Kinda need your help
How can i increase my output current stack is llama.cpp and flash attn enabled ive tried to optimise it as muc as i can but becof the ctx it keep offloading to ram and in getting 2-4 tok/s :/
How can i increase my output current stack is llama.cpp and flash attn enabled ive tried to optimise it as muc as i can but becof the ctx it keep offloading to ram and in getting 2-4 tok/s :/
This is the first model after 3.6 which was even able to register the Sharingan MCP which has basically most of the popular cyber security tools and take this is built based on Johnrizzo1 reverser and Kali linux. This was the first time running this on qwen code 3.827B and it was able to use all of the tools and also able to use it efficiently
Someone recommended to use the official harness idk why i didnt think of that earlier , the results were crazy good i added my own 91 pen test tools and boy this model is soo good like it genuinely beats any other ive tested by a MARGINNMMMM.