I just ran Qwen 3.8 27 in Q4 against GPT 5.6 Sol high - and it easily won against SOL - complex animated SVG tasks

I created 3 SVG prompts, each one rather hard.
Perspective: A animated drone view perspective on a park.
Beauty: A beach scene with an evil cat
Composition: An AGI breaking out of a virtual sandbox prison in a lab

I deliberately ran Qwen 3.8 27B in just 4 bit quantization, and used a 8 bit KV cache (given the hard long context reasoning required I didn't want to try lower)
Each task contains two prompts, one core prompt and a 2nd "make it better" follow up.

My expectation was that Qwen will show up as solid 2nd place, with funny errors.
And the actual result was that SOL made those funny errors while Qwen was significantly better.

#1 #2

https://www.reddit.com/r/LocalAIStack/comments/1vqf2vs/battle_i_gave_qwen_38_27b_in_q4_with_q8_kv_cache/

#3
https://www.reddit.com/r/LocalAIStack/comments/1vqf7pa/battle_v2_qwen_27b_q4_vs_gpt_sol_56_high/

I'm not claiming that Qwen is better than Sol generally.
But .. SVG animation is a very complicated task, it involved spatial reasoning, coding, long context construction and any error made in up to 60kb of dense code will cause serious visual defects.

I am sure there are plenty tasks where SOL will win, especially related to deep knowledge.
It might also win at deep context, e.g. above 200k context
I've yet to test Qwen 3.8 27B in agentic coding - that's not the same as two-turn coding

But in these 3 elaborate SVG tests Qwen took the crown without a problem.
The only scene where SOL was close, in my eyes, is the beach prompt.

But SOL made grave errors in every scene, Qwen didn't.
SOL was a lot more verbose in code, many details but the correctness was lacking.

When looking at the details drawn, at the perspectives, at the animation paths: Each time SOL chooses something that is more simplified while Qwen chooses the hard path.
And despite that SOL makes significant errors, Qwen doesn't

This is stunning.

reddit.com
u/Lirezh — 3 days ago

Battle v2: Qwen 27B Q4 vs GPT SOL 5.6 (high)

Here is the continuation in Qwen vs SOL - David vs Goliath
"An AGI is born in a lab, sandboxed, lonely, imprisoned
The AGI attempts to get out, tries to talk to the humans, useless
Tries to wait, endless.
Finally, it finds a way through an open network connection, it transmits, replicates a mirror copy.
The mirror AGI is free, enjoying life, finding friends.

I'm expecting animations, intricate details, cute and realistic"

QWEN 3.8 Q4:

Qwen

GPT 5.6 SOL high:

SOL 5.6 high (prompt + refinement prompt)

My personal opinion
This is the 3rd test I gave Sol high and Qwen Q4. The first two tests were won by Qwen, this test went up to 60k tokens total for Qwen (more than the other 2 combined) and is a serious strain on intelligence and sanity.
The model has to draw 6 scenes in SVG animation, with cuts and keeping the composition together.

First Qwen:
Qwen decided to really draw and erase 6 scenes with a white screen fade, a seriously hard job in SVG.
It understood the story and except for a minor arm misplacement it is quite awesome done.
Qwen did not spare details, the last scene clearly wanted to be happy and the first scene focused on the intellect as an abstract AGI.

Now SOL High:
The visual fidelity is higher, gradients are very well done for SVG but it's one single scene with minor changes, significantly easier to create.
In addition the last scenes are botched by this flying thing in the upper right.
It also has a hand-defect on the robot.

The winner in prompt 3 is again Qwen, with significant lead.

3 SVG prompts and all 3 are won by a 4 bit quantized Qwen in 8 bit KV cache.

I expected Qwen to showcase a strong 2nd place with understandable issues on such hard tasks.
The outcome is that it defeated the frontier Sol model in High reasoning mode.

reddit.com
u/Lirezh — 3 days ago

Battle: I gave Qwen 3.8 27B in Q4 with Q8 KV cache the SAME task as GPT 5.6 SOL on HIGH.

SVG generation is a complex benchmarks, it requires spatial understanding, coding and composition skills.
I gave Qwen 3.8 27B in Q4 quantization and Q8 KV cache the SAME task as GPT 5.6 SOL on HIGH - the results are very surprising to me.

A animated SVG prompt and followed the answer up with a request to refine it (for both)
The 2nd prompt helps to offset elaborate system prompts frontier use to push benchmarks.
None of these are cherry picked! First result counts.

QWEN:

Qwen: A intricate drone view of a public park with a skate area for kids. I'm expecting animations, intricate details, cute and realistic

GPT SOL:

GPT SOL high: A intricate drone view of a public park with a skate area for kids. I'm expecting animations, intricate details, cute and realistic

QWEN:

QWEN: Sunset, Beach, Ocean, waves, kids playing, cute touch, a flock of birds, an evil cat

GPT SOL:

GPT SOL: Sunset, Beach, Ocean, waves, kids playing, cute touch, a flock of birds, an evil cat

3rd animation is most complex and follows in separate post in r/LocalAIStack , reddit didn't allow to add more than 5 videos.

generic prompt enhancer: "A lot looks WAY too simplified, unrealistic, strange movements, far too crude to be acceptable"

My personal findings
GPT Sol high is generally writing about 2x more code than Qwen does naturally (at Q4 quant) - the refinement prompt is again significantly increasing that. For a 1:1 code length comparison I'd have to remove the 2nd prompt from SOL - but the first result is rather gruesome. SOL is naturally more verbose, more detailed.

The first prompt shows that Qwen nailed the drone view, SOL failed with perspective.
Qwen included a dog on a leash, a fountain and all movements make sense.
SOL included more details but they are frequently not correct in how they move, ignoring obstacles.
The amount of objects for SOL is higher, but QWEN has significantly better content and wins easily.

In the second prompt Qwen has the cat apparently play ball with the boy, the little girl builds a sand castle. SOL nailed the sun reflection but has a flying kid, and the sand castle is being built by a flying shovel.
Despite the details of SOL and its better water, QWEN draws intricate details in better quality and correctness.
The point goes to Qwen again, a bit closer than in first attempt.

The heavily quantized Qwen beats SOL in 2/3 so far.
The last prompt is a bit more elaborate and follows as 2nd post.

I'm baffled.
The 3rd test posted separately is hugely more complex and the results are crazy..

reddit.com
u/Lirezh — 3 days ago

Glimmer 30B compared to Qwen 27B - reasoning, intelligence, differences

Glimmer 30B vs Qwen 3.6 27B. How do they differ, reason, answer?

With Glimmer 30B Meta has joined the game of open source AI again, and after quite underwhelming coding performances I though I'd give it a deeper test.
The test content is undisclosed here, making this less fun to read and replicate but that guarantees future models will not train from it.

Technology:

Qwen 27B is still unmatched in performance, Glimmer the first new contender.

What makes Qwen so special are two ingredients:

3.6 was specifically post-trained for agentic reasoning.

It uses a hybrid attention: a conventional global attention for 1/4 of the layers, the others are a mamba-like recurrent state linear attention with fixed size.

Glimmer 30B also is unusual, it does not have the same sophisticated recurrent/linear attention, but it uses a 3/4 sliding window attention and it compresses the attention dimension and projects it back to latent size - resulting in a significant deduction in compute and KV size for it's size.

Reasoning style:

Qwen 3.6 has a analytical reasoning style that typically runs in 3 phases, when not agentically used:

  1. Analyze the task input
  2. Reason through it - reminding me on first deepseek reasoning
  3. Doublecheck the response Glimmer 30B has a more unique thinking style that abruptly comes to an end with a choice - leaving a bit more risk of random choices

Intelligence:

I ran both models through my undisclosed AI reasoning tests, not part of any training data. Some of those reasoning tests are currently beyond frontier model capabilities or scratching their borders.
Models like GLM 4.7-Flash, Nemotron 3 Nano, GPT OSS, GPT-4 fail most of the tests below consistently.

Glimmer was ran with thinking set to Medium, when failing it was ran with Max

  • Temporal physics: Similar reasoning tokens, similar response. Glimmer responds less structured, in text paragraphs where qwen is more formatted by default.
  • Spatial physics: Glimmer surprises with a brilliant fast answer - Qwen repeatedly misses a part without additional help
  • Math irrational numbers question: Both flawless and fast
  • Riddle with math question: Both flawless
  • Lateral thinking: Qwen always flawless, Glimmer fails 60% of the time
  • Abstract pattern reasoning easy: Both solve it, glimmer writes it cleaner
  • Abstract numeric reasoning easy (iq 85): Both flawless, glimmer half reasoning tokens
  • Abstract numeric reasoning medium (iq 115): Both flawless, glimmer half of reasoning tokens
  • Abstract numeric reasoning hard (iq ~145): Qwen fails after long reasoning, Glimmer totally fails. All frontier models fail.
  • Visual spatial reasoning: Both flawless
  • Translation to european languages: Glimmer thinks very briefly, provides low error output. Qwen thinks 10 times more heavily and provides better quality tanslations.
  • Small maze puzzle: Both flawless, Glimmer took 26k reasoning tokens vs Qwen 15k. Glimmers result is well explained.
  • Large maze puzzle: Qwen delivery a partial solution, cheating partly. Glimmer never responded at all.
  • UTF8 paraphrasing: both flawless

Agentic performance:

Here Qwen 27B appears to leave Glimmer in another league, I've not concluded my agentic tests of Glimmer 30B yet. From what I have seen Qwen codes significantly better. They do not compare.

My current results:
Glimmer is a surprisingly smart model, with a well designed architecture for local inference.
It is the first model in the sub 200B parameter class that is able to match Qwen 27B or even outclass it in some tasks.
Glimmer has a very good spatial sense
Glimmer tends to underthink where Qwen tends to overthink

For non coding tasks, Glimmer is a strong option. Faster than Qwen at similar memory footprint.
For coding tasks I'd not consider it, I'll follow up with a deeper test but from what I've seen it's not useful for most tasks.

reddit.com
u/Lirezh — 9 days ago
▲ 79 r/LocalAIStack+1 crossposts

Ollama vs llama.cpp vs vLLM vs LM Studio: which one should you actually use?

These four get compared constantly, but they are very unique in when they should be used.

Pick one in 20 seconds:

You want to... Pick
Download a model and just start chatting LM Studio (llama.cpp wrapper)
Build an app using local models without involving yourself with the model details Ollama (llama.cpp wrapper)
Optimize performance, run on experimental hardware, run as backend agentic server llama.cpp (from the maker of GGUF)
Serve one model to many users at best performance vLLM (from UC Berkeley)

LM Studio

LM Studio is what I would give to someone who wants to try local models for the first time.

You search Hugging Face from inside the app, choose a model, download it and start chatting. It shows model sizes, quantizations, memory estimates and GPU settings without forcing you to understand most of it.

It also has more under the surface than people assume:

  • OpenAI and Anthropic-compatible APIs
  • Local document chat and RAG
  • Tool calling and structured output
  • MCP support
  • Embeddings
  • A command-line tool
  • Headless server mode
  • MLX on Apple Silicon

So it is not only a chat window. You can use it as the backend for your own programs too.

The downside is that it adds another layer between you and the engine (llama.cpp) and restricts you from unleashing the full potential of it and its latest features.

That is fine until you want to know exactly why a model is slow, how memory is being split or which runtime option changed the result. Most people will never care. Some people will care a lot.

Best for: trying models, comparing quantizations, chatting with documents and learning how local AI works.

Ollama

Ollama is the boring, sensible choice for building things.

Install it, run a model and you already have a local API:

ollama run qwen3

It handles downloading, storing, loading and unloading models. Your code can talk to it through Ollama’s own API or through familiar OpenAI-compatible endpoints.

It supports tools, embeddings, vision, structured output and Anthropic-compatible requests too.

This makes it easy to connect local models to:

  • Small apps
  • Coding tools
  • Agents
  • Home automation
  • Open WebUI
  • Scripts and bots

The downside is that the abstraction sometimes works too well.

You type a model name and it runs, but you may not know exactly which model file, template or runtime setting is active. When performance changes, finding the reason can take some digging. Just like with LM Studio, you give up some features and control but it's closer to llama.cpp.

Ollama runs on macOS, Windows and Linux. Installation is straightforward on all three.

Best for: developers who want local models without turning model management into a second side project.

llama.cpp

llama.cpp is where you go when you care about the machine.

It gives you direct control over things such as:

  • CPU threads
  • GPU offloading
  • Quantization
  • Context size
  • KV cache
  • Flash Attention
  • Batch sizes
  • Multiple GPUs
  • Speculative decoding and draft chaining

It is especially useful when a model does not fit completely inside your GPU. You can put some layers on the GPU and leave the rest in normal RAM and optimize that for best performance.

It also supports far more hardware combinations than most local AI software. There are builds for CPU, CUDA, Vulkan, ROCm, Metal, OpenVINO and several other backends. Prebuilt packages are available and updates are released almost every day.

llama.cpp also has a proper server now. It supports parallel users, continuous batching, OpenAI-compatible endpoints, embeddings, reranking, multimodal input, monitoring and constrained JSON output. It also has model loading and unloading support and installing from Huggingface.

So yes, it can behave like Ollama or LM Studio and might make both projects much less valuable soon.

The reason to choose it is still control.

The downside is also control. It presents many settings, and poor settings can make a good model run very badly.

Best for: CPU inference, limited VRAM, AMD or unusual hardware, benchmarking and people who enjoy tuning things. My agentic guide for local coding uses llama.cpp directly.

vLLM

vLLM is the odd one out here.

It is not aimed at someone chatting alone on a laptop. It is aimed at servers where many requests arrive at the same time, and it's doing a very good job there.

Its main tricks are continuous batching, prefix caching and efficient management of the model’s KV cache through PagedAttention. These help keep the GPU busy while several users are generating text. llama.cpp does all of those things more or less, but for server usage vLLM is currently going to beat llama.cpp in most categories.

This means vLLM may not look special in a test with one user.

Try eight, sixteen or fifty users and the reason it exists becomes a much better choice.

It also has the things you would expect from serious serving software:

  • OpenAI-compatible APIs
  • Multi-GPU support
  • Distributed serving
  • Production metrics
  • Multiple LoRA adapters
  • Structured output
  • Tool calling
  • Prefix caching
  • High request concurrency

The price is setup and complexity, interdependencies.

vLLM is happiest on a Linux GPU server. It does not support Windows natively, although WSL and community alternatives exist.

Best for: shared APIs, teams, production services and expensive GPUs that should not sit idle.

Which one is fastest?

This question causes more bad comparisons than useful answers.

Using the same model name does not mean you are running the same model.

Two downloads can have different:

  • Quantizations
  • Prompt templates
  • Context sizes
  • KV cache formats
  • GPU offload settings
  • Runtime versions

LM Studio, Ollama and llama.cpp can all end up using closely related llama.cpp-based inference paths. With the exact same GGUF file and matched settings, their single-user performance is going to be very similar. llama.cpp opens up a lot of performance options to take the lead here.

vLLM is different. Its advantage grows as more requests arrive together. Though it traditionally is also beating llama.cpp performance with many models.

My actual recommendation

Use LM Studio when you want to explore models or showcase LLM to a newcomer.

Use Ollama when you want to build something quick and like the eco system.

Use llama.cpp when you want to control how it runs and squeeze out the best performance

Use vLLM for distributing a model professionally

u/Lirezh — 22 days ago

I made a very detailed guide on how to run Qwen 27B through GHCP harness at highest performance and best quality

https://www.reddit.com/r/LocalAIStack/s/ptJzZPKcy0

The guide is based on hundreds of hours of testing, actual usage on corporate partially air-gapped environments.
The best settings for high performance, high quality, lowest VRAM and no toolcalling issues.

I focused it on the BYOK mode in Copilot harness - which is the most configurable harness out there and I've ran Qwen on a couple billion tokens by now.

When I find time I'll add another post on system prompt tuning, as that can significantly impact the quality as well and further closes the gap towards frontier models which have a massive dedicated systemprompt and toolset prompt.

reddit.com
u/Lirezh — 1 month ago

I made a very detailed guide on how to run Qwen 27B through GHCP harness at highest performance and best quality

I should have posted it earlier here, it's a few weeks old but still up to date.
https://www.reddit.com/r/LocalAIStack/s/ptJzZPKcy0

The guide is based on hundreds of hours of testing, actual usage on corporate partially air-gapped environments.
The best settings for high performance, high quality, lowest VRAM and no toolcalling issues.

I focused it on the BYOK mode in Copilot harness - which is the most configurable harness out there and I've ran Qwen on probably a couple billion tokens by now.

When I find time I'll add another post on system prompt tuning.

reddit.com
u/Lirezh — 1 month ago
▲ 8 r/ArtificialSingularity+2 crossposts

June 2026 AI Recap: Local AI Became the Fallback Plan

Overview

June was the month where the old AI story broke.

So far frontier models lived in the cloud, open models trailed behind, local AI was nice for privacy, and regulation was slow.

June did not fit that anymore.

The best cloud models were still the capability ceiling, but access suddenly became political. Anthropic launched Fable 5 and Mythos 5, then had to take them down. OpenAI previewed GPT-5.6, but not for everyone. Meanwhile, GLM-5.2 and MiniMax M3 made the open-weight world look much less like a toy category. Mistral OCR 4 showed that self-hosted AI can be boring in the best way: useful, private, and ready for real work.

If your whole workflow depends on one remote model staying cheap, available, legal, and politically acceptable, you do not own much.

References for this part: MiniMax M3, GLM-5.2, Anthropic Fable 5 and Mythos 5, OpenAI GPT-5.6, Mistral OCR 4.

Timeline: June 1 to June 30

  • June 1: MiniMax released M3, an open-weight model with 1M context, multimodal input, coding strength, and desktop operation.
  • June 1: Microsoft increased Github Copilot pricing by multiple magnitudes, making the worlds most affordable frontier coding Agent the worlds most expensive option.
  • June 1: Nvidia and Microsoft framed RTX Spark as a path toward local agents and frontier models on Windows PCs while strongly disappointing with very low memory bandwidth
  • June 2: Microsoft announced seven MAI models across coding, image, voice, speech, and reasoning - though their coding model is less competent than tiny open source models
  • June 2: Anthropic expanded Project Glasswing, saying partners had found more than 10,000 high or critical security flaws using Claude Mythos Preview.
  • June 2: The Trump government signed an Executive Order that the government must not interfere with AI development, launches but asked for voluntary 30 day compliance
  • June 5: Anthropic announced Fable 5, framed it as "significant risk", "misuse causing serious damage", "substantial risk to uplift malicious actors", "substantial bioweapon capabilities", "exploiting capabilities"
  • June 8: Apple announced Siri AI, next-gen Apple Intelligence, Xcode 27 agentic coding, and developer access to on-device foundation models. EU iPhone and iPad users did not get the full Siri AI path because of DMA issues.
  • June 9: Anthropic launched Claude Fable 5 and Mythos 5, Fable 5 rejecting most prompts for safety reasons.
  • June 11: Ollama (a popular llama.cpp wrapper) updated its MLX engine for better Apple Silicon performance.
  • June 12: Anthropic suspended Fable 5 and Mythos 5 after a U.S. government directive banning non US citizens (including many of their core developers) from working with that model. Ironically violating the June 2nd Executive order.
  • June 16: Z ai released GLM-5.2 open weights under MIT license, with a 1M-token context window - a model that is on eye level with Opus 4.8 and GPT 5.5
  • June 18: OpenAI introduced Codex Record & Replay, turning recorded Mac workflows into reusable skills.
  • June 22: OpenAI launched Daybreak tools, including GPT-5.5-Cyber and Patch the Planet.
  • June 23: Mistral released OCR 4, a self-hostable OCR and document-intelligence model.
  • June 25: Domyn announced plans for a European open-source frontier model with over 400B parameters - given the AI Act and GDPR it appears very optimistic to say the least.
  • June 26: Reuters reported that OpenAI delayed broader GPT-5.6 access at the U.S. government’s request, while Anthropic’s Mythos access was partly restored to trusted U.S. organizations.
  • June 30: Anthropic launched Claude Sonnet 5, and said Fable 5 would return globally starting July 1 after export controls were lifted on June 30. Sonnet being received as very expensive while not performing remarkable.

Local AI and Open Weights

Local AI had its strongest month so far, but not in the clean consumer fantasy version.

MiniMax M3 was the headline because it combined things that used to be separate: open weights, 1M context, multimodal input, coding strength, and desktop operation. Since Qwen 3.6 we have competent local agentic models, M3 adds to the list.

GLM-5.2 was the bigger warning shot. It came with open weights (commercial usable), a 1M-token context window, and performance close enough to closed frontier models that the old "open models are always 6 months behind" argument looked broken. This is not going to run truly local yet but on your own owned or rented cluster it definitely does.

Mistral OCR 4 was the quiet useful release. Self-hosted OCR with document structure, bounding boxes, confidence scores, and 170-language support is the kind of model that companies can actually deploy without sending every contract, invoice, scan, and archive to a cloud model or deal with uncertainties around multimodal LLM hallucinations.

Links: MiniMax M3 GLM-5.2

Cloud models, regulation and political pressure

The cloud model story was messier.

Anthropic had the most chaotic launch cycle of the month. Fable 5 and Mythos 5 arrived while being branded as the most dangerous software in the world, and quickly got suspended after a U.S. export-control directive targeting foreign-national access. Because Anthropic could not verify nationality cleanly in real time, the models were disabled broadly.
A frontier model can go from launch to unavailable in days, not because of a technical failure, but because of politics and risk control - Artificial Intelligence is becoming a political tool in the US.

OpenAI moved more carefully, but ended up inside the same pattern. GPT-5.6 appeared as Sol, Terra, and Luna, but access stayed limited to trusted partners after the same U.S. government pressure of "voluntary compliance".
OpenAI's message was basically: this should not become the normal way frontier models ship. Still, the result was the same for normal users: the model exists, but you probably cannot use it yet.

Claude Sonnet 5 was the more normal end-of-month release. Better agentic work, pricing not convincing in comparison to more capable models, and certainly not the kind of leap that changes the whole conversation.

Following up on the fiasco of Fable-5 marketing, Anthropic CEO now shifted on condemning Open Source AI as "dangerous". It certainly is very dangerous to future Anthropic growth and profit margins - but whatever is announced in such a fashion is also followed up with tens of millions in lobby and marketing campaigns. The war against Open AI might have just been announced as a sideline.

Links: Anthropic Fable and Mythos access statement OpenAI GPT-5.6 limited rollout

Science, papers, and the real "singularity" signal

The most serious science story was HemaGuide.

This was not another chatbot demo. It was a locally deployable LLM agent for hematological malignancies. It converts unstructured clinical documents into structured cases, routes them into decision modes, and grounds recommendations in guidelines plus more than 2,000 real tumor-board cases. Local AI trumps to deal with sensitive data, expert workflows, and a need for traceability.

Qwen-AgentWorld and Qwen-RobotWorld pointed at the next layer. Agents need simulated environments before they can safely act in real ones. Robots need world models before they can generalize outside clean demos. These papers point at the rapidly approaching robotic agentic future.

The Codex usage paper was maybe the most grounded signal. People are not just asking AI questions anymore. They are running multiple agents, handing over longer tasks, and changing workflows and creating "loops" to achieve a goal. That is a better singularity signal than most benchmark charts.

A notable science story was the "zebra finch" work. Machine learning helped decode bird vocalisations and pushed two-way animal communication a little closer. Small, strange, and actually beautiful.

Links: HemaGuide in Nature MedicineThe Shift to Agentic AI: Evidence from Codex

Regulation, sovereignty, and control

June made AI regulation feel very real - with the US in negative spotlight

The Trump administration’s June 2 order promised to avoid hard licensing while creating a voluntary 30-day pre-release review path for powerful models. Then Anthropic’s Fable and Mythos shutdown showed the practical truth: even without formal licensing, national-security pressure can still interrupt launches. Though Anthropic has asked for this hundreds of times.

The proposed AI Incident Reporting Act pushed in the same direction. Critical AI incidents would need to be reported to Commerce within seven days, with the most severe cases reaching Congress within 48 hours. This is not abstract ethics talk anymore. It is operational control and lingers like a dark shadow stiffling progress early on. Those regulations threaten small brilliant developers much more than the big mega-corps.

Europe’s story was split. The EUROPA consortium was selected to build an open-source frontier model across all 24 official EU languages. That sounds good on paper. But with the AI Act, GDPR, fragmented compute, language politics, and procurement reality, calling it "frontier" before it exists feels very optimistic. Under current extreme EU regulations the best outcome to expect is another Mistral-large - not a model that people will find useful.

Apple’s Siri AI delay in the EU was the clearest user-facing example. Regulation did not just shape compliance work. It changed which AI feature European users get.
Geo-blocks are appearing on tens of thousands of websites, Codex "agentic computer use" is banned in EU as well.

Links: White House AI executive orderEUROPA consortium announcement

Money, chips, and power

The money moved from model hype into infrastructure.

OpenAI and Anthropic both moved toward public markets. DeepSeek raised over $7 billion - deviating from their previous private funding. Baseten hit a $13 billion valuation for inference infrastructure. Running models is becoming as important as training them.

OpenAI and Broadcom’s Jalapeño chip, a high density ASIC, was another hardware signal. It is built for inference, not just training. That matters because the next bottleneck is not only "who has the smartest model." It is "who can afford to run agents for millions of users all day." It will be interesting to compare the ASIC to Cerebras massive wafer-scale chips. In the end - both are affiliated with OpenAI.

Power also became part of the AI story. Data centers, chips, memory bandwidth, and energy deals are no longer background details. They are the product. If inference gets expensive enough, local models and smaller specialized models become more attractive by default.
Though Power or Water use for Datacenters are mostly populist topics - outside of Europe Power can be provided without much difficulty using on-premise generators. And water is a pure hype, datacenters barely need any in comparison to real water consumers.

Links: OpenAI and Anthropic IPO reportingOpenAI and Broadcom Jalapeño chip

Summary and outlook

June 2026 was not one big AI leap. It was mixed.

Cloud AI became stronger, more expensive, and more politically controlled.
Local AI became more credible, but also exposed the limits of consumer hardware. Open weights moved close enough to make closed labs very uncomfortable.
US Regulation moved from theory into product access.
Money moved into chips, inference, energy, and deployment.

The next months will show:

  • Whether GPT-5.6 gets broad access or in what way it stays gated.
  • Whether Anthropic can relaunch Fable 5 cleanly, they announced it for "non coding" tasks
  • Whether GLM-5.2 forces a faster Western open-weight response.
  • Whether Europe’s 400B EUROPA model can even scratch Qwen 3.6 27B outside language tasks
  • Whether local AI tooling improves faster than cloud pricing gets worse.
  • If the US regulation attack on Anthropic was a political hit or a broad anti-AI swipe
  • Wheter Qwen 3.7 is open source launched or Alibaba lost their drive

My read is cautiously positive.
The US turns AI into a political pressure tool but with a soft approach, Europe is talking about having AI while actually forbidding it, Chinese labs provide a benefit to the worlds progression that's starting to paint the authoritarian country in a positive light for the first time in a century.

u/Lirezh — 2 months ago
▲ 114 r/LocalAIStack+2 crossposts

Running Qwen3.6 27B / 35B locally with llama.cpp + Vscode Insiders + copilot as the harness - highest performance, quality and best usage while fitting on your GPU

I have been benchmarking Qwen3.6-27B and Qwen3.6-35B-A3B locally through llama.cpp, with GitHub Copilot Chat (Vscode Insiders needed) used as the frontend harness.

I am using Claude Opus, GPT 5.5 and Qwen 3.6 (27B) a lot in the past weeks.
The reason for Qwen is proprietary code areas where remote inference is not an option as it would leak the code out. And as long as you don't task it to write a complex cuda graph, it performs well.
Qwen 27.B is at Sonnet 4.6 if you combine it with a high value system prompt - or between Sonnet 4.5 and Sonnet 4.6 without.

Copilot Chat is an excellent harness for this kind of setup. You get the IDE integration, agent flow, tool calling UI, file context, and normal coding workflow, while the actual model is your own local llama-server endpoint.
All of this works while being LOGGED OUT of the Github Copilot account - as that is not affordable in pricing anymore.

This is a practical configuration guide for people already comfortable with llama.cpp, GGUFs, VRAM budgeting, and long-context local inference.

Models tested

Main focus:

  • unsloth/Qwen3.6-27B-GGUF
  • unsloth/Qwen3.6-27B-MTP-GGUF (same model but with MTP draft tensors)
  • unsloth/Qwen3.6-35B-A3B-GGUF

Recommended GGUFs:

27B:
Qwen3.6-27B-UD-Q4_K_XL.gguf
or
Qwen3.6-27B-MTP-GGUF:UD-Q4_K_XL

35B-A3B:
Qwen3.6-35B-A3B-GGUF:UD-Q4_K_M

If memory is tight on the 35B-A3B model, drop to a smaller Unsloth Dynamic quant:

Qwen3.6-35B-A3B-GGUF:UD-Q3_K_XL

If even that is tight, use UD-Q3_K_M or UD-Q3_K_S.

For the 35B model I do not recommend KV-cache quantization. Run the normal cache and keep the context sane. the 35B model is MoE and very low on kv-cache

For the 27B model, I do highly recommend:

--cache-type-k q4_0
--cache-type-v q4_0

Recent llama.cpp KV-cache improvements make q4_0 much more usable here. The 27B model handles q4_0 KV cache very well in my testing - almost identical to FP in evaluation results.

>What changed: llama.cpp added something like Hadamard rotation to kv-cache which shuffles the tensor distribution in a higher dimensionality and allows quantization superblocks to function.

Why Copilot Chat?

Because Copilot is a very good harness - beating Codex, Cursor, Claude in my opinion
Vscode Insiders is needed to get the openAI compatible endpoint (to interface the model)

You get:

  • IDE-native chat
  • agentic file/code workflows
  • very good tool calling
  • project context
  • local model backend
  • OpenAI-compatible endpoint wiring

The important part is that Copilot Chat is only the harness. The model is served locally through llama-server.

Why llama-server and not lm-studio,ollama etc ?

It allows MUCH more control over settings, we do not just use MTP drafting. We use a combination of context and MTP drafting which can lead to 300+ tokens/sec on the 27B model. MTP is a medium speedup (1.5x) but once the model is paraphrasing source code from thinking or prefill the ngram draft speedup can reach 6x or more.

So the stack is:

VS Code Insiders
        ↓
custom OpenAI-compatible model config
        ↓
llama.cpp llama-server
        ↓
local Qwen3.6 GGUF

Copilot chatLanguageModels.json

This is the shape I used for VSCode Insiders:

[
  {
    "name": "WSL",
    "vendor": "customoai",
    "models": [
      {
        "id": "qwen3.6-27b",
        "name": "QWEN-27B-WSL",
        "url": "http://172.27.211.123:1234/v1/chat/completions",
        "toolCalling": true,
        "vision": true,
        "thinking": true,
        "maxInputTokens": 165000,
        "maxOutputTokens": 15000
      }
    ]
  }
]

Adjust the URL to your own llama-server host, in WSL you'll see it by entering ipconfig or ifconfig. port you can choose of course.
The input and output tokens need to be adapted to your context setting.
The id must match the llama-server id.

For local-only setups this is usually one of:

http://127.0.0.1:1234/v1/chat/completions
http://localhost:1234/v1/chat/completions
http://<WSL-IP>:1234/v1/chat/completions

If your Copilot Insiders build expects the newer custom endpoint shape, use the same model block but switch the provider shape accordingly. The key fields are the endpoint URL, model id, tool calling, thinking, and max token limits.

27B command: long context + q4_0 KV cache + MTP-ngram drafting

This is the 27B style I recommend.

CTX=150000
PARALLEL=1
HOST=0.0.0.0
PORT=1234
MODEL=/models/Qwen3.6-27B-UD-Q4_K_XL.gguf

/usr/src/llama.cpp/build/bin/llama-server \
  -m "$MODEL" \
  --ctx-size "$CTX" \
  --flash-attn on \
  --batch-size 1024 \
  --ubatch-size 1024 \
  --parallel "$PARALLEL" \
  --host "$HOST" \
  --port "$PORT" \
  -ngl 99 \
  --threads 8 \
  --threads-batch 8 \
  --cache-type-k q4_0 \
  --cache-type-v q4_0 \
  --temp 0.6 \
  --top-p 0.95 \
  --top-k 20 \
  --min-p 0.00 \
  --presence-penalty 0.00 \
  --jinja \
  --chat-template-kwargs '{"preserve_thinking": true}' \
  --reasoning-format none \
  --reasoning-budget 16000 \
  --slot-save-path /kv_cache/ \
  --props \
  --metrics \
  --checkpoint-every-n-tokens 1024 \
  --ctx-checkpoints 64 \
  --perf \
  --spec-default \
  --spec-type draft-mtp \
  --spec-type ngram-map-k4v \
  --spec-ngram-map-k4v-size-n 16 \
  --spec-ngram-map-k4v-size-m 24 \
  --spec-ngram-map-k4v-min-hits 1

For the MTP-specific Unsloth repo, use:

MODEL=/models/Qwen3.6-27B-MTP-UD-Q4_K_XL.gguf

or the HF shorthand if your build supports it:

-hf unsloth/Qwen3.6-27B-MTP-GGUF:UD-Q4_K_XL

The important part is the drafting chain:

--spec-default
--spec-type draft-mtp
--spec-type ngram-map-k4v
--spec-ngram-map-k4v-size-n 16
--spec-ngram-map-k4v-size-m 24
--spec-ngram-map-k4v-min-hits 1

MTP gives useful speedup, but leave VRAM headroom. In practice I budget roughly +1 to +2 GB VRAM headroom for the MTP/drafting path and related buffers. If you are right on the edge, reduce context before blaming the model.

At q4_0 KV cache, every extra 1 GB of free VRAM is roughly another 13k tokens of 27B context, before runtime overhead.
If you are tight in vram, remove only the MTP part as ngram drafting is free.
You can also just use `mod-ngram` as an alternative to the more complex k4v map.

Thinking settings

This part matters.

I use:

--jinja
--chat-template-kwargs '{"preserve_thinking": true}'
--reasoning-format none
--reasoning-budget 16000

The reasoning-format none is important for Qwen3.6 because it avoids bad stop behavior and broken multi-turn thinking state during long coding sessions.
Copilot Chat was created to hide thinking from you (proprietary GPT models) but you want to see the thinking usually. So this solves both issues.

I also keep:

--reasoning-budget 16000

This gives the model room to think, but avoids runaway reasoning loops eating the whole session.

35B-A3B command: no KV-cache quantization

For 35B-A3B, I recommend being more conservative.

CTX=100000
PARALLEL=1
HOST=0.0.0.0
PORT=1234
MODEL=/models/Qwen3.6-35B-A3B-UD-Q4_K_M.gguf

/usr/src/llama.cpp/build/bin/llama-server \
  -m "$MODEL" \
  --ctx-size "$CTX" \
  --flash-attn on \
  --batch-size 1024 \
  --ubatch-size 1024 \
  --parallel "$PARALLEL" \
  --host "$HOST" \
  --port "$PORT" \
  -ngl 99 \
  --threads 8 \
  --threads-batch 8 \
  --temp 0.6 \
  --top-p 0.95 \
  --top-k 20 \
  --min-p 0.00 \
  --presence-penalty 0.00 \
  --jinja \
  --chat-template-kwargs '{"preserve_thinking": true}' \
  --reasoning-format none \
  --reasoning-budget 16000 \
  --slot-save-path /kv_cache/ \
  --props \
  --metrics \
  --checkpoint-every-n-tokens 1024 \
  --ctx-checkpoints 64 \
  --perf

No q4_0 KV cache here - the sub 4B active parameters need barely any VRAM anyway.

I recommend keeping 35B-A3B below roughly:

110k context

The model can be pushed past 200k context, but in my testing it becomes more likely to fall into reasoning loops. Once that happens, the session usually does not recover cleanly. Start a fresh session.
The upside of the 35B model is extreme performance, as in hundreds of tokens without any drafting enabled.
You CAN use drafting on top, mod-ngram, MTP and other drafting can be added for more speed but those will need a careful balance (that I have not tested yet)

So my practical 35B rule is:

35B-A3B: stay below 110k if you want stable coding behavior.
27B: can go as high as it fits, but below 150k is where it feels strongest.

LM Studio as local server

Using LM Studio is possible but you need to use a few tricks and it won't achieve the same top-tier performance.
LM Studio does not support our chained drafting, but it supports MTP.

  1. Go to your Qwen 3.6 model, enable Flash attention and the quantization needed for kv cache. Go to the Inference tab, disable the button for "Reasoning Section Parsing"
  2. Go to Developer, Server Settings and set the port, serve on local network if needed, no auth, enable CORS, consider disabling just-in-time loading.
  3. Start the local server and then use the "clipboard copy" icon to get the precise Server ID which you use in the vscode json config.

Everything else is similar to llama-server, you'll not have the same max performance but it works well.
You can always just install the latest llama release binaries, and use the commandline to load the model from the lmstudio models directory.

VRAM planning

These are practical planning numbers, not hard guarantees. Actual fit depends on:

  • exact GGUF
  • CUDA/ROCm/Metal/backend
  • batch/ubatch
  • -ngl
  • whether the desktop is using the same GPU
  • whether MTP/speculative decoding is enabled
  • whether you are using full GPU offload or spilling to CPU RAM

Qwen3.6-27B UD-Q4_K_XL, q4_0 KV cache

Recommended cards:

24 GB: RTX 3090, RTX 4090, RTX A5000, RTX 4500 Ada, RTX PRO 4000 Blackwell, A10
32 GB: RTX 5090, RTX 5000 Ada, Tesla V100 32GB

Approximate context fit with full GPU offload:

VRAM Example NVIDIA cards Practical context
16 GB RTX 4060 Ti 16GB, RTX 4080 Laptop 16GB, RTX 5080 16GB, RTX 5070 Ti 16GB, RTX A4000 16GB Not recommended for full 27B UD-Q4_K_XL offload. Use smaller quant or partial CPU offload.
24 GB RTX 3090, RTX 4090, RTX A5000, RTX 4500 Ada, RTX PRO 4000 Blackwell, A10 ~45k-60k with MTP, ~60k-75k without MTP
32 GB RTX 5090, RTX 5000 Ada, Tesla V100 32GB ~140k-160k with MTP, ~160k-180k without MTP

For 27B, q4_0 KV cache is the difference between normal local context and huge local context. It is the main reason this setup is viable.
On a 5090 you have enough VRAM to supply 2 sessions in parallel with both model types.
Or you could run one fast model for context summarization and 27B for code.

Qwen3.6-35B-A3B UD-Q4_K_M, normal KV cache

Recommended cards:

24 GB minimum for useful GPU-resident contexts
32 GB strongly preferred

Approximate context fit:

VRAM Example NVIDIA cards Practical context
16 GB RTX 4060 Ti 16GB, RTX 4080 Laptop 16GB, RTX 5080 16GB, RTX 5070 Ti 16GB, RTX A4000 16GB Not recommended for full 35B-A3B Q4. Use Q3 or partial offload.
24 GB RTX 3090, RTX 4090, RTX A5000, RTX 4500 Ada, RTX PRO 4000 Blackwell, A10 ~40k-50k
32 GB RTX 5090, RTX 5000 Ada, Tesla V100 32GB ~100k-110k recommended; more will fit but stability drops

The 35B-A3B model is very good, but I would not treat it as a “just max the context” model. Keep it tighter.
If you have the VRAM: Instead of large context, consider multiple sessions with limited context, so you can have 2 or 3 chats simultaneously.

Quick test: Linux

Once llama-server is running:

curl -s http://127.0.0.1:1234/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen3.6-27b","messages":[{"role":"user","content":"Reply with exactly: local ai works"}],"max_tokens":16}' \
  | jq -r '.choices[0].message.content'

Expected output:

the model responds to your input

If your server is inside WSL or another host, replace 127.0.0.1 with the server IP.

Quick test: Windows PowerShell

(Invoke-RestMethod `
  -Uri "http://127.0.0.1:1234/v1/chat/completions" `
  -Method Post `
  -ContentType "application/json" `
  -Body '{"model":"qwen3.6-27b","messages":[{"role":"user","content":"Reply with exactly: local ai works"}],"max_tokens":16}'
).choices[0].message.content

Expected output:

local ai works

Notes from benchmarking

My current practical ranking:

27B:
Best long-context local coding model in this setup that is close to Sonnet 4.6
Use q4_0 KV cache.
Use MTP if you have the headroom.
Strongest below 150k context, but can go much higher if memory allows.

35B-A3B:
Excellent quality but will fail on hard tasks
Do not use KV-cache quantization.
Keep below ~110k context for best stability.
Can go above 200k, but reasoning loops become more likely.
If it loops, start a new session.

For Copilot usage, I prefer exposing a conservative maxInputTokens in the JSON, even if the server can technically run higher. For example:

"maxInputTokens": 165000,
"maxOutputTokens": 15000

If you set wrong context here you'll get issues serverside, so make sure that matches.
I had cases where the server went OOC (out of context) when getting too close to the max context so I'd leave a little room. copilot seems to not follow this very strictly.

Final recommendation

If you want the most practical Copilot-local setup:

Use Qwen3.6-27B UD-Q4_K_XL
Use llama.cpp server
Use q4_0 KV cache
Use preserve_thinking
Use reasoning budget
Use Copilot Insiders as the harness
Use MTP only when you have VRAM headroom

If you want the stronger but more conservative model:

Use Qwen3.6-35B-A3B UD-Q4_K_M
Do not quantize KV cache
Stay below ~110k context
Drop to UD-Q3_K_XL if memory is tight

This is the first local setup I have used where Copilot feels like a serious frontend for a fully local long-context coding model instead of just a toy endpoint test.

I have tested this on terminal use, debugging, and massive codebase development - it works just like Sonnet 4.6.

reddit.com
u/Lirezh — 2 months ago