▲ 141 r/unsloth

Unsloth finally has Dark mode!

It's still a WIP but after 2 years of our website, we never had a darkmode except for our docs. Well after lots and lots of feedback, we finally have a darkmode. I can't believe it took this long but hey, it's here and hope you don't avoid going to our website anymore 😅

We'll also be doing a whole rebrand and also redesign of website in like a month or two or so, so what you're seeing is temporary :)

Thanks so much and if you have any feedback, be my guest!

Website: https://unsloth.ai/

u/yoracale — 3 days ago
▲ 126 r/unsloth

Unsloth is trending at #3 on GitHub today!

Hey guys just yesterday we were trending at #5 and now #3. Thanks so much guys! We couldn't believe it and it's literally all thanks to you guys! 🙏😭

We're currently adding many features and improvements especially regarding speed for UI/UX since some of y'all said it was laggy and also many extra new features including auto compaction etc

Feel free to star us on GitHub: https://github.com/unslothai/unsloth

u/yoracale — 4 days ago
▲ 304 r/unsloth

Unsloth is trending at #5 on GitHub!

Hey guys thanks so much for the support for Unsloth Desktop, we're currently trending at #5 overall on GitHub and the #2 overall package for Python! ❤️🦥

Feel free to Star us on GitHub: https://github.com/unslothai/unsloth

We are also adding a lot of new updates. If you have any feature requests or feedback, it would be best to make a GitHub issue, but also feel free to make a comment here.

Thanks so much once again for your support!

u/yoracale — 6 days ago
▲ 213 r/unsloth

Share your results from Qwen3.8-27B! 🔥

Here's 4-bit Qwen3.8-27B GGUF generating an interactive volcano simulation with geology, thermodynamics, fluid flow, buoyancy/drag, projectiles, cooling, groundwater interaction, and environmental effects in Unsloth Desktop.

What’s your favorite prompt to test a new model with? Full prompt in comment.

u/yoracale — 6 days ago
▲ 94 r/unsloth

New Unsloth Desktop Release v0.1.800-beta

Hey guys, we got some updates since we last updated Unsloth Desktop.

First of all we wanted to thank you for the support for our launch. It really means the world to us and it was amazing seeing you enjoy it as much as we loved making it.

Secondly, we took your feedback seriously. We dmed and messaged some of you individually for feedback and took it in. Currently we have 60+ new PRs which havent been merged yet but for now, we've addressed the most critical issues:

Highlights

  • Extra llama-server arguments allowed + custom VRAM toggle
  • External provider has tool calling + tool support + login with Codex
  • Fast FP8 10x faster MiniMax-H3 inference (3 minutes vs 30)
  • 10% faster inference for GGUFs + Bypass permissions fixed

Chat + tools

  • Connected AI providers can use their own Search or Unsloth Desktop's built-in Search and tools. Tool results are passed back to the model so it can continue multi-step tasks.
  • Sign in with a Codex subscription and use Codex tools inside Chat.
  • Chat shows live prompt and generation speeds, while long streaming replies use much less CPU.
  • Chat settings stay with the conversation across remote sessions.
  • Paste a YouTube link to attach its transcript, including the title, channel, duration, link, and caption language.
  • Save a full chat or reply into your project's sources while keeping its reasoning, tool calls, and citations.

MiniMax-H3

  • MiniMax-H3 can run on smaller supported GPUs by splitting large model parts into pieces that fit.
  • The model picker now hides H3 options that the current hardware cannot run instead of letting them fail after selection.
  • H3 options are labelled Fast FP8 or Slow, making the large speed difference clear before downloading.

Performance + hardware

  • Inference is up to 10% faster in supported cases, with lower VRAM use and a tunable memory limit.
  • Idle image and video models can optionally unload to free VRAM for Chat or Training.
  • Added better support for AMD RDNA 3, RDNA 4, and Strix Halo systems. VRAM checks no longer reserve extra GPU memory.
  • Multi-GPU ROCm device matching is safer.
  • Macs now choose context size from the memory that is actually free.
  • RAG document indexing uses the CPU by default, so it no longer leaves a large GPU memory block reserved.
  • Fixed GGUF image detection when choosing a model for the API.

Custom llama.cpp arguments

  • Model settings now include an Extra Arguments box for custom llama-server flags.
  • Unsloth checks flags against the installed build and saves valid ones per model for normal, startup, and API loads. Flags that could break model loading or app security are rejected with a clear message.

You can view the full release here: https://github.com/unslothai/unsloth/releases/tag/v0.1.800-beta

github.com
u/yoracale — 6 days ago
▲ 963 r/unsloth

Qwen3.8-27B is out now!

Qwen3.8-27B can now be run locally! ✨ Run on 17GB RAM via Unsloth Dynamic GGUFs and Unsloth Desktop. Qwen3.8-27B is by far the strongest model for its size.

Guide: https://unsloth.ai/docs/models/qwen3.8

We uploaded NVFP4 quants! GGUF: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF

We also did a new release for Unsloth Desktop with many new features and performance increases!

Unsloth Desktop Highlights:

  • External provider tool calling + Codex login
  • Extra llama-server arguments + custom VRAM controls
  • 10x faster MiniMax-H3 FP8 inference (3 min vs 30 min)
  • 10% faster GGUF inference + fixed bypass permissions

Unsloth Desktop Performance Improvements:

  • Up to 10% faster inference with lower VRAM use and adjustable memory limits.
  • Idle image/video models can unload to free VRAM for Chat or Training.
  • Improved support for AMD RDNA 3, RDNA 4, Strix Halo, and multi-GPU ROCm setups.
  • Added an Extra Arguments box for custom llama-server flags.
  • Macs now select context size based on available memory.
  • RAG indexing now uses CPU by default to reduce GPU memory usage.

Thank you!

u/yoracale — 7 days ago
▲ 60 r/unsloth+1 crossposts

1BIT Qwen 3.8 2.4T a95b (unsloth iQ1_S) (MEDIUM Reasoning)

Processing img az99qopcg8jh1...

So same as my prior post 1bit test... although this 1bit is a bit interesting you can read on unlsoth blog https://unsloth.ai/docs/models/qwen3.8 508 gigs being used I am using unsloth studio on the mac ultra 512. Im getting ~50pp and ~9.6 tgen drop to 5 at 50k tokens

Yes My computer i screaming at me. The things Im putting it through :'D

Flight Sim: the code for its is kind cool it is rendering gravity and physics rather then just a vector controls and force based dynamics which is more complex but doesn't show in the gameplay. I really like it even comparing to GLM 5.2 Not because of visuals because that we can just add later but because the controls feel kinda nice. (might be just the new toy bias). There is a bug with the water that its not rendering but Im just keeping whatever it output without telling it to fix it. I did run out of token head room 3 times, 1st i set 32k limit that didnt even finish the though process. then 65k it output the game and half output it. 96k it finished and tested to make sure it ran fixed a gravity bug where upward force was higher than gravity and accelerated upward. this is on Medium reasoning effort not xhigh. I dont have the token space for xhigh.

SVG Tests: All Medium Reasoning

Basic >_>' I know its dumb to test a 2.4T model. But Ill do the 3d renders later. just remember its almost like 1.5bit average

Panda: " generate an svg of a panda sitting having a picnic in Japan, make it beautiful"

https://preview.redd.it/szp3bxxrp8jh1.png?width=1660&format=png&auto=webp&s=7145485c1032cc8a1e6d4c0faf4996ff7de7b198

PS4 Controller: It obviously isn't fully right but its very coherent and things are in the right places. This is a difficult test for models

https://preview.redd.it/as31dw5yi8jh1.png?width=1182&format=png&auto=webp&s=5ba2137c8894789a988b69407eabe2f7abde0a5d

Capybara: "generate a capybara having a yuzu in an onsen as beautiful as possible" this is some other world drawing... it used python to generate the noise and other things like snow and textures and then generated the svg in the same tool call

https://preview.redd.it/9zswyz0cy8jh1.png?width=1990&format=png&auto=webp&s=5e6aea776a35b8f6165cf25e97b67daa4bec9a15

Pelican : "generate an svg of a pelican riding a bike" (This might have been on low thinking it only thought for 22 seconds I also didnt ask for it to be beautiful). But no missing parts, everything in the right place except the legs are on the same side of the bike.

https://preview.redd.it/weny3l4sv8jh1.png?width=1630&format=png&auto=webp&s=8e1c2e5636b93934ace4d76bb0e8dcbc408095e5

reddit.com
u/Ok_Technology_5962 — 7 days ago
▲ 89 r/unsloth

Unsloth Desktop launches on Product Hunt! ♥️

Hey guys, thanks so much for all the support for Unsloth Desktop yesterday and today, we appreciate every single of you and hope you are enjoying it as much as we made it.

Someone launched Unsloth Desktop on Product Hunt today,, so feel free to give us an upvote if you have any spare time: https://www.producthunt.com/products/unsloth/launches/unsloth-desktop

Also if you have any extra feedback or feature requests feel free to ask on GitHub or here!!

Thanks so much again guys! ❤️🦥

u/yoracale — 8 days ago
▲ 157 r/unsloth

Unsloth Desktop Features Breakdown

Hey guys thanks so much for all the love for Unsloth Desktop. Here's a breakdown of our 'main' features:

- Code execution: Unsloth improves tool calling with: 30–80% higher tool-call accuracy across models, more reliable call termination to reduce loops, better healing and deduplication to prevent XML leakage. Get 50% more accurate, self-healing tool calls and sandboxed code execution.

- Generate Images and Videos: For MiniMax-H3 on an NVIDIA B200, a 960×544, 124-frame, 8-step generation dropped from 70+ seconds to 13 seconds. Works with FLUX, Z-Image, LTX, Wan and fine-tuned LoRA adapters. Transform, inpaint, extend, upscale, reference and edit existing images.

- Serve and access your models anywhere: Serve your local or over HTTPS through Unsloth's free Cloudflare tunnel. Check a run from your phone, your laptop, or anywhere else you happen to be. Bind the app to your network with -H 0.0.0.0, or open a free Cloudflare tunnel for HTTPS:

- Connect your agent: Unsloth lets you connect Claude Code, Codex, Hermes, OpenCode, and other coding agents to a local model via the unsloth start command. The entire workflow can run offline on your own hardware. Unsloth Studio automatically configures the endpoint, API key, provider, model, and context length for each launch, so you can use your preferred agent without modifying them.

You can view details in our blog: https://unsloth.ai/docs/desktop

Available now on GitHub: https://github.com/unslothai/unsloth

Thanks so much!

u/yoracale — 9 days ago
▲ 478 r/unsloth+2 crossposts

Meet Unsloth Desktop - the first desktop app to run and train models

Hi guys, we're super excited to announce Unsloth Desktop today,
The first desktop app to run and train models locally.

  • Open-source and available on Mac, Windows, and Linux
  • Supports MLX, diffusion image/video models, audio models, and GGUF
  • Connect Claude Code and Codex to local LLMs
  • 50% more accurate with self-healing tool calls and sandboxed code execution
  • Supports CPU and multi-GPU setups across NVIDIA, AMD, Intel, and Mac
  • Train models 2× faster while using 70% less VRAM
  • Includes private web search, deep research, RAG, MCP, and exports (NVFP4, GGUF)
  • Use Unsloth’s OpenAI-compatible API with OpenAI and Anthropic cloud models
  • Securely deploy LLMs remotely and access them anywhere via Cloudflare HTTPS

Unsloth Desktop is now available on unsloth.ai and GitHub.

Thank you and we're here to answer any questions!

u/yoracale — 9 days ago
▲ 167 r/unsloth

2-bit Muse Glimmer GGUF made 100+ tool calls on 14GB VRAM.

What a powerful model even at 2-bit! Muse Glimmer did a complete repo bug hunt for 5 mins nonstop with: evidence, repro, fix, tests and a PR writeup.

You can now run and train it in Unsloth.

GitHub repo: https://github.com/unslothai/unsloth

u/yoracale — 11 days ago
▲ 429 r/unsloth+1 crossposts

DeepSeek-V4 now runs 2x Faster locally with DSpark!

Hey guys, DeepSeek-V4-Flash can now run 2× faster locally with DSpark! ⚡️

DSpark enables V4-Flash-0731 GGUFs to generate ~1.4–2× faster with no accuracy change.

DeepSeek-V4-Flash-0731 can reach at 120 tokens/s. DSpark is automatically enabled in Unsloth Studio!

GGUFs: https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF

Guide: https://unsloth.ai/docs/models/deepseek-v4

Thank you!

u/kwizzle — 15 days ago
▲ 894 r/unsloth

Qwen3.8-27B and Qwen3.8-Max announced!

Qwen just announced Qwen3.8-27B along with Qwen3.8-Max! 🔥

Qwen3.8-27B will run locally on 17GB RAM/VRAM setups and is expected to be the best performing model for its size.

We can't wait to support it at Unsloth AI. Qwen3.8-27B benchmarks are yet to be revealed, only Qwen3.8 -Max for now.

u/yoracale — 18 days ago