unsloth studio model loader settings ignored

hello first i want to say respect to unsloth devs they are doing a good job and i was impressed to see that unsloth studio has expanded to be more than fine tuning software,

i have an issue with the model loader it keep ignoring the context that i set and always load the default and when trying overriding from arguments sometimes it crash because i think it drops all the ui setting if you add only one argument, and i did make it work from argument but token generation dropped from 60 to 22 am i missing something or this is a bug

OS: kubuntu 24

Unsloth Version

v0.1.800-beta

Package Version

2026.8.18

Desktop App Version

0.1.800-beta

llama.cpp Version

b10360-mix-87da1a2

Hardware

GPU 0

NVIDIA GeForce RTX 5060 Ti · 16 GiB

GPU 1

NVIDIA GeForce RTX 5060 Ti · 16 GiB

CUDA

13.0

reddit.com
u/chocofoxy — 3 days ago

Qwen 3.8 27B is faster than expected

i ran this model on my two 5060 TI 16GB cards at Q4 in unsloth and LM studio ( i downloaded NVFP4 but didn't try it in vLLM ) i think it runs faster than expected it gives me 50 - 60 t/s with MTP.
this is surprising because it's a dense model and Qwen 3.6 was giving me 30t/s with MTP any one noticing this text generation speed peaking or it's a setup thing because i swapped from windows to linux last month and maybe Qwen 3.6 was fast but i had the wrong OS

reddit.com
u/chocofoxy — 3 days ago

The end of S2 gives me a headache

i feel season 2 logic is not in the room with us, i was mad because Dan has made an awful blender but continuing watching the season i felt like Peter was absolute mad man and he was right all the time, yes he made some mistakes with the ganging up and closing doors on other members but isn't that what Phaedra group is doing.

Phaedra did not give any names she was literally a female Dan , and her whole group was AFK until the end but they did not go after them and this thing with the loyalty in Phaedra group i did not understand maybe the show editing did not show it but i feel they made all the mistakes other people got banished for and they did get away with it

i am binge watching the show and season 1 made sense even wit Cirie risking it at the end, but season 2 the war at first made sense but at the end it went out the door because of a loyalty and emotions that i don't know about i literally do not know these people this is my first reality show beside cooking show

i feel if someone needed to win the season it's Peter group

reddit.com
u/chocofoxy — 14 days ago

The Traitors Season 2 episode 5

i started to watch the traitors this week and it been a good time until this dumb moment after the mission in season 2 episode 5 , when they paint Peter as a mastermind and have this master plan to not tell everyone who has the shields, BUT everyone knows who has the shields because the house group is the one who solved the sound of the bird ( the shield bird or something ) and Alan told everybody about it before the game starts if Dan and the traitors are going to fall for this i am going to be pissed more than when Cirie voted Christian out when they could've ended season 1 in episode 9

EDIT 1: I am officially pissed, especially because Dan literally said Peter was bluffing, and that killing him would ruin his own plan

EDIT 2: it was pointed out to me that the house group didn't know which bird was the shield bird, but still Dan was blind because he knew peter was bluffing and if he believes that he's bluffing now the shield is 50/50 chance in the other two pairs

reddit.com
u/chocofoxy — 14 days ago

How can you stop your model from looping

So i thought this is a small model issue but when i added a new gpu and i am able to run low mid model like Qwen 3.6 35b q4 or q5 this issue still exists now its not as much as small model but it does break when linking the model to copilot chat or Hermes the model mid task will start loop thinking or looping generating more than 40k token or generating a wrong tool call

reddit.com
u/chocofoxy — 3 months ago

running Qwen 3.6 35b A3B on 2x 5060TI

i ran Qwen 3.6 35b A3B two 5060TI 16gb ( 32 gb vram total also i have 32gb dram but i don't like offloading ) i used Q4 on LM Studio with full context and i get 90t/s any tricks to optimze this more to upgrade to Q6 or Q8 ?
thanks !

another thing if you recommend somthing for cooling because i am using 2 stacked gpus with 0 gap ( i have and mATX motherboard ) now the top gpu it not that hot but hotter then the bottom one

reddit.com
u/chocofoxy — 3 months ago

Got these two to add to my collection

I really like glenfiddich i trust it i tried the blue and orange ones bourbon casket and for glenlivet it was recommended a lot so i had to try it

u/chocofoxy — 3 months ago