303m parameter assistant model from scratch on local hardware
I trained a coherent assistant LLM completely from scratch on my own PC — TinyBrainBot-303M-Instruct
Hello, I've been messing with local AI for almost a year now, alongside a bunch of other projects, but my main goal the whole time has been to build a coherent assistant model entirely from scratch on my own hardware (2× Tesla P100). took like 300 hours but I'm done finally.
TinyBrainBot-303M-Instruct is a dense decoder-only transformer (LLaMA-style — RoPE, RMSNorm, SwiGLU):
- 303M params — 32k vocab, 1024 hidden, 24 layers, 16 heads (4 KV heads / GQA), 2816 FFN, 1024 context
It was pretrained on ~4.7B tokens of assorted data — FineWeb-Edu (10BT), English Wikipedia, Orca-Math-200k, OpenWebText2, and several distillation datasets I generated with Ling-2.6-Flash — then SFT'd on ~120M tokens including Dolly-15k, SmolTalk, UltraChat-200k, and more custom distillation sets.
The goal was a semi-coherent assistant with reasonable factual recall, and it got there. A few real examples (best of 5 samples, temperature tuned per prompt):
Q: What is the largest planet in the solar system?
>Jupiter.
Q: Explain what photosynthesis is in one or two simple sentences.
>Photosynthesis is how plants use sunlight, water, and air to make food (glucose) and release oxygen. It occurs in many types like green plants, algae, and trees.
Q: Write a short poem about the moon.
>Silver moon, slow and gold, hums low over the sleeping sky. It dreams in silver rivers, guarding dreams in morning light.
Q: List three healthy breakfast ideas.
>Greek yogurt with berries, whole wheat bread, cream cheese.
Q: Give me one tip for staying focused while studying.
>Use 25-minute focus blocks with 5-minute breaks, silence notifications, and a clear workspace to keep your mind fresh.
It's LLaMA-architecture, so it runs in llama.cpp / anything GGUF (F16 GGUF is up on the repo). Would love to hear your thoughts!
🔗 https://huggingface.co/nkthebass/TinyBrainBot-303m-instruct