49,296,896 parameters · trained from scratch · runs on your device
A language model that fits in fifty million parameters.
Trained on one GPU and running right here in your browser. Give it an opening line and watch it write.
13 layers fit
The word list costs 16.8% of the 50,000,000 cap before a single layer exists.
- Embeddings
- Layers
- Unused
Width 512, tied embeddings, 3,146,752 parameters per layer. Open the full explorer
embedding-tax-50m base model
- Preset
- New tokens
- Speed
- First token
- Stopped by
- Backend
A base model, not a chat assistant: it continues whatever you type. It writes fluently and is often wrong.
The model is being packaged.
This panel runs the real weights as soon as they land; nothing is simulated in the meantime.
The model could not run on this device. Try a current Chrome, Edge or Firefox.
Not sure where to start? Try these
Compare vocabularies
Runs on your device
No server runs this model.
The weights are one 50 MB file. Your browser downloads it once and keeps it. On a computer the page times WebAssembly against WebGPU, if your browser has it, and runs on the faster one; phones, tablets and low-memory devices run WebAssembly only, to stay within their memory.
The budget
Every parameter counts, including the word list.
The fifty million cap includes the table that turns tokens into numbers: 512 parameters for every token in the vocabulary. Each layer costs another 3,146,752, so a smaller word list leaves room for more layers.
-
GPT-2's 50,257 tokens7 layers fit
Word list: 51.5% of the cap
-
Ours, 16,384 tokens13 layers fit
Word list: 16.8% of the cap
- Word list
- Layers
- Unused
Same cap, same width of 512: 13 layers instead of 7. Open the explorer
Go deeper





