Every number here was measured,never claimed.
Move the cap and watch the layers fit,then read the run behind the model.
Open the calculator- 13 layers under the cap
- 1.9B tokens, one GPU
- lm-evaluation-harness
- Runs in your browser
Budget explorer
Fifty million, and the word list counts.
At our width of 512, GPT-2's 50,257-token vocabulary spends 51.5% of the cap before a single layer exists and leaves room for 7 layers. Ours has 16,384 tokens, spends 16.8% and fits 13. Everything in the calculator below is arithmetic on the cap, recomputed as you move it.
Try another word list
Start from a tokenizer, then move anything. This part is arithmetic on the cap, not a measurement.
Layers that fit
13 layers
vs 7 with GPT-2's vocabulary
Spent before any layer
49,296,896 of 50,000,000 parametersUnder the cap, 703,104 spare
The whole cap, drawn to scale
One block per transformer layer, on the word list's cost.
Rule 01 16,384 tokens × 512 wide = 8,388,608 before any layer exists. Smaller vocabularies split text into more tokens, so the same 1,024-token window holds less text.Enter a whole number between 1,024 and 65,536
One layer at this width costs 3,146,752.Enter a multiple of 64 between 256 and 1,024
Feed-forward is 2,097,152 of every layer.Enter a whole number between 2 and 8
Options
Rule 02 Tied: +3 layers over a separate head.
Rule 03 Rotary: saves 524,288, under one layer here.
Where every parameter goes
Every row is Each × Count. The groups are the parts of the drawing, and they add up to the total, checked against the cap.
| Part | Each | Count | Parameters |
|---|---|---|---|
| Before any layer16.8% of the cap | 8,388,608 | ||
| Token embedding16,384 tokens × 512 wide8,388,608 × 1 | 8,388,608 | 1 | 8,388,608 |
| Output headTied to the embedding: saves 8,388,6088,388,608 × 0 · tied | 8,388,608 | 0 | 0 |
| PositionsRotary: computed, not stored, saves 524,288524,288 × 0 · rotary | 524,288 | 0 | 0 |
| 13 layers3,146,752 each · 81.8% of the cap | 40,907,776 | ||
| Attention4d²: query, key, value and output1,048,576 × 13 | 1,048,576 | 13 | 13,631,488 |
| Feed-forward8d²: two matrices at 4× the width2,097,152 × 13 | 2,097,152 | 13 | 27,262,976 |
| Norms2d: two weight-only norms1,024 × 13 | 1,024 | 13 | 13,312 |
| Final normd: after the last layer512 × 1 | 512 | 1 | 512 |
| Total98.6% of the cap | 49,296,896 | ||
| Unused | 703,104 | ||
| Cap | 50,000,000 |
The cap counts the token embeddings and the output head.
Runs on this device
Try the model.
The same 49,296,896 parameters the calculator lands on, exported to run in your browser. Load once, then write.
It writes fluently and is often wrong. This checkpoint is from step a step not yet published of a training run that is still going.
One download, kept by your browser. Nothing you type leaves this page.
Enter to write · Shift+Enter for a new line
embedding-tax-50m base model
Pick an opening or type your own. The model's words follow yours in darker type as they arrive.
- Preset
- New tokens
- Speed
- First token
- Stopped by
- Backend
A base model, not a chat assistant: it continues whatever you type. It writes fluently and is often wrong.
The model is being packaged.
The weights are being exported for the browser. This panel will run them here as soon as they land; nothing is simulated in the meantime.
The model could not load on this device. Try a current Chrome, Edge or Firefox.