Every number here was measured,never claimed.

Move the cap and watch the layers fit,then read the run behind the model.

Open the calculator
  • 13 layers under the cap
  • 1.9B tokens, one GPU
  • lm-evaluation-harness
  • Runs in your browser
Trained from scratch on one RTX 4090.2026

Budget explorer

Fifty million, and the word list counts.

At our width of 512, GPT-2's 50,257-token vocabulary spends 51.5% of the cap before a single layer exists and leaves room for 7 layers. Ours has 16,384 tokens, spends 16.8% and fits 13. Everything in the calculator below is arithmetic on the cap, recomputed as you move it.

Try another word list

Start from a tokenizer, then move anything. This part is arithmetic on the cap, not a measurement.

Layers that fit

13 layers

vs 7 with GPT-2's vocabulary

Spent before any layer

16.8% of the cap

vs 51.5% with GPT-2's vocabulary

49,296,896 of 50,000,000 parametersUnder the cap, 703,104 spare

The whole cap, drawn to scale

One block per transformer layer, on the word list's cost.

At a 16,384-token vocabulary and width 512, 16.8% of the cap is spent before any layer exists, leaving room for 13 transformer layers. GPT-2's 50,257-token vocabulary with the same settings spends 51.5% and leaves 7.
tokens

Rule 01 16,384 tokens × 512 wide = 8,388,608 before any layer exists.

wide

One layer at this width costs 3,146,752.

× width

Feed-forward is 2,097,152 of every layer.

Options

Rule 02 Tied: +3 layers over a separate head.

Rule 03 Rotary: saves 524,288, under one layer here.

Where every parameter goes

Every row is Each × Count. The groups are the parts of the drawing, and they add up to the total, checked against the cap.

Parameter ledger for the current settings: each part's size, how many are paid for, and the parameters it costs.
PartEachCountParameters
Before any layer16.8% of the cap8,388,608
Token embedding16,384 tokens × 512 wide8,388,608 × 18,388,60818,388,608
Output headTied to the embedding: saves 8,388,6088,388,608 × 0 · tied8,388,60800
PositionsRotary: computed, not stored, saves 524,288524,288 × 0 · rotary524,28800
13 layers3,146,752 each · 81.8% of the cap40,907,776
Attention4d²: query, key, value and output1,048,576 × 131,048,5761313,631,488
Feed-forward8d²: two matrices at 4× the width2,097,152 × 132,097,1521327,262,976
Norms2d: two weight-only norms1,024 × 131,0241313,312
Final normd: after the last layer512 × 15121512
Total98.6% of the cap49,296,896
Unused703,104
Cap50,000,000

The cap counts the token embeddings and the output head.

Runs on this device

Try the model.

The same 49,296,896 parameters the calculator lands on, exported to run in your browser. Load once, then write.

It writes fluently and is often wrong. This checkpoint is from step a step not yet published of a training run that is still going.

embedding-tax-50m

Not loaded

One download, kept by your browser. Nothing you type leaves this page.

embedding-tax-50m base model

Pick an opening or type your own. The model's words follow yours in darker type as they arrive.

A base model, not a chat assistant: it continues whatever you type. It writes fluently and is often wrong.