Fifty million parameters, and the word list takes half

A smaller vocabulary bought this model six more layers

Validation loss

3.1039 / token

-45.5% vs. first measured point (5.69 at step 250)

The parameter cap Embeddings included 50,000,000

A fixed budget, a word list that quietly eats half of it, and one decision that buys six more layers of thinking.

01 · The budget

The cap counts the embeddings.

Fifty million parameters, and the rules include the token table and the output head. That one sentence decides the architecture before a single layer exists.

51.5%
Of the cap is a word list at GPT-2's vocabulary
16.8%
Of the cap is a word list at 16,384 tokens, ours

02 · The depth

Depth is what it buys.

The parameters a smaller vocabulary frees go straight into layers: seven becomes thirteen, at the same cap and the same width.

Open the budget explorer

The model in five numbers

Token table

16,384 tokens

A byte-level BPE fitted on the same corpus the model trains on.

Depth

13 layers

Where GPT-2's vocabulary would leave room for only seven.

Attention

8 attention heads

Four matrices a layer: query, key, value, and the output projection.

Positions

0 position parameters

Rotary, not learned: no position table to pay for.

Hardware

1 RTX 4090

From scratch: no pretrained weights, no fine-tuning, no distillation.

Card 1 of 5

03 · In use

Watch it write.

The 49,296,896-parameter model, running on your device. Load it once, give it an opening, and read what a model this size does with it.

It writes fluently and is often wrong. This checkpoint is from a step not yet published of a training run that is still going.

Try the model

Not a chat assistant: it continues whatever you type, word by word.

The model is being packaged.

The weights are being exported for the browser. This panel will run them here as soon as they land; nothing is simulated in the meantime. See the run so far

Focused · temperature 0.7 · top-k 40 · until done, at most 256 tokens · guards on, min-p 0.1 · seed 1337

Settings
PresetFocused

Uses Balanced's sampler but stops at the end of the first finished passage (about 55 tokens). Shorter, faster, and with less room to drift off topic.

Ranges are the ones tested with the guards on: temperature 0.6 to 1.4, top-k 40 to 200, up to 90 tokens (longer runs drift further from the prompt).

At most 256 new tokens. It stops sooner at the end of the first finished passage, when it starts repeating, or if a very long prompt fills its 1,024-token context. With the guards on, stopping at the end of the passage scored more coherent in blind tests than running to 90 tokens: Focused is Balanced with this on. Published with this on was not part of those tests, so it is not recommended there.

Guards: on. Every preset except Published adds these rules, and so do your own settings made from one of them; Published, with or without Unlimited, has none. They only change which token can come next; the model's words are never edited.

  • Stays on your prompt. Each step also runs the model with no prompt at all, then favours tokens your prompt made likelier (context-aware decoding, strength 1). Two passes per token, so about half the speed.
  • Skips long shots. Drops any token less than about a tenth as likely as the top choice (min-p, set from the temperature; stricter below 1).
  • Repeats less. Tokens already in the last 64, your prompt included, are made less likely (repetition penalty 1.2).
  • No blank start. The first 3 tokens cannot be a line break or the end of the text.
  • Breaks after sentences. A line break can only follow a finished sentence.

Weights and settings are read from the model manifest as it loads.

Fifty million parameters. Spend them on depth.

Slide the vocabulary and watch the word list trade against layers, or follow the 13-layer model through its training run.