16,384 tokens
A byte-level BPE fitted on the same corpus the model trains on.
A smaller vocabulary bought this model six more layers
3.1039 / token
A fixed budget, a word list that quietly eats half of it, and one decision that buys six more layers of thinking.
01 · The budget
Fifty million parameters, and the rules include the token table and the output head. That one sentence decides the architecture before a single layer exists.
02 · The depth
The parameters a smaller vocabulary frees go straight into layers: seven becomes thirteen, at the same cap and the same width.
50,257 tokens7 layers
16,384 tokens13 layers
03 · In use
The 49,296,896-parameter model, running on your device. Load it once, give it an opening, and read what a model this size does with it.
It writes fluently and is often wrong. This checkpoint is from a step not yet published of a training run that is still going.
Not a chat assistant: it continues whatever you type, word by word.
The model is being packaged.
The weights are being exported for the browser. This panel will run them here as soon as they land; nothing is simulated in the meantime. See the run so far
One download, kept by your browser. Nothing you type leaves this page.
49,296,896 parameters
The model could not load on this device. Try a current Chrome, Edge or Firefox.
The words the model writes are highlighted as they arrive.
Focused · temperature 0.7 · top-k 40 · until done, at most 256 tokens · guards on, min-p 0.1 · seed 1337
Uses Balanced's sampler but stops at the end of the first finished passage (about 55 tokens). Shorter, faster, and with less room to drift off topic.
Published keeps temperature, top-k and length at its fixed values. Pick another preset to change them.
Ranges are the ones tested with the guards on: temperature 0.6 to 1.4, top-k 40 to 200, up to 90 tokens (longer runs drift further from the prompt).
At most 256 new tokens. It stops sooner at the end of the first finished passage, when it starts repeating, or if a very long prompt fills its 1,024-token context. With the guards on, stopping at the end of the passage scored more coherent in blind tests than running to 90 tokens: Focused is Balanced with this on. Published with this on was not part of those tests, so it is not recommended there.
Guards: on. Every preset except Published adds these rules, and so do your own settings made from one of them; Published, with or without Unlimited, has none. They only change which token can come next; the model's words are never edited.
Weights and settings are read from the model manifest as it loads.
Slide the vocabulary and watch the word list trade against layers, or follow the 13-layer model through its training run.