49,296,896 parameters · trained from scratch · runs on your device

A language model that fits in fifty million parameters.

Trained on one GPU and running right here in your browser. Give it an opening line and watch it write.

PresetFocused

Uses Balanced's sampler but stops at the end of the first finished passage (about 55 tokens). Shorter, faster, and with less room to drift off topic.

Ranges are the ones tested with the guards on. Longer runs drift further from the prompt.

At most 256 new tokens. It stops sooner at the end of the first finished passage, when it starts repeating, or if a very long prompt fills its 1,024-token context. Focused is Balanced with this on. Published with this on was not tested, so it is not recommended there.

GuardsOn

Every preset but Published adds these rules. They only change which token can come next; the model's words are never edited.

  • Stays on your prompt. Each step also runs the model with no prompt, then favours tokens your prompt made likelier (context-aware decoding, strength 1). About half the speed.
  • Skips long shots. Drops tokens less than about a tenth as likely as the top choice (min-p).
  • Repeats less. Tokens in the last 64 are made less likely (repetition penalty 1.2).
  • No blank start, and line breaks only after a finished sentence.

Focused · temperature 0.7 · top-k 40 · until done, at most 256 tokens · guards on, min-p 0.1 · seed 1337

Not loaded

One 50 MB download, kept by your browser. Nothing you type leaves the page.

embedding-tax-50m base model

A base model, not a chat assistant: it continues whatever you type. It writes fluently and is often wrong.

Not sure where to start? Try these

Runs on your device

No server runs this model.

The weights are one 50 MB file. Your browser downloads it once and keeps it. On a computer the page times WebAssembly against WebGPU, if your browser has it, and runs on the faster one; phones, tablets and low-memory devices run WebAssembly only, to stay within their memory.

The budget

Every parameter counts, including the word list.

The fifty million cap includes the table that turns tokens into numbers: 512 parameters for every token in the vocabulary. Each layer costs another 3,146,752, so a smaller word list leaves room for more layers.

  • GPT-2's 50,257 tokens7 layers fit

    Word list: 51.5% of the cap

  • Ours, 16,384 tokens13 layers fit

    Word list: 16.8% of the cap

  • Word list
  • Layers
  • Unused

Same cap, same width of 512: 13 layers instead of 7. Open the explorer

Go deeper

Everything else, one click away