Seven hours on one GPU, and every number kept

Measured, not calculated. Everything below is a snapshot read from results.json, written from the training log and the evaluation output; the GPT-2 Small and chance rows are published reference figures kept in the same file. It never refreshes on its own, and the controls above change none of it.

The run so far

not measured yet

Results not loaded

Training results are read from results.json by this page's script. With scripts off, nothing below has been filled in.

Validation loss

not measured yet

Best so far

not measured yet

perplexity not measured yet per token

Held-out shard, scored on the same fixed sample at every evaluation. Lower is better. Perplexity is per token, so a smaller vocabulary lowers it without the model being any better: compare it only with models that use this exact tokenizer.

Tokens seen
not measured yet
Wall clock
not measured yet
Throughput
not measured yet
Steps
not measured yet
Peak VRAM
not measured yet

Corpus not measured yet

How it scores against GPT-2 Small

Scores load from results.json. Random chance is drawn on every row because at this size two of these tasks are expected to sit near chance.

Multiple choice, accuracy in %

Accuracy on four multiple-choice tasks for this model, GPT-2 Small and random chance, in percent.
Task Scale, 0 to 100 This model GPT-2 Small124M parameters Chance
Scores are read from results.json by this page's script. With scripts off, none are shown.

Text prediction, WikiText-2 test set

Word perplexity
not measured yetper word · lower is better
Bits per byte
not measured yetper byte, tokenizer-independent · lower is better

Every row runs 0 to 100. Both WikiText figures are divided by the text itself, per word and per byte, so the tokenizer cancels out and they compare fairly with any other model's. When they are published, lm-evaluation-harness scores them on the WikiText-2 test set, which WikiText-103 shares.

Published samples

Written once from the checkpoint at fixed settings and published unedited. For fresh text, use the model above; it runs on your device.

  1. Not generated yet

    Samples are written once from the finished checkpoint, at fixed settings, and published here unedited.

Only the architecture controls and the on-device model are live.