Measured, not calculated. Everything below is a snapshot read from results.json, written from the training log and the evaluation output; the GPT-2 Small and chance rows are published reference figures kept in the same file. It never refreshes on its own, and the controls above change none of it.
The run so far
Results not loaded published at step
Training results are read from results.json by this page's script. With scripts off, nothing below has been filled in.
Validation loss
not measured yet
Since the first evaluation
Best so far
not measured yet
perplexity not measured yet per token
- Tokens seen
- not measured yet
- Wall clock
- not measured yet
- Throughput
- not measured yet
- Steps
- not measured yet
- Peak VRAM
- not measured yet
Corpus not measured yet
How it scores against GPT-2 Small
Scores load from results.json. Random chance is drawn on every row because at this size two of these tasks are expected to sit near chance.
Multiple choice, accuracy in %
| Task | Scale, 0 to 100 | This model | GPT-2 Small124M parameters | Chance |
|---|---|---|---|---|
| Scores are read from results.json by this page's script. With scripts off, none are shown. | ||||
Text prediction, WikiText-2 test set
- Word perplexity
- not measured yetper word · lower is better
- Bits per byte
- not measured yetper byte, tokenizer-independent · lower is better
Every row runs 0 to 100. Both WikiText figures are divided by the text itself, per word and per byte, so the tokenizer cancels out and they compare fairly with any other model's. When they are published, lm-evaluation-harness scores them on the WikiText-2 test set, which WikiText-103 shares.
Published samples
Written once from the checkpoint at fixed settings and published unedited. For fresh text, use the model above; it runs on your device.
-
Not generated yet
Samples are written once from the finished checkpoint, at fixed settings, and published here unedited.
Only the architecture controls and the on-device model are live.