1am/small-data-gpt

What transfers to a GPT trained on 3.7M tokens, and what doesn't

from Advanced Machine Learning, ANU · Mayukh Das

For an Advanced Machine Learning project at the ANU, I trained a GPT from scratch on five-sentence stories: about 30 million parameters, and only 3.7 million training tokens. Chinchilla-style scaling would suggest around 600 million tokens for a model that size, so this is roughly 160 times less data than it "wants". That constraint turned every design choice into a question of efficiency rather than capacity, and it made several large-scale habits misbehave.

The models below are drawn to scale from the real configurations: volume is parameters.

1. The tokenizer was the biggest lever

the same model, two tokenizers
DRAG TO SPIN

19.3M of 31.69M parameters (61%) sit in the lookup table. Test perplexity 28.62.

The largest single gain didn't come from the architecture at all. With the GPT-2 tokenizer, the embedding table alone holds 50,257 × 384 ≈ 19.3M parameters: over 60% of the model, spent on a lookup table. Switching to a 32k vocabulary shrank it to about 12.3M and freed about 7M parameters for the layers that actually learn, and test perplexity fell from 28.62 to 23.71 on an otherwise identical model.

The graded evaluation required GPT-2 tokens, so the final model couldn't use it, but the lesson stands: when parameters are scarce, where they sit matters more than how many there are.

2. Depth beat width

four shapes, one budget (~31.7M)
DRAG TO SPIN

At matched parameter counts, deeper and narrower won every time: 7 layers scored 25.77, 11 scored 25.45, 16 scored 25.17, and 20 layers at width 272 scored 25.12. That contradicts the large-scale finding that depth versus width barely matters. My reading is that on small data each extra layer refines the previous one's features, while extra width adds parallel features the data is too small to teach.

3. Only two of three LLaMA components transferred

switch components on and off
DRAG TO SPIN

Test perplexity 25.33 · the baseline

RoPE was the biggest win (25.33 to 24.69) because it encodes position by rotating queries and keys, at no parameter cost: the learned position table disappears. RMSNorm helped modestly (25.01). SwiGLU's gates need enough data to learn meaningful values, and 3.7M tokens wasn't it: perplexity rose to 25.62 and training diverged after 6,500 steps. RoPE and RMSNorm together reached 24.67: complementary, but not additive.

4. Matching the data beat having more of it

test perplexity by data strategy · lower is better · the plane is the score to beat
DRAG TO SPIN

Pretraining on TinyStories, a much larger dataset of simple children's stories, was the worst result of the project (53.26), and adding synthetic stories also hurt. Fine-tuning on stories written by Gemini worsened perplexity too, but its stories read better. The one strategy that improved both was quality-filtering the original stories (25.33 to 24.92). Data that doesn't match the test distribution made things worse, regardless of quality or size.

5. Perplexity is not quality

the two finalists
DRAG TO SPIN

The submitted model (20 layers × 256, RoPE and RMSNorm) scored 24.67. A slightly wider sibling scored 24.84 and wrote noticeably better stories. Perplexity measures how well a model predicts the test set's exact words, not whether its stories are coherent. It is the right metric to optimise, and the wrong one to stop looking past.

real samples from the submitted model · same prompt
Training and validation loss of the submitted model over 12,000 steps
The submitted model's training (blue) and validation (orange) loss. Validation flattens after about 6,000 steps while training keeps falling.

Try the model

the live model, on Hugging Face

It may take a minute to wake up.

The thread through all of it

Recipes from large-scale training are tuned for large-scale problems. At small scale, some of them are free wins (RoPE), some are neutral, and some actively hurt (SwiGLU). The only reliable way to know which is to change one thing at a time and measure.