work/transformers-revenge

Transformers' Revenge: where "Were RNNs All We Needed?" breaks

A recent paper showed that a minimal recurrent model, minGRU, matches Transformers on language modelling while training in linear time. We asked whether that still holds when a task needs genuine algorithmic reasoning, and wrote down what we expected before running anything.

The question

Language modelling mostly rewards local statistical patterns, where even simple models do well. The tasks that separate architectures are the ones that need content-based retrieval: using what you see now as a query to find something specific from earlier. minGRU's gates depend only on the current input, which suggested it could not do this at all. We wanted to test that precisely, and to find where each architecture's limits are.

The setup

Predictions, made before the runs

  1. Transformer: succeeds on every task; attention can reach any earlier token by content.
  2. minGRU: fails every algorithmic task at every length; input-only gating cannot retrieve by content, whatever its size.
  3. gMLP: succeeds while the needed retrieval falls inside its window, and fails sharply beyond it.
  4. Everyone: similar results on Shakespeare, since that is the domain where the original claim was made.

What happened

Results of all 27 runs: green passes, red fails
All 27 runs. Shakespeare: perplexity. Copy: recall perplexity (1.0 perfect, about 26 random). Induction: accuracy (1.0 perfect, about 0.04 random).
PredictionObserved
minGRU fails copy and induction at every lengthchance level throughout (induction accuracy ≈ 0.04)confirmed
gMLP succeeds inside its window, fails beyondcopy-short perfect (perplexity 1.003); copy-medium and induction-long at chanceconfirmed
Transformer succeeds everywhereinduction 0.97–1.00 at every length; copy-long failed (perplexity 26.1)one miss
Similar Shakespeare perplexity for all3.7–4.4confirmed

So the original claim holds where it was tested, on language modelling, and breaks exactly where the architecture analysis said it would: any task that needs data-dependent retrieval.

The surprise: a phase transition

Transformer induction accuracy over training: flat for about 7,000 steps, then a jump to perfect
Transformer, induction at length 2048: near chance for about 7,000 steps, then perfect between steps 8,000 and 10,000.

The Transformer did not learn long induction gradually. It sat near chance for about 7,000 steps and then jumped to perfect accuracy within about 2,000 more, which is consistent with the attention heads suddenly forming an induction circuit. That also offers an explanation for our one missed prediction: copy-long only had 5,000 training steps, possibly too few for the same jump. We say so as a hypothesis, because the copy-long training curve itself shows no sign of a transition starting.

What was weak

My part

A team of four at the ANU (Deep Learning, 2026), with Vibhansh Gupta, Daniel Vaz and Adam Clark. I implemented the Transformer, designed the 27-run experiment grid and the comparison framework, and wrote the report.