Transformers' Revenge: where "Were RNNs All We Needed?" breaks
A recent paper showed that a minimal recurrent model, minGRU, matches Transformers on language modelling while training in linear time. We asked whether that still holds when a task needs genuine algorithmic reasoning, and wrote down what we expected before running anything.
The question
Language modelling mostly rewards local statistical patterns, where even simple models do well. The tasks that separate architectures are the ones that need content-based retrieval: using what you see now as a query to find something specific from earlier. minGRU's gates depend only on the current input, which suggested it could not do this at all. We wanted to test that precisely, and to find where each architecture's limits are.
The setup
- Three architectures at about 750–800K parameters, 4 layers each, with an identical optimiser and schedule: minGRU as published; a nanoGPT-based Transformer with the LLaMA recipe (RoPE, RMSNorm, no biases); and a causal gMLP, which we adapted from the original bidirectional design as a middle point that mixes across positions with fixed weights inside a 256-token window.
- Three tasks: character-level Shakespeare (language modelling), verbatim copying of a sequence, and induction heads (complete a pattern seen earlier in the sequence).
- Three sequence lengths per task, chosen to cross the gMLP's window. 3 × 3 × 3 = 27 runs.
Predictions, made before the runs
- Transformer: succeeds on every task; attention can reach any earlier token by content.
- minGRU: fails every algorithmic task at every length; input-only gating cannot retrieve by content, whatever its size.
- gMLP: succeeds while the needed retrieval falls inside its window, and fails sharply beyond it.
- Everyone: similar results on Shakespeare, since that is the domain where the original claim was made.
What happened
| Prediction | Observed | |
|---|---|---|
| minGRU fails copy and induction at every length | chance level throughout (induction accuracy ≈ 0.04) | confirmed |
| gMLP succeeds inside its window, fails beyond | copy-short perfect (perplexity 1.003); copy-medium and induction-long at chance | confirmed |
| Transformer succeeds everywhere | induction 0.97–1.00 at every length; copy-long failed (perplexity 26.1) | one miss |
| Similar Shakespeare perplexity for all | 3.7–4.4 | confirmed |
So the original claim holds where it was tested, on language modelling, and breaks exactly where the architecture analysis said it would: any task that needs data-dependent retrieval.
The surprise: a phase transition
The Transformer did not learn long induction gradually. It sat near chance for about 7,000 steps and then jumped to perfect accuracy within about 2,000 more, which is consistent with the attention heads suddenly forming an induction circuit. That also offers an explanation for our one missed prediction: copy-long only had 5,000 training steps, possibly too few for the same jump. We say so as a hypothesis, because the copy-long training curve itself shows no sign of a transition starting.
What was weak
- Single seeds per run, so small gaps (like the gMLP's slightly worse Shakespeare perplexity) are not conclusive.
- The copy-long budget of 5,000 steps leaves the Transformer's one failure unexplained by evidence; the next step is to train it longer.
- The tasks were chosen to expose limits. We think that is fair, since copy and induction are capabilities trained Transformers develop naturally, but it is a choice worth stating.
My part
A team of four at the ANU (Deep Learning, 2026), with Vibhansh Gupta, Daniel Vaz and Adam Clark. I implemented the Transformer, designed the 27-run experiment grid and the comparison framework, and wrote the report.