Predict the data, not the velocity, until you can't
Flow matching trains a network to turn noise into data. One design choice sounds like a detail: should the network predict the clean data (x-prediction) or the velocity that moves noise towards data (v-prediction)? Mathematically the two contain the same information. In practice, for an Advanced Machine Learning project at the ANU, they behaved completely differently.
The experiment
I used simple 2D datasets (a swiss roll, a ring of Gaussians, circles) and embedded them in 2, 8 and 32 dimensions, so the data always lies on a flat 2D sheet inside a bigger space. The same small network (5 layers, 256 wide) was trained with every combination of prediction target and loss.

What happened
At 2 dimensions, everything worked. At 8, v-prediction started to blur. At 32, both v-prediction setups failed completely, while both x-prediction setups still produced clean spirals, rings and clusters. The loss type mattered much less than the prediction target: matching the loss to the target helped a little, but never rescued v-prediction.
Why
The clean data still lives on a simple 2D sheet, however many dimensions it is embedded in, so x-prediction only has to learn that simple shape. The velocity, though, includes the noise, and the noise fills all 32 dimensions. So v-prediction has to model a target whose complexity grows with the full space.
Illustrative, with data and noise at the same scale: each bar is one direction of the space.
Same information, different difficulty: x-prediction's target stays low-dimensional, v-prediction's grows with the space.
Can v-prediction be rescued?
Yes, so the failure is about capacity, not impossibility. Widening the network from 256 to 1024 units gave clean spirals with v-prediction, but that means going from about 330K to 5.2M parameters: roughly 16 times the compute for what x-prediction did at the small size. Shifting the noise schedule, a trick from large-scale work, didn't help on its own.

So why do big image models predict velocity?
Because their world is different. Production models like SD3 and FLUX work in dense, learned latent spaces, where the data no longer sits on a simple low-dimensional sheet, so x-prediction loses its advantage. And v-prediction keeps the difficulty more even across noise levels, while x-prediction must reconstruct clean data from almost pure noise at the start. At scale, that steadier optimisation wins.
How many steps does sampling need?
Generating means following the learned flow from noise to data in small Euler steps. With the best setup (x-prediction at 32 dimensions), recognisable structure appeared by about 10 steps, and quality saturated between 20 and 50. Beyond that, extra steps bought almost nothing.
A footnote on one-step generation
I also implemented MeanFlow, which learns the average velocity over a whole interval so it can generate in a single step. It drew the swiss roll and circles in one step, but struggled with the separate Gaussian clusters, where averaging pulls the modes together. The paper's recommended timestep sampler made that worse here, which was a useful reminder in itself: recipes from large-scale work are tuned for failure modes that may not exist in a small problem.
