Dust Trains Transformers Without Backpropagation—at a Steep Compute Price
Q Labs' zeroth-order method perturbs token activations and uses large virtual populations to estimate updates. The experiments are provocative, but they do not make backpropagation obsolete.
Q Labs has published Dust, a zeroth-order optimization method that pretrains transformer language models without using a conventional backward pass. Instead of differentiating the loss through every layer, Dust adds independent noise to activations at each token, measures whether each perturbation improves the loss, and averages reward-weighted perturbations into an update estimate.
The launch reached technical communities quickly. The Hacker News submission had 140 points and 31 comments when checked at 1:33 p.m. Malaysia time on October 6, while GeekNews published a detailed Korean summary the same day. Those numbers establish attention, not scientific replication.
What happened
The paper’s central device is a virtual population. Traditional evolution strategies perturb weights and evaluate population members in separate forward passes. Dust perturbs hidden activations independently for many tokens, so one transformer forward pass evaluates thousands of local variations in parallel. The method then assigns each token a reward based on how the perturbation changed its own and later-token losses.
The researchers compare Dust with ordinary backpropagation and with a transformer implementation of EGGROLL, a weight-space evolution-strategy method. Their experiments use GPT-style models trained on FineWeb, including models from roughly 2 million to 243 million parameters. They also examine whether the estimated directions align with gradients from backpropagation at checkpoints trained on up to 1 billion tokens.
At larger virtual-population sizes, Dust approached backpropagation’s validation loss and exceeded it in some reported settings. The authors estimate that from 1 million training tokens upward it is roughly 1,000 to 10,000 times more population-efficient than their EGGROLL baseline. They also report an unexpected scaling pattern: larger models often used a given population more effectively than smaller ones.
Why it matters
Backpropagation is more than an algorithm; it is the design constraint around which modern neural-network software and accelerators evolved. It requires differentiable operations and stores information needed to send error signals backward. A useful alternative could make architectures with external programs, hard decisions, or very long recurrent loops easier to train.
Dust does not deliver that architecture freedom yet, but it provides a concrete transformer-scale experiment. The activation-space approach is important because it moves the search away from copying a complete set of weights for every population member. Token-level parallelism turns work a transformer already performs into a much larger search surface.
The model-size result also challenges a common intuition about zeroth-order methods. High dimensionality is normally expected to make noisy update estimates worse. In these experiments, overparameterization did not simply destroy the method’s population efficiency. That is a research clue worth testing independently.
Evidence
The official research page is the primary source for the method, experiments, tables, caveats, and related work. The linked repository provides code for inspection. Q Labs states that Dust’s estimate becomes more aligned with the backpropagation gradient as population grows, that the alignment persists across tested training checkpoints, and that Adam can improve both Dust and backpropagation.
Important qualifiers sit beside the headline results. Claims that Dust can beat backpropagation at very large populations sometimes depend on curve fitting or extrapolation rather than a directly observed infinite-population run. The authors acknowledge that comparable performance requires substantially more computation. Their paper explicitly says the method is not compute-efficient enough to replace backpropagation today.
The experiments are also small compared with frontier pretraining. The largest Dust model reported is 243 million parameters, not tens or hundreds of billions. Measuring gradient alignment at a 1-billion-token checkpoint is not the same as using Dust to pretrain a model for 1 billion tokens. The current work demonstrates a direction, not a production recipe.
Practical takeaway
For model researchers, the useful question is whether Dust opens a training regime where differentiability is the real blocker. Reproducing its loss curves, FLOP accounting, tuning burden, and hardware utilization should come before treating it as a faster optimizer. Comparisons should report total compute, wall-clock time, memory, tokens, model size, and the cost of hyperparameter searches—not only final loss.
For developers choosing how to train or fine-tune an ordinary model, backpropagation remains the practical default. Dust does not lower the cost of a standard training run, and the repository is research software. Its immediate value is conceptual: activation perturbations plus massive parallel evaluation may offer a bridge between gradient learning and search.
Limitations
The results come from the authors and have not yet accumulated broad independent replications. FineWeb language-model loss is a narrow outcome; it does not show downstream reasoning, safety, or data efficiency. The comparison with EGGROLL depends on implementations and extrapolated populations. Hardware optimized for forward-only search could change economics, but that hardware advantage has not been demonstrated here.
The responsible reading is therefore neither “backprop is dead” nor “zeroth-order training cannot scale.” Dust shows that one carefully constructed zeroth-order method can get surprisingly close on modest transformers when given a very large virtual population. Whether that becomes scientifically or commercially decisive remains open.
Sources
> Want more like this?
Get the best AI insights delivered weekly.
By subscribing, you agree to our Privacy Policy. You can unsubscribe at any time.
> Related Articles
The Next AI Efficiency Gain Is Amortizing Work Across the Whole System
Dust, Quail, and OpenRig attack different layers, but share one idea: useful AI systems get cheaper and more reliable when they reuse structure across training signals, data operations, and agent work.
Quail Makes the SQL Query Plan Control LLM Inference
CMU's open-source AI-SQL engine reuses KV cache across filters and joins, overlaps CPU preparation with GPU work, and reports faster execution than a request-at-a-time baseline.
Local-First AI Is Becoming a Control Stack, Not Just a Model Download
Strata, OpenMuse, and RemoveMacAI expose three layers of personal AI control: where inference runs, what tools an agent can use, and which operating-system capabilities remain enabled.
Tags
> Stay in the loop
Weekly AI tools & insights.