The Next AI Efficiency Gain Is Amortizing Work Across the Whole System
Dust, Quail, and OpenRig attack different layers, but share one idea: useful AI systems get cheaper and more reliable when they reuse structure across training signals, data operations, and agent work.
Three projects visible in developer communities on October 6 operate at different layers of the AI stack. Dust estimates transformer updates by evaluating many token-level perturbations in parallel. Quail plans LLM-backed SQL operators so document context can be reused across filters and joins. OpenRig organizes coding agents into persistent teams with named roles, shared context, queues, and direct addresses.
They are not competitors, and combining them would not automatically produce a good product. Their shared principle is more useful: stop treating every unit of AI work as an isolated call. Reuse structure across a larger execution plan, whether the unit is a training perturbation, a document prefix, or an agent’s responsibility and history.
From faster calls to less repeated work
AI performance discussions often begin with the model endpoint: tokens per second, time to first token, parameter count, or price per million tokens. Those measures matter, but a system can waste most of its budget before or after the endpoint.
Repeatedly encoding the same document wastes inference. Sending several agents the same large transcript wastes context. Training methods that materialize every candidate separately waste parallel structure. The system-level question is not only “How fast is one call?” It is “Which work can be shared, preserved, filtered, or delayed across the entire job?”
Traditional computing solved analogous problems with query optimizers, caches, compilers, schedulers, and durable processes. The emerging AI stack is rediscovering those disciplines under probabilistic workloads and long context windows.
Dust amortizes search across tokens
Dust’s virtual population uses token positions as parallel perturbation members. Instead of copying and evaluating a full weight set for each candidate, it perturbs activations and recovers an update from reward-weighted noise. One forward pass can therefore test many variations.
The gain is not that each perturbation becomes free. Dust currently needs far more computation than backpropagation for similar loss. The gain is that the algorithm maps a large search population onto work a transformer can execute in parallel. Its relevance lies in the shape of the computation and the possibility of training non-differentiable structures, not today’s cost advantage.
This is an important distinction. Amortization can make an otherwise impossible method investigable without making it economical. Engineering teams should separate those milestones.
Quail amortizes context across a query
Quail has a more immediate cost target. In AI-SQL, the same document may participate in several filters and joins. A generic request server sees separate prompts and may re-encode the prefix each time. Quail sees a query plan. It retains selected KV caches, rewinds question-specific portions, chooses reusable anchors, and overlaps CPU preparation with GPU work.
This makes context a query asset with a lifetime. The optimizer decides when preserving it costs less than recomputing it. That is familiar database reasoning applied to model memory.
The broader lesson is that an AI operation should expose enough structure for a scheduler to reason about reuse. An opaque remote call can be convenient, but it limits control over cache placement, batching, and evidence. Local engines offer more knobs while transferring operational burden to the user.
OpenRig amortizes coordination across agent work
OpenRig addresses a different bottleneck: humans and agents repeatedly reconstructing who owns what. It defines persistent teams in YAML, gives agents fixed roles and addresses, supports messaging and task queues, and keeps each seat’s context separate. A terminal-oriented interface exposes state across the network.
The repository had 5,291 stars and 388 forks at the October 6 check, with an Apache-2.0 license. That attention suggests demand for a layer above individual coding harnesses. But orchestration is not evidence that more agents improve outcomes. It can also multiply model cost, synchronization failures, and security exposure.
OpenRig’s documentation is unusually explicit that setup can modify trust settings, hooks, tmux configuration, and local state. That matters because persistent agents are not just chat participants. They are processes with tools, credentials, and workspaces. Reusing coordination safely requires durable ownership boundaries and inspectable permissions, not only shared memory.
A three-layer execution model
Together the projects suggest three reusable layers:
- Learning execution decides how many candidate update signals can share a forward computation.
- Data execution decides which model context survives across operators in a query.
- Work execution decides which agent retains responsibility, context, and permission across tasks.
Each layer converts a collection of calls into a plan. Plans allow the system to reason about dependencies, lifetimes, and failure. They also create new state that can become stale or unsafe. A wrong KV cache, an outdated agent assumption, or a biased perturbation estimate can be reused efficiently and spread the error farther.
Practical design rules
First, measure the unit being saved. For Dust that is population evaluation per forward compute. For Quail it is prefix encoding and GPU idle time. For agent systems it may be repeated briefing tokens, duplicated code inspection, or human coordination minutes. Without a baseline, “reuse” becomes a slogan.
Second, make cache and context lifetimes explicit. Document state should be tied to model and prompt versions. Agent memory should name its owner, source, and expiration. Training experiments should record seeds, tuning runs, and total FLOPs. Persistent state is valuable only when invalidation is designed.
Third, keep outcome verification outside the optimized path. A faster semantic join still needs labeled evaluation. A multi-agent patch still needs tests and independent review. A new training method still needs replication on downstream tasks. Optimization cannot supply correctness by itself.
Fourth, include coordination overhead. A four-agent team that repeats the same repository scan may cost more than one well-instrumented agent. A retained KV cache may consume memory that reduces batch throughput. A search population may increase total compute even when parallelism improves. End-to-end task success per dollar and per hour is more useful than component speed.
Limitations
The three sources report their own results and designs. Dust remains research at modest model scales. Quail supports a narrow hardware and model set, and its benchmark claims need reproduction. OpenRig demonstrates orchestration features, not a controlled productivity improvement. Their appearance on the same day’s feed is a useful analytic lens, not evidence of a coordinated industry standard.
There is also a boundary to amortization. Some tasks genuinely need fresh context, independent samples, or isolated execution. Reusing too aggressively can correlate errors and leak information between work units. The correct system reuses what is invariant while preserving independence where it protects quality or security.
Bottom line
The next efficiency frontier is not only a smaller or faster model. It is an execution stack that understands repeated structure. Dust, Quail, and OpenRig show that idea at training, data, and work layers. Their maturity differs sharply, but the direction is coherent: turn disconnected AI calls into plans, make state lifetimes explicit, and verify the final outcome independently.
The teams that benefit will not be those that add the most caches or agents. They will be the ones that can say exactly what work was avoided, what state was reused, when it was invalidated, and whether the finished task became more reliable.
Sources
> Want more like this?
Get the best AI insights delivered weekly.
By subscribing, you agree to our Privacy Policy. You can unsubscribe at any time.
> Related Articles
Dust Trains Transformers Without Backpropagation—at a Steep Compute Price
Q Labs' zeroth-order method perturbs token activations and uses large virtual populations to estimate updates. The experiments are provocative, but they do not make backpropagation obsolete.
Quail Makes the SQL Query Plan Control LLM Inference
CMU's open-source AI-SQL engine reuses KV cache across filters and joins, overlaps CPU preparation with GPU work, and reports faster execution than a request-at-a-time baseline.
Local-First AI Is Becoming a Control Stack, Not Just a Model Download
Strata, OpenMuse, and RemoveMacAI expose three layers of personal AI control: where inference runs, what tools an agent can use, and which operating-system capabilities remain enabled.
Tags
> Stay in the loop
Weekly AI tools & insights.