AI NEWS 7 min read

Solar Mini 4 Targets Repetitive Agent Work With 3B Active Parameters

Upstage's new sparse model combines a 35B-parameter footprint with 3B active parameters per token, structured output, parallel tool calls, and a one-H100 quantized deployment target.

By EgoistAI ·
Solar Mini 4 Targets Repetitive Agent Work With 3B Active Parameters

Upstage has released Solar Mini 4, a mixture-of-experts language model designed for high-volume tasks such as information extraction, classification, routing, structured decisions, and tool use. The architecture contains about 35 billion parameters but activates roughly 3 billion for each token. Upstage says a quantized deployment can fit on one 80 GB H100, while its published throughput example uses two H100s for 32 simultaneous requests.

That distinction matters. “Three billion active” describes computation per token; it does not mean the model stores only three billion parameters. Teams evaluating self-hosting still need to account for the full weights, the key-value cache, concurrency, quantization, serving software, and operational headroom.

What happened

The official model documentation lists a 512,000-token context window, JSON-schema structured output, parallel tool calling, and adjustable reasoning effort. Upstage positions the model below its Pro line: Mini is intended to make repetitive workflow steps cheap enough to run at scale, while harder planning or synthesis can be escalated.

The public API list price is $0.10 per million input tokens, $0.01 per million cached input tokens, and $0.40 per million output tokens. OpenRouter also lists the model, giving developers a second public place to inspect availability and price. Upstage offers API access, a chat interface, and enterprise deployment options.

The company reports 22.3 percent on AutomationBench-AA and 47.2 on τ³-Banking. It also cites an Artificial Analysis Intelligence Index score of 24.1 among models with a similar active-parameter range. These are vendor-selected comparisons and should be read as starting points for evaluation, not a guarantee on private workloads.

Why it matters

Most production agent traffic is not a frontier reasoning problem. A support workflow may classify a request, retrieve two records, apply a policy, and return a typed object. A document system may extract the same fields from thousands of pages. Paying a large model for every step can make a useful automation uneconomic even when its accuracy is acceptable.

Solar Mini 4 is an explicit bet on routing. A smaller sparse model handles well-specified, repeated operations; a larger model receives ambiguous or high-risk cases. The combination can lower average cost without pretending every task is easy.

Single-GPU feasibility also changes who can test private deployment. It does not make deployment simple: an H100 is expensive, and a reliable service still needs monitoring, batching, access controls, upgrades, and failure recovery. But one-device minimum capacity is a materially different procurement problem from a multi-node cluster.

Evidence

Upstage’s documentation is the primary source for architecture, context, features, price, and deployment claims. The GeekNews post surfaced the release and made the Korean-language community response visible. OpenRouter confirms a third-party served endpoint and its posted commercial terms.

The evidence remains mostly first-party. The benchmark configurations do not represent every language, tool schema, or latency target. Upstage’s two-H100 throughput result uses a 4K input and 1K output without prompt caching; it should not be treated as the throughput of the one-H100 minimum configuration. Long-context support also does not imply that half a million tokens will be fast, cheap, or equally accurate across the window.

Practical takeaway

Evaluate Solar Mini 4 on a narrow workflow before treating it as a general agent brain. Build a held-out set of real documents and tool calls. Measure schema-valid responses, field-level accuracy, tool-selection error, latency at expected concurrency, and the rate at which cases must escalate. Test both the ordinary path and adversarial inputs that place instructions inside retrieved content.

For self-hosting, measure memory under the chosen quantization and maximum context, not just the weight-file size. Include idle capacity, container overhead, cache growth, and failover in the cost comparison. For API use, add output tokens, retries, and cached-input assumptions to the headline input price.

The useful question is not whether a 3B-active model “beats” a large model on a composite chart. It is whether it completes one bounded job reliably enough that the larger model can be reserved for exceptions.

Limitations

Independent replications were limited at publication time. Prices and served features can change. Upstage has not disclosed every training-data or operational detail needed for a full risk assessment. A supported context length is not a quality guarantee, and a model suitable for document routing is not automatically suitable for legal, medical, financial, or autonomous high-impact decisions.

Share this article

> Want more like this?

Get the best AI insights delivered weekly.

By subscribing, you agree to our Privacy Policy. You can unsubscribe at any time.

> Related Articles

Tags

Solar Mini 4Upstagesmall language modelsAI agentson-premises AI

> Stay in the loop

Weekly AI tools & insights.