Radar · 07/08/2026 · happened on 06/08/2026 · models

Meta's Muse Spark 1.2 reaches the frontier at $0.69 per test: price becomes the competitive criterion

Muse Spark 1.2, the update to Meta’s coding model launched in July, enters the top 5 of the Vals Index at $0.69 per test. Three times cheaper than Kimi, ten times cheaper than Claude Opus, Fable and GPT-5.6 Sol. It’s also the first model above 60% on Finance Agent v2, at $0.77 against Opus 5’s $5.12.

Meta reports gold-medal results across five STEM olympiads, with perfect scores in theoretical physics tests. It achieved these through pure reasoning and parallel multi-agent orchestration, without calling external tools.

For those using AI in their work, the signal matters more than individual numbers. Price per correct task is becoming the discriminator among frontier models. When performance is similar, cost decides. It’s the same pattern we saw with DeepSeek V4-Flash and the 80% price cut on Luna.

Meta also offers a “contributor” version at $0.10 per million input tokens and $0.20 output, if you accept that your data will be used to train their models. The standard version costs $1.25/$4.25.

In detail

What came before. Meta arrived late to the coding agent race. Muse Spark 1.1, in July, had aggressive pricing but performance still below the frontier tier. The 1.2 jump comes from scaling up training on coding: Meta increased compute on coding tasks, diversified training environments, and co-trained the model with Muse Code, its proprietary coding agent. The idea is that the model performs best when paired with its specific harness.

Numbers on the Vals Index. 1.2 goes from “unranked” to top 5 at $0.69 per test. On Finance Agent v2, it’s the first above 60% accuracy at $0.77 per attempt, against Opus 5’s $5.12, and at double speed. Artificial Analysis recorded one of the largest score jumps for 1.2 after updating evaluation criteria (patch v4.1.1). How much of the jump is real and how much is changed grading remains unclear, but the pattern is consistent across two independent benchmarks.

Olympiads without tools. Meta participated in five STEM olympiads (physics, mathematics, chemistry, mathematical reasoning) achieving gold-medal level in all of them, three under official competition conditions with official evaluation. The detail that sparked debate is the “no tools”: no search, no code, no calculator. Just reasoning, with parallel multi-agent orchestration. François Chollet and others discussed how much of the result is the model and how much is the surrounding harness. The underlying question, which resurfaces every time a lab publishes such numbers, is whether we’re measuring model capacity or the quality of the infrastructure around it.

Price as a competitive weapon. Two model IDs with the same engine. The standard version, muse-spark-1.2, costs $1.25 per million input tokens and $4.25 output, in line with Gemini 3.6 Flash ($1.50/$7.50). The contributor version, muse-spark-1.2-contributor, drops to $0.10/$0.20, close to GPT-5.6 Luna ($0.20/$1.20) and Gemini 3.1 Flash-Lite ($0.25/$1.50). The difference is data consent: in the contributor version, Meta uses your conversations for training. It converts privacy into discount: those who give up their data pay an order of magnitude less.

Limitations. Olympiad results are self-reported by Meta, though three were evaluated under official competition conditions. The “no tools” claim is under academic scrutiny. Prices are public, but actual serving capacity remains an open question. There are no independent tests yet on production tasks, only benchmarks. For those evaluating the model, the usual principle applies: benchmarks measure averages, not your case.

Type to search across course, playbooks, skills, papers…