
Meta’s newest AI model Muse Spark 1.3, unveiled yesterday, is faster and more performant on third-party benchmarks than its predecessor — with a caveat. The company’s strongest results come from a configuration still completing safety testing, leaving developers with a broadly available version that is very good but not the benchmark leader.
What the benchmarks actually show
Muse Spark 1.3 makes significant gains over last month’s 1.2 release, particularly on long-running agent tasks. The version developers can access now uses Meta’s previously available reasoning settings, including xhigh. Meta’s strongest benchmark results come from its max reasoning configuration, which the company says will arrive “shortly” after completing additional safety testing. Artificial Analysis evaluated max in a limited partner preview and currently lists no API provider for the configuration at all.
Related: OpenAI Launches GPT-6 Astra, Ushering AGI Era
Meta does disclose results for both configurations in its underlying evaluation report. On GDPval-AA v2, the company reports scores of 1,754 Elo for max versus 1,709 for xhigh. On OSWorld 2.0, max scores 66.9 versus xhigh at 57.2. The gap narrows on some tests: DeepSearchQA is tied at 89.4, while xhigh edges ahead on Terminal-Bench 2.1 with 89.2 versus max at 88.8.
Artificial Analysis scores Muse Spark 1.3 max at 62 on its Intelligence Index and the shipping xhigh version at 61. The xhigh variant ties GPT-5.6 Sol max, Grok 4.6 high and Claude Opus 5 high. Anthropic still occupies the top of the leaderboard: Claude Fable 5.1 reaches 66 at max and 65 at xhigh, while Claude Opus 5 reaches 63 at both settings.
Muse Spark 1.3 xhigh is legitimately in the frontier cluster, but it is not currently setting the frontier. That still marks a substantial change from 1.2, which generally trailed Anthropic’s best model on the coding comparisons Meta presented. With 1.3, Meta is no longer merely showing up in that contest — on several coding and agentic evaluations, it is trading wins with OpenAI and Anthropic.
Related: Prompt injection tops security threat lists
Token pricing stayed the same
Muse Spark 1.3 did not receive an API price cut. Meta kept Standard pricing exactly where it was for 1.2: $1.25 per million input tokens, $4.25 per million output tokens and $0.15 per million cached input tokens. That makes Zuckerberg’s “almost too cheap to meter” line less a statement about lower token prices than about what Meta believes developers can accomplish with those tokens.
For enterprises paying for thousands or millions of agent loops, behavioral improvements could matter more than another leaderboard point. Meta says the underlying model has become easier to operate: it maintains multiple workflows in a long thread, gathers context with tools, detects gaps in its own plans, asks users for clarification when necessary and confirms before consequential actions. In Meta engineers’ internal comparisons, it used roughly 20% fewer tool calls and 25% fewer tokens than 1.2 during coding work.
Meta also retains its Contributor tier at $0.10 per million input tokens and $0.20 per million output tokens in exchange for permission to use prompts and completions for training. That may be attractive for prototyping but creates a materially different data-governance calculation for enterprises working with proprietary code or sensitive internal information.
Related: Quectel unveils Android 16 IoT modules
How it stacks up against Google’s Gemini
Meta chief AI officer Alexandr Wang was considerably less qualified in celebrating the release. After Artificial Analysis posted its Muse Spark results, Wang reposted them on X, adding: “i really hate to say it, but… gemini who? 😱💨” The shade was particularly pointed because Google released Gemini 3.8 Flash on the same day, pitching it at almost exactly the same class of workload: long-horizon software engineering, autonomous agents and multi-step professional reasoning.
Independent numbers give Wang something to work with, though hardly a knockout. Artificial Analysis gives Muse Spark 1.3 xhigh a 61 Intelligence Index score at $0.55 per task, compared with 59 and $0.58 for Gemini 3.8 Flash at high reasoning. Meta edges Google on both intelligence and task cost at those particular settings. Google wins decisively on throughput: Artificial Analysis measures Gemini 3.8 Flash high at about 305 output tokens per second, versus 235 for Muse Spark — roughly 30% faster.


