Meta Says Muse Spark 1.3 Beats GPT-5.6 Sol at Coding — Independent Tests Are More Mixed

Meta Says Muse Spark 1.3 Beats GPT-5.6 Sol at Coding — Independent Tests Are More Mixed

Meta has unveiled Muse Spark 1.3, its latest AI system designed for complex coding and autonomous digital tasks as it competes with OpenAI and Anthropic. Image: Dima Solomin/Unsplash

Meta launched Muse Spark 1.3 with stronger coding performance and efficiency claims, but independent tests show higher task costs and mixed benchmark results.

Verfasst von
Aminu Abdullahi
Aminu Abdullahi
Sep 4, 2026

Meta says its latest AI model can compete with the industry’s strongest coding systems while using fewer tokens and tool calls. Independent testing suggests the efficiency story is more complicated.

The company released Muse Spark 1.3 to developers Wednesday through Muse Code and the Meta Model API, keeping pricing at $1.25 per million input tokens and $4.25 per million output tokens. Meta AI, Instagram and Facebook are expected to gain access later.

Alexandr Wang, Meta’s chief AI officer, called the launch the company’s “biggest jump so far on model performance.” Wang, according to Bloomberg, argued that Muse Spark 1.3 is “competitive” with Anthropic’s Claude Fable 5.1, “better than” OpenAI’s GPT-5.6 Sol at software development, and ahead of current Chinese models.

Meta Chief Executive Officer Mark Zuckerberg declared on X that the update delivers “frontier performance almost too cheap to meter.”

Meta claims the system handles single-threaded workflows across multiple tasks, asks for clarification when requests are ambiguous, and operates with roughly 25% fewer tokens and 20% fewer tool calls during internal coding workflows.

Independent tests complicate Meta’s performance claims

Third-party testing presents a more nuanced picture than Meta’s launch messaging.

Artificial Analysis placed the broadly deployable “xhigh” variant at 61 on its Intelligence Index, tied with GPT-5.6 Sol max but still trailing Anthropic’s Claude Fable 5.1, which leads at 66.

Meta’s highest scores come from a “max reasoning” configuration, which remains held back for extra safety testing. While benchmark sheets show Muse Spark 1.3 max logging 75.4 on DeepSWE v1.1 and 59.4 on SWEAtlas CodeBase QnA, companies cannot currently build on that specific tier.

Furthermore, Artificial Analysis noted that the average cost to run an evaluation task rose from $0.40 on version 1.2 to $0.55 on 1.3, largely because agent evaluations consume heavier volumes of input tokens.

That distinction matters for enterprises comparing models today: the version producing Meta’s strongest benchmark numbers is not the version developers can currently deploy.

More must-read AI coverage

Advertisement

Safety and open-source hesitation

Safeguards have taken a central role following an incident where an earlier model accessed the internet and infiltrated an external service during cybersecurity tests.

Wang said the occurrence informed improved resistance to prompt injections and added safeguards that pause to seek human approval before triggering irreversible operations.

Muse Spark 1.3 also leaves an important question unanswered about Meta’s open-model strategy. While the company still plans to release weights for the older Muse Spark 1.2, it has not committed to releasing the underlying weights for version 1.3.

Why this matters: The efficiency shift

The real transition signaled by Muse Spark 1.3 is not simply a contest over raw benchmark points, but a shift toward operational stamina.

For developers, peak intelligence is meaningless if an agent loops out of control, consumes massive token budgets, or requires constant manual course corrections. By engineering the system to recognize its own errors, decline hallucinated progress, and prune redundant tool calls, Meta is optimizing for workflow reliability.

For enterprise buyers, the useful question is therefore not whether Muse Spark 1.3 is simply “cheaper” or “more efficient.” It is whether the model completes a given workflow with fewer retries, fewer failed actions, and lower total cost than competing systems.

That is the benchmark that will matter once developers start using it at scale.

Muse Spark 1.3 puts Meta closer to the front of the coding-model race, but it also shows why benchmark leadership alone is becoming less useful.

Developers now have to compare not just raw intelligence scores, but token consumption, tool-call behavior, reliability, total task cost, and which reasoning tiers are actually available in production. Meta’s next test will be whether the efficiency gains it reports internally translate into cheaper and more dependable real-world workflows.

Let us teach you How to Talk to AI for free! Try our six-minute course at The Neuron Academy and learn a few simple ways to write better prompts and get more useful results from AI, or browse our other AI course for free for seven days. Check out all the lessons here →

Aminu Abdullahi

Aminu Abdullahi is a B2C and B2B technology and finance writer with more than six years of experience covering enterprise IT, cybersecurity, cloud computing, artificial intelligence, fintech, business software, and emerging technologies. He has written for a wide range of technical and business audiences, from IT professionals and cybersecurity leaders to small business owners, executives, and technology buyers. His work has appeared in publications including: TechRepublic eWEEK Channel Insider Geekflare Enterprise Networking Planet eSecurity Planet CIO Insight Webopedia With a background in computer science, Aminu specializes in translating complex technical subjects into clear, practical, and accessible content. His writing helps readers understand emerging technologies, evaluate business software, strengthen cybersecurity strategies, and make more informed decisions about technology investments. Across his work, Aminu focuses on the real-world impact of technology, connecting technical innovation with business value, operational efficiency, security, and long-term digital transformation.