TNS
VOXPOP
As a JavaScript developer, what non-React tools do you use most often?
Angular
0%
Astro
0%
Svelte
0%
Vue.js
0%
Other
0%
I only use React
0%
I don't use JavaScript
0%
NEW! Try Stackie AI
AI / AI Agents / AI Models / Large Language Models

“Google was ahead only a few hours”: Muse Spark 1.3 edges out Gemini as Meta claims its biggest coding leap yet

Sep 3rd, 2026 11:37am by
Featued image for: “Google was ahead only a few hours”: Muse Spark 1.3 edges out Gemini as Meta claims its biggest coding leap yet

It has been a big week for Meta in the AI sphere, formally launching its Muse Code coding agent out of beta with a triumvirate of new subscription plans. Then late on Wednesday, the company unveiled Muse Spark 1.3, the latest version of its low-cost reasoning model, with Meta claiming its biggest gains yet in coding and agentic tasks.

Available through Muse Code and the Meta Model API, Muse Spark 1.3 is the latest iteration of the reasoning model Meta first introduced in April, and arrives less than a month after the release of Muse Spark 1.2. Meta CEO Mark Zuckerberg took to social media to hype the new release, describing it as delivering “frontier performance almost too cheap to meter.”

“This is the biggest jump we’ve made so far on coding and agentic work.”

“This is the biggest jump we’ve made so far on coding and agentic work,” Zuckerberg writes on X.

Open-weight release ‘coming soon’

Zuckerberg also confirmed that the open-weight releases of Muse Spark will be “coming soon,” a move that could let developers download and run the model on their own infrastructure, bypassing Meta’s hosted API.

Exactly how permissive that will be, however, isn’t clear, as Meta has yet to publish the licence terms that will accompany the release. Its previous open-weight models have carried varying restrictions on use, so those details will determine how freely Spark can be modified, redistributed or deployed.

Notably, Zuckerberg also teased Meta’s much-hyped next model, codenamed Watermelon, using the somewhat apt watermelon emoji.

That model is understood to be a much larger system than Spark and was reportedly still in training in July, according to a Business Insider report at the time. There is still no firm word on when Watermelon will see the light of day, though the implication from Zuckerberg is that it won’t be long.

For now, though, Meta is making some sizeable claims for Spark 1.3 itself. Zuckerberg accompanies his post with a benchmark table pitting the model against OpenAI’s GPT-5.6 Sol and Anthropic’s Claude Opus 5 across coding, agentic, computer-use and long-context tests.

Among the standout figures, Muse Spark 1.3 scored 75.4% on the DeepSWE coding benchmark and 98.1% on the 512K-1M version of the MRCR long-context test.

Muse Spark 1.3: Benchmarked
Muse Spark 1.3: Benchmarked

These are Meta-assembled results, however. The company says its Spark 1.3 scores were generated through the Meta Model API, while comparison figures are drawn from a mixture of Meta’s own evaluations, official leaderboards and results reported by rival model providers. Meta also describes its testing of third-party models as “best-effort,” meaning the table shouldn’t be read as a single independent head-to-head test conducted under identical conditions

There is another caveat to these results: Meta ran Spark 1.3 at its new “max” reasoning level for the headline comparisons, while the highest reasoning level generally available to developers today is “xhigh.” Max remains in limited preview while Meta completes additional safety testing.

“Gloves are off”: How Spark 1.3 stacks up

Independent testing by Artificial Analysis does offer a useful outside perspective. The San Francisco-based company, which specializes in benchmarking AI models and providers, gives the publicly available Muse Spark 1.3 xhigh a score of 61 on its Intelligence Index, four points ahead of Spark 1.2 and level with GPT-5.6 Sol max, Grok 4.6 high and Claude Opus 5 high. It was also able to test the limited-preview max version, which scored 62, placing it behind only Claude Fable 5.1 and Claude Opus 5 among the models in its comparison at launch.

Artificial Analysis Intelligence Index (credit: Artificial Analysis)
Artificial Analysis Intelligence Index (credit: Artificial Analysis)

Alex Volkov, AI evangelist at cloud infrastructure company CoreWeave, points to the chart as evidence of how quickly Meta has closed the gap with the leading models.

“Damn, gloves are off!,” Volkov writes on X, calling the showing “quite the statement” from Meta while predicting “busy weeks ahead of us!”

“Damn, gloves are off!”

Cost is another part of Meta’s pitch. Artificial Analysis describes Spark 1.3 xhigh as the “most cost-efficient model” at its level of measured intelligence, with its combination of a 61 Intelligence Index score and comparatively low per-task cost placing it on the company’s Pareto line.

Intelligence Index vs. Cost per Intelligence Index Task (Credit: Artificial Analysis)
Intelligence Index vs. Cost per Intelligence Index Task (Credit: Artificial Analysis)

Artificial Analysis calculates that Spark 1.3 xhigh costs around $0.55 per Intelligence Index task, the lowest of any model scoring 59 or higher on its index. GPT-5.6 Sol max and Grok 4.6 high, both tied with Spark at 61, came in at $0.95 and $0.94 respectively.

Spark 1.3 was still more expensive per task than Spark 1.2’s $0.40, with Artificial Analysis attributing that increase largely to the newer model consuming around 57% more input tokens on agentic evaluations.

Cost per Intelligence Index Task (Credit: Artificial Analysis)
Cost per Intelligence Index Task (Credit: Artificial Analysis)

Muse meets Gemini: A ‘playground slapfight’

It’s worth noting that shortly before Meta unveiled Muse Spark 1.3, Google released a new low-cost model of its own: Gemini 3.8 Flash, its third Flash release in just six weeks. Google also pitched its model as its best reasoning and coding Flash model yet, while keeping introductory pricing at $0.75 per million input tokens and $3.75 per million output tokens.

Artificial Analysis initially placed Gemini 3.8 Flash (high) on its Intelligence-versus-Cost Pareto frontier — essentially the group of models for which there is no alternative that is both more capable and cheaper. Gemini scored 59 on the Intelligence Index at a cost of $0.58 per task.

However, just a few hours later, Muse Spark 1.3 xhigh arrived at 61 and $0.55 per task, beating Gemini 3.8 Flash on both measures and pushing it off that frontier.

“Google was ahead only a few hours.”

This point was not lost on many in the AI community. Indeed, AI researcher Benjamin Marie notes on X that “Google was ahead only a few hours.

Florian Brand, a research engineer at AI infrastructure and research company Prime Intellect, also sums up the turnaround succinctly.

“Gemini 3.8 held a spot at the pareto frontier for *checks notes* 3.5 hours,” Brand writes on X.

And none other than Meta’s chief AI officer himself Alexandr Wang weighed in, taking the time to cast shade at Google off the back of the Artificial Analysis report.

However, Corey Quinn, co-founder and chief cloud economist at cloud and AI cost management company Duckbill, is quick to mock the exchange, suggesting that neither Meta nor Google are meaningfully setting the pace in the AI arms race.

“Meta casting shade at Google in AI is a playground slapfight outside a MMA championship,” Quinn writes on X.

Muse Code enters the scene

For context, Muse Spark 1.3 is the fourth version of the model Meta has shipped since April, and follows Muse Spark 1.1 in July and 1.2 the month after. Arguably the more important part of Meta’s push, though, is the agent harness it’s building around the model.

That comes in the form of Muse Code, Meta’s terminal-based coding agent, which arrived in beta alongside Spark 1.2 in early August. It officially launched on Tuesday, replete with new subscription plans starting at just $5 per month and a new SDK also hitting developer preview.

So while Meta is clearly trying to compete with the frontier labs on model quality, evidenced by the gains in the Muse Spark lineup, it’s also chasing them a level up, where Anthropic’s Claude Code and OpenAI’s Codex have set the pace for how developers actually work with an agent day to day.

And this explains why so much of the 1.3 release language centers on coding and agentic work specifically. The model is a key piece of a broader developer platform Meta is now pushing hard on price as much as capability.

📢 แจ้งเตือน: muse spark gemini - Kapook Update 2025 Created with Sketch.
TNS owner Insight Partners is an investor in: OpenAI, Anthropic.
TNS DAILY NEWSLETTER Receive a free roundup of the most recent TNS articles in your inbox each day.