Meta Releases Muse Spark 1.3 Model for Longer Tool-Based Work

Meta released its Muse Spark 1.3 model in Muse Code and the Meta Model API for longer work that uses external tools.

TL;DR
  • Developer Release: Meta released its Muse Spark 1.3 model in Muse Code and the Meta Model API for longer work that uses external tools.
  • Same Unit Price: Standard rates remain $1.25 per million input tokens and $4.25 per million output tokens, matching Muse Spark 1.2.
  • Mixed Results: Artificial Analysis found the available higher-effort xhigh mode improved from 1.2 on two tool-use tests but regressed on long-document and knowledge tests.
  • Mode Limit: Meta says its more compute-intensive max mode will remain limited until additional safety testing is complete.

Meta released Muse Spark 1.3, a multimodal reasoning model, on September 2 for longer work that uses external tools. Developers can access it through Muse Code, Meta’s terminal coding agent, and the Meta Model API, which lets applications call the model; standard prices remain unchanged from 1.2. Muse Spark 1.3 is designed to keep longer coding and research tasks on track, ask the user when requirements are unclear and seek confirmation before consequential actions.

Developers can now test the new model in the two surfaces established for Muse Spark 1.2 in August. Meta did not make every configuration from its benchmark table generally available. The xhigh reasoning setting is available, while the more compute-intensive max mode remains limited to partners until Meta completes additional safety testing.

Muse Spark began in April as the first model from Meta Superintelligence Labs. Meta initially offered developers a private API preview. Muse Spark 1.2 and the Muse Code beta created the developer baseline that 1.3 now updates.

What Developers Can Use Now

The standard API rates remain $1.25 per million input tokens, $0.15 per million cached input tokens and $4.25 per million output tokens. Those units match Muse Spark 1.2, but they do not fix the price of a completed job. A workload’s bill still depends on its input, output, cache use, reasoning effort and tool activity.

Muse Spark 1.3 has a one-million-token combined context window. The model accepts text, images and video as input and produces text. Exact account and geographic eligibility remain unspecified.

The model remains proprietary. Meta has put an open-weights Muse Spark release on its roadmap, but it has not identified the checkpoint, terms or date.

What Agentic Work Means Here

Meta uses “agentic” for a model that can work through a multi-step task using tools, retain requirements during a long exchange, gather information from different sources and revise its plan when it encounters a gap. The model can also pause to ask the user for clarification or help and request confirmation before an irreversible action.

Those behaviors matter in Muse Code because the product is intended to plan, write and validate changes across a software repository. Meta says the model can incorporate an added requirement during an interrupted task instead of beginning the work again. The company also says it trained 1.3 on more long-horizon coding work, but it does not disclose enough about those training tasks to turn the design claim into a general success rate.

Meta reports that its engineers saw about 20% fewer tool calls and 25% fewer tokens when comparing 1.3 with 1.2. The company has not published the task count, accuracy pairing or variance for that internal comparison, so the figures describe Meta’s test rather than every coding workload.

Independent Tests Show Mixed xhigh Results

Artificial Analysis measured the available xhigh mode against its xhigh predecessor under the same evaluation framework. On Tau3-Bench Banking, which tests an agent’s ability to complete tool-mediated banking tasks, the pass rate rose from 35% to 47%. On Terminal-Bench 2.1, a set of command-line tasks checked by executable verifiers, it rose from 80% to 85%.

Meta AI Muse Spark 1.3 Artificial Analysis Index

The same independent testing also found regressions. The score on a 100-question long-document test fell from 83% to 79%, while accuracy on a 6,000-question knowledge test slipped from 45% to 42%. Artificial Analysis linked the latter decline to more abstentions, which reduced wrong answers counted by its separate hallucination measure. Together, the results support improvement on selected tool-heavy tasks, not an across-the-board advance.

Meta’s broad comparison table used 1.3 in max mode but compared it with 1.2 in xhigh mode, and some rows came from provider reports or public leaderboards rather than one common test run. Those results describe Meta’s max evaluation package, not the generally available xhigh configuration on release day.

Meta AI Muse Spark 1.3 Benchmarks

Stable token rates also do not guarantee a lower bill. Artificial Analysis measured the cost of its composite index task rising from about $0.40 for 1.2 to $0.55 for 1.3 because the newer model used more input tokens in that evaluation. Developers therefore have to judge both task success and token use on their own workload.

Safety Claims Meet the Release Boundary

Meta says Muse Spark 1.3 is more resistant to adversarial inputs and prompt injection, and better calibrated about actions that cannot be undone. The public release gives no release-specific rates, sample sizes or independent safety measurements for those claims. They establish Meta’s safety position, not superiority over another model.

Meta says general access to max will follow additional safety testing, but it gives no date or completion criteria. Until that changes, developers can use xhigh while the strongest reasoning mode in the release’s scorecard remains outside general access.

Meta presents Muse Spark as progress toward personal superintelligence, but the September 2 release is a developer model rather than a shipping personal agent or a new consumer rollout. The same unit prices leave a concrete question for developers: whether 1.3 completes their own workloads more successfully without using enough extra tokens to make each successful task more expensive.

Markus Kasanmascheff
Markus Kasanmascheff
Markus has been covering the tech industry for more than 15 years. He is holding a Master´s degree in International Economics and is the founder and managing editor of Winbuzzer.com.
Subscribe
Notify of
guest
0 Comments
Newest
Oldest Most Voted