Muse Spark 1.3 reaches GPT-5.6 Sol level on “xhigh”
Muse Spark 1.3 is now publicly available through the Muse Code coding agent, which runs on macOS and Linux; it can also be accessed across platforms through the Model API. The independent benchmarking platform Artificial Analysis gives the available Reasoning tier “xhigh” 61 points on the Intelligence Index, putting it on par with GPT-5.6 Sol “max”.
Meta is targeting the new version at software development and general agentic tasks. Compared with Muse Spark 1.2, Spark 1.3 is designed to follow complex instructions more reliably, handle multiple tasks simultaneously more effectively and improve collaboration with users. The model processes text, images and video; its context window remains unchanged at one million tokens.
Meta’s own benchmarks also show progress. At the same “xhigh” Reasoning tier, Spark 1.3 clearly beats its predecessor on GDPVal-AA v2 and Terminal-Bench 2.1. On GDPVal-AA v2, it is virtually on par with GPT-5.6 Sol “max”, but trails Claude Opus 5 “max”; on Terminal-Bench 2.1, it narrowly beats both. However, the comparison graphic in Meta’s blog post consistently pits Spark 1.3 with “max” against Spark 1.2 with “xhigh” — at the same tier, some of the gains are smaller.
Artificial Analysis confirms the increase from 57 to 61 points and sees the biggest gains in agentic knowledge work. The higher “max” tier scores 62 points but is not yet generally available at launch. For the additional point, it requires about 62% more reasoning tokens than “xhigh” on GDPVal-AA, according to Artificial Analysis; Meta plans to make “max” available after further safety testing.
For developers, the practical change is that they can already use a more capable model for coding tasks through Muse Code or the API, without waiting for the announced release of open model weights. The regular price remains $1.25 per million input tokens and $4.25 per million output tokens. The Contributor plan costs $0.10 for input, $0.20 for output and $0.002 for cached inputs — with the clear trade-off that Meta may use prompts and responses to train future models.
Comments
Loading the thread…
Sign in to leave a comment. Sign in