Muse Spark 1.1 vs Grok 4.5: Benchmarks & Pricing
Maya Chen
Lead AI Researcher

TLDRMuse Spark 1.1 costs $1.25/$4.25 per M tokens with a 1M context window; Grok 4.5 costs more per task but wins on some game-code tests. Full comparison.
Muse Spark 1.1 vs Grok 4.5: Benchmarks, Pricing, and How They Compare
Muse Spark 1.1 matches Grok 4.5 on the Artificial Analysis Intelligence Index (both score 51) and beats it on cost per task ($0.26 vs $0.31), while Grok 4.5 still edges ahead on some real-world game-code generation tests — the right pick depends on whether your workload is agentic reasoning at the lowest price or physics-heavy code that must ship on the first run.
Key Takeaways
- Intelligence tie: Both models score 51 on the Artificial Analysis Intelligence Index, placing them in the same frontier tier.
- Muse Spark 1.1 is cheaper per task: $0.26 vs $0.31 on the AA Intelligence Index; API pricing is $1.25 / $4.25 per million input/output tokens.
- Grok 4.5 wins on game-code robustness: In a third-party test of three arcade-style code prompts, Grok 4.5 nailed physics and playback on all three where Muse Spark 1.1 lagged badly on the Crossy Road clone.
- Muse Spark 1.1 wins on theoretical CS: It outperformed Grok 4.5, Opus, and Gemini on a new finite model theory eval, per Meta's Alexandr Wang.
- Context window: Muse Spark 1.1 ships with a documented 1M-token window; Grok 4.5's window is not in the public sources cited here.
- Head-to-head coverage is thin. Only a handful of direct comparisons exist as of mid-July 2026, and most benchmark chatter comes from Meta-affiliated authors — read with that in mind.
Muse Spark 1.1 vs Grok 4.5 at a Glance
Sources: Artificial Analysis model page, Meta's official Muse Spark 1.1 announcement.
Benchmarks: Where Each Model Wins
The best head-to-head data point comes from Artificial Analysis, which measured Muse Spark 1.1 (xhigh) at 69 on the Coding Agent Index in the Opencode harness — landing "just below GPT-5.5 (medium) in Codex (71) and ahead of Claude Opus 4.8 (medium) in Claude Code (67)." Cost per task is roughly $1.4, among the lowest of frontier coding agents, with the tradeoff of higher time per task.
Source: @ArtificialAnlys
On the broader Intelligence Index, both models tie at 51. Muse Spark 1.1's cost per task is $0.26 versus Grok 4.5's $0.31 — a ~16% cost advantage on the same workload.
DeepSeek researcher Teortaxes noted that "Muse Spark 1.1 is surprisingly close to Grok 4.5 on many high-signal evals" and topped CritPT at the time of posting on July 11, 2026.
Source: @teortaxesTex
Meta's Alexandr Wang has posted a series of narrower wins for Muse Spark 1.1:
- Finite model theory / theoretical CS eval: beats Opus, Grok 4.5, and Gemini.
- Radiologists Last Exam Handover Readiness Index (RadLE-H): SOTA, nearing human expert performance.
- HealthBench Professional: SOTA.
- Debate Benchmark: #3 behind Fable 5 and Opus 4.7, ahead of GPT-5.6 Sol.
These are single-source claims from a Meta executive. Treat them as directional until third-party leaderboards confirm.
Where Grok 4.5 wins: a third-party test from AI/ML API ran three game-generation prompts (Fruit Ninja slicer, Angry Birds fort collapse, Crossy Road clone) and reported that "Grok nailed all three: clean physics, smooth playback, good visuals." Muse Spark 1.1 was cheapest per run ($1.08 vs Grok's $2.47), but "its Crossy Road lagged badly." The verdict: "Cheap stops being cheap when it doesn't work." This is a small n=3 test, but it's the clearest concrete failure mode documented so far.
For more on Meta's model, see our Muse Spark 1.1 deep dive.
Pricing and Cost-Efficiency
Muse Spark 1.1 API pricing, per Meta's official documentation:
- Input: $1.25 per 1M tokens
- Output: $4.25 per 1M tokens
- Cache hit price: $0.15 per 1M tokens (88% discount)
- Free credit: $20 to start
Grok 4.5's list API pricing is not surfaced in the sources gathered for this page. What is measurable is normalized cost per Intelligence Index task: Muse Spark 1.1 at $0.26, Grok 4.5 at $0.31 (per Artificial Analysis). At this scale, the difference is meaningful for high-volume agentic workflows but not decisive for occasional use.
The full evaluation cost for Muse Spark 1.1's Intelligence Index run was $548.07. Output verbosity is above average at 94M tokens for the full index (versus a 60M average), which erodes some of the input-price advantage on long generations.
Context Window and Modalities
Muse Spark 1.1 ships with a documented 1M-token context window, plus text and image input (text output only). Meta emphasizes that the model "actively manages" its context, compacting earlier work while keeping steps needed later — a claim aimed at long-running agentic sessions.
Grok 4.5's context window is not disclosed in the sources gathered here. If you need a large-context agent on public numbers alone, Muse Spark 1.1 has the confirmed spec. For xAI-side details, see our Grok 4.5 leak analysis.
Availability
Muse Spark 1.1 is available through the new Meta Model API (public preview as of July 9, 2026), in "Thinking" mode inside the Meta AI app, at meta.ai, and — according to a community post — via a listing on OpenCode. Regional availability is limited; not all developer keys are available in all regions.
Grok is available through xAI's own API and inside X. Both models operate on paid-API terms similar to Anthropic and OpenAI.
Which One Should You Use?
Choose Muse Spark 1.1 if:
- You're running high-volume agentic workflows and want the cheapest cost per task at frontier-tier intelligence.
- You need a documented 1M-token context window for long-horizon tasks.
- Your work is quantitative-but-legible: multi-step calculation over structured documents, procedural review, report drafting. A third-party enterprise benchmark writeup found Muse Spark 1.1 particularly strong on this class of work.
- You want text + image input in one API call.
Choose Grok 4.5 if:
- You're generating self-contained code that must run correctly on the first shot — particularly graphics, physics, or game logic where a failed output cascades into rework.
- You're already in the xAI ecosystem or need Grok-specific integrations.
- You value the specific personality and behavior tuning xAI has invested in over Meta's more research-lab framing.
If cost isn't the constraint and the workload is coding-heavy, note that Artificial Analysis ranks Claude Fable 5 highest on the Intelligence Index (60), ahead of both models compared here.
Frequently Asked Questions
Is Muse Spark 1.1 better than Grok 4.5?
Muse Spark 1.1 and Grok 4.5 tie on the Artificial Analysis Intelligence Index at 51, with Muse Spark 1.1 winning on cost per task ($0.26 vs $0.31) and losing on some game-code generation tasks. Neither model dominates across the board.
Is Muse Spark 1.1 cheaper than Grok 4.5?
Yes. Artificial Analysis measured $0.26 per task for Muse Spark 1.1 (xhigh) versus $0.31 for Grok 4.5 (high) on the Intelligence Index. Muse Spark 1.1 API pricing is $1.25 per million input tokens and $4.25 per million output tokens.
Which is better for coding, Muse Spark 1.1 or Grok 4.5?
Muse Spark 1.1 scores 69 on the Artificial Analysis Coding Agent Index in the Opencode harness, sitting between Claude Opus 4.8 (67) and GPT-5.5 (71). Grok 4.5's Coding Agent Index score is not published in comparable form, but Grok 4.5 outperformed Muse Spark 1.1 on three game-generation prompts in a third-party test.
What is the context window difference between Muse Spark 1.1 and Grok?
Muse Spark 1.1 has a 1 million token context window, per Meta's official documentation. Grok 4.5's context window is not documented in the sources available for this comparison.
When was Muse Spark 1.1 released compared to Grok 4.5?
Muse Spark 1.1 was released on July 9, 2026, one day after Grok 4.5's July 8, 2026 release, according to model tracker Vals.ai.
Can I use Muse Spark 1.1 through an API like Grok?
Yes. Muse Spark 1.1 is available through the new Meta Model API in public preview, launched alongside the model on July 9, 2026. Grok is available through xAI's own API. Meta offers $20 in free credits to start.
Does Muse Spark 1.1 beat Grok 4.5 on any benchmarks?
Muse Spark 1.1 outperformed Grok 4.5 on a finite model theory / theoretical CS evaluation and is reported to be surprisingly close to Grok 4.5 on many high-signal evals, including topping CritPT at the time of measurement. Grok 4.5 leads on some game-code generation tasks.
What to Watch Next
Three signals will decide how this comparison shapes up over the next month: (1) whether independent leaderboards confirm the Meta-authored eval wins on RadLE-H, HealthBench Professional, and Debate Benchmark; (2) whether Grok 4.5 publishes a Coding Agent Index score in the same Opencode harness for an apples-to-apples code comparison; and (3) whether Meta ships the rumored "Watermelon" mega-sized variant referenced in Reddit chatter, which would reset the frontier tier entirely.
Building similar chat and agentic workflows? On Uptech API you can try Grok 4.5, GPT-5.6, and Claude Opus 4.8.
About Maya Chen
Maya tracks AI model releases, benchmarks, and developer adoption signals across the open and closed model landscape.
About Maya Chen
Maya tracks AI model releases, benchmarks, and developer adoption signals across the open and closed model landscape.
