Muse Glimmer Benchmarks

A 30B model scoring within striking distance of 70B+ models — and running on hardware you already own.

Benchmark Scores

Muse Glimmer measured against GPT-5.6 Sol and Qwen3.6 across six major evaluations.

Benchmark Muse Glimmer Comparison
IFBench 77.0 GPT-5.6 Sol: 76.0 | Qwen3.6: 70.8
AIME 2026 94.7 GPT-5.6 Sol: 89.2 | Qwen3.6: 94.1
GPQA Diamond 83.5 GPT-5.6 Sol: 85.7 | Qwen3.6: 84.2
Humanity's Last Exam 22.0 GPT-5.6 Sol: 23.6 | Qwen3.6: 23.1
AA-LCR 80.0 GPT-5.6 Sol: 68.3 | Qwen3.6: 73.3
Beam 128K 65.1 GPT-5.6 Sol: 58.2 | Qwen3.6: 63.0

Scores reflect Meta's published results at launch. Higher is better for all benchmarks listed.

What the Numbers Mean

77.0
IFBench — Beats GPT-5.6 Sol

Instruction following is where Muse Glimmer outshines the competition. At 77.0 vs GPT-5.6 Sol's 76.0, it handles complex multi-constraint instructions better than models running in the cloud at scale.

94.7
AIME 2026 — Elite Math

Competition-level math on a single GPU. Scoring 94.7 puts Muse Glimmer at the frontier for mathematical reasoning — beating GPT-5.6 Sol at 89.2 and matching Qwen3.6 at 94.1.

80.0
AA-LCR — Long Context

Long-context retrieval is 12 points ahead of GPT-5.6 Sol (68.3). For document analysis, RAG pipelines, and code review over large codebases, this gap matters enormously in practice.

Benchmark FAQ

How can a 30B model beat GPT-5.6 Sol on some benchmarks?
Post-training matters enormously. Meta focused Muse Glimmer's training on the specific capabilities that matter for local AI: instruction following, agentic tool use, and long context. It sacrifices breadth for depth in its target use cases.
What is IFBench?
Instruction Following Benchmark — measures how accurately a model follows complex, multi-constraint instructions. Muse Glimmer scores 77.0, beating GPT-5.6 Sol at 76.0.
What is AA-LCR?
A long-context retrieval benchmark. Muse Glimmer scores 80.0, significantly beating GPT-5.6 Sol (68.3) and Qwen3.6 (73.3) — notable because long context on local hardware is a key use case.

Ready to see how it compares head-to-head or run it yourself?