Muse Glimmer Benchmarks
A 30B model scoring within striking distance of 70B+ models — and running on hardware you already own.
Benchmark Scores
Muse Glimmer measured against GPT-5.6 Sol and Qwen3.6 across six major evaluations.
| Benchmark | Muse Glimmer | Comparison |
|---|---|---|
| IFBench | 77.0 | GPT-5.6 Sol: 76.0 | Qwen3.6: 70.8 |
| AIME 2026 | 94.7 | GPT-5.6 Sol: 89.2 | Qwen3.6: 94.1 |
| GPQA Diamond | 83.5 | GPT-5.6 Sol: 85.7 | Qwen3.6: 84.2 |
| Humanity's Last Exam | 22.0 | GPT-5.6 Sol: 23.6 | Qwen3.6: 23.1 |
| AA-LCR | 80.0 | GPT-5.6 Sol: 68.3 | Qwen3.6: 73.3 |
| Beam 128K | 65.1 | GPT-5.6 Sol: 58.2 | Qwen3.6: 63.0 |
Scores reflect Meta's published results at launch. Higher is better for all benchmarks listed.
What the Numbers Mean
Instruction following is where Muse Glimmer outshines the competition. At 77.0 vs GPT-5.6 Sol's 76.0, it handles complex multi-constraint instructions better than models running in the cloud at scale.
Competition-level math on a single GPU. Scoring 94.7 puts Muse Glimmer at the frontier for mathematical reasoning — beating GPT-5.6 Sol at 89.2 and matching Qwen3.6 at 94.1.
Long-context retrieval is 12 points ahead of GPT-5.6 Sol (68.3). For document analysis, RAG pipelines, and code review over large codebases, this gap matters enormously in practice.
Benchmark FAQ
How can a 30B model beat GPT-5.6 Sol on some benchmarks?
What is IFBench?
What is AA-LCR?
Ready to see how it compares head-to-head or run it yourself?