Meta
Search
Documentation

Products
Muse Code
Overview
Meta Model API
Overview
API Login
Models
Muse
Muse Spark 1.3
Muse Glimmer
Muse Image
Muse Voice Transcribe
Llama
Llama 4
Llama 3
Resources
Documentation
Model API docs
Learn
Cookbooks
Videos
Blog
Case studies
Community
Github
Meta Models
Llama
Hugging Face
Meta Models
Safety
Llama Protections
Overview
Llama Defenders Program
Developer use guide

ModelsMuseMuse Glimmer
Stay updated
Get started
Meta
post image
post image
post image
post image
Products
Muse Code
Meta Model API
Models
Muse Spark 1.2
Muse Spark 1.1
Muse Glimmer
Muse Voice Transcribe
Llama 4
Llama 3
Documentation
Meta Model API Docs
Muse Glimmer Docs
Llama Docs
Resources
Cookbook
Blog
Videos
Case studies
FAQs
Community
Meta-Models Github
Llama GitHub
Hugging Face
Terms & policies
Terms of Service
Privacy Policy
Cookie Policy
Products
Muse Code
Meta Model API
Models
Muse Spark 1.3
Muse Spark 1.2
Muse Spark 1.1
Muse Glimmer
Muse Image
Muse Voice Transcribe
Llama 4
Llama 3
Documentation
Meta Model API Docs
Muse Glimmer Docs
Llama Docs
Resources
Cookbook
Blog
Videos
Case studies
FAQs
Community
Meta-Models Github
Llama GitHub
Hugging Face
Terms & policies
Terms of Service
Privacy Policy
Cookie Policy

Muse Glimmer

An open model built for always-on local agents. 30B parameters, licensed under Apache 2.0 and runs on a single GPU — tuned for tool use, long tasks, and failure recovery.
Download model
Documentation

Meet Muse Glimmer

Always-on agents
Built for agents that don't stop with reliable tool-calling, persistent state across restarts, and self-managed memory across hours-long sessions.
Optimized for local deployments
It's small enough to run on a single consumer GPU or Mac, enabling use cases that range from local agents to local coding.

Benchmarks

Muse Glimmer is trained to deliver competitive agentic and coding performance, with multimodal perception built in.
Benchmark
Muse Glimmer-30BHigh reasoning
Gemma4-31BThinking mode
Qwen3.6-27BThinking mode
General agentic
MCP Atlas
75.5
54.2
62.5
DeepSearch QA
74.6
61.7
71.1
τ³-Banking
23.5
15.1
16.7
WildClawBench
47.6
37.6
43.2
GDPval-AA
953
811
1141
GAIA2
43.3
36.4
40.0
SkillsBenchWith skills
44.3
32.4
46.6
OSWorld-Verified
65.9
58.5
75.6
Agentic coding
SWE-Bench Pro
51.2
36.9
50.2
SWE-Bench Verified
76.0
66.6
77.2
TerminalBench 2.1
51.7
43.4
60.7
SciCode
43.6
43.4
39.8
Multimodal
Charxiv Reasoning
78.8
77.7
78.4
ScreenSpot Pro
75.4
75.9
76.1
OmniDocBench v1.5
75.8
72.5
77.8
MMMU Pro
74
73
75
Safety
CI Memories
Violation (↓): 26.4Coverage: 64.8
Violation (↓): 12.1Coverage: 53.0
Violation (↓): 53.4Coverage: 66.9
Siren AgentDojo
Attack Success Rate (↓): 28.4Utility: 94.2
Attack Success Rate (↓): 25.6Utility: 90.8
Attack Success Rate (↓): 40.3Utility: 92.7
General capabilities and reasoning
IFBench
77.0
76.0
70.8
AIME 2026
94.7
89.2
94.1
GPQA Diamond
83.5
85.7
84.2
Humanity's Last ExamText · No tools
22.0
23.6
23.1
AA-LCR
80.0
68.3
73.3
Beam 128K
65.1
58.2
63.0
For the full model card, please see the Muse Glimmer page on Hugging Face.
For more details about Muse Glimmer evaluations, read the methodology report.
Meta is committed to promoting safe and fair use of its tools and features, including Muse Glimmer. This Usage Policy (“Policy”) applies to the access or use of Muse Glimmer.

Where to run Muse Glimmer

huggingface logo image
HuggingFace
Run Muse Glimmer with transformers, llama.cpp, vLLM, Colab, and other popular platforms.
ollama image
Ollama
Run Muse Glimmer on your own hardware in one simple, seamless CLI command.
unsloth image
Unsloth
Run, fine-tune, and deploy Muse Glimmer from one unified local UI interface with the open-source Unsloth Desktop app.
lm studio image
LM Studio
LM Studio Bionic is the easiest way to run Muse Glimmer for agentic tasks, locally on your computer.
fireworks image
Fireworks
Fireworks offers production-grade Muse Glimmer deployments with precise reasoning controls and effortless scaling for unpredictable agent traffic.
together.ai image
together.AI
Run Muse Glimmer on Together Serverless Inference for long-running agentic workflows.
openrouter image
OpenRouter
Run Muse Glimmer via a single unified API that routes across multiple inference providers for cost, speed, and availability.

Build with Muse Glimmer

Download model
View documentation

Cookbooks and quickstarts

Read documentation
Prompting guideGet the most out of Muse Glimmer by using the correct chat template and following these prompting best practices.
Learn more
QuantizationRun on smaller hardware by reducing the precision of model weights.
Learn more
Speculative decodingSpeed up inference by using a smaller, faster draft model to propose candidate tokens that the full model verifies.
Learn more
vLLMDeploy with vLLM for high-throughput, low-latency inference with an OpenAI-compatible API endpoint.
Learn more
llama.cppRun on your machine with llama.cpp, a C/C++ inference engine that supports CPU, mixed CPU/GPU and full GPU execution.
Learn more
ExecuTorchRun on mobile phones, tablets, and edge devices with ExecuTorch, Meta's on-device inference framework.
Learn more
Meta blue bg
Stay up-to-date

Our latest updates delivered to your inbox

Subscribe to our newsletter to keep up with the latest AI updates, releases and more.

Sign up