Meta AI Releases Muse Glimmer: A 30B Open-Weights Agentic Model That Runs on One Consumer GPU
Meta has released Muse Glimmer , a 30-billion-parameter multimodal model distilled from Muse Spark.

AI Releases Muse Glimmer: A 30B Open-Weights Agentic Model That Runs on One Consumer GPU">
Meta has released Muse Glimmer , a 30-billion-parameter multimodal model distilled from Muse Spark. It is tuned for always-on local agent workflows, and ships under Apache 2.0. A 30B model normally needs over 55 GB of memory at full precision. Meta compresses it to roughly 4-bit, then adds block-level speculative decoding so it answers fast enough to sit inside a real agent loop. The result runs on one consumer GPU or a Mac, with no network call.
Yes, the weights are open under Apache 2.0 . The Hugging Face collection carries BF16 weights, GGUF k-quants, ExecuTorch builds, and the DFlash drafter. Self-hosting is the day-one path.
Muse Glimmer is a dense causal transformer with a dedicated perception encoder. Total parameters are roughly 30B, including the vision tower. Grouped-query attention uses 32 query heads and 2 KV heads. Attention repeats a [Local, Local, Local, Global] pattern with a 2,048 sliding window. RoPE is applied to local layers only, with theta 500,000. The vision side is a ~1.8B ViT-G/14 perception encoder accepting up to 4,096 visual tokens per image. Context length is 131,072+, vocabulary is 202,048 tokens, and the knowledge cutoff is January 4, 2026. Input is text and image; output is text. Audio is not supported, and video is processed as individual frames.
At full precision the model needs over 55 GB of memory. Meta compresses weights to approximately 4-bit precision, which brings the language model under 20 GB. That leaves headroom inside a 24 GB or 32 GB envelope. The KV cache, perception encoder, and drafter share it. Two quantized builds ship. K-Quant-Dynamic targets 32 GB VRAM at 0.2% average degradation. K-Quant-17GB targets 24 GB VRAM at 1.0%. Degradation is averaged over accuracy metrics across 15 common benchmarks.
Generation speed comes from DFlash , a block-diffusion drafter that predicts 16 tokens in one forward pass. The main model verifies the block in parallel. The drafter uses 5 layers, sliding-window attention at 2,048, and 32 query / 8 KV heads. Meta measured K-Quant-17GB at batch size 1 with greedy decoding. On an RTX 5090, throughput rises from 74.9 to 233.4 tok/s, a 3.1x speedup. Apple M5 Max moves from 26.6 to 50.2 tok/s, and M4 Max from 23.7 to 37.8 tok/s.
Meta compares Muse Glimmer against Gemma4-31B and Qwen3.6-27B in thinking mode. It leads on MCP Atlas at 75.5, against 54.2 and 62.5. It also leads on DeepSearch QA at 74.6, Gaia2 at 43.3, and SWE-Bench Pro at 51.2. Reasoning scores follow: AIME 2026 at 94.7, IFBench at 77.0, AA-LCR at 80.0. Qwen3.6-27B stays ahead on OSWorld-Verified, 75.6 versus 65.9. It also leads TerminalBench 2.1 at 60.7 and SWE-Bench Verified at 77.2. The pattern is consistent. Muse Glimmer wins on agentic orchestration and reasoning. It trails on computer-use and terminal work.
Source: MarkTechPost