The GenAI Field Guide · Trending
A 30B dense open-weight model from Meta, under Apache 2.0, built to run agent loops on one consumer GPU or a high-memory Mac.
Muse Glimmer is a 30 billion parameter dense model that Meta Superintelligence Labs released with open weights on 10 August 2026, distilled from its larger Muse Spark model and tuned for tool use, long tasks and failure recovery. A roughly 4-bit build fits in 24 to 32 GB of GPU or unified memory, and the licence is plain Apache 2.0 rather than the Llama community licence. If you want an agent or coding model that runs on your own hardware with no data leaving it, it belongs on your shortlist, but its benchmarks are Meta's own and its prompt injection numbers are not reassuring.
Meta describes Glimmer as a model for always-on local agent workflows: something that sits on a laptop or workstation and drives tools, coding agents and personal assistants without a cloud call. Mark Zuckerberg announced it as a 30B dense model that can run locally, as reported by VentureBeat. VentureBeat frames it as Meta returning to open source and reports that Zuckerberg also promised to open the weights of the larger Muse Spark 1.2 soon, which had not happened in anything we found.
The licence is the news for many teams. Apache 2.0 allows commercial use, modification and redistribution with no user cap, where the older Llama community licence carried conditions such as the 700 million monthly user ceiling. The weights are on Hugging Face as meta-models/Muse-Glimmer-30B, with GGUF k-quants, ExecuTorch builds and a separate speculative decoding drafter.
It takes text and images and returns text. The model card says video is processed as individual frames and that the model is not optimised for video; the Hugging Face launch blog lists video as an input, so treat video as frames rather than a real capability. The card gives a 131,072 token context window, a knowledge cutoff of 4 January 2026 and training on more than 100 languages, with quality that varies by language. Reasoning effort is set in the system prompt as low, medium, high or xhigh.
Under the marketing, the useful claim is that a 30B model can be a competent tool caller. Meta's own table puts it at 75.5 on MCP Atlas against 62.5 for Qwen3.6-27B and 54.2 for Gemma4-31B, 74.6 on DeepSearch QA and 51.2 on SWE-Bench Pro. It trails Qwen3.6-27B on OSWorld-Verified (65.9 against 75.6) and, per DataCamp's reading of the table, on Terminal-Bench 2.1 (51.7 against 60.7). These are vendor-run numbers; we found no independent reproduction.
It is a dense causal transformer, not a mixture of experts, so all 30B parameters are active on every token. The model card puts the total at about 29.6B including a 1.8B perception encoder for images (the Hugging Face blog rounds the encoder to 2B and the decoder to 28B). The decoder has 52 layers in a repeating pattern of three sliding-window attention layers with a 2,048 token window followed by one full-attention layer without positional embeddings, and uses grouped-query attention with 32 query heads sharing 2 key-value heads. The sliding layers keep the KV cache small, which is what lets a long context fit next to the weights on a 24 GB card.
Training ran in three stages, per Meta's announcement: pre-training by logit distillation from Muse Spark, mid-training on longer-context, agent-heavy data with reasoning traces and multimodal inputs, and post-training that mixes supervised fine-tuning, on-policy distillation and reinforcement learning across general, reasoning, coding and agentic tasks.
Local deployment rests on two tricks. First, quantisation to about 4 bits takes the model from more than 55 GB in BF16 to under 20 GB; Meta ships a K-Quant-Dynamic build aimed at 32 GB with about 0.2 percent degradation and a K-Quant-17GB build aimed at 24 GB with about 1 percent, by its own measure. Second, a block-diffusion drafter called DFlash proposes 16 tokens per forward pass for speculative decoding; Marktechpost reports 74.9 to 233.4 tokens per second on an RTX 5090 (3.1 times), 26.6 to 50.2 on an M5 Max and 23.7 to 37.8 on an M4 Max.
Meta assessed it under its Advanced AI Scaling Framework and says it does not meet the framework's frontier definition, rating chemical and biological, cyber and loss-of-control risk as moderate or lower. The model card recommends extra guardrails before letting it take real-world actions.
Anyone can download it; there is no gate on the Hugging Face repository and no Meta-hosted API. For local use you need about 24 GB of GPU memory or 32 GB of unified memory on a Mac for the 4-bit builds, or a single 80 GB H100 for BF16. If you would rather not host it, Together AI, Fireworks AI and OpenRouter serve it as a paid API.