The GenAI Field Guide · Trending
OpenAI's cheaper tiers below GPT-6 Astra: Sol for coding and professional agent work, Luna for fast high-volume tasks, and a GPT-6.1 Sol refresh one week later.
GPT-6 Sol and GPT-6 Luna (22 September 2026) are the mid and budget models under GPT-6 Astra, priced at about half their GPT-5.6 predecessors. GPT-6.1 Sol (29 September, at DevDay) is a separate model ID that OpenAI pitches as near-Astra performance at a lower cost, with cheaper cache reads and a beta multi-agent mode in the Responses API. For most teams these, not Astra, are the models that decide the monthly bill, so the practical work is re-running your own evals and re-checking your cost per task.
OpenAI now ships the GPT-6 family in three tiers. Astra (3 September) is the flagship for the hardest end-to-end work. Sol targets complex coding and professional work at lower cost, and Luna is described in OpenAI's own docs as its most efficient model for focused, high-volume tasks such as summarization, extraction and focused coding. Press coverage reports that OpenAI built Sol and Luna with methods similar to Astra's, and that both are API-only with no weights to self-host.
There are three model IDs, not two, and the naming trips people up. gpt-6-sol and gpt-6-luna launched on 22 September 2026. gpt-6.1-sol launched on 29 September 2026 as a separate model; OpenAI's help docs list GPT-6 Sol as a previous-generation option that stays available during the 6.1 rollout, and the Codex CLI moved its default to GPT-6.1 Sol. There is no GPT-6.1 Luna as of 2 October 2026.
Under the marketing, the story is price per unit of capability. Sol and Luna list at 2 and 10 US dollars, and 0.10 and 0.50 dollars, per million input and output tokens, about half the GPT-5.6 rates reported by press. OpenAI's benchmark claims are framed the same way: GPT-6 Sol is said to make about half as many mistakes as GPT-5.6 Sol, to beat Claude Opus 5 on AutomationBench at a fraction of the cost, and GPT-6.1 Sol is said to nearly match Astra on the DeepSWE v1.1 coding benchmark at about one fifth of the cost. These are vendor numbers relayed by press and blogs; treat them as hypotheses to test on your own workload.
All three are reasoning models that take text and image input and return text, with a 1,050,000 token context window, up to 922,000 input tokens and 128,000 output tokens in OpenAI's model docs. Amazon Bedrock's model card lists GPT-6.1 Sol as 1M context with 131,072 output tokens, a small difference worth checking if you run near the limits. Knowledge cutoffs per OpenAI's model pages are 20 April 2026 for GPT-6 Sol, 30 April 2026 for GPT-6.1 Sol and 18 May 2026 for GPT-6 Luna.
Reasoning effort is the main dial. GPT-6 Sol and Luna accept none, low, medium (the default), high, xhigh and max. GPT-6.1 Sol, like Astra, drops none and minimal, so low is its floor; OpenAI's migration guide says to keep GPT-6 Sol where you need none. When effort is above none, temperature, top_p and log probabilities must be removed. Tool calling with GPT-6.1 Sol requires the Responses API, while GPT-6 Sol and Luna support function calling in Chat Completions only at none effort. DataCamp's write-up reports that effort can be changed mid-conversation without invalidating the prompt cache.
Caching is where 6.1 Sol differs most on price. Cache reads cost 0.20 dollars per million tokens on GPT-6 Sol and 0.10 on GPT-6.1 Sol, with a separate cache write charge of 2.50 dollars per million on both; Luna reads cost 0.01 and writes 0.125. Requests over 272,000 input tokens are billed at long-context rates for the whole request: OpenAI's GPT-6 Sol page and Bedrock's GPT-6.1 Sol card both show twice the input and cache rates and one and a half times the output rate. Batch and Flex run at half the standard rate on OpenAI's API.
GPT-6.1 Sol also brings Multi-agent, a beta in the Responses API (also available for all GPT-5.6 models). You enable it per request with a betas flag; a root agent named /root spawns subagents with their own contexts, which share the request's model and tools, and the root synthesizes their results. The docs list default concurrency of three subagents, no fixed limit on tree depth, automatic compaction per agent, and warn that item schemas may change during the beta. They do not state how multi-agent runs are billed.
Any paid OpenAI API account can call gpt-6-sol, gpt-6-luna and gpt-6.1-sol through the Responses API or Chat Completions. In ChatGPT, the models appear in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu, with Luna also on Free and Go in the desktop app; for Enterprise and Edu, GPT-6.1 Sol stays off until an admin enables it. The example below is OpenAI's own multi-agent quickstart.