The GenAI Field Guide · Trending
Google's mid-priced Flash model retuned for long-running coding and agent work, now the default brain of the Antigravity managed agent in the Gemini API.
Gemini 3.8 Flash is a post-training update on Gemini 3.7 Flash, aimed at long-horizon software engineering and autonomous agents, at the same introductory price of 0.75 US dollars per million input tokens and 3.75 per million output until 31 December 2026, doubling on 1 January 2027. Two weeks later Google moved its hosted agent harness, the Antigravity managed agent, onto it with a new version that replaces the May harness, which shuts down on 5 October 2026. Teams on the Gemini API should check their token bills and migrate any code pinned to the old agent this week.
Google released gemini-3.8-flash on 2 September 2026 and calls it its most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents and complex enterprise workflows. The DeepMind model card is direct about what it is under the name: it is based on Gemini 3.7 Flash, which reached general availability only on 13 August, and refers readers to the 3.7 card for architecture. So this is a new post-training pass on an existing base, not a new foundation model. Google's own frontier safety section says it found no meaningful new capabilities compared with 3.7 Flash.
Alongside it Google announced Gemini 3.8 Flash Cyber, a security-tuned variant for vulnerability discovery and automated patching. It is not generally available: access goes through a new Fairwind Program limited to what Google calls trusted defenders, such as government authorities and critical infrastructure operators. Google claims a 47.2 percent pass@1 on CWE-Bench patching and that Chrome Security got 2.6 times more correct patches than from commercial alternatives; these are the maker's numbers, not independent results.
The second half of the story is the harness. Managed Agents arrived in the Gemini API in preview on 19 May 2026: you call an agent instead of a model and Google runs the loop, the tools and an isolated Linux sandbox for you. On 17 September Google released antigravity-preview-09-2026, which replaces and deprecates antigravity-preview-05-2026, runs on Gemini 3.8 Flash by default, and brings the agent's tools closer to those in Google's Antigravity coding product. The old version shuts down on 5 October 2026.
As a model, 3.8 Flash takes text, images, audio and video with a context window of about one million tokens (1,048,576) and returns text up to 64K tokens (65,536). Reasoning is controlled by a thinking_level string with low, medium and high; medium is the default and minimal is not supported. Google's migration notes for Gemini 3 say to replace thinking_budget with thinking_level, drop temperature, top_p, top_k and candidate_count, and use the server-side previous_interaction_id for multi-turn state through the Interactions API. The model card's knowledge cutoff is March 2026.
As an agent, the Antigravity managed agent provisions a Linux sandbox per interaction. Default tools are code execution in Bash, Python and Node.js, Google Search and URL fetching, and filesystem tools switch on when you pass an environment. Reusing an environment ID keeps files between calls, and an environment can be seeded from inline content, a Git repository or Cloud Storage. You add your own function calls or remote MCP servers, cap spend with max_total_tokens (the interaction returns an incomplete status if it hits the cap), and can switch the underlying model to an older or lighter Gemini Flash model.
Custom managed agents are the same base agent plus your own instructions and files: a system_instruction, an AGENTS.md file under .agents/, skills folders with SKILL.md files, and workspace data. Outbound network access goes through a domain allowlist, and stored credentials are injected by a proxy at request time so secrets do not sit in the agent definition or the prompt. You can save up to 1,000 named agents per project and call them by ID; each run forks a fresh sandbox.
The September harness changed the tool interface itself. The changelog says tool parameters moved from snake_case to PascalCase and file edits now use line-range replacements instead of full rewrites, and Google's docs describe new native code search tools. Google's docs cite improved prompt caching, and press coverage reports a 22 percent cache hit-rate improvement on long multi-step tasks plus a new Files API and Credentials API. Any prompt, eval or parser that names the old tools needs updating.
Anyone with a Google AI Studio account can call the model on the Gemini API free tier, where Google says inputs and outputs may be used to improve its products; the paid tier opts you out of that. The managed agent is a preview available to free and paid projects through the Interactions API and AI Studio. You need an API key and the google-genai SDK.