[AINews] Google I/O 2026: Gemini 3.5 Flash, Omni (NanoBanana for Video), Spark (background agents), and Antigravity 2.0
Google I/O 2026 unveiled Gemini 3.5 Flash (GA with 1M context, 65k output, 4 thinking levels, thought preservation), Gemini Omni (multimodal video generation/editing), Antigravity 2.0 (desktop, CLI, SDK, managed agents), and generative UI and information agents in Search. Google claims 3.2 quadrillion tokens/month processed (7x YoY) and 900M+ monthly Gemini users.
The full keynote livestream was 2 hours, but as usual, The Verge has the best supercut down to 30 mins, which is very worthwhile to get a narrative sense:
The mainline Gemini 3.5 Flash is GA today (very nice compared to some staged rollouts) and is sold as a decent step up even compared to 3.1 Pro, with 3.5 Pro coming next month. Perhaps more impressive were the Gemini Live (Voice) and Omni (Video) and Google Pics/Flow (Images/VFX/music) modalities, where Google demonstrated industry leading capabilities and latency, all presumably made possible by industry leading hardware and models.
Per longstanding tradition at every bigtech keynote these days, Google also showed off some smart glasses tech, which seems a little more likely to be seen on the street than many prior iterations from both Google and their peers.
AI News for 5/18/2026-5/19/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies!
AI Twitter Recap
Google used I/O to reposition Gemini as both a consumer AI surface and a developer/agent platform, with three core technical announcements: Gemini 3.5 Flash for fast agentic/coding workloads, Gemini Omni for multimodal generation/editing starting with video, and a broader Antigravity agent stack spanning desktop/CLI/SDK/API. Official posts emphasized scale — Google says it now processes over 3.2 quadrillion tokens/month, up 7x YoY from 480T/month, while the Gemini app has 900M+ monthly users and is available in 230+ countries and 70+ languages (Google, Google, GeminiApp). The most technically substantive release was Gemini 3.5 Flash, framed by Google as its strongest agentic/coding model yet, GA immediately, with 1M-token context, 65k max output, 4 thinking levels (“minimal/low/medium/high”), and “thought preservation” across turns (GoogleDeepMind, Google, _philschmid). Google paired that with Gemini Omni, a new family combining Gemini reasoning with generative media, initially via Omni Flash, capable of taking text/image/video/audio inputs and producing video edits/generation in Gemini, Flow, Shorts, and later APIs (GoogleDeepMind, Google, GeminiApp). Around those models, Google launched or expanded Antigravity 2.0 desktop, CLI, SDK, Managed Agents in the Gemini API, Search-native generative UI/coding, Gemini Spark background agents on cloud VMs, and a long list of Gemini-app/Workspace/commerce/media integrations (Google, Google, Google).
Facts vs. opinions
Facts / directly claimed by official or third-party benchmark sources
Google says it now processes 3.2 quadrillion tokens/month, up from 480 trillion a year earlier (Google).
Google says Gemini has 900M+ monthly users (Google).
Google says Gemini 3.5 Flash is GA today across Gemini app, Search AI Mode, Gemini API, AI Studio, Antigravity, Android Studio, and enterprise surfaces (Google, GeminiApp).
Google says Gemini 3.5 Flash has 1M context, 65k max output, 4 thinking levels, and “thought preservation” across turns ( _philschmid).
Google says 3.5 Flash beats Gemini 3.1 Pro on Terminal-Bench 2.1, GDPval-AA, and MCP Atlas (GoogleDeepMind, Google).
Google says 3.5 Flash runs 4x faster than comparable frontier models, and up to 12x faster in Antigravity (Google, JeffDean).
Independent benchmarker Artificial Analysis reports Gemini 3.5 Flash scores 55 on its Intelligence Index, +9 vs Gemini 3 Flash, at >280 output tok/s, with MMMU-Pro 84%, GDPval-AA Elo 1656, and pricing of $1.50 / $9.00 per 1M input/output tokens; it also reports the model is 5.5x costlier to run than Gemini 3 Flash on its suite and 75% costlier than Gemini 3.1 Pro (ArtificialAnlys).
Arena reports Gemini 3.5 Flash reached #9 overall in Text Arena and #9 in Code Arena: Frontend, scoring 1507, a +70 jump over Gemini 3 Flash, and becoming the top score in its price tier (arena).
Google says Gemini Omni Flash is available in Gemini/Flow today for paid users, in Shorts/Create starting this week for free, and via APIs in coming weeks (Google).
Google says Spark runs on dedicated Google Cloud virtual machines, allowing long-running tasks while user devices are closed (Google).
Google claims an Antigravity + Gemini 3.5 Flash demo built a functioning OS in 12 hours using 93 parallel sub-agents, 15k+ model requests, 2.6B tokens, and 280 output tok/s
Some discussion cited ~867 tok/s in Antigravity-specific optimized serving (scaling01, scaling01)
Third-party evaluation:
Artificial Analysis says 3.5 Flash is the leader on the intelligence-vs-speed Pareto frontier, but the economics are notably worse than prior Flash:
Intelligence Index 55
+9 over Gemini 3 Flash
Hallucination rate reduced to 61%, a 31-point drop vs Gemini 3 Flash on its omniscience setup
GDPval-AA 1656 Elo
5.5x costlier than Gemini 3 Flash to run on its benchmark suite
75% costlier than Gemini 3.1 Pro on the same suite (ArtificialAnlys)
Arena:
#9 Text Arena
#9 Code Arena: Frontend
1507 score, +70 over Gemini-3 Flash
Better than Gemini 3.1 Pro across categories in its frontend coding eval (arena, arena)
Implications
The notable shift is that Google appears to be using a “Flash” label for a model that, in prior cycles, would have been described more like a high-end product model optimized for deployment rather than simply a cheap lightweight tier. Several posters called this out directly, arguing Flash is becoming more expensive and possibly absorbing former Pro territory (enricoros, simonw).
The strongest technical signal is not “best absolute benchmark model,” but:
material agentic gains
extreme serving speed
deep integration into product surfaces
tooling built around subagents and long-horizon execution
That makes 3.5 Flash strategically important even if some competitors still win on raw price-adjusted intelligence in certain third-party comparisons.
Gemini Omni: multimodal generation/editing as “create anything from any input”
What Google announced
Google introduced Gemini Omni as a new family merging Gemini reasoning/world knowledge with Google’s generative media stack, starting with video creation and editing. Official messaging described it as “create anything from any input,” but current rollout is narrower:
Inputs: text, images, audio, video
Initial output emphasis: video
Product availability: Gemini app, Flow, YouTube Shorts/Create, later APIs
Current shipping model: Gemini Omni Flash (GoogleDeepMind, Google, Google)
Google/DeepMind claims:
Better world understanding
More robust physics
Multi-turn editing where scene/character consistency is retained
Ability to “reimagine” user video footage with conversational edits (Google, Google)
Rollout specifics:
Paid Gemini users globally in app/Flow “today”
YouTube Shorts/Create rolling out “starting this week” at no cost
APIs for developers/enterprise in coming weeks (Google, GeminiApp)
Perspectives
Supportive: users and Google employees described Omni as a major quality step, especially for video editing and consistency (joshwoodward, fofrAI, osanseviero).
Strategic interpretation: several posters framed Omni as evidence Google is investing in world models and embodied/physical priors, not just text/code competition (demishassabis, jparkerholder, kimmonismus).
Skepticism: some UI/output examples drew criticism for looking like “B-tier video game interface” or too polished/template-like (teortaxesTex, shlomifruchter).
Context
Omni matters less as “yet another video model” and more as Google’s attempt to unify:
multimodal understanding,
media editing,
world grounding,
agent interfaces,
and eventually any-input/any-output generation.
This aligns with DeepMind’s long-running world-model agenda and Google’s product distribution advantage.
Antigravity: Google’s agent OS, not just a coding assistant
A major underappreciated I/O theme was that Google is no longer presenting agents as a thin wrapper around a chat model. Antigravity is becoming the execution substrate.
What launched / expanded
Antigravity 2.0 desktop app: agent-first desktop with core conversations, artifacts, multi-agent orchestration (Google, Google)
Antigravity CLI (Google, Google)
Antigravity SDK (Google)
Managed Agents in Gemini API: single API call gives an agent plus hosted Linux sandbox; supports Bash/Python/Node, files, browsing, custom markdown-defined skills, repo/GCS mounts (Google, GoogleAIStudio, _philschmid)
Integrations with AI Studio, Android, Firebase, Workspace, web (Google, Google)
One-click export from AI Studio to Antigravity (Google)
Native Android app generation in AI Studio / Android support in Antigravity (Google, AndroidDev)
Technical signaling
Google’s own demos centered on parallel sub-agents, hosted execution, high-frequency iterative loops, and artifact-oriented workflows. Jeff Dean explicitly described 3.5 Flash as a strong engine for “deploy sub-agents that collaborate, run high-frequency iterative loops, and solve real-world problems at scale” (JeffDean).
The marquee proof point:
OS built in 12h
93 parallel sub-agents
15k+ requests
2.6B tokens
< $1K credits (Google)
Even if this is mostly a stage-managed benchmark/demo, it reveals the architecture Google wants developers to adopt: many fast agents over one slow monolithic run.
Reactions
Positive: this is Google’s answer to Codex/Claude Code/OpenClaw/Hermes-style workflows, with a stronger infra story (iScienceLuvr, theo).
Critical: branding and product sprawl remain confusing; some users aren’t sure whether they should use Gemini CLI or Antigravity CLI, and Google’s design choices drew complaints (kchonyc, zachtratar, teortaxesTex).
Search, Gemini app, and consumer agents
Search
Google announced a redesigned AI-powered Search box, multimodal query support, and the most ambitious consumer-facing move: Search generating custom visual tools and simulations on the fly using Antigravity + Gemini 3.5 Flash (Google, Google).
It also previewed information agents in Search:
persistent monitoring tasks
web/news/social/real-time signals
synthesized updates with links and actions
rolling out to Pro/Ultra this summer (Google, Google)
This is a notable strategic shift: Search moves from retrieval/ranking to background agentic monitoring + generated applets.
Gemini app
Consumer Gemini updates included:
new “Neural Expressive” design language (Google)
inline/instant Gemini Live voice (Google)
Daily Brief personalized digest from inbox/calendar/tasks (Google, GeminiApp)
Gemini Spark as a 24/7 personal AI agent on cloud VMs, checking with users before major actions (Google, GeminiApp)
macOS app + upcoming Spark/voice desktop workflows (Google, GeminiApp)
Pricing / subscriptions
Google introduced a new pricing ladder:
new $100/month plan
top-tier Ultra cut from $250 to $200/month (Google, GeminiApp)
This reads as a more aggressive bid for premium power users, especially coders and creators.
Trust, provenance, and standards
Google pushed SynthID across Search, Gemini, Chrome, and hardware/media surfaces, and announced partnerships with OpenAI, NVIDIA, Kakao, and ElevenLabs to bring SynthID to their generated content (Google, Google).
That is one of the more consequential standards moves from I/O:
it gives Google a shot at owning part of the provenance layer for generative media;
notably, OpenAI separately announced support for checking OpenAI-generated images via SynthID watermark + C2PA credentials (OpenAI).
This was less flashy than Omni/3.5 Flash, but likely more durable if provenance becomes mandatory infrastructure.
Google’s science and world-model angle
Several I/O items reinforced that Google does not want to compete only on coding/chat:
Gemini for Science: Literature Insights, Hypothesis Generation, Computational Discovery (GoogleDeepMind, Google)
Nature publication links around ERA / Co-Scientist (Google
[truncated for AI cost control]