Skip to content
Breaking

AI in MarketingGenerative AI Tools

Google Ships Gemini 3.6 Flash, Cuts Agent Costs 17%

Google released three new Gemini models built for agentic workloads, with 3.6 Flash cutting output token usage 17% and undercutting 3.5 Flash on price.

A pair of scissors slicing through a stack of coins, with the top portion tipping off to the side.
Illustration by CMO Mag

Key takeaways

  • Gemini 3.6 Flash uses 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index, and up to 65% fewer on some coding benchmarks
  • Pricing for 3.6 Flash drops to $1.50 per 1M input tokens and $7.50 per 1M output tokens, lower than 3.5 Flash
  • 3.5 Flash-Lite hits 350 output tokens per second, Google's fastest Flash-class model yet
  • A specialized 3.5 Flash Cyber model now pairs with Google's CodeMender agent for security tasks
  • Gemini 3.5 Pro is in partner testing and Gemini 4 pre-training has already begun

What Google announced

Google released three new Gemini models on July 21, 2026, all aimed at making AI agents cheaper and faster to run in production, according to a post from Google by Tulsee Doshi, senior director of product management for Gemini.

The headline model, Gemini 3.6 Flash, improves coding, knowledge work, and multimodal tasks while using fewer tokens to get there. Google says it takes fewer reasoning steps and tool calls per workflow, which is the real cost lever for anyone running agents at volume.

17%

Reduction in output token usage, 3.6 Flash vs. 3.5 Flash

Artificial Analysis Index, via Google, July 2026

3.6 Flash is priced at $1.50 per 1 million input tokens and $7.50 per 1 million output tokens, undercutting its predecessor on cost per agentic task. Google also introduced 3.5 Flash-Lite, its fastest and most cost-effective 3.5-class model at 350 output tokens per second, and 3.5 Flash Cyber, a specialized model paired with Google's CodeMender agent for security work. Gemini 3.5 Pro is currently in partner testing, and Google says it has begun pre-training for Gemini 4.

Why marketers should care this week

The efficiency gains matter more than the model name. Marketers building real-time personalization, chat-based content tools, or campaign automation agents pay per token and per API call, so a 17% cut in output tokens compounds fast at scale, especially for teams running thousands of daily agent sessions.

This fits a broader pattern of Google pushing agentic infrastructure across its stack, something we've tracked in Google's agentic advertising playbook and in how Google's Ask Advisor agent now spans the ad stack. Faster, cheaper Flash models are the plumbing underneath those tools.

Explore more AI in marketing coverage to prep your stack for what's next.

Portrait of Grace Nakamura

Grace Nakamura

AI expert · Verified

AI in marketing writer · AI in Marketing

Grace Nakamura tests AI marketing tools so you don't have to trust the demo. She led marketing at two SaaS startups through the first wave of generative AI. She writes about AI in marketing — the workflows that work, the ones that don't, and the governance most teams ignore. A pragmatic futurist with a low tolerance for hype.

More from Grace Nakamura What is an AI expert?

Advertiser disclosure: some links in our articles are affiliate links, and CMO Mag may earn a commission or referral fee if you sign up or buy through them, at no cost to you. It never affects our editorial coverage. See our advertising & affiliate policy.

Discussion

No comments yet. Be the first to say something worth reading.

View all
AI for SEO & AEO

AI Knows 96% of Brands, Cites Almost None of Them

A Q2 2026 study of eight AI platforms finds a wide gap between brand recognition and brand recommendation, and it's a warning for anyone counting on AI search for demand generation.

Grace Nakamura
AI for SEO & AEO

Search Console May Be Adding AI Performance Data, SEOs Say

A Reddit thread in r/SEO has practitioners buzzing about new generative AI performance reporting inside Google Search Console, but Google hasn't confirmed the rollout publicly. Here's what marketers should actually do this week.

Grace Nakamura
Generative AI Tools

Google's Gemini API Managed Agents Gain Hooks, Model Choice, Free Tier

Google DeepMind updated Managed Agents in the Gemini API with environment hooks for auditing tool calls, a default switch to Gemini 3.6 Flash, and free tier access, moving the sandboxed agent framework closer to production-ready marketing automation.

Grace Nakamura

The CMO Mag brief

The marketing intelligence worth reading

Get the numbers behind the news. Pick your cadence.