AI in MarketingGenerative AI Tools
Google Ships Gemini 3.6 Flash, Cuts Agent Costs 17%
Google released three new Gemini models built for agentic workloads, with 3.6 Flash cutting output token usage 17% and undercutting 3.5 Flash on price.

Key takeaways
- Gemini 3.6 Flash uses 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index, and up to 65% fewer on some coding benchmarks
- Pricing for 3.6 Flash drops to $1.50 per 1M input tokens and $7.50 per 1M output tokens, lower than 3.5 Flash
- 3.5 Flash-Lite hits 350 output tokens per second, Google's fastest Flash-class model yet
- A specialized 3.5 Flash Cyber model now pairs with Google's CodeMender agent for security tasks
- Gemini 3.5 Pro is in partner testing and Gemini 4 pre-training has already begun
What Google announced
Google released three new Gemini models on July 21, 2026, all aimed at making AI agents cheaper and faster to run in production, according to a post from Google by Tulsee Doshi, senior director of product management for Gemini.
The headline model, Gemini 3.6 Flash, improves coding, knowledge work, and multimodal tasks while using fewer tokens to get there. Google says it takes fewer reasoning steps and tool calls per workflow, which is the real cost lever for anyone running agents at volume.
Reduction in output token usage, 3.6 Flash vs. 3.5 Flash
Artificial Analysis Index, via Google, July 2026
3.6 Flash is priced at $1.50 per 1 million input tokens and $7.50 per 1 million output tokens, undercutting its predecessor on cost per agentic task. Google also introduced 3.5 Flash-Lite, its fastest and most cost-effective 3.5-class model at 350 output tokens per second, and 3.5 Flash Cyber, a specialized model paired with Google's CodeMender agent for security work. Gemini 3.5 Pro is currently in partner testing, and Google says it has begun pre-training for Gemini 4.
Why marketers should care this week
The efficiency gains matter more than the model name. Marketers building real-time personalization, chat-based content tools, or campaign automation agents pay per token and per API call, so a 17% cut in output tokens compounds fast at scale, especially for teams running thousands of daily agent sessions.
This fits a broader pattern of Google pushing agentic infrastructure across its stack, something we've tracked in Google's agentic advertising playbook and in how Google's Ask Advisor agent now spans the ad stack. Faster, cheaper Flash models are the plumbing underneath those tools.
Explore more AI in marketing coverage to prep your stack for what's next.
Advertiser disclosure: some links in our articles are affiliate links, and CMO Mag may earn a commission or referral fee if you sign up or buy through them, at no cost to you. It never affects our editorial coverage. See our advertising & affiliate policy.
More in AI in Marketing
View allAI Knows 96% of Brands, Cites Almost None of Them
A Q2 2026 study of eight AI platforms finds a wide gap between brand recognition and brand recommendation, and it's a warning for anyone counting on AI search for demand generation.
Search Console May Be Adding AI Performance Data, SEOs Say
A Reddit thread in r/SEO has practitioners buzzing about new generative AI performance reporting inside Google Search Console, but Google hasn't confirmed the rollout publicly. Here's what marketers should actually do this week.
Google's Gemini API Managed Agents Gain Hooks, Model Choice, Free Tier
Google DeepMind updated Managed Agents in the Gemini API with environment hooks for auditing tool calls, a default switch to Gemini 3.6 Flash, and free tier access, moving the sandboxed agent framework closer to production-ready marketing automation.




Discussion
No comments yet. Be the first to say something worth reading.