NEW

Google releases two faster Gemini models

Gemini 3.6 Flash and 3.5 Flash-Lite are built for tasks where response time and cost are important. The models are now available for developers, businesses and in the Gemini app, while the security-focused Flash Cyber has much more limited availability.

Publicerad 28 July 2026, 11.06

AI-generated illustration of four professionals comparing two abstract AI flows for speed and cost.

Google has released Gemini 3.6 Flash and Gemini 3.5 Flash-Lite. There are two AI models for tasks where many answers need to be produced quickly, such as document sorting, code work, and AI agents that use multiple tools.

Gemini 3.6 Flash is the more capable of the two. It can receive text, image, video, audio and pdf files and returns text as its output. The model supports, among other things, file searching, code execution and function calls. Computer control is available as a preview version, which means that the feature can still be changed.

The price is $1.50 per million input tokens and $7.50 per million output tokens. A token is a small piece of text that the model reads or writes. The cost therefore depends both on how much documentation is sent in and how long the answers the model gives.

Gemini 3.5 Flash-Lite is aimed at larger amounts of simpler work. In global traffic, it costs $0.30 per million input tokens and $2.50 per million output tokens. Google charges a slightly higher price for some traffic that does not run globally.

Both models are available in the Gemini API via Google AI Studio and in Google's platform for enterprise agents. Google also states that the models are in the Gemini app. Flash-Lite is also being rolled out in Google Search, but the company isn't publishing a full country or language list in the launch post.

At the same time, Google presented Gemini 3.5 Flash Cyber, a special model for finding and fixing security flaws in code. It is not generally available. Google is planning a limited pilot for government agencies and trusted partners through the CodeMender security tool.

The comparisons of quality, speed and resilience are primarily based on Google's own tests and selected customer data. Some speed measurements come from Artificial Analysis, but they do not say what the models cost or perform in a certain Swedish business's own flows.

Därför spelar det roll

For businesses that build AI support, the choice of model can affect both response time and running costs. A cheaper model may suit simple, high-volume tasks, while a more capable model may be needed when the surface is mixed or the work has many steps. The price list does not replace a test with real data, safety requirements and human review.

Det här kan du göra

  1. Measure cost, response time and quality of the same representative data before changing models.
  2. Use the lighter model for simple flows first and only move more difficult cases to a more expensive model.
  3. Limit tools and permissions when a model is allowed to control the computer or other systems, especially while the feature is in preview.