LOADING

Google Gemini 4 has launched, with three consecutive releases in one night

Three in one breath, Google is really not installing this wave.

Today, Google DeepMind launched a series of big moves, presenting three trump cards at once:

Gemini 3.6 Flash
Gemini 3.5 Flash-Lite
Gemini 3.5 Flash Cyber

A stronger main force, a faster lightweight model, and a security special forces specialized in fixing vulnerabilities.

The three Geminis point to the same thing: making AI agents running in production environments faster, smarter, and cheaper.

 

On the same day of its release, Google also casually threw a bigger thunderbolt——

The most aggressive pre training in Google DeepMind’s history has begun, with the goal of the next generation model Gemini 4.

 

On one side is Gemini Triple Launch, and on the other side is Gemini 4 Run.

The signal conveyed by Google is crystal clear, and the ultimate game of AI is just beginning.

Google’s late night amplification strategy
Three Gemini models released simultaneously
Three Gemini models, to be disassembled one by one.

Gemini 3.6 Flash, Violent Province 65% Token
Let’s first take a look at the real ‘C-bit’ – Gemini 3.6 Flash.

It focuses on a highlight this time, the Giant Province Token.

According to the Artificial Analysis Index, compared to the previous generation 3.5 Flash, its output token is 17% less.

On coding benchmarks like DeepSWE, the maximum savings can be up to 65%.

 

Not only does it speak less, but it also requires fewer reasoning steps and tool calls to complete a multi-step task, with shorter detours, which is both fast and cost-effective.

Not only that, Gemini 3.6 Flash not only saves tokens, but also has a lower price:

Input $1.5 and output $7.5 per million tokens, which is cheaper than 3.5 Flash.

More economical, cheaper, and capable of fighting, that’s what sincerity is all about. 3.6 Flash has surpassed its predecessor in several key tests——

DeepSWE: Programming testing 37% ➔ 49%
MLE Bench: Machine Learning Research Test 49.7% ➔ 63.9%
OSWorld Verified: Allow models to be directly operated on computers for testing 78.4% ➔ 83%
GDPVal-AA v2: Achieved 1421 knowledge-based job tasks, more than 70 points higher than the previous generation

In the following demo, 3.6 Flash showcases its multi-agent orchestration skills, which not only easily accomplish complex code migration tasks, but also achieve dimensionality reduction in response speed and code quality for 3.5 Flash.

 

With the help of Gemini Canvas and 3.6 Flash, the 3D workflow is seamlessly integrated, allowing for the creation of a professional photography grade texture extraction tool in just a few minutes.

It can also transform into an “AI interaction designer” and easily handle immersive theme studios.

 

 

3.5 Flash-Lite, Defeated Big Brother
The second Gemini 3.5 Flash Lite takes a different path: it focuses on being fast and cheap enough.

The output speed of 3.5 Flash Lite reaches 350 tokens per second, which is the fastest in the entire 3.5 series.

The price is also outrageously low, with an input of $0.3 and an output of $2.5 per million tokens.

Fast and cheap, naturally prepared for high-frequency and high-volume scenarios such as “massive document processing” and “agent search”.

The most impressive thing is that it actually stepped on its own ‘big brother’ with its backhand.

In multiple programming and agent tests, this lightweight model surprisingly outperformed the larger Gemini 3 Flash——

SWE-Bench Pro:54.2% vs 49.6%
OSWorld-Verified:74.0% vs 65.1%

3.6 Flash is used as the “main brain” for disassembly tasks, while Flash Lite is used as a “clone” for batch work.

One person is responsible for thinking, while the other is responsible for running, easily smoothing out high concurrency pressure.

In the official demonstration, Flash Lite combines the two to come up with 25 sets of highly available web design solutions in one go, with a speed that is almost as fast as “just thinking about it and coming up with a batch”.

Flash Lite is now available on the Gemini App and will gradually enter Google search.

Flash Cyber: Can fix vulnerabilities, but not everyone can use it
Gemini 3.5 Flash Cyber, A hardcore model specializing in network security.

Its task is only one, to find and fix vulnerabilities.

The awkward reality at present is that the speed at which AI searches for vulnerabilities has exceeded the speed at which existing systems fix them.

 

Google has integrated this model into the CodeMender code security agent, relying on multiple Cyber agents to collaborate and refresh SOTA on the security benchmark CyberGym, at a lower cost than larger models——

Use a lighter model to perform the tasks that require the most precision.

However, this model is currently inaccessible to ordinary people.

Gemini 4 starts running
The most radical pre training in history
Although Gemini 3.5 Pro is not yet available, Gemini 4 has also started training.

This triple play is just the appetizer, but the official announcement at the end is the real highlight.

Google has confirmed that it has initiated the most aggressive pre training in its history, targeting Gemini 4.

 

The big shot said that according to Google’s regular 6-month training pace, maybe Gemini 4 can be seen by the end of the year.

 

For those of us who spend real money to buy tokens and build agents, today’s release actually has only one word – ‘decrease’.

The price has decreased, token consumption has decreased, and the cost of running an agent has decreased.

As for the delayed flagship and the recently launched Gemini 4, they are still “futures”.

What can be used and is cheap is what can be put into your wallet today. No matter how strong the flagship is, it must first land.

© 版权声明

相关文章