Gemini 4 Pro has another checkpoint out!
Two brand-new checkpoints (internal codename “barium-b” or Checkpoint 2) were spotted by developers—and hands-on tests say Gemini 4 Pro is terrifyingly strong.
One developer wrote: “After grinding 400 hours on Opus 5.5 and GPT-6, Gemini 4 Pro still made my scalp tingle!”
Google’s next heavy hitter shows fierce 3D physics simulation, code generation, and long-context reasoning.
And Google leadership is blunt about it.
Google DeepMind’s new chief Koray Kavukcuoglu publicly confirmed for the first time that Gemini 4 is already in early post-training.
Training is far ahead of schedule: early pretraining basically wrapped in about two months.
Buckle up—a wave of hard demos is coming.
Core demos: seven hard hits
Overnight it felt like everyone could use Gemini 4 Pro.
From shared demos, its 3D generation, physics-engine simulation, and long-form code builds look stunning.
The peacock SVG test showed a huge leap too.
Seaplane takeoff: crushing Opus 5.5
Reviewer @AI_Screening asked it and Opus 5.5 to code a 3D floatplane taxiing and taking off from water.
Opus 5.5’s earlier clip had stunned crowds—but this time Gemini 4 Pro clearly surpassed it.
Opus 5.5’s plane moves, but the fuselage shakes hard as if about to break apart, and water feedback looks stiff.
Gemini 4 Pro is far smoother: steady accel, clear reflections, and splash rendering that feels real.
On fluid logic and visual cleanliness, it rubs Opus 5.5 into the ground.
16 minutes: a hand-rolled 3D ISS
Developer @thtbee_ dropped a bomb.
With no 3D assets imported, Gemini 4 Pro took just 16 minutes of pure code to build an interactive 3D International Space Station orbiting Earth.
Surprisingly, it not only modeled the station in Three.js but also pulled a correct Earth texture from the web so latitude/longitude and colors were accurate.
That “create from nothing” modeling still wows people.
Some still roast the UI as ugly, with a rear perspective bug remaining.
From lighthouse to rocket launch (Three.js)
User @HarshithLucky3 gave a hard prompt: “Use Three.js to turn a lighthouse into a rocket launch—no UI elements.”
Gemini 4 Pro followed perfectly.
Testers praised its “taste”: rich textures on every element, 3D quality absurdly high.
High-fidelity 3D Wii remote
Developer @LuminaBench tested industrial design modeling with a 3D Nintendo Wii remote.
Detail grasp on 3D objects was strong, and generation was “unbelievably fast.”
The imperfect bit: it over-ornaments prompts, packing detail even when not asked.
Taste across parts is uneven—some brilliant, some slightly off.
Tests say Checkpoint 2’s detail density is excellent, far more precise—and currently the fastest frontier model, unmatched.
Mechanical hummingbird: crushing GPT-6 Sol
Blogger @TimJayas threw his heirloom mechanical-hummingbird prompt at complex structure and motion.
Verdict: “Gemini 4 Pro was simply fantastic! Honestly, quality rivals Fable and far beats GPT-6 Sol!”
Sometimes GPT-6 Astra still wins, of course.
Full site in 14 minutes; games from one prompt
Engineering chops are strong too.
One developer generated a full creative product site—front and back—in 14 minutes. “Really, a site in minutes.”
Even more: many report a single prompt can spit out a complete interactive web game.
Leaked specs: 10M context + autonomous agents
The new checkpoint is fresh, but Gemini 4 Pro’s specs had already leaked.
Reports claim a 10-million-token context window and 256k output cap—if true, goodbye “code cuts off halfway,” and full large-software codebases from one shot.
Permanent cross-session memory, web access without an API, and real agent ability are also claimed.
It may even control robots—acting as a brain for physical machines.
An earlier benchmark chart showed Gemini 4 Pro beating GPT-6 Astra and Claude Fable 5.1 on coding, agents, and reasoning.
On DeepSWE v1.1 (agent coding) it hit 88%; on Terminal-bench 2.1 coding it soared to 95.3%.
Someone sniffed RSI in the air.
Google is urgent: trained in 2 months, all-in on 4.0
Why did Gemini 4 Pro drop from the sky? Google’s floor got hit.
After trailing GPT-6 and Opus 5.5 for a while, the giant struck back.
On Wednesday, new DeepMind chief Koray Kavukcuoglu publicly “opened the books.”
Per The Information, Koray said Gemini 4’s “most ambitious pretraining ever,” started in July, basically finished early pretraining in about two months.
Gemini 4 is now fully in the critical post-training phase.
“Our intention is to release an early post-training version as fast as possible,” Koray said, barely hiding excitement, “because we’ve seen the results and we’re very excited!”
That is why a “half-baked but already fierce” test node suddenly appeared on the arena.
And after finding Gemini 3.5 Pro not enough to crush rivals, Google cut deep—throwing all compute at Gemini 4.
Asked sharply whether Google fears falling behind peers, Koray said flatly: “I have 100% confidence in my team. We are destined to stay at the frontier forever.”
He said Gemini 4 is already used internally by engineers on Antigravity AI—a retort to rumors Google staff coded with Claude.
He stressed they no longer cling to the AGI label: “The real question is whether we can build agents we can fully trust.”
Gemini 4’s R&D, he added, is helping hardware engineers design “the next two or three generations of TPUs”—an algorithms-to-silicon moat that keeps NVIDIA wary.
Finally Koray set a clock: he hopes Gemini 4’s formal release comes “well before year-end.”
Some say Gemini 4 Pro could ship around China’s National Day holiday week.
Performance bombs aside, Google may also start a price war. Leaked pricing is aggressive:
Input: $2.25 / 1M tokens
Output: $11.25 / 1M tokens
Versus rivals charging tens of dollars for flagship models, that is “bone-fracture pricing.”
No wonder people shout: “This value is insane—if 4 Pro beats expectations, topping up membership now is bottom-fishing!”
And Google is sending AI chips to space!
Next week a Falcon 9 will loft a satellite carrying four TPUs.
Someday a vast Starlink-style AI data-center net may really hang in the sky.
Looks like LMSYS’s gemini-3.8-flash was only the tip of “very early post-training.”
As one tester put it: “Checkpoint 2 is a huge visible leap over checkpoint 1—Google is truly cooking a big dish this time.”
How will OpenAI and Anthropic answer?
Could Gemini 4 Pro change the endgame of the ASI finals?