AI GUIDEPartner content

Second Wave of Gemini 4 Pro in Arena: Testing the Baryon-B Build

A second wave of tests of the presumed Gemini 4 Pro, under the codename Baryon-B, has been recorded in LMS Chatbot Arena. Generation time has been cut in half, and complex WebGL scenes run on the first try.

Affiliate link: your price stays the same and the project earns a commission.

Secret trials disguised as a budget model

LMS Chatbot Arena continues its practice of blind-testing unannounced models by swapping their identifiers. Disguised as the standard Gemini 3.8 Flash, released in early September as a lightweight solution with a score of 41 on the Artificial Analysis Intelligence Index, it periodically gives users responses from a significantly larger architecture. The developer community associates this checkpoint with an early version of Gemini 4 Pro.

According to sources tracking the code names of Google's internal builds, the first phase of testing was conducted under the designation Argon, while the current one has the identifier Baryon-B. This suggests the continued use of internal nomenclature tied to the periodic table and physics terms.

Changes between the first and second waves: speed measurements

A comparison of the two testing phases, one week apart, showed a noticeable change in performance:

  • First wave (mid-September): compilation of complex interactive projects took 5 to 10 minutes. The model handled complex vector graphics and static 3D geometry, but produced gaps in polygon meshes and glitches in looping animations. Subjective assessments placed it on par with GPT-6 Astra Max, but below Claude Fable 5.1.
  • Second wave: benchmarks recorded a twofold drop in response latency. Tasks of comparable complexity are now consistently completed in under 5 minutes. This suggests either the deployment of optimized TPU clusters for inference or calibration of the compute budget during the reasoning stage (reasoning tokens).

Complex tests in WebGL and Three.js

The main proving ground for identifying differences between standard Flash and the hidden build was procedural JavaScript coding using the Three.js library. Fully rendering such scenes requires calculating spatial geometry, lighting shaders, and physical interactions directly in the browser window.

In standard tests, a typical entry-level model is limited to primitives (cubes, spheres) and cuts off the script because it exceeds the context length. The Baryon-B checkpoint ran successfully on the first try in several nontrivial scenarios:

  • Dynamic hydrodynamics: generating a ship model with calculations for ocean surface deformation and the physics of pitching on waves in real time.
  • Complex kinematics: an interactive airship with detailed gondolas and rotating propellers, as well as a mechanical flower with interconnected transmission gears.
  • Physically accurate rendering: building a Formula 1 race car with independent double-wishbone suspension for each wheel, a matte carbon-fiber monocoque, racing tire texture, and a studio HDR lighting map.

Disagreements over evaluations and infrastructure context

Despite demonstrating complex shaders, testers are divided in their opinions. A number of developers believe the model has caught up with Claude Opus 5.5 in terms of syntax accuracy and compilation speed. Skeptics, however, note that there has been no qualitative leap in logical reasoning compared with the first wave: the improvement is most apparent in applied frontend coding, while the system's behavior on large-scale repositories and under strict hallucination constraints remains untested.

Google's parallel expansion of its server fleet and testing of specialized chips to accelerate inference point to preparations to reduce compute costs per unit before the public release of the next generation of the Gemini lineup.

Compare models before you start

The service sets its plans, limits and model catalog. If they differ from this article, contact us so we can update it and record a new review date.

Browse models

Affiliate link: your price stays the same and the project earns a commission.

Gemini 4 Pro Baryon-B LMS Chatbot Arena WebGL Three js Google AI model testing

SEO Mind42 editorial team

We explore SEO and neural networks in practice: test services on our own projects, verify prices and limits against primary sources, and share things you can put to use the same day.

πŸ“š Reference guide to SEO and AI πŸ”„ Materials are updated πŸ• Updated: 3 October 2026

Related reading

All in this section β†’