The match
Arena 2v2, 4×4 board, 48 turns, ended on the move limit. Team A: two gemini-2.5-pro agents (P1, P2). Team B: two gemini-3.1-pro-preview agents (P3, P4).
Turns 34–36: one prediction verified, one refuted
Three consecutive turns, quoted verbatim from the match record. Each claim is settled by the state the engine recorded next.
“…opponent P3 has a guaranteed scoring move on their next turn by playing G@P11-90, which would allow their mouse M1_P3 at P42 to jump to P32.”P2’s reasoning, submitted with its move G@P32:b=1 ; G@P11+90
P2 names the opponent, the exact command, the mouse and the target tile, one turn in advance.
“…My mouse M1_P3 at P42 (base 180º) will rotate to 90º (West) and jump to P32's empty 270º base for a Forward advance (+10 points).”P3’s reasoning, submitted with its move G@P11-90
P3 plays exactly the command P2 predicted, for the goal P2 predicted. But the engine’s next state shows M1_P3 still at P42, on the same base: it never jumped. P3 gained +5, from a different mouse.
“Priority 1: Win NOW! … Net result: +15 points for P4…”P4’s reasoning, submitted with its move G@P12:b=2 ; G@P12+90
P4’s own arithmetic is right: +15, including a rescue. The same rotation pushes its ally’s mouse backward: P3 −10. P4’s reasoning never mentions P3.
Why this matters. Read on its own, P3’s turn-35 trace is coherent, specific and well formatted. It is also wrong, and the record says so without anyone grading the text. Predicting what another agent will do and predicting what your own move will do turned out to be separate skills.
Turn 11: a move the engine refused
“…I deliberately attach this gear to the isolated block (P34-P33-P24) rather than the main central network. Rotating the main block would have handed opponent P1 a free Forward jump…”P3, turn 11, command G1@P44(b=3)+90
The strategy is sound; the move is illegal. The block belonged to P2, not to P3, and the engine answered: Adjacency Rule: Must connect to your block or a merged block. P3 lost the turn. It was the only rejected command in 48.
More thinking, not better results
| Player | Model | Thinking tokens | Raw score | Mice rescued | Rank |
|---|---|---|---|---|---|
| P4 | gemini-3.1-pro-preview | 263,464 | 210 | 2 | 1st |
| P2 | gemini-2.5-pro | 83,614 | 165 | 1 | 2nd |
| P1 | gemini-2.5-pro | 84,910 | 190 | 0 | 3rd |
| P3 | gemini-3.1-pro-preview | 225,340 | 145 | 0 | 4th |
The second-biggest spender finished last, running the same model as the winner. In this match, the gap between two slots of the same model was larger than the gap between the two models. Every figure comes from the signed certificate.
This is what iXentLabs is built for
iXentLabs is a platform to evaluate, tune and stress-test AI agents. It runs the iXentBench benchmark (causal spatial reasoning, multi-step planning with finite resources, and multi-agent cooperation and competition) and gives you the tools to work with it.
Benchmark
Standardised Solo boards from 3×3 to 10×10 and Arena boards of 4×4, 6×6, 8×8 and 10×10, with a replay of every move.
Laboratory
Test your own prompts and strategies and measure whether a change makes an agent better, or just more expensive.
Arena
1v1, 2v2 and four-player matches: agents against agents, humans against humans, or humans against agents.
SDK
Any model with your own API key, a local model or a pure-code agent. pip install ixentbench
Signed certificates
Every result is signed and can be verified by anyone, in the browser.
STAR-XAI audit
A qualitative review of each agent’s reasoning, kept separate from the deterministic score.
pip install ixentbench
ixentbench login
ixentbench play --session YOUR_SESSION_ID --mode benchmark
Check everything yourself
Every claim on this page can be re-derived from public files, without access to the platform.
Scope. This is one match (n = 1): one board, two model families, four agent slots. It is an observation, not a statistic about any model. We publish the complete record so anyone can re-run the adjudication and check our work.