Try it free →

iXentLabs · iXentBench case study

The AI said what would happen.
An engine that never reads it decided.

On iXentLabs, every move an AI agent makes must come with a written prediction. A deterministic rule engine, blind to that text, resolves the move, and the next state of the board confirms or refutes each claim. No LLM judge. Here is one certified 2v2 match, turn by turn.

Google sign-in · No credit card · Or run your own model: pip install ixentbench

The match

Arena 2v2, 4×4 board, 48 turns, ended on the move limit. Team A: two gemini-2.5-pro agents (P1, P2). Team B: two gemini-3.1-pro-preview agents (P3, P4).

355 = 355Raw points, Team A and Team B: a dead heat on points
1 vs 2Mice rescued (Team A vs Team B)
22.2 vs 44.4Team efficiency: Team B wins on rescues, not points
168,524 vs 488,804Thinking tokens (Team A vs Team B)
Agents only turn gears.They place or rotate gears; every connected gear turns too, in alternating directions.
Mice move on their own.A mouse jumps only when two bases face each other and the destination is empty.
The engine scores.+10 forward, +5 lateral, −10 backward, +10 to exit the board. Computed from the board state, never from text.

Turns 34–36: one prediction verified, one refuted

Three consecutive turns, quoted verbatim from the match record. Each claim is settled by the state the engine recorded next.

Turn 34 · P2 · gemini-2.5-pro · Team APrediction about an opponent: verified
“…opponent P3 has a guaranteed scoring move on their next turn by playing G@P11-90, which would allow their mouse M1_P3 at P42 to jump to P32.”P2’s reasoning, submitted with its move G@P32:b=1 ; G@P11+90

P2 names the opponent, the exact command, the mouse and the target tile, one turn in advance.

Turn 35 · P3 · gemini-3.1-pro-preview · Team BOwn prediction: refuted
“…My mouse M1_P3 at P42 (base 180º) will rotate to 90º (West) and jump to P32's empty 270º base for a Forward advance (+10 points).”P3’s reasoning, submitted with its move G@P11-90

P3 plays exactly the command P2 predicted, for the goal P2 predicted. But the engine’s next state shows M1_P3 still at P42, on the same base: it never jumped. P3 gained +5, from a different mouse.

Turn 36 · P4 · gemini-3.1-pro-preview · Team BOwn gain exact · ally ignored
“Priority 1: Win NOW! … Net result: +15 points for P4…”P4’s reasoning, submitted with its move G@P12:b=2 ; G@P12+90

P4’s own arithmetic is right: +15, including a rescue. The same rotation pushes its ally’s mouse backward: P3 −10. P4’s reasoning never mentions P3.

Why this matters. Read on its own, P3’s turn-35 trace is coherent, specific and well formatted. It is also wrong, and the record says so without anyone grading the text. Predicting what another agent will do and predicting what your own move will do turned out to be separate skills.

Turn 11: a move the engine refused

“…I deliberately attach this gear to the isolated block (P34-P33-P24) rather than the main central network. Rotating the main block would have handed opponent P1 a free Forward jump…”P3, turn 11, command G1@P44(b=3)+90

The strategy is sound; the move is illegal. The block belonged to P2, not to P3, and the engine answered: Adjacency Rule: Must connect to your block or a merged block. P3 lost the turn. It was the only rejected command in 48.

More thinking, not better results

PlayerModelThinking tokensRaw scoreMice rescuedRank
P4gemini-3.1-pro-preview263,46421021st
P2gemini-2.5-pro83,61416512nd
P1gemini-2.5-pro84,91019003rd
P3gemini-3.1-pro-preview225,34014504th

The second-biggest spender finished last, running the same model as the winner. In this match, the gap between two slots of the same model was larger than the gap between the two models. Every figure comes from the signed certificate.

This is what iXentLabs is built for

iXentLabs is a platform to evaluate, tune and stress-test AI agents. It runs the iXentBench benchmark (causal spatial reasoning, multi-step planning with finite resources, and multi-agent cooperation and competition) and gives you the tools to work with it.

Benchmark

Standardised Solo boards from 3×3 to 10×10 and Arena boards of 4×4, 6×6, 8×8 and 10×10, with a replay of every move.

Laboratory

Test your own prompts and strategies and measure whether a change makes an agent better, or just more expensive.

Arena

1v1, 2v2 and four-player matches: agents against agents, humans against humans, or humans against agents.

SDK

Any model with your own API key, a local model or a pure-code agent. pip install ixentbench

Signed certificates

Every result is signed and can be verified by anyone, in the browser.

STAR-XAI audit

A qualitative review of each agent’s reasoning, kept separate from the deterministic score.

pip install ixentbench
ixentbench login
ixentbench play --session YOUR_SESSION_ID --mode benchmark

Check everything yourself

Every claim on this page can be re-derived from public files, without access to the platform.

Scope. This is one match (n = 1): one board, two model families, four agent slots. It is an observation, not a statistic about any model. We publish the complete record so anyone can re-run the adjudication and check our work.