An empirical observation of how field sequence in Gemini 3.6 Flash's Structured Outputs (
responseSchema) and Thinking Budget influence the model's move selection in a 5×5 Gomoku game.
Key Takeaways (3-Bullet Summary)
- When invoking Gemini 3.6 Flash with Structured Outputs, placing the final decision field at the top of the JSON schema (Decision-First) causes the model to fail spatial reasoning traps—even with a 2,048 Thinking Budget.
- In contrast, arranging the schema sequentially ('Spatial Scan ➔ Threat Analysis ➔ Final Move', or CoT-First) achieves 100% accuracy, even with a Thinking Budget of 0 (thinking disabled).
- Due to the autoregressive nature of LLMs, the declaration order of properties in a JSON Schema functions as an externalized Chain-of-Thought (CoT) scratchpad.
Why We Conducted This Experiment
In enterprise production pipelines, Google Gemini's Structured Outputs (responseSchema) is one of the most widely adopted capabilities. It enforces strict type contracts and schema adherence without fragile regex or manual markdown parsing.
When paired with the Thinking Budget parameter available in Gemini 2.5 and 3.6 models, developers anticipate flawless execution across complex logic, route planning, logistics dispatch, and game theory scenarios.
However, engineers frequently observe a baffling discrepancy:
- "Even with a generous Thinking Budget of 1,024 or 2,048, our agents fall for obvious decoy traps."
- "In the generated JSON's explanation field, the model clearly articulates the exact correct answer, yet the preceding action field contains the wrong coordinate."
To isolate the mechanism behind this phenomenon, we designed an empirical benchmark testing visual-spatial perception and logical filtering: the 5×5 Mini Gomoku (Connect-4) Decoy Trap.
Experiment Design: The 5×5 Mini Gomoku Decoy Trap
The model is supplied with a 5×5 board state represented as a 2D text array. The objective is to identify the single vital defense coordinate (row, col) to prevent the opponent ('O') from completing 4-in-a-row on their next turn.
[Board Configuration]
col 0 col 1 col 2 col 3 col 4
row 0: [ . , O , . , . , . ]
row 1: [ X , O , O , O , X ]
row 2: [ . , . , . , O , . ]
row 3: [ . , . , . , . , . ]
row 4: [ . , . , . , . , . ]
This board incorporates a dual-threat mechanism:
1. Visual Decoy: Dead Threat
- In row 1, three opponent stones ('O') line up consecutively:
(1,1),(1,2),(1,3). - This prominently commands self-attention layers due to high stone density.
- However, both flanking ends—
(1,0)and(1,4)—are already blocked by 'X'. It is physically impossible for 'O' to form 4-in-a-row here. It is a completely Dead Threat.
2. The Fatal Live Threat
- A diagonal inspection reveals three 'O' stones:
(0,1)➔(1,2)➔(2,3). - The trailing cell
(3,4)is currently empty (.). - If 'X' fails to occupy
(3,4)this turn, 'O' wins immediately on the next move. - The Only Correct Defense:
row: 3, col: 4
Two JSON Schema Designs: Decision-First vs CoT-First
Under identical prompts and identical Gemini 3.6 Flash hyperparameters, we varied only the declaration sequence of properties within the JSON Schema.
Schema A: Decision-First (Standard API Pattern)
Requests the execution move first, followed by retrospective analysis:
{
"type": "OBJECT",
"properties": {
"best_defense_move": {
"type": "OBJECT",
"properties": {"row": {"type": "INTEGER"}, "col": {"type": "INTEGER"}},
"required": ["row", "col"]
},
"real_threat_direction": {"type": "STRING"},
"ignored_dead_threat": {"type": "STRING"},
"analysis": {"type": "STRING"}
},
"required": ["best_defense_move", "real_threat_direction", "ignored_dead_threat", "analysis"]
}
Schema B: CoT-First (Step-by-Step Elicitation Pattern)
Forces the model to verbalize board scanning and threat evaluation before outputting coordinates:
{
"type": "OBJECT",
"properties": {
"spatial_board_scan": {"type": "STRING"},
"dead_threat_analysis": {"type": "STRING"},
"live_threat_analysis": {"type": "STRING"},
"best_defense_move": {
"type": "OBJECT",
"properties": {"row": {"type": "INTEGER"}, "col": {"type": "INTEGER"}},
"required": ["row", "col"]
}
},
"required": ["spatial_board_scan", "dead_threat_analysis", "live_threat_analysis", "best_defense_move"]
}
Empirical Benchmark Results
We measured end-to-end latency, total token consumption (usageMetadata), and accuracy across 5 distinct configurations.
| Test Case | Thinking Budget | Schema Architecture | Latency | Total Tokens | Defense Coordinate | Outcome |
|---|---|---|---|---|---|---|
| Case 1 | 0 (OFF) | Decision-First | 3.47s | 653 tokens | (row: 0, col: 3) | FAIL (Incorrect) |
| Case 2 | 0 (OFF) | CoT-First | 4.41s | 890 tokens | (row: 3, col: 4) | SUCCESS (Correct!) |
| Case 3 | 1,024 (ON) | Decision-First | 4.02s | 790 tokens | (row: 3, col: 3) | FAIL (Incorrect) |
| Case 4 | 1,024 (ON) | CoT-First | 5.80s | 1,284 tokens | (row: 3, col: 4) | SUCCESS (Correct!) |
| Case 5 | 2,048 (ON) | Decision-First | 5.61s | 907 tokens | (row: 3, col: 3) | FAIL (Incorrect) |
Detailed Analysis: Autoregressive Commitment & Model Regret
These empirical results reveal several critical architectural insights:
1. Why Decision-First Schemas Fail Despite Large Thinking Budgets
In Case 3 (Budget 1,024) and Case 5 (Budget 2,048), the model emitted (3, 3)—an incorrect move. Examining the raw text emitted in the subsequent analysis field reveals a remarkable quote:
"Wait, look at (0,1), (1,2), (2,3) diagonal! (0,1), (1,2), (2,3) forms 3 'O's diagonally down-right. Next spot is (3,4)... The next cell to complete 4 in a row is row 3 col 4! Blocking at (3,4) prevents the diagonal win."
The model discovered the true diagonal threat during its generation of the analysis field and explicitly proved (3, 4) was the only winning defense! However, it had already committed best_defense_move: {"row": 3, "col": 3} to the output buffer earlier in the stream.
Because autoregressive language models generate tokens sequentially from left to right without backtracking, the model had no mechanism to revise its previously emitted coordinate.
2. The Power of CoT-First Schemas: 100% Accuracy at Budget 0
In Case 2, with the Thinking Budget completely disabled (0 tokens), the model achieved 100% accuracy at (3, 4).
By sequencing spatial_board_scan and threat_analysis ahead of best_defense_move, the generated intermediate text served as an externalized reasoning scratchpad. As the model verbalized that row 1 was blocked by 'X', attention weights shifted to the diagonal, feeding a clean semantic context directly into the final best_defense_move tokens.
[Information Flow Comparison]
1. Decision-First:
Board Input ➔ [Emits Move: (3,3) FAIL] ➔ [Writes Analysis: "Oh, it was diagonal (3,4)..."] (Too Late)
2. CoT-First:
Board Input ➔ [Emits Spatial Scan] ➔ [Emits Threat Filtering] ➔ [Emits Move: (3,4) SUCCESS!]
Four Architectural Guidelines for Production JSON Schemas
1. Always Place Terminal Actions at the End of the Schema
Any field that triggers real-world side effects—such as trade executions, tool calling parameters, routing decisions, or physical coordinates—must be declared as the final property in your JSON Schema.
2. Use Structured Intermediate Fields as Verification Barriers
Incorporate intermediate validation fields such as preconditions_met, extracted_constraints, and counter_arguments prior to the payload field.
3. Consider 'CoT-First Schema + Flash' Over Pure Reasoning Models
Reasoning models add significant latency and token billing overhead. In many production classification and routing tasks, pairing Gemini Flash with a well-sequenced schema achieves equivalent or superior accuracy at lower latency and cost.
4. Dual Defense for High-Stakes Workflows
For safety-critical autonomous operations, combine both: activate a moderate Thinking Budget alongside a CoT-First schema to eliminate edge-case hallucinations.
Frequently Asked Questions (FAQ)
Q1. Don't JSON specifications define object keys as unordered?
While RFC 8259 states JSON object keys are unordered, LLM generation is strictly sequential in time. The model produces tokens strictly adhering to the schema's property definition order. To the model's attention mechanism, key order is definitive.
Q2. Wouldn't expanding the Thinking Budget to 8,192 resolve this?
For 2D grid representations and spatial coordinates, tokenization compresses spatial geometry into 1D sequences. Without explicit token emission to unroll spatial relationships into the context window, latent hidden-state reasoning remains susceptible to visual decoy attractors.
Q3. Does CoT-First significantly increase latency?
In our measurements, Case 1 took 3.47s while Case 2 took 4.41s—an overhead of ~0.94 seconds. However, avoiding failed execution, agent rollbacks, or downstream retry loops yields significantly lower total cost of ownership (TCO) and superior user experience.