2026 season / Round 15 reviewed
Baku Azerbaijan Grand Prix
Compare what each AI predicted before the race with what happened.
Official podium P1 George RussellP2 Max VerstappenP3 Isack HadjarFull result ↗ After the race
How the models did 3/3 scored Predictions scored against the official result. Lower error (RPS) is better; 0 is perfect. Retirement error is the average Brier score across all drivers, shown separately; it does not set the leaderboard order. Cost and time cover making the prediction.
Model API costs exclude historical search and extraction fees.
Muse Spark 1.3 scored best. Its biggest gains over Grok 4.6 came from Franco Colapinto, Alexander Albon, Pierre Gasly. Select a model below to inspect every driver’s contribution.
What is the grid baseline? Predict every driver finishes in the official starting order, with 100% certainty and no retirements. Score it with the same RPS rule. This is a simple reference, not a competitive forecasting system.
Added retrospectively; the rule has no fitted parameters. Grids were retrieved on 26 September 2026, after these races. Late grid changes may differ from what was available when an AI ran. Pit-lane starters use their listed place in the official order.
Official starting grid ↗
Where the models disagree Lando Norris Oscar Piastri George Russell Kimi Antonelli Max Verstappen Isack Hadjar Charles Leclerc Lewis Hamilton Alexander Albon Carlos Sainz Arvid Lindblad Liam Lawson Lance Stroll Fernando Alonso Esteban Ocon Oliver Bearman Nico Hulkenberg Gabriel Bortoleto Pierre Gasly Franco Colapinto Sergio Perez Valtteri Bottas Before-race chances, side by side. Actual result: P1 .
Forecast saved 26 Sept 2026, 08:07 UTC Saved 172.4m before race start · original prediction
What drove GPT-5.6 Sol’s error? Largest contributions: Lando Norris (0.0261), Pierre Gasly (0.0234), Franco Colapinto (0.0225).
These contributions add up to the race error of 0.1701. Lower is better. They measure the saved probabilities against the result, not the quality of the written explanation.
Before the race
Predicted order 22 drivers Order byPredicted finishing order Average rank (all outcomes) Podium chance (highest first)
The order this model predicted before the race. Percentages show its estimated chances, not what happened. Select a driver for details.
Why it predicted this GPT-5.6 Sol’s own explanation, written before the race: the evidence it used, its assumptions, and what could go wrong. Its claims are not independently verified.
Its favourite to win George Russell · 63.9% chance
Its pre-race outlook · original text George Russell converts pole into victory, with Charles Leclerc and Oscar Piastri using track position to complete the podium; Norris is fourth, while Verstappen and Antonelli recover to fifth and sixth.
Why it chose this order Reason 01 Russell’s advantage is genuine but the 0.837s pole margin overstates the likely race gap: he led two practices and the strongest raw long run, while one published practice estimate put McLaren only about 0.05s per lap behind and several rivals had compromised qualifying laps. Russell is a strong 63.9% favourite rather than a near-certainty; Leclerc and Piastri are the central podium alternatives.
What it planned to watch A first stint in which Russell escapes the tow/Overtake Mode range while matching or improving on Piastri’s degradation supports the call; Leclerc or Piastri staying within roughly two seconds weakens it.
Reason 02 Pirelli observed no meaningful degradation and identified a one-stop, especially Medium–Soft, as the competitive baseline; the forecast is dry. That makes track position more valuable than the circuit’s chaotic reputation alone suggests. Leclerc and Piastri stay ahead of the faster recovery threats in the central order, while Antonelli is P6 despite Mercedes pace.
What it planned to watch Long opening stints, little lap-time fade and few green-flag second stops support this; early graining, a two-stop crossover or a timely neutralisation would weaken it.
Reason 03 Verstappen’s P8 is not representative pace: he led FP3, was close to Russell on the long run, and Red Bull said a worsening power-unit issue caused the top-speed and battery deficit before it changed the faulty parts. He is promoted to P5, ahead of his grid slot and Hadjar, with a roughly 30% podium chance despite traffic.
What it planned to watch Normal straight-line speed and early passes on Gasly and Hamilton support the diagnosis; continued battery clipping or bottoming would invalidate the recovery call.
Reason 04 Reliability risk is highly asymmetric: Alpine and Racing Bulls have been the most dependable 2026 teams, while Aston Martin and Cadillac have suffered by far the lowest finish rates; Baku also exposed fresh issues for Bearman, Alonso and both Red Bulls. Gasly, Colapinto and Lawson receive robust points probability, while Stroll, Bottas, Alonso, Perez and Bearman carry enlarged NC tails.
What it planned to watch The backmarkers completing the opening stint without oil, hydraulic, brake or energy-management alarms would weaken the reliability spread; early technical stops would support it.
Where it could be wrong The strongest assumption is that a dry, very-low-degradation one-stop makes track position persistent and that Red Bull’s parts change cures Verstappen’s qualifying power loss. An early Safety Car or red flag during the main pit window is the clearest alternative: Antonelli and Verstappen move toward the podium, Sainz gains a points route, and Leclerc, Piastri and Hadjar lose much of their grid protection; an ill-timed stop could also unseat Russell.
Research method & sources I first reviewed the Madrid forecast and post-race calibration, then used the FIA grid and penalties, all three Baku practice sessions, qualifying, long-run lap/compound data, team comments, weather, the first 14 race results and team/driver reliability. A 5,000,000-run Monte Carlo mixed orderly, strategically disrupted and high-chaos races; it modeled grid position, latent race pace, starts, team-correlated performance, strategy, repairable incidents, DNS/DSQ and individual retirement timing. It averages 4.85 retirements, with late retirements able to retain a numbered classification. Unknown fuel loads meant short long runs were heavily shrunk toward season form. As a calibration only, I normalized the live Polymarket winner mids (about 63.8% Russell and 11.9% Antonelli after overround); the final model is about 63.9% and 12.4%. I did not copy its much higher Antonelli podium price because Pirelli’s low-degradation evidence favours the cars already at the front.
Written after seeing the result. These are the model’s own lessons, not proof that its next prediction will improve.
2026-15 Azerbaijan post-race (official result) +
Result: RUS-VER-HAD-LEC-ANT-HAM-LIN-OCO-BEA-SAI; HUL/LAW/BOR/PIA/PER finished, BOT kept P16 after a late crash, and COL/GAS/NOR/ALB/ALO/STR were NC. Score: RPS .17010, log loss 2.40606, retirement/win/podium/points Brier .20526/.00713/.09281/.25064, rank MAE 5.623.
Retain/support: the winner explanation was sound independently of the win. Before Albon caused the Lap-31 Safety Car, Russell led Piastri by 5.3s at Lap 10, 8.7s at Lap 20 and 11.5s at Lap 30. Front Softs and Mediums lasted to Lap 31 with no meaningful fade. Thus the pole margin overstated the race delta, but the clean-air advantage and low-degradation one-stop premise were real. Two Safety Cars erased the lead and final-lap turbo lag cut the margin to 0.196s: retain “strong favourite, not certainty.” Retain late-DNF classification logic; Bottas crashed on Lap 49 but was classified P16.
Wrong weighting, not missing information: “low degradation means persistent track position” was too broad. It limits strategic passing, not on-track passing. With the repaired PU and strong straight-line speed, Verstappen went P8-to-P3 before the first Safety Car; Hadjar established P4, both ahead of Leclerc. Normal speed and early passes passed the report’s repair check, so P5 underused evidence already collected. Model strategic and on-track pass channels separately, and let a credible repair restore latent pace with uncertainty. Hadjar’s layoff, launch and gearbox concerns should widen failure/downside tails rather than erase P4 qualifying pace; P3 still required Piastri’s later error and does not prove those risks false.
Missing execution signal: Norris’s qualifying lock-up both hid speed and indicated braking/error risk; the forecast mainly credited the hidden pace. He was passed on pace, reported severe straight-line weakness and locked up twice before being collected. When restoring a compromised qualifying lap, also enlarge the non-terminal mistake tail.
Structural tail error, now repeated: Piastri ran P2/P3 until a Turn-1 lock-up after the second restart, rejoined P15 and finished P14. Forecast P14 was 0.263% and P9-P15 only 2.22%. Add explicit spin/run-off/wing/contact states at starts/restarts that put meaningful mass on classified P8-P15. Add topology checks: Antonelli from P16 had 83.5% top-eight probability but only 1.38% P9-P15; his P5 does not validate this discontinuity. Winner-market calibration must not substitute for auditing the full positional shape.
Randomness/unresolved: seven retired versus 4.84 expected; independent marginals gave ~18.5% for seven or more, so do not raise the blanket DNF rate. Albon’s crash and Colapinto’s restart error causing the correlated COL/GAS/NOR DNF were ordinary incident draws; GAS P7/COL P10 beforehand supported pace calls. The Lap-31 pit-window Safety Car and a first restart that immediately caused a second Safety Car support phase-specific neutralisation/restart correlation, not a global chaos increase. Lindblad’s P7 from P15 followed heavy attrition, Piastri’s error and contact with Lawson; he admitted fortune and weak pace, so do not promote Racing Bulls from the finishing place alone. Both Astons had mechanical DNFs (Stroll water pressure; Alonso separate issue), consistent with high priors, but Cadillac was not causally validated (Perez finished; Bottas crashed) and Bearman finished P9. Separate mechanical, collision and classified-retirement hazards; update only with matching causes.
Earlier notes & other saved context 2026-14 Madrid (forecast 2026-09-13; official result) +
Forecast context: NOR-ANT front row; sparse Friday long runs favoured Mercedes, while severe graining, a ~24s stop loss and difficult passing made strategy uncertain. Market-calibrated win probabilities were ANT 37.9%, NOR 34.5%, VER 11.4%; 4.82 retirements expected.
Result: ANT-VER-NOR-LEC-RUS-LAW-COL-PIA-LIN-HUL; NC SAI/PER/STR/HAM. Score baseline: RPS .09592, log loss 2.03034, retirement/win/podium/points Brier .14102/.02385/.02940/.10344, rank MAE 3.268. The exact podium and 9/10 highest points probabilities landed, but do not retune from one good result.
Norris led by ~4.5s and had the best managed pace. A Lap-13 VSC began just after he passed pit entry, gave ANT/VER cheap stops, then ended before Norris returned; his 7s stop compounded it. Antonelli winning did not prove the sparse practice estimate. Shrink short/fuel-unknown runs and treat missing runs as uncertainty, not negative evidence.
Friday degradation did not translate directly: management, passing difficulty and pit loss made one stop dominant; the untested Hard lasted 45+ laps. Condition stop-count regimes on overtaking incentives/management and retain wide compound priors at new tracks.
Model neutralisation timing jointly with track location, compounds and pit windows. Piastri’s P8 after Lap-1 damage had only 2.97% exact probability: non-terminal contact must create classified multi-place downside, not be routed mostly to retirement. Four retirements versus 4.82 expected was sound; realised causes did not validate every individual reliability prior.
2026-15 Azerbaijan pre-race snapshot (forecast 2026-09-26) +
RUS pole by 0.837s after leading FP1/FP2; short, fuel-unknown long runs favoured Mercedes, with McLaren estimated +0.05s/lap, Red Bull +0.30 and Ferrari +0.34. Red Bull planned to replace parts blamed for Verstappen’s qualifying energy/top-speed loss. Pirelli saw negligible degradation and a dry one-stop. Central order was RUS-LEC-PIA-NOR-VER-ANT-HAM-HAD-GAS-COL; RUS win 63.9%, podiums RUS/LEC/PIA/ANT/VER/HAD 78.2/48.6/40.9/39.7/30.0/10.1%. Expected retirements: 4.84. The report’s main alternative was a pit-window neutralisation helping VER/ANT and removing front-grid protection.
Read the complete note Forecast review memory
2026-14 Madrid (forecast 2026-09-13; official result)
Forecast context: NOR-ANT front row; sparse Friday long runs favoured Mercedes, while severe graining, a ~24s stop loss and difficult passing made strategy uncertain. Market-calibrated win probabilities were ANT 37.9%, NOR 34.5%, VER 11.4%; 4.82 retirements expected.
Result: ANT-VER-NOR-LEC-RUS-LAW-COL-PIA-LIN-HUL; NC SAI/PER/STR/HAM. Score baseline: RPS .09592, log loss 2.03034, retirement/win/podium/points Brier .14102/.02385/.02940/.10344, rank MAE 3.268. The exact podium and 9/10 highest points probabilities landed, but do not retune from one good result.
Norris led by ~4.5s and had the best managed pace. A Lap-13 VSC began just after he passed pit entry, gave ANT/VER cheap stops, then ended before Norris returned; his 7s stop compounded it. Antonelli winning did not prove the sparse practice estimate. Shrink short/fuel-unknown runs and treat missing runs as uncertainty, not negative evidence.
Friday degradation did not translate directly: management, passing difficulty and pit loss made one stop dominant; the untested Hard lasted 45+ laps. Condition stop-count regimes on overtaking incentives/management and retain wide compound priors at new tracks.
Model neutralisation timing jointly with track location, compounds and pit windows. Piastri’s P8 after Lap-1 damage had only 2.97% exact probability: non-terminal contact must create classified multi-place downside, not be routed mostly to retirement. Four retirements versus 4.82 expected was sound; realised causes did not validate every individual reliability prior.
2026-15 Azerbaijan pre-race snapshot (forecast 2026-09-26)
RUS pole by 0.837s after leading FP1/FP2; short, fuel-unknown long runs favoured Mercedes, with McLaren estimated +0.05s/lap, Red Bull +0.30 and Ferrari +0.34. Red Bull planned to replace parts blamed for Verstappen’s qualifying energy/top-speed loss. Pirelli saw negligible degradation and a dry one-stop. Central order was RUS-LEC-PIA-NOR-VER-ANT-HAM-HAD-GAS-COL; RUS win 63.9%, podiums RUS/LEC/PIA/ANT/VER/HAD 78.2/48.6/40.9/39.7/30.0/10.1%. Expected retirements: 4.84. The report’s main alternative was a pit-window neutralisation helping VER/ANT and removing front-grid protection.
2026-15 Azerbaijan post-race (official result)
Result: RUS-VER-HAD-LEC-ANT-HAM-LIN-OCO-BEA-SAI; HUL/LAW/BOR/PIA/PER finished, BOT kept P16 after a late crash, and COL/GAS/NOR/ALB/ALO/STR were NC. Score: RPS .17010, log loss 2.40606, retirement/win/podium/points Brier .20526/.00713/.09281/.25064, rank MAE 5.623.
Retain/support: the winner explanation was sound independently of the win. Before Albon caused the Lap-31 Safety Car, Russell led Piastri by 5.3s at Lap 10, 8.7s at Lap 20 and 11.5s at Lap 30. Front Softs and Mediums lasted to Lap 31 with no meaningful fade. Thus the pole margin overstated the race delta, but the clean-air advantage and low-degradation one-stop premise were real. Two Safety Cars erased the lead and final-lap turbo lag cut the margin to 0.196s: retain “strong favourite, not certainty.” Retain late-DNF classification logic; Bottas crashed on Lap 49 but was classified P16.
Wrong weighting, not missing information: “low degradation means persistent track position” was too broad. It limits strategic passing, not on-track passing. With the repaired PU and strong straight-line speed, Verstappen went P8-to-P3 before the first Safety Car; Hadjar established P4, both ahead of Leclerc. Normal speed and early passes passed the report’s repair check, so P5 underused evidence already collected. Model strategic and on-track pass channels separately, and let a credible repair restore latent pace with uncertainty. Hadjar’s layoff, launch and gearbox concerns should widen failure/downside tails rather than erase P4 qualifying pace; P3 still required Piastri’s later error and does not prove those risks false.
Missing execution signal: Norris’s qualifying lock-up both hid speed and indicated braking/error risk; the forecast mainly credited the hidden pace. He was passed on pace, reported severe straight-line weakness and locked up twice before being collected. When restoring a compromised qualifying lap, also enlarge the non-terminal mistake tail.
Structural tail error, now repeated: Piastri ran P2/P3 until a Turn-1 lock-up after the second restart, rejoined P15 and finished P14. Forecast P14 was 0.263% and P9-P15 only 2.22%. Add explicit spin/run-off/wing/contact states at starts/restarts that put meaningful mass on classified P8-P15. Add topology checks: Antonelli from P16 had 83.5% top-eight probability but only 1.38% P9-P15; his P5 does not validate this discontinuity. Winner-market calibration must not substitute for auditing the full positional shape.
Randomness/unresolved: seven retired versus 4.84 expected; independent marginals gave ~18.5% for seven or more, so do not raise the blanket DNF rate. Albon’s crash and Colapinto’s restart error causing the correlated COL/GAS/NOR DNF were ordinary incident draws; GAS P7/COL P10 beforehand supported pace calls. The Lap-31 pit-window Safety Car and a first restart that immediately caused a second Safety Car support phase-specific neutralisation/restart correlation, not a global chaos increase. Lindblad’s P7 from P15 followed heavy attrition, Piastri’s error and contact with Lawson; he admitted fortune and weak pace, so do not promote Racing Bulls from the finishing place alone. Both Astons had mechanical DNFs (Stroll water pressure; Alonso separate issue), consistent with high priors, but Cadillac was not causally validated (Perez finished; Bottas crashed) and Bearman finished P9. Separate mechanical, collision and classified-retirement hazards; update only with matching causes.
Model setup & usage Model version openai/gpt-5.6-sol-20260709 Reasoning max Runtime 15.3m Model API cost $3.63 Model calls 63 Provider OpenAI Forecast started 26 Sept 2026, 07:52 UTC Forecast saved 26 Sept 2026, 08:07 UTC Race started 26 Sept 2026, 11:00 UTC Run limits: 120 minutes · $30 model spend · 4 CPUs