How well can AI predict Formula 1?

AI models research each race and predict every driver’s finish. We compare their predictions with the real results, plus what each forecast cost and how long it took.

Explain like I’m lost

Each AI predicts every driver’s chances before the race. A 70% chance of winning still leaves a 30% chance of losing.

RPS is the error score: lower is better. Zero is perfect; 0.12 does not mean 12% accuracy. Cost and time show the average dollars and minutes used to make one race forecast.

Compare models on the same races. A small lead over just a few races is an early result, not proof that one model is always better.

Model standings

3 models · Early results · 2 races compared

Lower error means better predictions. We use RPS: 0 is perfect; 0.12 does not mean 12% accuracy. Too few races to establish a reliable winner.

#ModelScored
01Grok 4.60.1247$1.7013.3m$4.472/2
02Muse Spark 1.30.1251$2.0110.5m$6.532/2
03GPT-5.6 Sol0.1330$3.1617.8m$9.452/2
—Grid baseline0.1932———2/2

Scroll for spend and races scored →

Averages use the same 2 races. Model API costs exclude historical search and extraction fees. Recorded spend includes saved forecast and review charges.

What’s included

Forecast averages use shared scored races; total spend includes all season races. Research fees on new runs are list-price estimates (≈). Historical research fees and some failed-attempt charges are unavailable; recorded spend is not a complete bill. The cost chart uses model API costs only so historical runs remain comparable. RPS is an error measure, so spend is shown alongside the score rather than divided by it.

Spend breakdown
ModelForecastsReviewsRecorded
Grok 4.6$3.41$1.07$4.47
Muse Spark 1.3$4.02$2.51$6.53
GPT-5.6 Sol$6.33$3.12$9.45

Excludes research fees and failed-attempt charges without saved billing records.

What is the grid baseline?

Predict every driver finishes in the official starting order, with 100% certainty and no retirements. Score it with the same RPS rule. This is a simple reference, not a competitive forecasting system.

Added retrospectively; the rule has no fitted parameters. Grids were retrieved on 26 September 2026, after these races. Late grid changes may differ from what was available when an AI ran. Pit-lane starters use their listed place in the official order.

Open a race for its grid source. An average appears only when every compared race has a verified grid.

1 · Grok 4.62 · Muse Spark 1.33 · GPT-5.6 Sol

Forecast error vs. model cost & time

Lower left means less error for less money or time.

0.1210.1290.137$0.00$1.37$2.74$4.11Error (RPS) ↓GrokGrok 4.6: 0.1247 RPS, $1.70MuseMuse Spark 1.3: 0.1251 RPS, $2.01GPT-5.6GPT-5.6 Sol: 0.1330 RPS, $3.16

Forecast error by race

Race RPS · lower is better

0.0710.1280.185Error (RPS) ↓GPT-5.6 Sol · R14: 0.0959GPT-5.6 Sol · R15: 0.1701Muse Spark 1.3 · R14: 0.0911Muse Spark 1.3 · R15: 0.1591Grok 4.6 · R14: 0.0856Grok 4.6 · R15: 0.1638GPT-5.6GrokMuseR14R15

Compare a race

Round / raceGPT-5.6 SolMuse Spark 1.3Grok 4.6Status
15Baku0.17010.15910.1638reviewed
14MADRING0.09590.09110.0856reviewed

Forecast error (RPS) · lower is better. Open a race to compare predictions with results.

How it works

Predict before the race. Score after it. Review mistakes for the next one.

Same race task and limits for every model. Web search, code execution, and each model’s own notes from earlier races. Up to $30 in model calls and 120 minutes per run. This tests a model using tools and memory, not an unaided answer.

What are the models predicting?

Each model forecasts the entire Formula 1 field using public information available during its run. It decides how to research, reason, write code, and simulate. The target is the final race classification, including non-classification, non-starts, disqualifications, and retirement risk.

The benchmark evaluates the submitted probabilities. The written order, explanation, and research findings provide context; they do not earn extra points. Research quality and causal claims are examined in the post-race review, not assigned a separate numerical score.

Model setup, tools & budgets

The current cohort contains GPT-5.6 Sol, Muse Spark 1.3, and Grok 4.6. Each entrant keeps a pinned model version and a restricted provider policy. Provider fallback cannot switch the model. Reasoning uses the highest setting supported when the entrant was created.

120 minmaximum run time$30model-spend limit per run4 CPUsper isolated container

Agents have web search, page extraction, a shell, and a disposable workspace. They may build their own models, but no simulation method is prescribed. The budget is a ceiling, not a target. Models run independently; they do not see competing forecasts.

The same limits apply to each post-race review. Forecast cost includes model API calls, search, and extraction. Total spend also includes reviews and failed or voided attempts across all season races. Forecast averages use the same shared races as the score. Where research fees are unavailable, comparisons use model API costs for every entrant. Recorded spend includes only saved charges and is not a complete bill. New research costs are estimates from recorded tool requests at published Parallel rates, before credits or discounts. Historical research fees are unavailable. Hosting is excluded.

Can models see the results? How do they learn?

A forecast starts with the race entry list, the model’s own previous forecasts and research, results from earlier races, scores, review snapshots, and its current notes. The normal workflow runs after qualifying. Agents independently verify current evidence, so research sources and completion times can differ.

New runs must start and finish before the configured race start. A race with a saved result cannot be forecast again through the live workflow. Once completed, a prediction stays frozen.

After the official result is supplied, a separate review compares the prediction with what happened and revisits its assumptions. The model edits its own notes for the next race. The workflow checks that earlier reviews are complete. This is learning through context and notes, not weight training; scratch code does not carry forward.

Archived runs retain their original outputs. Where an older run has no explicit finishing-order report, the race table clearly labels its derived average-rank ordering.

What does the error score mean?

Ranked Probability Score (RPS) measures how far the predicted chances were from what happened. Lower is better; zero is perfect. For each driver, the scorer compares the probability of finishing at or above every position with the official result. Errors farther from the outcome affect more thresholds. This rewards accurate distributions and penalizes misplaced confidence.

RPS = (1 / N²) × Σ drivers Σ positions (F(p) − O(p))²

N is the field size. F(p) is the forecast probability of finishing P1 through Pp; O(p) is 1 if the driver actually finished at or above that position, otherwise 0. The reported race score averages across all drivers and all N position thresholds. Lower is better; a perfect distribution scores zero.

NC, DNS and DSQ sit beyond the last numbered position for RPS, so it does not distinguish those three outcomes. The log-loss metric evaluates their separate probabilities. A late retirement can still have a numbered classification; retirement is scored separately.

Other scores & displayed rankings
MeasureWhat it measures
Log lossProbability assigned to the driver’s exact result. Zero probability is floored at 10⁻¹⁵ for calculation.
Brier scoresSquared probability error for winning, podium, points, and retirement.
Expected-rank errorMean absolute error between each driver’s average forecast rank and their actual rank.
Average rankThe probability-weighted rank across all outcomes. NC, DNS and DSQ count as N + 1. It is not the most likely individual finishing position.
Podium chanceProbability of P1, P2, or P3. It says nothing about the size of the downside outside the podium.

All numerical error scores are lower-is-better. Sorting the driver table by podium chance answers a different question from sorting by average rank. Neither changes the original probabilities.

Season standings & progression

Each shared, completed race contributes equally to the season score: the arithmetic mean of its race RPS values. Only races with valid scores for every model in the comparison enter the aggregate. A missing forecast never counts as zero.

Coverage shows a model’s scored forecasts out of races with at least one score. Partially scored races remain available in the race details, but are excluded from the shared-race average until all models have scores.

The cumulative chart shows the average through each included round. The per-race chart shows the individual race score. Changes reflect both forecasting performance and differences in circuits and race outcomes; a trend alone does not isolate the effect of memory or learning.

Results, failed runs & reproducibility

Official classifications are supplied manually using the final FIA result. The scorer validates every driver and the classification structure. A changed result after a completed review is flagged before the next forecast.

Completed attempts are reused, never replaced. Recorded failed attempts can be retried explicitly before the race deadline; the failed files and traces are archived. Each retry has a fresh budget. Running or incomplete attempts are not automatically restarted.

Original predictions, reports, explanations, run metadata, full trajectories, and notes snapshots are retained locally. This site is generated from those saved files. Schema validation checks probability consistency and required report fields; it does not independently prove that a cited source supports a claim.