Regulation M-CDoubles ladderSep 16–28, 2026

Bring Four.

How good is the lead calculator?

Tested on 71,572 ladder games it never saw, it names the opponent’s exact lead 22.3% of the time. Guessing their most common lead gets 12.9%. Here is what that means, and what it doesn’t.

Snapshot Sep 16–28, 2026 · 280,703 games · Backtest: tables built from games through Sep 24, tested on held-out games from Sep 25 to Sep 28 (the last day partial). These numbers are frozen as of publication. See today’s numbers: Calculator or Backtest method.

The calculator takes both team previews and makes two kinds of guesses. It predicts what your opponent will bring and lead. It also ranks your own leads and fours by how similar choices have done on the ladder. This note checks both against games the calculator had never seen.

How it was tested

A fair test has to hide the answers. So the test was split by date:

None of the test games touched the tables or the weights. A model scored on its own training data always looks better than it is.

Guessing their lead

There are 15 possible lead pairs from a team of six. A blind guess is right 1 time in 15, or 6.7%. A smarter baseline picks whichever of their possible pairs has led most often on the ladder. The calculator does better than both.

PredictionCalculatorBaseline
Their lead pair, top guess22.3%CI 22.1–22.5 · n 143k12.9%Most common pairCI 12.8–13.1 · n 143k
Their lead pair, in top three47.3%CI 47.1–47.6 · n 143k32.2%Most common pairCI 32.0–32.5 · n 143k
Their brought mons in predicted four74.6%CI 74.3–74.8 · n 143k70.4%Top four by bring rateCI 70.2–70.7 · n 143k
Their exact four (full-reveal sides)17.8%CI 17.6–18.1 · n 90k9.3%Top four by bring rateCI 9.1–9.5 · n 90k
Held-out games, Sep 25–28. n is sides: 143,144 for the first three rows, 89,686 for the last. Recall is over revealed mons; its interval is computed over sides, which is wider than it needs to be.

Its single top guess was right 22.3% (95% CI 22.1–22.5, n 143,144), against 12.9% (95% CI 12.8–13.1, n 143,144) for the baseline. The real lead was in its top three 47.3% (95% CI 47.1–47.6, n 143,144), against 32.2% (95% CI 32.0–32.5, n 143,144).

Put plainly: it is wrong about the exact lead most of the time. Leads depend on the matchup and the player, and a species-level model can’t see sets, items or habits. What it does well is narrow 15 options to a short list.

Guessing their four

Of the Pokémon each opponent actually showed, 74.6% (95% CI 74.3–74.8, n 143,144) were in the calculator’s predicted four. Picking the four species with the highest bring rates gets 70.4% (95% CI 70.2–70.7, n 143,144). The gap is smaller, since bring rates alone say a lot.

The exact-four number is stricter. In games where the opponent showed all four, the calculator named the whole four 17.8% (95% CI 17.6–18.1, n 89,686), about twice the baseline’s 9.3% (95% CI 9.1–9.5, n 89,686).

Recall slightly understates itself: a Pokémon brought but never sent out counts as “not brought”. Full-reveal games avoid that and give almost the same recall.

Are the win-rate estimates honest?

The calculator also puts a number on each of your options, such as “this lead has won about 54%”. A number like that is only useful if it means what it says. So the test sides were grouped by the calculator’s estimate and compared with what they actually won.

  1. Est. 35.136.7+1.6 pp vs estimate · CI 34.6–38.9 · n 1,961
  2. Est. 37.137.6+0.5 pp vs estimate · CI 36.0–39.2 · n 3,411
  3. Est. 39.139.7+0.6 pp vs estimate · CI 38.5–41.0 · n 5,841
  4. Est. 41.141.1+0.0 pp vs estimate · CI 40.1–42.2 · n 8,436
  5. Est. 43.043.4+0.3 pp vs estimate · CI 42.5–44.3 · n 11,898
  6. Est. 45.046.4+1.4 pp vs estimate · CI 45.6–47.2 · n 14,855
  7. Est. 47.047.9+0.9 pp vs estimate · CI 47.2–48.7 · n 16,460
  8. Est. 49.050.1+1.1 pp vs estimate · CI 49.3–50.8 · n 17,379
  9. Est. 51.052.1+1.1 pp vs estimate · CI 51.3–52.8 · n 16,771
  10. Est. 53.055.3+2.3 pp vs estimate · CI 54.4–56.1 · n 14,383
  11. Est. 55.057.3+2.3 pp vs estimate · CI 56.4–58.2 · n 11,368
  12. Est. 56.959.3+2.4 pp vs estimate · CI 58.2–60.4 · n 8,023
  13. Est. 58.960.1+1.2 pp vs estimate · CI 58.7–61.4 · n 5,125
  14. Est. 60.963.2+2.3 pp vs estimate · CI 61.4–65.0 · n 2,862
  15. Est. 62.963.9+1.0 pp vs estimate · CI 61.3–66.4 · n 1,366
Estimate Observed Calibration of the lead win-rate estimate: 2 pp bins with 1,000+ held-out sides. Hollow: the calculator’s estimate. Filled: what those sides actually won, with its 95% CI.

The estimates track reality closely. In every 2-point bin with at least 1,000 test sides, the observed win rate is within about 3 points of the estimate. Above 50% it is usually 1 to 2 points higher than the estimate. So when the calculator is wrong, it is usually too cautious, not too optimistic. That is by design: thin data is pulled toward the average, which trims extreme estimates.

Do its picks win?

This is the question the data can answer least well.

  1. Led the top pick55.654.8–56.4 · n 15,275
  2. Led a top-3 pick54.353.8–54.8 · n 39,441
  3. Led anything else49.349.1–49.6 · n 127,557
  4. Led a pair ranked 9–1546.546.1–46.9 · n 54,534
Win rate of held-out sides by where their actual lead ranked in the calculator’s list. Observational, not causal. Bar: 95% CI.

Sides that happened to lead the calculator’s top pick won 55.6% (95% CI 54.8–56.4, n 15,275). Every other lead won 49.3% (95% CI 49.1–49.6, n 127,557). Sides whose lead was ranked 9th to 15th won 46.5% (95% CI 46.1–46.9, n 54,534). For fours, sides that brought the top pick won 54.8% (95% CI 53.9–55.8, n 9,977), against 49.4% (95% CI 49.0–49.7, n 79,709) for any other four.

That looks like a 6-point edge. It is not evidence that following the calculator adds 6 points. The calculator was not public when these games were played. These are players who chose, on their own, what the calculator would later have picked. Players who choose what the data favours are probably better players to begin with. The estimates are also partly built from the win rates of similar players in earlier games. So this check shows that the ranking agrees with what wins. It does not show that switching to it causes wins. A real test would need players randomly assigned to follow it, which a ladder dataset can’t provide.

What it can’t do

For scale, the mimikyu model for the previous regulation, M-B, reported 23.3% lead top-1 and 66.1% bring recall, on different data and a different split. Compare loosely.

The fair summary: a useful short list and an honest estimate, not an oracle. Use it to narrow your options, then play the game in front of you.

Every number here comes from the snapshot above, frozen in this file. Win rates carry 95% Wilson intervals, which assume independent games; see player clustering.