CompStrength blends recent patch-weighted pro play with solo queue performance, plus champion-pair synergy and lane matchup history, into a logistic model of blue-side win probability — and, when you optionally select the two teams, adds team Elo, per-player Elo, and player-champion comfort from the inferred starting fives. The numbers below come from holding out games the model never trained on.
Current season (16.x patches, 7,642 held-out games)
With teams: 65.9% (log-loss 0.620) · Draft only: 55.5% (log-loss 0.685) · baseline 53.7%
This slice matches what predictions face today: current-season games, with all prior history available for training. The headline metrics above average over the full multi-season walk-forward (including early folds where the model had little history), so they run lower.
With teams vs. draft only (same held-out games)
| Inputs | Accuracy | Log loss | Brier |
|---|---|---|---|
| Teams + draft | 64.8% | 0.631 | 0.220 |
| Draft only | 55% | 0.687 | 0.247 |
Team strength carries most of the predictable signal in pro play; the draft refines it. “Teams + draft” uses three history signals that all ride on the team selection: team Elo over the game history, per-player Elo of the five starters (tracks roster moves the team rating smooths over), and each player's record on the champion they're drafting (comfort picks). Select both teams on the draft builder to get the top row's model — the starting five is pre-filled from each team's most recent game and you can edit any seat to swap in a substitute.
Team Elo scores how decisively a game was won, not just who won it. Classic Elo moves ratings by K × (result − expected), which treats a 12-minute demolition and a 45-minute base race as the same event — throwing away the most informative thing about a pro game. Each win now moves ratings in proportion to how decisive it was, measured by the winner's gold lead as a fraction of the loser's gold (relative, so it compares fairly across game lengths), with diminishing returns on huge margins and a damping term so favourites who were *expected* to stomp don't run away. Held out on the walk-forward backtest this lifted current-season accuracy 65.7% → 66.2% and log-loss 0.622 → 0.618. The margin only ever affects the post-game update — every prediction still uses strictly pre-game ratings.
A stronger-looking version of this was rejected. Replacing the win/loss outcome entirely with a continuous dominance score scored better on the pooled metric (66.4%/0.613) — but it compressed the rating spread by two-thirds, which flattens Elo's expected-score term and quietly removes the opponent-strength correction. The ratings degenerated into “average gold margin against whoever you happened to play”: minor-league and academy sides floated to the top of the table (majors held 1 of the top 10 slots, down from 6), and accuracy on cross-region games actually dropped61.9% → 60.7%. The pooled number hid it because international games are a small slice of the total. The shipped version moves the ratings the other way, making them more separated than before (spread +10%).
Aren't multi-season team ratings stale, since rosters change? We tested exactly that. Rebuilding team strength from the current season ONLY (resetting every rating at the year boundary) scores 62.9% — 2.3 points WORSE than keeping prior-season history. The reason is subtle: teams churn, but players don't. When both team and player ratings are reset, accuracy craters; keep the players' individual histories and it recovers almost entirely. Last season stays informative because the same people are still playing — which is the whole point of the per-player Elo feature. A mild discount on old team ratings is optimal (the shipped carryover keeps 70% of a team's prior-season deviation), but throwing last season away is strictly worse.
Fold 1/8 skipped: insufficient train/test data (train_games=0, test_games=2210). Walk-forward validation on real Oracle's Elixir pro-match data (15,473 held-out games; newest patch 16.x = 2026 season). Each fold refits the entire pipeline using only games strictly before that fold, so there is no lookahead leakage. Draft-only prediction (champions picked, no in-game state) is a genuinely weak signal at the pro level -- treat accuracy a few points above the pick-majority baseline as expected, not a bug.
It's tempting to think a hard counter-pick (Renekton into Shen, Ashe into Teemo) should let a draft-only model predict much better than ~55%. We tested that directly, with a leak-free walk-forward harness (each model trained only on games strictly before the games it's scored on), and none of these beat the shipped model on both accuracy and calibration:
The draft-only mode now has its own dedicated fit. Previously the no-teams prediction reused the full model's coefficients with the team terms zeroed — but those weights were fit with team features present, and without them the pairwise-synergy term (which is estimated in-sample) gets over-trusted. Measured on the walk-forward harness, adding synergy to a draft-only fit makes it worse than champion strength alone (55.7% vs 56.0% current-season), while the lane-matchup term helps. The shipped draft-only fit therefore uses champion strength + lane matchups with synergy excluded: 56.2% vs 56.0% current-season and 55.1% vs 54.6% overall against the old zeroed-coefficients behaviour, with better calibration on both slices. (The draft builder's synergy-pick suggestion still ranks through the full model's synergy weight — it's a relative comparison among candidates, not a calibrated probability.) We also tested a champion “dominance” rating — margin-of-victory averaged per champion, the champion-level analogue of the team-Elo margin weighting. Rejected: log-loss identical to four decimals in every configuration; dominance alone predicts 55.2%, nearly as well as win rates do, confirming the two measure the same underlying champion strength.
The honest conclusion: at the pro level, which champions get drafted is a genuinely weak predictor once you can't see who's piloting them. That's exactly why the “teams + players” model jumps ~10 points to ~66% — the predictable signal lives in team and player identity, not the champion select screen. Disaster drafts are real but rare (a handful per season), so even predicting every one perfectly moves aggregate accuracy under a point.
Should it train on the current season only? We measured that too. Training on this season alone (walk-forward within it) scores 63.7% with teams, ~2 points BELOW the shipped model that also uses last season as history — the early-season games starve without it. More recent-weighted data doesn't beat the current recency decay; less data just loses signal.
International events (MSI / Worlds / EWC) are the hardest. Held-out accuracy on cross-region events runs well below regional play — ~57% aggregate (EWC 66%, Worlds 58%, MSI ~52% on a tiny 80-game sample) vs 65% inside a single league. The cause is structural: team Elo only bridges regions through the handful of inter-region games, so when the best of two regions meet the rating gap is genuinely uncertain. We help it where we can — inter-region games (an international league, or any game where the two teams' home leagues differ) move Elo 3× as much, since they're the only games that calibrate strength across regions. That lifted held-out international accuracy 55.7% → 56.5% and improved its calibration (log-loss 0.715 → 0.714) with no cost to regional games — a real but modest gain; cross-region prediction stays fundamentally data-starved. The draft builder flags these matchups so you can weight them accordingly.
What about wombo combos, all-AD/all-AP teams, or a low-damage/all-squishy comp? This is a real gap: the model's only inter-champion signal is pairwise synergy (exactly 2 champions at a time, shrunk toward 0) — there is no feature that looks at the 5-champion whole. We built the best measurable proxies available and tested each: team damage output and tankiness (from real per-game damage/tanking stats), damage-share concentration (one carry vs a balanced comp), and AD/AP balance and engage-champion count (from champion attributes, since Oracle's Elixir has no crowd-control-time column to measure a true “wombo” signal from). Draft-only, one candidate (AD/AP lean) showed a small, consistent gain across every test fold. But once combined with team and player Elo — the model actually used once you pick teams — that gain vanished into noise (currentSeason 65.71% → 65.76%, well within run-to-run variance) while overall accuracy and log-loss both got marginally worse. Team/player identity is simply too dominant a signal for these subtler structural effects to move the needle on real pro outcomes. None of these features ship.
Why not just use Riot's official Global Power Rankings? We tested it properly rather than assuming. The GPR history is real and, importantly, not retroactively rewritten — we verified 7,032 overlapping historical observations fetched through different query windows and found zero conflicts, so using it wouldn't leak future results into past predictions. Added as a feature it changed held-out log-loss from 0.6125 to 0.6125: nothing. The reason is that GPR is itself an Elo-style rating computed from the same match results our own Elo already consumes — and ours covers all 612 teams in the data, where GPR ranks only ~58 tier-1 teams (about a quarter of games, with no academy or regional sides at all). It is a less complete version of a signal the model already has.
What about form, streaks, rest days, or head-to-head records? Also tested, also negative. Recent form correlates 0.75with team Elo — it is a noisier restatement of what the rating already encodes. Win/loss streaks actively hurt (−0.5 points of current-season accuracy). Days of rest and games-in-the-last-week have essentially no relationship with the outcome at all (correlation ~0.015). Head-to-head history, series game number, playoffs-vs-regular-season, and roster changes all landed inside the measurement noise. The honest summary: once you know who is playing and how dominant they have been, schedule and momentum narratives add nothing measurable.
Every prediction shows fair decimal odds (1 ÷ probability) for each side. If a bookmaker offers longer odds than the fair number on a side, the model sees positive expected value on that side — before costs.
Be honest about the bar: a bookmaker's implied probabilities contain a built-in margin (typically ~4–7% across both sides), and closing lines on major leagues are sharp. To profit you need the model's calibration edge to exceed that margin consistently — check the log-loss and calibration table above, not just accuracy, and treat small-sample leagues (see the breakdown below) with extra skepticism. One structural caveat: team Elo is anchored across leagues only by the few inter-league games (MSI, Worlds, EWC), so ratings are most trustworthy for matchups WITHIN a league or at international events — an isolated league's ratings can drift high or low as a block. Nothing on this page is betting advice; it's a measured, walk-forward-validated probability estimate with known error bars.
17,683real professional games, drawn from the leagues and patches below. More recent patches are weighted exponentially more (the “weight” column is each patch's share of full weight); older games still contribute, just less. On top of that, premier leagues (LCK + LPL) are up-weighted to carry ~70% of the total training weight, and international events (MSI/Worlds/EWC) get their own boost — so the champion statistics reflect the highest level of play even though minor leagues supply more raw games.
| Patch | Games | Weight |
|---|---|---|
| 16.16 | 92 | 100% |
| 16.15 | 691 | 50% |
| 16.14 | 546 | 25% |
| 16.13 | 377 | 13% |
| 16.12 | 27 | 6% |
| 16.11 | 383 | 3% |
| 16.10 | 795 | 2% |
| 16.09 | 747 | 1% |
| 16.08 | 700 | 0% |
| 16.07 | 730 | 0% |
| 16.06 | 326 | 0% |
| 16.05 | 306 | 0% |
| 16.04 | 320 | 0% |
| 16.03 | 623 | 0% |
| 16.02 | 546 | 0% |
| 16.01 | 433 | 0% |
| 15.24 | 78 | 0% |
| 15.23 | 56 | 0% |
| 15.22 | 19 | 0% |
| 15.21 | 60 | 0% |
| 15.20 | 192 | 0% |
| 15.19 | 378 | 0% |
| 15.18 | 180 | 0% |
| 15.17 | 622 | 0% |
| 15.16 | 613 | 0% |
| 15.15 | 637 | 0% |
| 15.14 | 598 | 0% |
| 15.13 | 385 | 0% |
| 15.12 | 166 | 0% |
| 15.11 | 445 | 0% |
| 15.10 | 652 | 0% |
| 15.09 | 875 | 0% |
| 15.08 | 796 | 0% |
| 15.07 | 820 | 0% |
| 15.06 | 464 | 0% |
| 15.05 | 214 | 0% |
| 15.04 | 371 | 0% |
| 15.03 | 569 | 0% |
| 15.02 | 536 | 0% |
| 15.01 | 315 | 0% |
| League | Games |
|---|---|
| LPL | 1,382 |
| LCK | 977 |
| LCKC | 972 |
| LJL | 729 |
| EM | 723 |
| LAS | 716 |
| LEC | 622 |
| AL | 609 |
| PRM | 566 |
| LFL | 555 |
| LCP | 552 |
| CD | 526 |
+ 40 more leagues
The same walk-forward held-out predictions, split by patch and by league. “Edge” is accuracy minus that segment's own pick-majority baseline — a positive edge means the model beat simply guessing the more common outcome there. Per-segment numbers are noisier the fewer games the segment has.
| Patch | Games | Accuracy | Baseline | Edge | Log loss |
|---|---|---|---|---|---|
| 16.16 | 92 | 59.8% | 68.5% | -8.7pp | 0.642 |
| 16.15 | 691 | 63.7% | 57.5% | +6.2pp | 0.620 |
| 16.14 | 546 | 67.2% | 53.8% | +13.4pp | 0.622 |
| 16.13 | 377 | 66.8% | 56.5% | +10.3pp | 0.625 |
| 16.11 | 383 | 62.1% | 57.4% | +4.7pp | 0.645 |
| 16.10 | 795 | 65.3% | 51.2% | +14.1pp | 0.621 |
| 16.09 | 747 | 67.5% | 53% | +14.5pp | 0.605 |
| 16.08 | 700 | 70.3% | 53.3% | +17.0pp | 0.589 |
| 16.07 | 730 | 68.2% | 54.7% | +13.6pp | 0.604 |
| 16.06 | 326 | 72.1% | 55.8% | +16.3pp | 0.581 |
| 16.05 | 306 | 65% | 54.9% | +10.1pp | 0.626 |
| 16.04 | 320 | 62.8% | 51.9% | +10.9pp | 0.631 |
| 16.03 | 623 | 65% | 53.5% | +11.6pp | 0.632 |
| 16.02 | 546 | 63% | 51.8% | +11.2pp | 0.628 |
| 16.01 | 433 | 63.3% | 51.7% | +11.5pp | 0.666 |
| 15.24 | 78 | 39.7% | 55.1% | -15.4pp | 0.847 |
| 15.23 | 56 | 66.1% | 55.4% | +10.7pp | 0.711 |
| 15.21 | 60 | 73.3% | 63.3% | +10.0pp | 0.570 |
| 15.20 | 192 | 59.9% | 50% | +9.9pp | 0.632 |
| 15.19 | 378 | 61.6% | 55.6% | +6.1pp | 0.674 |
| 15.18 | 180 | 58.9% | 50.6% | +8.3pp | 0.657 |
| 15.17 | 622 | 57.9% | 53.4% | +4.5pp | 0.686 |
| 15.16 | 613 | 65.7% | 52.4% | +13.4pp | 0.625 |
| 15.15 | 637 | 63.7% | 52.7% | +11.0pp | 0.619 |
| 15.14 | 598 | 66.6% | 52% | +14.5pp | 0.630 |
| 15.13 | 385 | 61% | 54.3% | +6.8pp | 0.668 |
| 15.12 | 166 | 60.8% | 53.6% | +7.2pp | 0.734 |
| 15.11 | 445 | 59.8% | 51.7% | +8.1pp | 0.692 |
| 15.10 | 652 | 63% | 55.7% | +7.4pp | 0.630 |
| 15.09 | 875 | 65.8% | 52.2% | +13.6pp | 0.624 |
| 15.08 | 796 | 68% | 55.3% | +12.7pp | 0.603 |
| 15.07 | 820 | 65.6% | 53.5% | +12.1pp | 0.620 |
| 15.06 | 255 | 66.3% | 54.5% | +11.8pp | 0.632 |
| Other (3 smaller) | 50 | 58% | 50% | +8.0pp | 0.670 |
| League | Games | Accuracy | Baseline | Edge | Log loss |
|---|---|---|---|---|---|
| LPL | 1,202 | 62.5% | 52.8% | +9.7pp | 0.667 |
| LCK | 868 | 67.2% | 52.5% | +14.6pp | 0.623 |
| LCKC | 865 | 60% | 52% | +8.0pp | 0.677 |
| LAS | 682 | 69.6% | 52.5% | +17.2pp | 0.569 |
| EM | 649 | 63.2% | 56.2% | +6.9pp | 0.667 |
| LJL | 597 | 73.4% | 50.8% | +22.6pp | 0.541 |
| LEC | 531 | 65% | 56.7% | +8.3pp | 0.636 |
| AL | 516 | 66.1% | 54.8% | +11.2pp | 0.604 |
| NACL | 496 | 63.3% | 55.2% | +8.1pp | 0.650 |
| LFL | 492 | 63.8% | 54.9% | +8.9pp | 0.660 |
| PRM | 487 | 63% | 57.9% | +5.1pp | 0.649 |
| CD | 466 | 58.6% | 50.6% | +7.9pp | 0.672 |
| LCP | 463 | 65.7% | 54.4% | +11.2pp | 0.608 |
| ROL | 377 | 68.2% | 55.4% | +12.7pp | 0.593 |
| HLL | 352 | 72.2% | 54% | +18.2pp | 0.562 |
| NLC | 349 | 70.5% | 54.7% | +15.8pp | 0.575 |
| TCL | 328 | 65.9% | 56.4% | +9.5pp | 0.622 |
| EBL | 323 | 66.9% | 53.3% | +13.6pp | 0.565 |
| LIT | 319 | 62.1% | 50.2% | +11.9pp | 0.637 |
| RL | 313 | 63.9% | 50.8% | +13.1pp | 0.633 |
| LRN | 304 | 68.1% | 50% | +18.1pp | 0.595 |
| LRS | 303 | 59.1% | 53.5% | +5.6pp | 0.659 |
| HC | 299 | 63.5% | 56.9% | +6.7pp | 0.644 |
| HM | 295 | 70.5% | 52.5% | +18.0pp | 0.542 |
| PCS | 276 | 71% | 51.4% | +19.6pp | 0.589 |
| EWC | 265 | 63.8% | 54% | +9.8pp | 0.645 |
| Asia Master | 240 | 54.6% | 54.6% | +0.0pp | 0.776 |
| LPLOL | 238 | 65.1% | 52.1% | +13.0pp | 0.624 |
| CBLOL | 215 | 64.2% | 58.6% | +5.6pp | 0.642 |
| VCS | 203 | 62.1% | 52.7% | +9.4pp | 0.638 |
| LCS | 195 | 66.2% | 55.4% | +10.8pp | 0.598 |
| LTA S | 192 | 55.7% | 51.6% | +4.2pp | 0.691 |
| LTA N | 188 | 62.8% | 52.7% | +10.1pp | 0.641 |
| LVP SL | 183 | 63.9% | 61.2% | +2.7pp | 0.654 |
| LES | 170 | 77.6% | 50.6% | +27.1pp | 0.496 |
| LFL2 | 157 | 59.9% | 55.4% | +4.5pp | 0.671 |
| MSI | 151 | 60.3% | 58.9% | +1.3pp | 0.663 |
| NL | 116 | 81.9% | 54.3% | +27.6pp | 0.563 |
| HW | 103 | 67% | 58.3% | +8.7pp | 0.640 |
| WLDs | 96 | 57.3% | 51% | +6.3pp | 0.671 |
| NEXO | 95 | 63.2% | 55.8% | +7.4pp | 0.616 |
| DCup | 78 | 39.7% | 55.1% | -15.4pp | 0.847 |
| PRMP | 73 | 54.8% | 53.4% | +1.4pp | 0.637 |
| KeSPA Cup | 61 | 47.5% | 60.7% | -13.1pp | 0.866 |
| KeSPA | 56 | 66.1% | 55.4% | +10.7pp | 0.711 |
| CT | 50 | 66% | 56% | +10.0pp | 0.651 |
| FST | 45 | 68.9% | 66.7% | +2.2pp | 0.601 |
| ASI | 42 | 57.1% | 59.5% | -2.4pp | 0.737 |
| CCWS | 40 | 57.5% | 60% | -2.5pp | 0.688 |
| Other (3 smaller) | 69 | 68.1% | 56.5% | +11.6pp | 0.641 |
For predictions bucketed by predicted win probability, how often did the predicted side actually win?
| Predicted bucket | Predicted mean | Actual win rate | Count |
|---|---|---|---|
| 0.0-0.2 | 13% | 20.6% | 889 |
| 0.2-0.4 | 31.5% | 34.1% | 2925 |
| 0.4-0.6 | 50.5% | 49.3% | 5300 |
| 0.6-0.8 | 69% | 66.3% | 4704 |
| 0.8-1.0 | 86.9% | 82.4% | 1655 |