Introduction: What the Findings Mean
"He always lights up this team" is one of the most common reasons given for a player prop. It is also hard to test by eye: a quarterback sees a non-division opponent about once every three or four years, and twenty plate appearances against one pitcher carry about ±.100 of noise in batting average. A run of good games against one team is exactly what chance produces somewhere in a league of hundreds of players.
We asked a narrow question. Suppose a forecast already knows the player's recent and career form, the strength of tonight's opposing defense, and whether he is home or away. Does adding his own history against this particular opponent, or his own home/away split, make the forecast more accurate on games it has not seen? We answered it the same way in three sports, using only information that would have been available before each game.
The answer is no. In most markets the tuning seasons gave the opponent history and the personal venue split zero or very small weight; where they kept some, it did not improve accuracy on the test seasons. The streak check is the clearest. NHL skaters who had out-shot their forecast by 49% against an opponent came in at 1.025 times the forecast the next time they met; those who had fallen 42% short came in at 1.021; everyone with that much history averaged 1.015. In baseball, batters who had hit 62% above expectation against a pitcher came in at 1.053, and those 49% below at 1.025, against 1.003 for all pairs with that much history: a gap of 111 points in the history shrank to 2.8 points. Aaron Judge, against pitchers he had already faced fifteen or more times, had 53 hits in 222 plate appearances in 2023–25, against 54.1 forecast without any matchup history.
What does work is the mechanism behind a matchup, measured on large samples. The home/away effect is real but small for yardage, and our fixed NFL multipliers overstated it. Handedness matters in baseball. And a player's record against the model's own earlier forecasts carries information the hockey model was missing. We show the past meetings in our player snapshots as context, labeled as not a model input.
Data and Evaluation Windows
Each sport uses its own public record. NHL player-games come from the frozen history behind model nhl-v2.3 (official NHL per-game skater and team reports, 2022–23 through 2025–26) [1]. NFL player-games come from nflverse weekly statistics, 2012 through week 4 of 2026, joined to the nflverse schedule for home, away and neutral sites [2]. MLB plate appearances come from Retrosheet event files for 2012–2025 through the Chadwick Bureau mirror [3]; the parser reproduces published season lines exactly (Aaron Judge, 2024: 180 hits, 58 home runs, 171 strikeouts, 142 walks and hit-by-pitches). Regular seasons only.
| Sport | Source | Unit | History from | Tuning | Test | Test rows |
|---|---|---|---|---|---|---|
| NHL | nhl-v2.3 archive (official NHL game logs) | player-game | 2022–23 | 2024–25 | 2025–26 | 47,230 |
| NFL | nflverse weekly player stats | player-game | 2012 | 2022–2023 | 2024 to 2026 wk 4 | 14,793 (7 markets) |
| MLB | Retrosheet event files | plate appearance | 2012 | 2021–2022 | 2023–2025 | 549,195 |
Every history is point in time. A game counts toward a player's history only on later dates: the next day for NHL and MLB, later weeks for NFL. Tuning and test seasons never overlap.
Method: Shrunk Histories over a Fixed Forecast
Baseline forecasts
The NHL baseline is the production player model nhl-v2.3 (projected ice time × recency-weighted production per 60 minutes, scaled by the opponent's shots or goals allowed; negative-binomial counts), refitted on seasons before each scored season exactly as its own evaluation does. The NFL baseline rebuilds the production recipe from public data: a four-game recency average blended with a decayed career mean shrunk to the position average, times the production opposing-defense rating and the production home/away multiplier. The MLB baseline is a plate-appearance model built for this study: the batter's and pitcher's decayed, shrunk rates combined by odds ratio against the league rate, times a shrunk park factor and the league home factor.
Adding a history
For a player i and a set of his earlier games S (all of them, those at the same venue, or those against tonight's opponent), let A be his actual output and E the baseline's expected output over S. The history enters as an empirically shrunk ratio [4]:
Here k is a prior strength in the stat's own units (expected shots, games of output, or plate appearances) and m is the value the ratio shrinks toward. For the player's overall ratio m = 1. For a venue split or an opponent split m is the player's overall ratio, so the split only moves the forecast by how much the player differs from himself, not from the league. A history of n units therefore gets weight n ÷ (n + k).
The player's overall ratio is a control. Without it, an opponent split would also pick up any general bias the baseline has for that player, and look useful for the wrong reason.
Choosing k and scoring
For each component, k is chosen from a fixed grid on the tuning seasons by mean log loss (NHL, MLB) or squared error (NFL), in order: the player's overall ratio first, then each split given it. "Not used" means the grid's top value, infinite k, won. The test seasons are then scored once. Differences are paired on the same rows, with 95% intervals from resampling whole games [5].
| Market | Player vs this opponent | Player's own home/away split |
|---|---|---|
| NHL shots | 1000 expected events of prior | not used (zero weight) |
| NHL scoring | not used (zero weight) | 1000 expected events of prior |
| NFL passing yards | not used (zero weight) | 300 games of prior |
| NFL pass attempts | not used (zero weight) | 300 games of prior |
| NFL completions | not used (zero weight) | 100 games of prior |
| NFL rushing yards | 30 games of prior | not used (zero weight) |
| NFL rush attempts | 100 games of prior | not used (zero weight) |
| NFL receiving yards | not used (zero weight) | 300 games of prior |
| NFL receptions | 100 games of prior | not used (zero weight) |
| MLB hit | 1000 PA of prior (batter vs pitcher) | not used (zero weight) |
| MLB strikeout | 300 PA of prior (batter vs pitcher) | not used (zero weight) |
| MLB walk or hbp | 300 PA of prior (batter vs pitcher) | not used (zero weight) |
| MLB home run | 1000 PA of prior (batter vs pitcher) | not used (zero weight) |
Player Versus Opponent, and Personal Home/Away Splits
Table 3 compares each split with the same forecast without it. Most weights were zero or small: 2% of the evidence after eight NHL meetings, and 3% (hits, home runs) to 9% (strikeouts, walks) after thirty plate appearances against one pitcher. Two NFL cases kept more: personal venue splits in four markets (12% to 29% after forty games) and opponent history for rushing yards (about 21% after eight meetings). None of them improved the test seasons; the rushing-yard opponent history made squared error worse on average, and the batter-versus-pitcher changes are a few hundred-thousandths of a nat per plate appearance.
| Market (metric) | Player vs this opponent | Player's own home/away split |
|---|---|---|
| NHL shots (log loss) | +0.00001 [−0.00003, +0.00005] | +0.00000 [+0.00000, +0.00000] |
| NHL goals (log loss) | +0.00000 [+0.00000, +0.00000] | +0.00002 [−0.00000, +0.00004] |
| NHL assists (log loss) | +0.00000 [+0.00000, +0.00000] | −0.00002 [−0.00005, +0.00001] |
| NHL points (log loss) | +0.00000 [+0.00000, +0.00000] | −0.00000 [−0.00004, +0.00004] |
| NFL passing yards (squared error) | +0.00 [+0.00, +0.00] | +0.28 [−15.83, +13.05] |
| NFL pass attempts (squared error) | +0.00 [+0.00, +0.00] | +0.14 [−0.03, +0.29] |
| NFL completions (squared error) | +0.00 [+0.00, +0.00] | +0.19 [−0.04, +0.39] |
| NFL rushing yards (squared error) | +5.26 [−0.40, +11.11] | +0.00 [+0.00, +0.00] |
| NFL rush attempts (squared error) | −0.01 [−0.05, +0.04] | +0.00 [+0.00, +0.00] |
| NFL receiving yards (squared error) | +0.00 [+0.00, +0.00] | +1.43 [−0.01, +2.68] |
| NFL receptions (squared error) | +0.00 [−0.00, +0.01] | +0.00 [+0.00, +0.00] |
| MLB hit (log loss ×10⁻⁴) | +0.0 [−0.0, +0.1] | +0.0 [+0.0, +0.0] |
| MLB strikeout (log loss ×10⁻⁴) | −0.3 [−0.4, −0.2] | +0.0 [+0.0, +0.0] |
| MLB walk or hbp (log loss ×10⁻⁴) | −0.2 [−0.3, −0.0] | +0.0 [+0.0, +0.0] |
| MLB home run (log loss ×10⁻⁴) | +0.0 [−0.0, +0.1] | +0.0 [+0.0, +0.0] |
Table 4 asks the bettor's question directly. It takes players whose earlier record against an opponent was far above or below the forecast, and checks how they did the next time against that same opponent. The "hot" groups did not keep their edge. In hockey and football they landed within about one point of the reference group of everyone with that much history, and NFL rushing yards reversed (the hot group under-performed and the cold group over-performed). In baseball a small part of the earlier gap carried over, consistent with the few percent of weight batter-versus-pitcher history earned in tuning.
| Market | Group | Next meetings | Earlier ratio | Next-meeting ratio |
|---|---|---|---|---|
| NHL shots | everyone with 4+ prior games vs this opponent (reference) | 34,964 | 0.967 | 1.015 |
| NHL shots | beat expectation by 30%+ vs this opponent (4+ prior games) | 4,929 | 1.492 | 1.025 |
| NHL shots | fell 23%+ short vs this opponent (4+ prior games) | 9,170 | 0.581 | 1.021 |
| NHL points | everyone with 4+ prior games vs this opponent (reference) | 34,964 | 0.954 | 0.994 |
| NHL points | beat expectation by 30%+ vs this opponent (4+ prior games) | 8,443 | 1.715 | 0.989 |
| NHL points | fell 23%+ short vs this opponent (4+ prior games) | 13,492 | 0.410 | 0.986 |
| NFL receiving yards | everyone with 4+ prior games vs this opponent (reference) | 718 | 1.110 | 0.946 |
| NFL receiving yards | beat expectation by 15%+ vs this opponent (4+ prior games) | 306 | 1.394 | 0.953 |
| NFL receiving yards | fell 13%+ short vs this opponent (4+ prior games) | 159 | 0.702 | 0.941 |
| NFL rushing yards | everyone with 4+ prior games vs this opponent (reference) | 223 | 1.083 | 1.014 |
| NFL rushing yards | beat expectation by 15%+ vs this opponent (4+ prior games) | 88 | 1.339 | 0.936 |
| NFL rushing yards | fell 13%+ short vs this opponent (4+ prior games) | 43 | 0.721 | 1.077 |
| NFL passing yards | everyone with 4+ prior games vs this opponent (reference) | 290 | 1.028 | 0.964 |
| NFL passing yards | beat expectation by 15%+ vs this opponent (4+ prior games) | 41 | 1.254 | 0.959 |
| NFL passing yards | fell 13%+ short vs this opponent (4+ prior games) | 26 | 0.814 | 1.038 |
| MLB hits | every pair with 20+ PA (reference) | 11,731 | 1.055 | 1.003 |
| MLB hits | 20+ PA vs this pitcher, hits 40%+ above expectation | 1,925 | 1.622 | 1.053 |
| MLB hits | 20+ PA vs this pitcher, hits 30%+ below expectation | 1,888 | 0.512 | 1.025 |
Personal home/away splits fared no better. Where tuning gave a player's own venue split any weight, it did not improve the test seasons. A player who has looked much better at home is, on the evidence, mostly showing the league-wide home effect plus noise.
What Did Help: Venue, Platoon and Player Bias
League home/away effects of the right size (NFL)
The NFL model had applied fixed home/away multipliers since 2025: ±6% for passing and receiving yards, ±4% for rushing yards, ±3% for receptions. We measured the gap directly: actual over expected output at home and away, each divided by the overall ratio so only the venue gap remains, with neutral sites excluded. On 2024–26 games held out from fitting, gaps fitted on 2012–2021 replaced the fixed values with lower squared error for passing yards (−99.6 [−198.9, −7.8]) and receiving yards (−5.3 [−9.6, −1.2]). Refitted on every completed season, the passing-yard home factor is +1.4%, not +6%.
| Market | Previous fixed | Fitted 2012–2025 | Home 95% interval | Holdout change (fitted − fixed) | Holdout |
|---|---|---|---|---|---|
| Passing yards | +6.0% / −6.0% | +1.4% / −1.4% | +0.6% to +2.2% | −99.56 [−198.88, −7.78] | better |
| Pass attempts | +0.0% / +0.0% | −0.2% / +0.2% | −1.1% to +0.6% | +0.05 [−0.12, +0.22] | no detectable change |
| Completions | +0.0% / +0.0% | +0.8% / −0.8% | −0.1% to +1.6% | −0.04 [−0.11, +0.03] | no detectable change |
| Rushing yards | +4.0% / −4.0% | +2.4% / −2.4% | +0.9% to +3.8% | −1.77 [−3.75, +0.33] | no detectable change |
| Rush attempts | +2.0% / −2.0% | +1.7% / −1.7% | +0.7% to +2.7% | −0.00 [−0.03, +0.02] | no detectable change |
| Receiving yards | +6.0% / −6.0% | +1.4% / −1.4% | +0.5% to +2.2% | −5.32 [−9.62, −1.18] | better |
| Receptions | +3.0% / −3.0% | +0.7% / −0.7% | −0.1% to +1.4% | −0.01 [−0.02, +0.00] | no detectable change |
| Passing TDs | +6.0% / −6.0% | +4.6% / −4.6% | +2.8% to +6.4% | −0.00 [−0.00, +0.00] | no detectable change |
| Interceptions | −5.0% / +5.0% | −1.8% / +1.8% | −5.1% to +1.2% | −0.00 [−0.00, +0.00] | no detectable change |
| Anytime TD | +5.0% / −5.0% | +5.4% / −5.3% | +3.8% to +6.8% | +0.00 [−0.00, +0.00] | no detectable change |
Platoon (MLB)
The production MLB model has no handedness input. In plate-appearance data the effect is large and stable: same-handed matchups produce fewer home runs and walks and more strikeouts (Table 6). Adding the league platoon factor improved walk and home-run forecasts on the test seasons; a batter's own platoon split, shrunk heavily, added a little more for strikeouts only.
| Outcome | Same-hand ÷ opposite-hand | Home ÷ away | Adding league platoon | Adding batter's own platoon split |
|---|---|---|---|---|
| Hit | 0.963 | 1.024 | −0.2 [−0.5, +0.0] | +0.0 [+0.0, +0.0] |
| Strikeout | 1.081 | 0.959 | −0.4 [−0.9, +0.3] | −1.0 [−1.4, −0.6] |
| Walk or HBP | 0.867 | 1.073 | −1.8 [−2.4, −1.2] | +0.0 [+0.0, +0.0] |
| Home run | 0.881 | 1.029 | −0.9 [−1.2, −0.5] | +0.0 [+0.0, +0.0] |
Player bias against the model's own forecasts (NHL)
In the hockey backtest the control itself (the player's overall ratio of actual to forecast output) improved shots and points clearly. That is not a matchup effect: it says the production model was consistently over- or under-projecting some players. Because it changes a published model, it went through the hockey model's own locked protocol as version nhl-v2.4, with a point-in-time ledger of each player's actual output and the opportunity means his features gave before each game, a prior strength fitted on training rows only, selection on the 2023–24 and 2024–25 folds, and a single paired final test. The final-test change in shots log loss against v2.3 is −0.00220 [−0.00298, −0.00141]; calibration error at the reference lines falls by about half.
| Market | v2.3 | v2.4 | Difference [95% interval] | ECE |
|---|---|---|---|---|
| shots | 1.54065 | 1.53846 | −0.00220 [−0.00298, −0.00141] | 0.0167 → 0.0060 |
| goals | 0.45255 | 0.45237 | −0.00018 [−0.00037, +0.00001] | 0.0056 → 0.0044 |
| assists | 0.64402 | 0.64364 | −0.00037 [−0.00064, −0.00010] | 0.0126 → 0.0062 |
| points | 0.83855 | 0.83799 | −0.00056 [−0.00093, −0.00016] | 0.0201 → 0.0114 |
Changes to the Published Models
- NFL: the home/away multipliers now hold the gaps refitted on 2012–2025 (Table 5). Pass attempts and completions keep no venue factor; their measured gaps are under 1% and indistinguishable from zero.
- NHL: model nhl-v2.4 adds each player's record against our earlier forecasts (Table 7). Opponent history and venue splits are not inputs.
- MLB: platoon is the strongest candidate improvement. It is not yet in the model: the production model trains daily on official box scores, and the change will be compared with and without platoon on that data before it ships.
- Player snapshots: the Form tab shows the player's average against tonight's opponent, at home and away, and how often he cleared tonight's line, labeled as history that the model does not use.
Threats to Validity
- These are tests of predictive accuracy, not betting returns. Archived prop prices for these seasons are not available, so no return is claimed for any change.
- The NFL baseline rebuilds the production recipe from public data rather than replaying production's saved weekly forecasts, and the MLB baseline is a plate-appearance model built for this study, not the production game-level model. A weaker baseline would make matchup history look more useful, not less.
- "Especially at home" against one opponent could not be tested reliably: splitting already thin matchup samples by venue leaves almost nothing, and the matchup effect alone is about zero.
- Matchup histories span roster, coaching and scheme changes. That is inherent to the claim being tested, but a scheme-specific matchup (a receiver against one coverage) could still matter where this test cannot see it.
- The NHL test season, 2025–26, had served as the final test for earlier hockey models, and this backtest scored it too before v2.4 was evaluated. v2.4's validation folds alone selected it and improved every market in both folds, but the final test is not untouched for that idea.
- NFL neutral-site games are excluded from the fit but, at forecast time, still take the listed home team's factor.
Conclusion for Readers
A player's record against one team is a story, not a signal: once form, the opposing defense and the venue are known, it does not predict the next meeting. Use the "against" history in our snapshots as context, and weigh the model's inputs (the player's form, the opponent's defense, the venue, and in baseball the handedness matchup) above it. Where a factor is real, it should be measured rather than assumed; our own fixed NFL venue multipliers were four times too large for yardage.
Reproduction
The backtests are research scripts; no production model reads their output. From a checkout with the data in local folders:
python scripts/research/matchups/nhl_matchups.py --features features.pkl --out reports/matchups/nhl.json
python scripts/research/matchups/nfl_matchups.py --data nflverse/ --out reports/matchups/nfl.json
python scripts/research/matchups/nfl_venue.py --data nflverse/ --out reports/matchups/nfl_venue.json
python scripts/research/matchups/mlb_matchups.py --retro retrosheet/ --out reports/matchups/mlb.json
python scripts/research/matchups/report.py
python3 scripts/research/build_matchup_paper.py
The NHL v2.4 evaluation follows the hockey model's own commands in reports/nhl-v2.4/README.md. The evidence JSON lists the SHA-256 of every evidence file used here.
References and Research Materials
- National Hockey League. NHL Statistics. Team and skater per-game reports, archived with model nhl-v2.3 and later.
- nflverse. nflverse-data: weekly player statistics (stats_player) and schedules. Retrieved October 6, 2026.
- Retrosheet event files, via the Chadwick Baseball Bureau mirror. The information used here was obtained free of charge from and is copyrighted by Retrosheet. Interested parties may contact Retrosheet at www.retrosheet.org.
- Efron, B., and Morris, C. (1975). Data Analysis Using Stein's Estimator and Its Generalizations. Journal of the American Statistical Association, 70(350), 311–319.
- Gneiting, T., and Raftery, A. E. (2007). Strictly Proper Scoring Rules, Prediction, and Estimation. Journal of the American Statistical Association, 102(477), 359–378.
- Fourth & Value Research Group (2026). Backtest scripts, results and the NHL v2.4 evaluation.
Research status: experimental forecasts. No betting advantage is established for any change described here. This report is an internal technical publication, not external peer review.