Fourth & Value Research · Technical Report FV-2026-03

Player Matchups and Home Field:
Do Head-to-Head History and Venue Splits Predict the Next Game?

Fourth & Value Research Group

fourthandvalue.com

October 6, 2026 · Version 1.0 · NFL, NHL and MLB player forecasts

Technical report; not peer reviewed.

Paper, numerical evidence, and reproducible implementationDownload PDFEvidence JSON
Abstract

Bettors often say a player "owns" an opponent, or is far better at home. We test both claims as forecasting inputs. Each history is added to a forecast that already knows the player's form, the opponent's defense and the venue, as an empirically shrunk ratio of actual to expected output over strictly earlier games. The shrinkage strength is chosen on tuning seasons, and each sport is scored once on later test seasons: 47,230 NHL player-games (2025–26), NFL games from 2024 through week 4 of 2026 in seven markets, and 549,195 MLB plate appearances (2023–25) from Retrosheet, where batter-versus-pitcher history is exact. Player-versus-opponent history and personal home/away splits added nothing measurable in any sport: the tuning seasons gave them zero or small weight, where they kept some weight it did not improve the test seasons, and players who had beaten their forecast against an opponent did not keep that edge the next time. Three things did help. League-wide home/away effects of the right size: the NFL's fixed ±6% multipliers for passing and receiving yards were about four times the measured gap. Platoon (batter side against pitcher hand) in baseball. And, in hockey, correcting for players the model consistently over- or under-projected, which became model nhl-v2.4 after a separate locked evaluation.

Keywords: player props; head-to-head; batter versus pitcher; home-field advantage; platoon; empirical Bayes shrinkage; chronological validation; calibration.

Introduction: What the Findings Mean

"He always lights up this team" is one of the most common reasons given for a player prop. It is also hard to test by eye: a quarterback sees a non-division opponent about once every three or four years, and twenty plate appearances against one pitcher carry about ±.100 of noise in batting average. A run of good games against one team is exactly what chance produces somewhere in a league of hundreds of players.

We asked a narrow question. Suppose a forecast already knows the player's recent and career form, the strength of tonight's opposing defense, and whether he is home or away. Does adding his own history against this particular opponent, or his own home/away split, make the forecast more accurate on games it has not seen? We answered it the same way in three sports, using only information that would have been available before each game.

The answer is no. In most markets the tuning seasons gave the opponent history and the personal venue split zero or very small weight; where they kept some, it did not improve accuracy on the test seasons. The streak check is the clearest. NHL skaters who had out-shot their forecast by 49% against an opponent came in at 1.025 times the forecast the next time they met; those who had fallen 42% short came in at 1.021; everyone with that much history averaged 1.015. In baseball, batters who had hit 62% above expectation against a pitcher came in at 1.053, and those 49% below at 1.025, against 1.003 for all pairs with that much history: a gap of 111 points in the history shrank to 2.8 points. Aaron Judge, against pitchers he had already faced fifteen or more times, had 53 hits in 222 plate appearances in 2023–25, against 54.1 forecast without any matchup history.

What does work is the mechanism behind a matchup, measured on large samples. The home/away effect is real but small for yardage, and our fixed NFL multipliers overstated it. Handedness matters in baseball. And a player's record against the model's own earlier forecasts carries information the hockey model was missing. We show the past meetings in our player snapshots as context, labeled as not a model input.

Data and Evaluation Windows

Each sport uses its own public record. NHL player-games come from the frozen history behind model nhl-v2.3 (official NHL per-game skater and team reports, 2022–23 through 2025–26) [1]. NFL player-games come from nflverse weekly statistics, 2012 through week 4 of 2026, joined to the nflverse schedule for home, away and neutral sites [2]. MLB plate appearances come from Retrosheet event files for 2012–2025 through the Chadwick Bureau mirror [3]; the parser reproduces published season lines exactly (Aaron Judge, 2024: 180 hits, 58 home runs, 171 strikeouts, 142 walks and hit-by-pitches). Regular seasons only.

Table 1. Data, history and evaluation windows. Tuning chooses how much weight each history gets; the test season is scored once.
SportSourceUnitHistory fromTuningTestTest rows
NHLnhl-v2.3 archive (official NHL game logs)player-game2022–232024–252025–2647,230
NFLnflverse weekly player statsplayer-game20122022–20232024 to 2026 wk 414,793 (7 markets)
MLBRetrosheet event filesplate appearance20122021–20222023–2025549,195

Every history is point in time. A game counts toward a player's history only on later dates: the next day for NHL and MLB, later weeks for NFL. Tuning and test seasons never overlap.

Method: Shrunk Histories over a Fixed Forecast

Baseline forecasts

The NHL baseline is the production player model nhl-v2.3 (projected ice time × recency-weighted production per 60 minutes, scaled by the opponent's shots or goals allowed; negative-binomial counts), refitted on seasons before each scored season exactly as its own evaluation does. The NFL baseline rebuilds the production recipe from public data: a four-game recency average blended with a decayed career mean shrunk to the position average, times the production opposing-defense rating and the production home/away multiplier. The MLB baseline is a plate-appearance model built for this study: the batter's and pitcher's decayed, shrunk rates combined by odds ratio against the league rate, times a shrunk park factor and the league home factor.

Adding a history

For a player i and a set of his earlier games S (all of them, those at the same venue, or those against tonight's opponent), let A be his actual output and E the baseline's expected output over S. The history enters as an empirically shrunk ratio [4]:

r(S) = (A + k·m) ÷ (E + k). (1)

Here k is a prior strength in the stat's own units (expected shots, games of output, or plate appearances) and m is the value the ratio shrinks toward. For the player's overall ratio m = 1. For a venue split or an opponent split m is the player's overall ratio, so the split only moves the forecast by how much the player differs from himself, not from the league. A history of n units therefore gets weight n ÷ (n + k).

forecast = baseline × rall × r(venue) ÷ rall × r(opponent) ÷ rall. (2)

The player's overall ratio is a control. Without it, an opponent split would also pick up any general bias the baseline has for that player, and look useful for the wrong reason.

Choosing k and scoring

For each component, k is chosen from a fixed grid on the tuning seasons by mean log loss (NHL, MLB) or squared error (NFL), in order: the player's overall ratio first, then each split given it. "Not used" means the grid's top value, infinite k, won. The test seasons are then scored once. Differences are paired on the same rows, with 95% intervals from resampling whole games [5].

Table 2. Prior strength chosen on tuning seasons. A history of n units gets weight n ÷ (n + prior); "not used" means the tuning seasons gave it zero weight.
MarketPlayer vs this opponentPlayer's own home/away split
NHL shots1000 expected events of priornot used (zero weight)
NHL scoringnot used (zero weight)1000 expected events of prior
NFL passing yardsnot used (zero weight)300 games of prior
NFL pass attemptsnot used (zero weight)300 games of prior
NFL completionsnot used (zero weight)100 games of prior
NFL rushing yards30 games of priornot used (zero weight)
NFL rush attempts100 games of priornot used (zero weight)
NFL receiving yardsnot used (zero weight)300 games of prior
NFL receptions100 games of priornot used (zero weight)
MLB hit1000 PA of prior (batter vs pitcher)not used (zero weight)
MLB strikeout300 PA of prior (batter vs pitcher)not used (zero weight)
MLB walk or hbp300 PA of prior (batter vs pitcher)not used (zero weight)
MLB home run1000 PA of prior (batter vs pitcher)not used (zero weight)

Player Versus Opponent, and Personal Home/Away Splits

Table 3 compares each split with the same forecast without it. Most weights were zero or small: 2% of the evidence after eight NHL meetings, and 3% (hits, home runs) to 9% (strikeouts, walks) after thirty plate appearances against one pitcher. Two NFL cases kept more: personal venue splits in four markets (12% to 29% after forty games) and opponent history for rushing yards (about 21% after eight meetings). None of them improved the test seasons; the rushing-yard opponent history made squared error worse on average, and the batter-versus-pitcher changes are a few hundred-thousandths of a nat per plate appearance.

Table 3. Test-season change from adding each history to the same forecast without it. Mean per row [95% interval, resampling whole games]; negative is better; 0 [0, 0] means tuning gave it zero weight.
Market (metric)Player vs this opponentPlayer's own home/away split
NHL shots (log loss)+0.00001 [−0.00003, +0.00005]+0.00000 [+0.00000, +0.00000]
NHL goals (log loss)+0.00000 [+0.00000, +0.00000]+0.00002 [−0.00000, +0.00004]
NHL assists (log loss)+0.00000 [+0.00000, +0.00000]−0.00002 [−0.00005, +0.00001]
NHL points (log loss)+0.00000 [+0.00000, +0.00000]−0.00000 [−0.00004, +0.00004]
NFL passing yards (squared error)+0.00 [+0.00, +0.00]+0.28 [−15.83, +13.05]
NFL pass attempts (squared error)+0.00 [+0.00, +0.00]+0.14 [−0.03, +0.29]
NFL completions (squared error)+0.00 [+0.00, +0.00]+0.19 [−0.04, +0.39]
NFL rushing yards (squared error)+5.26 [−0.40, +11.11]+0.00 [+0.00, +0.00]
NFL rush attempts (squared error)−0.01 [−0.05, +0.04]+0.00 [+0.00, +0.00]
NFL receiving yards (squared error)+0.00 [+0.00, +0.00]+1.43 [−0.01, +2.68]
NFL receptions (squared error)+0.00 [−0.00, +0.01]+0.00 [+0.00, +0.00]
MLB hit (log loss ×10⁻⁴)+0.0 [−0.0, +0.1]+0.0 [+0.0, +0.0]
MLB strikeout (log loss ×10⁻⁴)−0.3 [−0.4, −0.2]+0.0 [+0.0, +0.0]
MLB walk or hbp (log loss ×10⁻⁴)−0.2 [−0.3, −0.0]+0.0 [+0.0, +0.0]
MLB home run (log loss ×10⁻⁴)+0.0 [−0.0, +0.1]+0.0 [+0.0, +0.0]

Table 4 asks the bettor's question directly. It takes players whose earlier record against an opponent was far above or below the forecast, and checks how they did the next time against that same opponent. The "hot" groups did not keep their edge. In hockey and football they landed within about one point of the reference group of everyone with that much history, and NFL rushing yards reversed (the hot group under-performed and the cold group over-performed). In baseball a small part of the earlier gap carried over, consistent with the few percent of weight batter-versus-pitcher history earned in tuning.

Table 4. Players who had beaten or missed the forecast against one opponent, and how they did the next time (actual ÷ forecast, test seasons). The reference row is everyone with that much history.
MarketGroupNext meetingsEarlier ratioNext-meeting ratio
NHL shotseveryone with 4+ prior games vs this opponent (reference)34,9640.9671.015
NHL shotsbeat expectation by 30%+ vs this opponent (4+ prior games)4,9291.4921.025
NHL shotsfell 23%+ short vs this opponent (4+ prior games)9,1700.5811.021
NHL pointseveryone with 4+ prior games vs this opponent (reference)34,9640.9540.994
NHL pointsbeat expectation by 30%+ vs this opponent (4+ prior games)8,4431.7150.989
NHL pointsfell 23%+ short vs this opponent (4+ prior games)13,4920.4100.986
NFL receiving yardseveryone with 4+ prior games vs this opponent (reference)7181.1100.946
NFL receiving yardsbeat expectation by 15%+ vs this opponent (4+ prior games)3061.3940.953
NFL receiving yardsfell 13%+ short vs this opponent (4+ prior games)1590.7020.941
NFL rushing yardseveryone with 4+ prior games vs this opponent (reference)2231.0831.014
NFL rushing yardsbeat expectation by 15%+ vs this opponent (4+ prior games)881.3390.936
NFL rushing yardsfell 13%+ short vs this opponent (4+ prior games)430.7211.077
NFL passing yardseveryone with 4+ prior games vs this opponent (reference)2901.0280.964
NFL passing yardsbeat expectation by 15%+ vs this opponent (4+ prior games)411.2540.959
NFL passing yardsfell 13%+ short vs this opponent (4+ prior games)260.8141.038
MLB hitsevery pair with 20+ PA (reference)11,7311.0551.003
MLB hits20+ PA vs this pitcher, hits 40%+ above expectation1,9251.6221.053
MLB hits20+ PA vs this pitcher, hits 30%+ below expectation1,8880.5121.025

Personal home/away splits fared no better. Where tuning gave a player's own venue split any weight, it did not improve the test seasons. A player who has looked much better at home is, on the evidence, mostly showing the league-wide home effect plus noise.

What Did Help: Venue, Platoon and Player Bias

League home/away effects of the right size (NFL)

The NFL model had applied fixed home/away multipliers since 2025: ±6% for passing and receiving yards, ±4% for rushing yards, ±3% for receptions. We measured the gap directly: actual over expected output at home and away, each divided by the overall ratio so only the venue gap remains, with neutral sites excluded. On 2024–26 games held out from fitting, gaps fitted on 2012–2021 replaced the fixed values with lower squared error for passing yards (−99.6 [−198.9, −7.8]) and receiving yards (−5.3 [−9.6, −1.2]). Refitted on every completed season, the passing-yard home factor is +1.4%, not +6%.

Table 5. NFL home / away factors. Earlier fixed values, values refitted on 2012–2025, the 95% interval for the home factor, and the 2024–26 holdout change in squared error when gaps fitted on 2012–2021 replace the fixed values.
MarketPrevious fixedFitted 2012–2025Home 95% intervalHoldout change (fitted − fixed)Holdout
Passing yards+6.0% / −6.0%+1.4% / −1.4%+0.6% to +2.2%−99.56 [−198.88, −7.78]better
Pass attempts+0.0% / +0.0%−0.2% / +0.2%−1.1% to +0.6%+0.05 [−0.12, +0.22]no detectable change
Completions+0.0% / +0.0%+0.8% / −0.8%−0.1% to +1.6%−0.04 [−0.11, +0.03]no detectable change
Rushing yards+4.0% / −4.0%+2.4% / −2.4%+0.9% to +3.8%−1.77 [−3.75, +0.33]no detectable change
Rush attempts+2.0% / −2.0%+1.7% / −1.7%+0.7% to +2.7%−0.00 [−0.03, +0.02]no detectable change
Receiving yards+6.0% / −6.0%+1.4% / −1.4%+0.5% to +2.2%−5.32 [−9.62, −1.18]better
Receptions+3.0% / −3.0%+0.7% / −0.7%−0.1% to +1.4%−0.01 [−0.02, +0.00]no detectable change
Passing TDs+6.0% / −6.0%+4.6% / −4.6%+2.8% to +6.4%−0.00 [−0.00, +0.00]no detectable change
Interceptions−5.0% / +5.0%−1.8% / +1.8%−5.1% to +1.2%−0.00 [−0.00, +0.00]no detectable change
Anytime TD+5.0% / −5.0%+5.4% / −5.3%+3.8% to +6.8%+0.00 [−0.00, +0.00]no detectable change

Platoon (MLB)

The production MLB model has no handedness input. In plate-appearance data the effect is large and stable: same-handed matchups produce fewer home runs and walks and more strikeouts (Table 6). Adding the league platoon factor improved walk and home-run forecasts on the test seasons; a batter's own platoon split, shrunk heavily, added a little more for strikeouts only.

Table 6. MLB plate appearances. Rate ratios from 2012–2019, and test-season (2023–25) change in log loss ×10⁻⁴ per plate appearance. Negative is better.
OutcomeSame-hand ÷ opposite-handHome ÷ awayAdding league platoonAdding batter's own platoon split
Hit0.9631.024−0.2 [−0.5, +0.0]+0.0 [+0.0, +0.0]
Strikeout1.0810.959−0.4 [−0.9, +0.3]−1.0 [−1.4, −0.6]
Walk or HBP0.8671.073−1.8 [−2.4, −1.2]+0.0 [+0.0, +0.0]
Home run0.8811.029−0.9 [−1.2, −0.5]+0.0 [+0.0, +0.0]

Player bias against the model's own forecasts (NHL)

In the hockey backtest the control itself (the player's overall ratio of actual to forecast output) improved shots and points clearly. That is not a matchup effect: it says the production model was consistently over- or under-projecting some players. Because it changes a published model, it went through the hockey model's own locked protocol as version nhl-v2.4, with a point-in-time ledger of each player's actual output and the opportunity means his features gave before each game, a prior strength fitted on training rows only, selection on the 2023–24 and 2024–25 folds, and a single paired final test. The final-test change in shots log loss against v2.3 is −0.00220 [−0.00298, −0.00141]; calibration error at the reference lines falls by about half.

Table 7. NHL v2.4 under the locked protocol: 2025–26 final test, paired against v2.3 on the same 47,230 player-games. Count log loss; 95% interval from resampling whole games; calibration error (ECE) at the reference line.
Marketv2.3v2.4Difference [95% interval]ECE
shots1.540651.53846−0.00220 [−0.00298, −0.00141]0.0167 → 0.0060
goals0.452550.45237−0.00018 [−0.00037, +0.00001]0.0056 → 0.0044
assists0.644020.64364−0.00037 [−0.00064, −0.00010]0.0126 → 0.0062
points0.838550.83799−0.00056 [−0.00093, −0.00016]0.0201 → 0.0114

Changes to the Published Models

Threats to Validity

Conclusion for Readers

A player's record against one team is a story, not a signal: once form, the opposing defense and the venue are known, it does not predict the next meeting. Use the "against" history in our snapshots as context, and weigh the model's inputs (the player's form, the opponent's defense, the venue, and in baseball the handedness matchup) above it. Where a factor is real, it should be measured rather than assumed; our own fixed NFL venue multipliers were four times too large for yardage.

Reproduction

The backtests are research scripts; no production model reads their output. From a checkout with the data in local folders:

python scripts/research/matchups/nhl_matchups.py --features features.pkl --out reports/matchups/nhl.json
python scripts/research/matchups/nfl_matchups.py --data nflverse/ --out reports/matchups/nfl.json
python scripts/research/matchups/nfl_venue.py --data nflverse/ --out reports/matchups/nfl_venue.json
python scripts/research/matchups/mlb_matchups.py --retro retrosheet/ --out reports/matchups/mlb.json
python scripts/research/matchups/report.py
python3 scripts/research/build_matchup_paper.py

The NHL v2.4 evaluation follows the hockey model's own commands in reports/nhl-v2.4/README.md. The evidence JSON lists the SHA-256 of every evidence file used here.

References and Research Materials

  1. National Hockey League. NHL Statistics. Team and skater per-game reports, archived with model nhl-v2.3 and later.
  2. nflverse. nflverse-data: weekly player statistics (stats_player) and schedules. Retrieved October 6, 2026.
  3. Retrosheet event files, via the Chadwick Baseball Bureau mirror. The information used here was obtained free of charge from and is copyrighted by Retrosheet. Interested parties may contact Retrosheet at www.retrosheet.org.
  4. Efron, B., and Morris, C. (1975). Data Analysis Using Stein's Estimator and Its Generalizations. Journal of the American Statistical Association, 70(350), 311–319.
  5. Gneiting, T., and Raftery, A. E. (2007). Strictly Proper Scoring Rules, Prediction, and Estimation. Journal of the American Statistical Association, 102(477), 359–378.
  6. Fourth & Value Research Group (2026). Backtest scripts, results and the NHL v2.4 evaluation.

Research status: experimental forecasts. No betting advantage is established for any change described here. This report is an internal technical publication, not external peer review.