Our first full week of 2026 produced a losing prop shortlist, an especially rough run in rushing attempts, and a useful reminder that good prices cannot rescue every forecast. Favorites and game-total overs had a better week than our selections. The strongest thing our totals model did was accept how little weight its own predictions deserved.
The numbers below come from the pages we actually published before kickoff. We recovered the September 9 board and Top Picks from Git, matched them to completed Week 1 results, and kept the original probabilities and prices. The Week 2 model already incorporates Week 1 outcomes; using it here would answer the wrong question.
The headline: 815 gradable offers on the archived shortlist went 365–450, returning −7.2% at one unit risked per offer. Those offers repeat many of the same opinions across sportsbooks and alternate lines. A second, less repetitive comparison—one ticket per player, game, and market—went 97–135, −27.43 units, and −11.8% ROI on 232 graded tickets.
These are hypothetical flat-stake audits, not a record of bets our readers placed. The second comparison uses a rule defined for this review: retain the archived offer with the highest published expected return in each player/game/market, without looking at results. Neither comparison accounts for sportsbook limits, later price availability, promotional injury protection, or individual execution.
Ten of the 242 selected tickets remain ungraded because the saved results do not establish a matching player outcome or sufficient participation evidence. Missing players were not quietly turned into winning unders. Thirty-three of the original 848 shortlist offers remain unresolved for the same reason.
Where we made money, and where we lost it
The table uses the one-ticket comparison throughout. One unit means one unit risked, with winning returns calculated from the actual archived American odds. Pending tickets are excluded from the ROI denominator.
| Market | Graded W–L | Pending | Net units | ROI |
|---|---|---|---|---|
| Passing yards | 2–0 | 0 | +3.70 | +185.0% |
| Receiving yards | 3–5 | 1 | −0.60 | −7.5% |
| Receptions | 48–68 | 6 | −10.78 | −9.3% |
| Rushing yards | 18–22 | 2 | −5.05 | −12.6% |
| Passing attempts | 6–7 | 0 | −1.72 | −13.3% |
| Passing completions | 5–7 | 1 | −2.00 | −16.7% |
| Rushing attempts | 15–26 | 0 | −10.97 | −26.8% |
Rushing attempts were our worst market by both net loss and return on stake. Receptions were almost as expensive in units because we selected so many of them. Together those two markets lost 21.75 units, about 79% of the portfolio's net loss. Passing yards were the only profitable market in this comparison, but that result rests on two correlated tickets in one game.
The split by side is just as revealing. Our overs went 40–48 for −3.44 units and −3.9% ROI. Our unders went 57–87 for −23.98 units and −16.7%. Unders accounted for about 87% of the net loss. That tells us where the damage accumulated; it does not establish that every under was a bad forecast or that next week's answer is to reverse every pick.
What worked—and how much credit it deserves
The passing-yard winners were Drake Maye under 199.5 and Sam Darnold under 199.5, both available at +185 in the archived shortlist. Maye finished with 178 passing yards; Darnold finished with 13. Risking one unit on each produces the table's +3.70 units.
Those were useful prices on winning outcomes. They were not evidence that we had mastered quarterback projections. Every one of the board's 436 modeled passing-yard offers carried a 50.0% probability. The model treated those two tickets as coin flips. At +185, that alone generates an estimated return of 42.5%.
Darnold also left the opener early with a hip injury, which the Seahawks confirmed during the game. Our stat-based audit counts the under as a win because he participated; actual settlement can differ where a book offers injury protection. More fundamentally, an early exit is not proof that a pregame estimate of passing production was accurate. We should take the favorable outcome without inventing a forecasting accomplishment around it.
There were other successes. Adam Trautman over 1.5 receptions at +188 won with two catches. Mack Hollins over 1.5 at +162 won with four. The Tampa Bay–Cincinnati selections went 9–5 for +5.14 units; Dallas–New York went 7–3 for +5.10. Those are the strongest game-level pockets in the one-ticket audit.
Price shopping helped even in the losing portfolio. Keeping the same selected players, lines and sides, our ticket prices improved realized returns by 3.95 units versus the median payout available across the saved books at those exact lines. That reduced the loss from 31.38 to 27.43 units. It is a concrete benefit of comparison shopping, although it was not large enough to overcome the selections' losses.
The highest estimated-return bucket also deserves a fair accounting. The 50 graded tickets with a published expected return of at least 20% went 24–26 but earned +7.62 units, or +15.2%, because the winning prices were large enough. The lower three buckets all lost money. This week does not support the claim that our biggest displayed edges necessarily performed worst. It does support keeping hit rate and profit separate: fewer than half of that top bucket won, yet the bucket was profitable.
These are descriptive slices chosen for the review, not independently validated strategies. We cannot find a favorable subgroup after the games and promote it as next week's proven selection rule.
The biggest struggle: volume and an under-heavy portfolio
Rushing attempts should help us understand how much work a player will get. In Week 1, our published mean missed by 3.84 carries on average across the 55 comparable player markets. The representative sportsbook line missed by 3.21. Our forecasts were low by 1.57 carries on average.
Rushing yards showed a similar pattern: a model error of 18.20 yards versus 15.25 for the market line, with a model bias of −5.41 yards. Receptions were closer, but still behind: 1.69 catches of average error versus 1.61 for the line. Those are comparisons on identical player outcomes, not different samples selected to make one side look better.
New Orleans–Detroit makes the failure concrete. We selected Jahmyr Gibbs under 18.5 carries at −102 and under 84.5 rushing yards at −110. He finished with 29 carries and 156 yards. We also selected Tyler Shough under 33.5 attempts; he threw 56 times. Receptions unders on Amon-Ra St. Brown, Chris Olave, Jahmyr Gibbs and Travis Etienne all lost.
That game finished 31–30 in overtime. It is reasonable to investigate how workload, pace and game state affected these misses. It would be stronger than the evidence to assign a precise share of the error to overtime or any single feature. The observable fact is that several selections depended on quieter workloads, and those assumptions failed together.
San Francisco–Los Angeles was the worst game for our ticket portfolio: 4–13, −8.27 units. New Orleans–Detroit lost 6.45 units, Baltimore–Indianapolis lost 6.32, and Chicago–Carolina lost 6.20. Multiple player markets in one game are multiple ways to express related assumptions. A long list of picks does not automatically create a diversified portfolio.
The Chicago–Carolina game was an extreme example. It produced 96 points against our 47-point market snapshot, while our raw totals forecast was only 46.3. Our prop selections there went 5–11. A weekly review should show that concentration instead of treating each losing ticket as an unrelated surprise.
Did our probabilities improve on the sportsbooks?
Profit matters, but one week of returns can reward an inaccurate forecast and punish a reasonable one. We therefore also compared the published probabilities with a market baseline.
For each player/game/market, we selected the exact line offered by the most distinct books. Ties were resolved by proximity to the median of distinct book/line pairs, then the lower line. We scored one Over probability at that line, using only outcomes with both a model estimate and an exact-line, de-vigged market consensus. That leaves 530 player/game/market observations covering 192 player-game identities and 16 games.
Brier score measures squared probability error; lower is better. A constant 50% forecast scores 0.2500 on settled binary outcomes.
| Market | Outcomes | Model Brier | Market Brier |
|---|---|---|---|
| Passing attempts | 31 | 0.2588 | 0.2469 |
| Passing completions | 31 | 0.2516 | 0.2473 |
| Passing yards | 31 | 0.2500 | 0.2498 |
| Receptions | 152 | 0.2495 | 0.2478 |
| Receiving yards | 153 | 0.2500 | 0.2499 |
| Rushing attempts | 55 | 0.2529 | 0.2409 |
| Rushing yards | 77 | 0.2510 | 0.2489 |
| All comparable markets | 530 | 0.2509 | 0.2479 |
The market had a lower Brier score in every modeled category. Log loss gave the same overall ordering: 0.6950 for our model against 0.6890 for the market. Our raw mean projections also had higher absolute error than the representative line in all seven categories.
The difference is small, and the uncertainty matters. Resampling whole games 5,000 times produced an approximate 95% interval of −0.00043 to +0.00639 for model Brier minus market Brier. That interval crosses zero. Week 1's point estimates favor the market; this one slate does not establish a permanent or statistically conclusive ranking.
The confidence pattern is less comforting than the small average difference. At representative lines, the 68 forecasts with directional confidence between 55% and 60% averaged 57.1% confidence but were right 48.5% of the time. The seven forecasts between 60% and 70% were right three times. These buckets are too small for a stable recalibration, but they offer no reason to increase trust simply because the displayed percentage is higher.
Calibration also compressed many forecasts into a few repeated values. Passing yards were exactly 50%; rushing yards used only three distinct probabilities; receiving-yard probabilities ranged from roughly 48.4% to 51.6%. That may restrain an overconfident raw model, but it makes a large displayed edge at a long price especially important to inspect. The edge can be driven by the price combined with an almost flat probability curve.
Across the 232 graded selected tickets, the model expected about +28.45 units and produced −27.43. Their average predicted win probability was 53.5%; their realized win rate was 41.8%. A single week can move a long way from its expectation, particularly with correlated outcomes. Still, those are the numbers the next model review must confront.
Game totals: shrinking the model helped, but did not beat the market
Our totals page published both the raw forecast and a version pulled toward the sportsbook consensus. Week 1 justified that restraint.
| Forecast | Mean absolute error | Root mean squared error |
|---|---|---|
| Raw model | 12.47 points | 17.65 points |
| Published model after shrinkage | 11.79 points | 16.68 points |
| September 9 market snapshot | 11.78 points | 16.66 points |
| Recorded closing total | 11.69 points | 16.49 points |
The closing total has the benefit of later information, so the September 9 snapshot is the more direct comparison with our forecast. Shrinkage reduced error by about 0.68 points per game relative to the raw model. It essentially brought us back to the market, with no incremental gain over the same-time baseline.
Blindly following the raw model's direction on every total would have gone 6–9–1 and lost 3.34 units at the saved best prices on the consensus line. The calibrated forecast retained the same directional lean, but made the disagreements much smaller. Only New England–Seattle had a positive estimated return after pricing the calibrated probabilities. That over 44.5 lost when the game finished with 23 points.
In our earlier “The Edge Was the Price” article, we highlighted over 44.5 at −102 and a 50.9% model probability. The better price was real; the outcome was a loss. An estimated edge of less than one dollar per $100 risked should never have been read as a strong prediction of a high-scoring opener.
The raw model did have useful individual calls, including the under in Arizona–Los Angeles and the over in Buffalo–Houston. Its misses included overs in the opening two games and unders in several high-scoring Sunday games. Removing the 96-point Chicago–Carolina outlier still leaves the raw model with 9.99 points of average error against the market's 9.30. That one game enlarged the errors without creating the entire ranking.
All 16 games: saved total, forecast and result
| Game | Final (away–home) | Saved total | Raw model | Published model | Raw lean | Result |
|---|---|---|---|---|---|---|
| NE @ SEA | 10–13 | 44.5 | 51.1 | 44.8 | Over | Loss |
| SF @ LA | 27–7 | 48.5 | 51.2 | 48.6 | Over | Loss |
| ATL @ PIT | 13–20 | 42 | 46.8 | 42.2 | Over | Loss |
| BAL @ IND | 41–23 | 48 | 46.8 | 48.0 | Under | Loss |
| BUF @ HOU | 36–31 | 45 | 51.0 | 45.3 | Over | Win |
| CHI @ CAR | 59–37 | 47 | 46.3 | 47.0 | Under | Loss |
| CLE @ JAX | 10–34 | 40.5 | 47.0 | 40.8 | Over | Win |
| NO @ DET | 30–31 | 50 | 45.7 | 49.8 | Under | Loss |
| NYJ @ TEN | 23–10 | 38.5 | 39.2 | 38.5 | Over | Loss |
| TB @ CIN | 27–33 | 50.5 | 49.7 | 50.5 | Under | Loss |
| ARI @ LAC | 26–14 | 47.5 | 41.9 | 47.3 | Under | Win |
| GB @ MIN | 22–39 | 46 | 44.1 | 45.9 | Under | Loss |
| MIA @ LV | 13–27 | 40.5 | 39.7 | 40.5 | Under | Win |
| WAS @ PHI | 22–24 | 44.5 | 46.2 | 44.6 | Over | Win |
| DAL @ NYG | 20–28 | 48 | 46.8 | 48.0 | Under | Push |
| DEN @ KC | 10–31 | 43.5 | 43.4 | 43.5 | Under | Win |
What last week's odds got right
This is a separate question from how our picks performed. A market can be difficult for our model while still producing a profitable week for a simple favorite or Over strategy.
Using the saved schedule's closing lines, favorites won 12 of 16 games outright and went 9–6–1 against the spread. Home-designated teams went 10–6, although that category includes the neutral-site Rams game. Game-total overs went 9–7.
At the recorded moneyline prices, risking one unit on each favorite returned +2.62 units, or +16.4%; betting every underdog lost 4.32 units, or −27.0%. Over bets at the recorded closing prices returned +1.34 units, or +8.4%, while unders lost 2.65 units, or −16.6%. These are bettor-side results, not estimates of sportsbook profit: we do not have the books' stakes, liabilities, promotions or betting mix.
Favorites against the spread returned +2.09 units, or +13.0%, at their recorded closing prices; underdogs against the spread lost 3.45 units, or −21.6%. Those ROI denominators include the Seattle push as a stake returned. Favorites were the clearest broad game-market winners in this slate, both outright and against the handicap.
The league scored 791 points against 724.5 implied by our September 9 totals. Chicago–Carolina alone exceeded its saved total by 49 points, accounting for about 74% of that aggregate excess. High total scoring does not mean almost every Over won. The distribution mattered.
Timing mattered too. Against our September 9 snapshot, overs went 8–7–1 rather than 9–7. Dallas–New York finished with 48 points: a push at our saved 48, an Over win at the recorded closing 47.5. Always betting the snapshot Over at its best available same-line price returned +0.58 units, or +3.6%, across 16 stakes. Different snapshots can produce different honest records for the same week.
Seattle supplied the corresponding spread lesson. A three-point win pushes a Seattle −3 ticket and loses a −3.5 ticket; both handicaps appeared in our pregame book comparison. The exact number belongs in the record alongside the price.
Which prop markets rewarded bettors generally?
For this comparison we ignored our model's preferences. We used one representative offered line per player/game/market and the best archived price at that exact line, separately for Over and Under. This covers our saved board, not every wager available throughout the week. Prices were captured September 9; these are not closing prop odds.
| Market | Graded player markets | Over W–L | Over ROI | Under ROI |
|---|---|---|---|---|
| Passing touchdowns | 31 | 16–15 | +7.4% | −7.9% |
| Interceptions thrown | 31 | 16–15 | +6.9% | −13.2% |
| Receptions | 160 | 85–75 | +4.2% | −10.3% |
| Rushing yards | 79 | 43–36 | +3.8% | −14.4% |
| Passing completions | 31 | 16–15 | −1.5% | −6.8% |
| Receiving yards | 161 | 80–81 | −6.0% | −4.5% |
| Passing attempts | 31 | 15–16 | −8.8% | −2.0% |
| Rushing attempts | 56 | 27–29 | −8.9% | −3.1% |
| Passing yards | 31 | 14–17 | −14.7% | +3.6% |
The profitable sides were passing-touchdown Overs, interception Overs, receptions Overs, rushing-yard Overs and passing-yard Unders. This is a retrospective description, not a reason to bet them indiscriminately next week.
Receptions offer the sharpest contrast with our own selections. The general best-price Over comparison earned 4.2%, while our selected receptions portfolio lost 9.3%. Rushing attempts are equally instructive: indiscriminately betting the representative Under lost 3.1%, but our selected market portfolio lost 26.8%. The market was not handing everyone an easy profit; our selection of players, sides and lines made the week worse.
Receiving yards were almost perfectly balanced—80 Overs and 81 Unders—yet both flat-stake sides lost money. Passing-completion Overs won 16 of 31 and still lost money. Those results are why a bare win percentage is an incomplete betting analysis.
Passing touchdowns and interceptions had no published model probabilities, so their market-level results are not model successes. Anytime, first and last touchdown markets are excluded from this settlement audit. We did not have a validated model for them, and the result matching here is not designed to settle first-scorer or last-scorer wagers. Coverage is part of the report, not something to fill in after the fact.
What should change after this week
First, review the calibration curves before trusting the size of the displayed edges. A flat 50% passing-yard forecast can identify favorable prices relative to that assumption, but it does not establish that every alternate line is a coin flip. The next validation should compare a market-only baseline, the raw model, the calibrated model and a blend on future, untouched weeks. The curve needs to preserve useful distinctions across lines while earning any departure from market prices.
Second, prioritize workload forecasts. Rushing attempts had the worst selection results and the largest relative Brier shortfall. Receptions created about half of the selected tickets. Review projected roles, carries and targets, and test improvements chronologically. Changing coefficients until they explain Gibbs's Week 1 would be fitting the answer we already know.
Third, make the public record easier to interpret. Preserve the whole offer archive, but maintain a separately defined selection policy with one reproducible ticket choice and explicit game-level exposure limits. Decide those rules before the next slate. A page containing hundreds of qualifying offers should not imply hundreds of independent opportunities.
Fourth, preserve quote timing and improve settlement evidence. We can prove these prices and probabilities existed before the games. We cannot prove that Wednesday's prop prices remained available on Sunday, or compute prop closing-line value without a closing archive. Capture later snapshots, starting-role information, participation and book-specific void treatment so future reviews can distinguish stale quotes, unavailable players and genuine forecast misses.
Finally, retain the restraint that already helped. Shrinking the totals model reduced its errors. Suppressing unsupported touchdown estimates avoided presenting guesses as player-specific probabilities. Neither earns fictional winning bets, but both improve the quality of what we publish.
Week 1 leaves us with useful wins, measurable failures and a clear priority: our probability estimates have to earn their place next to the prices. For now, the evidence favors the sportsbook baseline. The next week's record should be built to tell us whether that changes.
Sources and audit notes
The primary forecast evidence is the archived props board and archived Top Picks page, built September 9 at approximately 6:36 PM Eastern. Totals use the saved 6:52 PM Eastern snapshot. The audit retains the source hashes and timestamps in the downloadable summary.
Results come from the project's saved 2026 Week 1 schedule and player-stat files, using the nflverse player-stat release and nflverse schedule dataset as the upstream sources. Game identity, both teams and Eastern game dates are checked before matching players. Name punctuation and generational suffixes are normalized; ambiguous identities stop the audit. The old unseasoned grades_week1.csv file was not used.
Only players with a matching game record and evidence of an attempt, carry, target or reception are graded. Zero-usage cases remain unresolved without better participation evidence, which limits how representative the graded sample is. The audit assumes ordinary action after participation, retains pushes as returned stakes, and excludes unresolved stakes from ROI. There were no pushes in the selected prop tickets. The reported returns describe saved offers, not verified cash settlements.
Repeated books and alternate lines remain in the full-offer audit but are removed from the one-ticket comparison. Markets for the same player and players in the same game remain correlated. The game-block bootstrap acknowledges that dependence, but 16 games cannot establish season-long skill. All calculations use frozen Week 1 forecasts; no Week 2 model output or private user betting records enter the review.