AI models have benchmarks. So do ours. Every ranking model published on ScratchOffsNY is graded here against one thing only: what actually happened. Not projected odds, not our own math — the prize money New Yorkers really claimed, per dollar they really spent, in the three weeks after each ranking went live. Updated daily as new outcomes resolve.
Quick answer: As of 2026-09-10, the ScratchOffsNY model whose top-5 picks returned the most per dollar is Tally (82.5% realized payout over the following 21 days). The model best at surfacing $5,000+ wins is Tally (58.8 big prizes per million tickets among its top 5). For reference, buying the ticket with the best printed odds returned 76.3%, buying the highest published payout rate returned 72.3%, and a random pick returned 65.6%.
Best return: Tally — its top-5 picks returned 82.5% per dollar. Best for big wins: Tally — 58.8 $5,000+ prizes per million tickets among its top 5.
| Model | Top-5 realized payout | vs a random ticket | Top-5 $5k+ wins / M tickets | Big-Win IC | Payout IC | Beat Smart Score | Days graded |
|---|---|---|---|---|---|---|---|
| #1 Tally | 82.5% | +16.9¢ | 58.8 | 0.605 | 0.484 | 67% | 105 (walk-forward) |
| #2 Apex | 80.7% | +15.1¢ | 45.9 | 0.381 | 0.511 | 81% | 105 (walk-forward) |
| #3 Smart Score | 78.9% | +13.3¢ | 28.0 | 0.300 | 0.424 | — | 105 (walk-forward; 161 live) |
| #4 Verity | 70.0% | +4.4¢ | 36.1 | 0.341 | 0.467 | 58% | 45 (walk-forward; 52 live) |
| Two-Stage (shadow) shadow | — | — | — | — | — | — | 0 (walk-forward) |
| Reference: best printed odds (on the ticket) baseline | 76.3% | +10.7¢ | 28.4 | 0.475 | 0.283 | 16% | 161 |
| Reference: highest published payout rate baseline | 72.3% | +6.7¢ | 42.6 | 0.490 | 0.492 | 80% | 161 |
| Reference: random pick baseline | 65.6% | — | 17.8 | 0.000 | 0.000 | 1% | 161 |
Top-5 realized payout answers the only question most players have: “if I bought the model’s five best games, what did that cohort return per dollar?” Realized payout is the prize value actually claimed divided by the dollars of tickets actually sold over the following 21 days. Every scratch-off pays back less than a dollar by design — a random New York ticket returned 65.6% per dollar on this board — so there is no 100¢ ticket and a number in the seventies or eighties means the picks beat the market. The vs a random ticket column makes that explicit. Top-5 $5k+ wins is the same cohort’s rate of big hits per million tickets; it is rarer and noisier than payout and moves more month to month. Big-Win IC and Payout IC are the technical ordering metrics: the Spearman rank correlation between a model’s full 60-plus-game ranking and each outcome (1.0 perfect, 0 coin flip). A model can order the whole field well yet still pick a weaker top 5, or vice versa — which is exactly why both are shown. Beat Smart Score is the share of days a model’s Payout IC exceeded the incumbent’s.
The gray rows are not models and are not recommendations. They are what a benchmark needs to mean anything: the strategies a real player already has without any analytics, drawn the way an AI benchmark draws a dashed “human baseline.” Neither of the two informed references accounts for sales or changing odds. Best printed odds is the everyday heuristic — read the “1 in 3.61” on the back of the ticket and buy the lowest number. Printed odds are fixed the day a game is printed and never change; because every $30 game prints roughly the same odds, the rule amounts to “buy $30 tickets,” which is why it returns well and why it cannot tell a $30 game with three top prizes left from one with a single top prize and 93% sold. Highest published payout rate is the informed player’s heuristic — the remaining-prize percentage every odds site shows for each game. It does update as prizes are claimed and orders the field well, but it cannot tell a great return from a great buy (on launch day its #1 was a nearly sold-out $1 ticket with a $5,555 top prize). Random pick is a blind grab. The models, by contrast, re-price every game every day from live sales and claims. A model that cannot clear these lines on the player metrics has no business being offered, and we say so on this page when it happens.
Smart Score is the production model: a factor-weighted composite whose weights are tuned by a daily optimizer, anchored to realized outcomes, and versioned in the changelog. Its honest claim is narrow and real: the best return per dollar on its top five picks, by a few cents. Verity uses the same factors but takes its weights straight from the raw realized-payout fit each day. Apex starts from the published payout rate (half its score) and adds only the factors with a measured track record since March 2026 at predicting realized payout and $5,000+ wins. Tally (public since September 4, 2026) is a different kind of model: no factors and no learned coefficients. It reads the prize ledger the Lottery publishes and asks, per ticket still out there, how much prize money under $5,000 is left, how much of everything except the jackpot is left, and how many $5,000+ prizes are left. Leaving the jackpot out of the payout side is the whole idea — jackpots almost never resolve inside three weeks, so a game whose remaining value is mostly jackpot looks rich and pays out like everyone else. Its three blend weights are re-fit every day on all graded history and nudged, not jumped, toward the new fit, so it improves as outcomes accumulate without lurching. In its pre-registered window it produced the best top-5 realized payout on this board and the best big-win ordering, while ordering the middle of the field a little worse than Apex; its blend weights were chosen on data through mid-June, inside that window, so treat the exact figures as optimistic until the live rows fill in. All four models are selectable on the rankings page. Shadow rows are candidates being graded before any public release; they are listed below the public models and never ranked.
ScratchBench is a public benchmark that grades every ScratchOffsNY ranking model (Smart Score, Verity, Apex, Tally) on realized outcomes: the prize money actually claimed per dollar of tickets actually sold, and the number of $5,000+ prizes actually hit, over the 21 days after each ranking was published. It updates daily.
As of 2026-09-10, the ScratchOffsNY model whose top-5 picks returned the most per dollar is Tally (82.5% realized payout over the following 21 days). The model best at surfacing $5,000+ wins is Tally (58.8 big prizes per million tickets among its top 5). For reference, buying the ticket with the best printed odds returned 76.3%, buying the highest published payout rate returned 72.3%, and a random pick returned 65.6%.
Realized payout = prize value actually claimed ÷ dollars of tickets actually sold, per game, over a 21-day window, using official New York Lottery claim and sales data. It is not the odds printed on the ticket and not a model projection. Every scratch-off pays back less than a dollar by design, so read the number as distance above what a random ticket returns (about 66 cents).
IC (information coefficient) is the Spearman rank correlation between a model’s ranking of all active games on a given day and each game’s realized outcome over the next 21 days. 1.0 is a perfect ordering, 0 is a coin flip. Payout IC uses realized payout; Big-Win IC uses $5,000+ win density.
Not a model and not a recommendation. It is what a player gets by reading the payout percentage the Lottery publishes on every game page and buying the highest, with no analytics. It is a strong reference because that number already captures most of the next three weeks, and every model here uses it as an ingredient. It cannot tell a great return from a great buy: its top pick is often a nearly sold-out $1 game with a tiny top prize. A model that cannot clear this line on the player metrics has no business being offered.
Scores are read from archived snapshots of the day they were published and are never recomputed. Outcomes come from official data, not from any model. New models enter with a pre-registered walk-forward backtest and switch to live results after 10 graded days. No model is removed for scoring badly. Limitations: one state (New York), roughly six months of history, and claim-lag noise on short horizons.
ScratchOffsNY. “ScratchBench: NY Scratch-Off Model Leaderboard.” Updated 2026-09-10. https://scratchoffsny.com/methodology/benchmarks. Data CC BY 4.0. Machine-readable results: /feeds/scratchbench.json.
Data sourced from nylottery.ny.gov and NY Open Data. Independent, non-commercial project. Play responsibly — NY Problem Gambling · 1-877-8-HOPENY.