ScratchBench: We Benchmarked Our Own Scratch-Off Models Against Real Payouts

Quick answer: ScratchBench is ScratchOffsNY’s public benchmark that grades its three scratch-off ranking models — Smart Score, Verity and Apex — on what actually happened over the following 21 days. On the opening board, Apex’s top-5 picks returned the most per dollar (80.7¢) and produced the most $5,000+ wins (45.9 per million tickets), Smart Score returned 78.9¢, and Verity 70.0¢. For reference, buying the ticket with the best printed odds also returned 80.7¢ — but printed odds never change and that rule amounts to “buy $30 tickets,” so it hit big prizes at barely half of Apex’s rate; sorting by remaining payout % returned 75.2¢; a random pick returned 66.2¢. All three models are one click apart on the rankings page.

Every AI lab publishes benchmark scores. It is how you know whether a new model is actually better or just newer. We think a site that ranks lottery tickets owes its readers the same thing — so today we are publishing ScratchBench, a live leaderboard that grades every ranking model on ScratchOffsNY against one standard: what actually happened.

Not our own projections. Not the odds printed on the back of the ticket. The prize money New Yorkers really claimed, per dollar they really spent, in the three weeks after each ranking went live.

Try the models: scratchoffsny.com/rankings — switch between Smart Score, Verity and Apex at the top of the list. Every score you see there is graded the way this post describes.

Why we built a benchmark for lottery models

Since March we have published a daily Smart Score for every active New York scratch-off. In July we added Verity, an alternate model users can switch to. Today we are adding a third, Apex. Three models means three opinions about which ticket is best — and without a shared scorecard, “which one is right?” becomes a matter of taste.

A benchmark fixes that. It also keeps us honest. A model that looks brilliant in a backtest can quietly fall apart in production; a public board where bad scores stay posted is the only cure we know of.

How ScratchBench works

The design borrows from how quantitative funds evaluate signals, adapted to lottery data:

  1. Scores are locked. Each model’s score for every game is recorded the day it is published and never changed afterward, so there is no way to peek at the answer.
  2. Outcomes are real. Twenty-one days later we measure, for each game, the realized payout: prize value actually claimed (from official New York Lottery claim data) divided by dollars of tickets actually sold. We also count how many $5,000+ prizes were actually hit per million tickets sold.
  3. Models are graded on ordering. The core metric, which we call IC, measures how closely a model’s ranking matched the order in which games actually paid out. A perfect model would score 1.0; a coin flip scores 0. (More on IC below.)
  4. Top-5 is what you would actually buy. We also report the average realized payout and big-win rate of each model’s five highest-ranked games — the practical “if I followed this list” number.
  5. Reference lines are mandatory. Three no-model strategies sit on the board the way an AI benchmark draws a dashed “human baseline”: buy the ticket with the best printed odds, buy the highest published payout rate, or pick at random. They are not recommendations. They are the bar. A model that cannot clear them is not earning its place.

The opening scoreboard

ScratchBench chart: Apex, Smart Score and Verity compared on top-5 return, top-5 big wins and whole-field ordering, with a dashed reference line for a random ticket
Click the chart to enlarge. Dashed line = a random ticket. Same 105 days for every bar.

Every row below was graded on the same 105 days from April through August 2026, always on days the models had not yet seen: for each month, a model could only use what it knew before that month began. Bold green is the best model on each column.

ModelTop-5 realized payoutMore money back than a random ticketTop-5 $5k+ wins per M ticketsPayout ICBig-Win IC
#1 Apex80.7¢+22%45.90.5110.381
#2 Smart Score78.9¢+19%28.00.4240.300
#3 Verity70.0¢+6%36.10.4670.341

Benchmark reference lines. These are not models and not strategies we recommend — they are the bar the models have to clear, the way an AI benchmark draws a dashed “human baseline.” Neither of the first two responds to sales or claims. Explained in full below.

Reference (no model)Top-5 realized payoutvs random ticketTop-5 $5k+ wins per MPayout ICBig-Win IC
Buy the lowest printed odds (odds never change; in practice “buy $30 tickets”)80.7¢+22%28.20.3010.464
Sort by remaining payout % only (the “Payout” column on any odds site)75.2¢+14%43.50.4880.365
Random ticket (the no-skill floor)66.2¢18.00.0000.000

Verity launched July 2, so its row covers the 45 days of the window after that date. Apex figures are for the version released today, which leans on the published payout rate for half of its score (see below).

How to read “top-5 realized payout” (and why nobody scores 100¢)

This is the column that matters most, so it is worth being precise about it. For each day, take a model’s five highest-ranked games. Over the next 21 days, add up the prize money New Yorkers actually claimed on those five games and divide by the money they actually spent on them. That is the realized payout: cents back per dollar, measured after the fact from official claim and sales data. It is not the odds printed on the ticket and it is not our projection.

Two things about the scale. First, every scratch-off pays back less than a dollar — that is how lotteries work. Across all active New York games in our test window, a random ticket returned 66.2 cents per dollar. There is no 100-cent ticket, so a number in the seventies or eighties is not a failing grade; it is a ticket that beat the market. Second, the useful comparison is therefore distance above 66. That is what the “more money back than a random ticket” column shows. Apex’s picks returned 22% more than a random ticket. Smart Score’s returned 19% more — a strong result, and within two cents of the best model. The gap between Smart Score and Apex on this column is about a tenth of the gap between either of them and a random ticket.

Where Smart Score genuinely lags is the next column: big wins. Its top-5 picks hit $5,000+ prizes at 28 per million tickets, versus 46 for Apex and 44 for the payout-rate reference. Big wins are rarer and noisier than payouts, so that column moves more from month to month — but the gap has been consistent, and it is the specific thing Apex was built to fix. If you have been following Smart Score, you have been getting close to the best available return; you have not been getting the best odds of a life-changing hit.

What the numbers say

Apex is the only model that clears every reference line on every metric. Its top-5 picks returned 80.7 cents per dollar and hit $5,000+ prizes at 45.9 per million tickets — more return than any model and more big wins than any model or reference. It also ordered the whole field better than anything else on the board.

Smart Score is strong on return and weak on big wins. Its top-5 returned 78.9 cents per dollar — 19% more than a random ticket, clearly ahead of the payout-rate reference, and within two cents of the best model. On big wins it trails: 28 per million versus Apex’s 46. That is the specific gap Apex was built to close, and the reason we are publishing all of this rather than just the flattering column.

Verity is now the weakest of the three on the metric players care about. It orders the field well (0.467) but its top-5 return of 70.0 cents trails every reference except random. It stays on the board and on the rankings page because that is the rule, but we would not point a new player at it today.

The reference lines are strong, and we are publishing that. Reading the printed odds off the back of the ticket returns as much on its top 5 as Apex does — because it amounts to buying $30 tickets, and $30 games pay back the most. Sorting by remaining payout % orders the field nearly as well as Apex. Neither is a model. Neither responds to sales or claims: printed odds never change at all, and the payout sort cannot tell a great return from a great buy. The section below explains what each one misses and why they are the right bar for us to clear. The one-line version: for pure money back, “just buy $30 tickets” is a strong strategy, and the value of our models is knowing which ones.

The reference lines: what they are and why they matter

A benchmark with no reference is just a beauty contest between our own models. So three no-model strategies sit on the board, drawn the way AI benchmarks draw a dashed “human baseline” line. They are not recommendations and they are not things we built. They are what a New Yorker already has without us:

The rule we hold ourselves to: a model that cannot clear these lines on the player metrics has no business being offered, and the board will say so.

What “IC” means, and why the payout-rate reference scores so high on it

IC stands for information coefficient, a term borrowed from quantitative finance. It compares two orderings: how a model ranked all 60-plus active games on a given day, and how those same games actually ranked by realized payout over the next 21 days. A score of 1.0 means the model’s order was perfect; 0 means it was no better than shuffling the deck; negative means it was backwards. In noisy real-world data, anything above about 0.3 is a genuinely strong signal, and 0.5 is excellent.

The crucial thing to understand is that IC grades the whole field — the 40 games nobody would buy count just as much as the top 5. That is why the payout-rate reference does so well on it. Realized payout over three weeks is, mechanically, very close to “remaining prize money divided by remaining tickets” — and that is almost exactly what the published payout rate is. Sorting by it correlates with the outcome nearly by definition. It is a thermometer predicting tomorrow’s temperature by saying “about like today”: right across all 365 days, and useless at telling you when to bring a coat.

The models, unlike either reference, re-price every game every day from live sales and claims — how many tickets have actually sold, which prizes have actually been claimed, how fast, and what is left. They deliberately deviate from the published payout rate — penalizing games down to their last top prize, adjusting for where a game is in its life, rewarding real access to big prizes. Every deviation costs a little whole-field correlation and buys something at the top of the list, which is the only part of the list a player actually uses. Launch day shows the trade exactly: the payout-rate reference’s #1 was that $1 Instant Take 5 ticket; Smart Score’s and Apex’s top picks were $30 games with $10 million top prizes still in play. The thermometer ranked the deck well and then handed you a $1 ticket.

So read the board this way: top-5 realized payout and top-5 big wins are the player metrics — what following the list actually returned. IC is the engineering metric — how much of the whole field a model understands. A model that wins on all of them, as Apex does, is getting better in a way that should hold up. A reference that wins only on IC is measuring something true but not something you can spend.

What Apex is, and why it exists

When we reviewed six months of outcomes, we found that Smart Score had been tuned to keep its rankings steady from day to day rather than to predict payouts — and as a result it was leaning on some signals that, in hindsight, did not predict anything. Apex is the direct answer. It looks at the same live game data, but trusts each signal only as much as that signal’s track record since March at predicting two things at once: realized payout and $5,000+ wins. Then it leans on the game’s published payout rate for half of its score, because that number turned out to be very hard to beat — and the honest way to beat it is to start from it and add only what has proven to help.

That last point is worth a confession. When we first built Apex this week, it leaned on the payout rate for only 30% of its score, and in testing it beat the payout-rate reference on return by a statistically unconvincing two and a half cents while losing to it on big wins. Raising the payout-rate share to 50% fixed both: the gap on return grew to five and a half cents and Apex moved ahead on big wins. We made that change before publishing this post, and the numbers above reflect it. One caveat we owe you: that 50% was chosen from three values tested on the same days, so the exact margin is probably a little optimistic. The direction — trust the published number more, not less — is not in doubt.

Apex was chosen from ten candidate designs by testing each on months of data it had never seen; several runners-up looked better on paper and worse in practice. It is available now as a third option on the rankings page. Smart Score and Verity are unchanged.

What we got wrong along the way

Benchmarks are only useful if you publish the misses. Two from this round:

Rules of the board

The scores are re-graded every day as new outcomes come in, and the rankings page always shows you the current pick from each model. If a new model cannot beat what is already there, it will not be offered.

Frequently asked questions

What is ScratchBench?

A public, daily-updated benchmark that grades every ScratchOffsNY ranking model against realized outcomes: prize money actually claimed per dollar of tickets actually sold, and $5,000+ prizes actually hit, over the 21 days after each ranking was published.

Which model should I use?

Apex, on today’s numbers: it returned the most on its top five (80.7¢ per dollar), hit the most $5,000+ prizes (45.9 per million tickets), and ordered the whole field best. Smart Score is a close second on return (78.9¢) but well behind on big wins. Verity trails on return. All three are one click apart on the rankings page, and the live board will tell you if this changes.

How is “realized payout” different from the odds on the ticket?

The odds on the ticket are the game’s printed design. Realized payout is measured after the fact: the prize value New Yorkers actually claimed divided by the dollars they actually spent, per game, from official claim and sales data. It captures how the remaining prize pool actually paid out in the real world.

Why not just buy the best odds on the ticket, or the highest payout rate?

Both sit on the board as reference lines, and neither accounts for sales or changing odds. Printed odds are fixed the day a game is printed; the rule amounts to “buy $30 tickets,” which returned as much as Apex (80.7¢) but hit big prizes at barely half the rate (28.2 vs 45.9 per million) because it cannot see which $30 games still have prizes. Sorting by remaining payout % does update, but it cannot tell a great return from a great buy and returned less (75.2¢). The models re-price every game daily from live sales and claims. Apex clears both references on every metric; that is the whole point of it.

What does “top-5 realized payout” mean, and why is it under 100¢?

Take a model’s five highest-ranked games on a given day. Over the next 21 days, divide the prize money actually claimed on them by the money actually spent on them. Every scratch-off pays back less than a dollar by design — a random New York ticket returned about 66¢ per dollar in our test — so read the column as distance above 66. Apex’s picks returned 22% more than a random ticket; Smart Score’s 19% more.

What does IC mean?

Information coefficient: the rank correlation between a model’s ordering of all active games and their realized payout over the next 21 days. 1.0 is perfect, 0 is a coin flip, above 0.3 is strong. It grades the whole field, not just the top picks — which is why the payout-rate reference scores well on IC while its top-5 picks return less than Apex’s.

Related reading

Outcome data: official New York Lottery prize claim reports and ticket sales. Numbers in this post reflect September 2, 2026; the live leaderboard supersedes them.

S
ScratchOffsNY Data Team
Model research and validation

The team builds and grades the Smart Score, Verity and Apex models daily from official NY Lottery prize reports and sales data. Every model is scored publicly on ScratchBench.