Evidence

Where our simulated field misses

When you run a tournament simulation here, you are playing against a field we generated. This is that field measured against 329 real DraftKings contests, on a standard we wrote down before the run. Three measures come back clearly wrong and two are things we do not model at all. They are on this page because you should know them before you trust a simulated finish.

6of 18 measures calibrated
7drifting
3miscalibrated
2not modelled at all

The standard, written first

It is easy to publish a scorecard once you know you passed. So the grading was fixed before any field existed. For each measure we take the gap between our field and the real one and divide it by how much that measure naturally varies from contest to contest. Being $200 off on salary means something different from being 20 ownership points off, and dividing by the real spread puts them on one scale.

Half a standard deviation or less is calibrated. Up to one and a half is drifting. Beyond that is miscalibrated. A thing our generator cannot produce at all is reported as not modelled, which is worse than a wrong number, and is deliberately not allowed to hide as one.

The generator is fed each contest's own salaries and projections, so what is being graded is the field model and the ownership model, not whether we can project football.

Where it is wrong

Three measures are miscalibrated.

The salary gap is the one to understand, because the number looks small and the grade looks severe. Our field spends a couple of hundred dollars less than a real one, but real fields are remarkably consistent about spending nearly all of the cap, so a small dollar gap is a large gap in the units that matter. A field that underspends is a slightly weaker opponent than the one you will actually face, and the same is true of it projecting fewer points.

Two things we do not model.

Every lineup our field generates is unique. In a real contest the same lineup gets entered over and over, which splits prizes between the people holding it. Our own study of 329 contests found duplication is one of the few things that genuinely hurts at the top of a tournament, so a field without it is missing a real effect rather than a cosmetic one.

Ownership, player by player

Ownership is the part we measure best and it is still not perfect. Across the same 329 contests, our average error on a single player is 0.78 points. The chalkiest player on a slate we get right on average — ours 46.0%, real 45.1% — but that average hides a 7.6 point typical error on which player it is and how high.

The clearest way to see the shape of the problem:

Slates with a player owned above…RealOurs
70%429
80%014

In 329 real contests, no player was ever owned above 80%. Our field produced one on 14 slates. We over-concentrate at the very top, which makes the simulated field look more chalk-locked than the real one.

The full card

MeasureRealOursGap (SD)Verdict
Cumulative ownership of a lineup139%113%-1.00drifting
Spread of cumulative ownership37%39%+0.21calibrated
The most contrarian lineups (5th pct)77%53%-1.48drifting
The chalkiest lineups (95th pct)199%178%-0.60drifting
Salary spent$49,883$49,657-7.97miscalibrated
Projected points of a lineup124.71115.47-1.94miscalibrated
Games represented in a lineup5.695.73+0.17calibrated
Most players taken from one game3.163.07-0.51drifting
Ownership of a lineup's most-owned player37%35%-0.24calibrated
How many entries share the same lineup2.041.00-0.13not modelled
Lineups using a bring-back48%42%-0.79drifting
Lineups within $200 of the cap86%95%+2.23miscalibrated
Entries that are a one-of-a-kind lineup95%100%+1.37not modelled
Lineups with no stack12%13%+0.20calibrated
Lineups stacking 1 teammate42%45%+0.70drifting
Lineups stacking 2 teammates39%35%-0.63drifting
Lineups stacking 3 teammates7%7%+0.13calibrated
Lineups stacking 4 or more1%1%-0.28calibrated

What this does not tell you

The archive carries no Vegas totals, so every game in this run was handed to the model as a neutral 45-point game. Anything in our ownership or field construction that leans on game environment was running blind, and any gap that causes is ours to own rather than the archive's to explain.

This also grades the field, not the outcome. It says how closely our opponents resemble real opponents. It does not say you will win.

What would change these numbers

The salary and projection gaps point the same way and probably have one cause: our field is more spread out than a real one, so it wanders further from the obvious, expensive, high-projection lineups that most entrants actually submit. Concentrating it is the open work. Duplication is a separate build, not a tuning problem.

When we re-run this, the new numbers replace these ones on this page, including if they are worse. That is the only version of a scorecard worth reading.