Draft night is an AI problem
What the engine does today, what I have tried to replace it with, and what came of that. A line below marks where the shipped product stops.
What ships today
NineDraft is a deterministic engine — a calculator, and I mean that as a compliment. One person built it, it is live in beta, and it has been used in a real draft. What it does:
- Before the draft it prices every player by points above what you could have had instead for nothing, not by projected points. That is why a running back projected for fewer points than a quarterback can be worth several times as much at auction. (At quarterback, kicker and defense, inside the league shapes that measurement covers, the "for nothing" bar is the better of two things: the last drafted starter, and what streaming the waiver wire is really worth, measured off real football history. Outside those shapes, and at every other position, it is still the blend it started as — the last starter at the position and the last player worth rostering. I measured tight end against the wire and the reading was too close to call. Running back and receiver I have not measured against it at all.)
- During the auction it watches every sale, measures how hard the room is paying position by position, and moves what it expects each remaining player to cost.
- It tracks the whole room — every team's remaining budget, roster and needs. When the live feed is not faithful enough to trust, it suppresses opponent-specific claims instead of guessing at them.
- With a player on the block it gives a max bid or a pass, weighing what your roster still needs, what you can afford, and which comparable players are left to pivot to.
- Every recommendation carries its reason in plain English. You know things the engine does not, and one glance is enough to overrule it.
- When the data goes thin — a broken feed, a signal it cannot measure — it says so and labels the estimate.
Same inputs, same advice, every time. No learning loop, no black box. (The valuation math underneath keeps moving: the 2026 changes so far are measured calibration work, one piece at a time, and the bar above is one of them.)
Below this line is research, not the product.
What I have tried, what it measured, and what is still only planned — none of it in NineDraft. The shipped engine appears below only as the thing being compared against.
Why draft night is an AI problem
Nobody in an auction announces their maximum bid. There is a clock, and an auctioneer counting down is a bad place to do careful arithmetic. You are making a sequence of decisions against people who each have their own budget and their own temper, and every dollar you spend on one player is a dollar that cannot chase another.
A draft is also small, which is what makes it worth pointing a learning agent at: one league, a fixed number of roster spots, a few hundred players worth bidding on. A simulator plays a complete draft, nomination by nomination, in a fraction of a second, so a research machine can play millions of them. Games with that shape have a known playbook, and I have spent the past year trying to apply it.
What I have tried, and what it did
That attempt runs in a separate lab, and the notes are public: The judge comes first. Nothing from it is in the product. The strongest drafter in the lab still runs on hand-written valuation with simulation at decision time, and no learned agent has cleared the bar to replace it — some have won individual fields, none has been promoted.
The obvious recipe is self-play: leagues of AI agents drafting against each other millions of times, finding nomination bluffs and budget traps nobody wrote down for them. Self-play has a known failure mode, agents that are excellent against copies of themselves and strange against people, and in this game it showed up as a collapsed market. Seat ten copies of the strongest brain at one table and nobody bids: across a thirty-draft probe it deployed 31% of its budget against itself, while still filling a legal starting lineup 97.7% of the time. So the plan trains against a population of realistic drafter types instead — rankings followers, people who spend early and hard, bargain hunters, and the manager who drafts straight down a printed list. That same probe left one thing I never settled: with every seat playing the identical strategy, the best seat reached a mean win probability of 0.164 and the worst 0.061, but the permutation test came back p=0.079 on thirty drafts a seat, so it stays a suspicion.
Small networks distilled from the search all refused to buy quarterbacks, and so did the search itself when I gave it far more thinking time. I replayed recorded drafts with the search instrumented and found the cause in its model of how a draft finishes: that model assumes teams buy starters only, so the moment league-wide starter demand at a position is met it prices what is left at replacement, about a dollar, while a simulated room full of managers hoarding backups goes on clearing them near forty-eight. The candidate wanted quarterbacks and was willing to pay for them. It lost the ones it valued early, then declined the rest all draft. A conservative patch cut the failure roughly in half on the same block the defect was diagnosed on, which is the only block it was measured on, and it still left the candidate short of the promotion bar. Most of the work around those small networks turned out to be regularization, and I got there late — my first read of the regressions was that the network was too small, so I made it bigger, which was exactly backwards.
What is still planned
Today's engine plans against one expected future: this player will probably cost about that much. The planned replacement runs thousands of simulated futures — price swings and the runs on a tier that follow them — finishes your roster in each one, and reads the distribution. Advice would stop being "this price seems fair" and start being "bidding here wins in 68 of 100 simulated futures; passing wins in 79", with illustrative numbers there, because nothing produces real ones yet.
A draft only matters through the season it produces, so the plan scores a finished roster with a season simulator — injuries, byes, lineup calls, the waiver wire — over many simulated seasons, instead of scoring it on how it looks the night it was drafted. The lab already made that swap in its own scoring, for a mundane reason: one sampled season was noisy enough that two identical rosters could score a win and a loss.
At decision time the pattern would be the AlphaZero one. A trained model supplies the first instinct and a bounded search over simulated futures checks it, thinking longer exactly when the call is close.
The rule that never bends
No learned strategy gets near a real draft until it can explain itself in plain English. "Bid — 62% confidence" is a shrug with decimals. Every call the shipped engine makes already arrives with its reason, and that requirement carries into anything learned: if the model cannot say why — the tier is drying up, two opponents still need the position and can outbid you later — then it does not get to advise.
The same applies to the numbers underneath. Anything the system leans on is measured against real football history, and where the measurement does not exist the engine says the estimate is an estimate. The waiver-wire bar above is the working example: measured at three positions, unmeasured at two, and the page says which.
Where this stands
The whole plan rests on one unglamorous thing: the simulator has to be realistic and the scoring has to be trustworthy, or a strong model learns the wrong lesson efficiently. Most of the lab's past year went into that, and not into training runs. I rebuilt the arena in August around absolute measurement: a candidate and the incumbent draft the same scenario in separate rooms, and the difference in the rosters they build is the reading. I commit every evaluation to that repository, inconclusive ones included.
So, plainly: today there is a calculator I trust, and it is the product. The agent is a research project with a public record, and nothing in it has yet been promoted over the hand-built drafter. I will update this page when something is.