Two lessons ago we earned a beautiful number: reject the first 37%, then leap at the first option that beats everything you saw. It felt like the final word. It is not. It is the answer to one very specific, slightly deranged question — and that question is almost never the one you’re actually asking.
Here’s the buried assumption. The 37% rule maximises the probability of getting the single best option in the whole pool. Not a great option. Not a top-five option. The best. The literal maximum. And it scores your outcome like a psychopath: if you land the #1 candidate, you win; if you land the #2 candidate — a hair behind, brilliant, would-have-been-perfect — you lose exactly as hard as if you’d hired the worst person in the stack. Second place and last place are the same score: zero.
That is a bizarre way to value a hire, a house, or a partner. This lesson swaps in the objective most decisions actually have — the best expected value (expected quality, or expected rank) of your pick — and watches everything change: the optimal look shrinks below 37%, and the bar you demand falls as your runway runs out. If you remember one upgrade to the whole model, make it this one. Picking the right objective matters more than any formula.
Before you read — take a guess
The classic '37% rule' is proven optimal for one specific goal. Which goal?
The 37% rule optimises a harsh, all-or-nothing goal
Imagine a talent show where the grand prize is a golden trophy for first place only, and every other finisher — second through hundredth — goes home with nothing. In that world, a contestant who finishes 98th-percentile-brilliant but not #1 has failed. Their reward is identical to the person who finished dead last. That’s the world the 37% rule is built for. It’s an all-or-nothing objective (also called a best-only or win-or-lose objective): the payoff is 1 if you grab the single best and 0 otherwise.
Formally, the classic secretary problem uses what’s called a relative-rank or cardinal 0/1 payoff: you only care about being right about the top. It does not care how good the person you hired was — only whether they were The One.
Now say that out loud about hiring. You interview 100 people, you hire a candidate who’s in the top 2% of everyone you saw — a genuinely excellent engineer — but unbeknownst to you there was one person marginally stronger back in the stack. Under the best-only objective, you lost. Your outcome is scored the same as if you’d hired the least qualified applicant in the entire pool.
Nobody runs a company like that. A 98th-percentile hire is not “a failure identical to the worst hire.” It’s nearly as good as perfect, and in the real world you’re thrilled. The harshness is the whole problem: the best-only objective throws away all the information in “how close did I get,” and that information is exactly what you care about.
The single most common misuse of the whole model
People memorise “look at 37%, then leap” and apply it to everything — hiring, apartments, dating. But that rule is optimal only when second place is genuinely worthless. For almost every real decision, near-misses are valuable, and the 37% rule is solving the wrong problem. Getting the objective wrong dwarfs any error in the arithmetic.
Switching to expected value changes the optimum
So let’s fix the objective. Instead of “probability of the single best,” we want to maximise the expected value of whatever we end up with — the expected quality, or equivalently minimise the expected rank, of our pick. Now a 98th-percentile option is worth about 98, a 50th-percentile option is worth about 50, and near-misses get the credit they deserve.
Watch what this does to strategy. Under the best-only rule, holding out is cheap: passing on a merely-excellent option costs you nothing if it wasn’t literally #1, because excellent and terrible score the same anyway. Under the expected-value rule, holding out is expensive: every time you reject a 95th-percentile option hoping for a 99th, you’re gambling a near-certain 95 against a small chance of a slightly-better number — and risking ending up with far less. Perfectionism has a price now, and the price is steep precisely because excellent is almost as good as perfect.
The consequence: the optimal look fraction shrinks below 37%. When you’re chasing expected value with known, uniformly-distributed option values, the math (this is the flavour of the Chow–Robbins and cardinal-payoff results) says look at far fewer than 37% before you start accepting — the optimal sampling region drops toward the sqrt(N) neighbourhood rather than N/e. With 30 options, best-only says look at about 11; the expected-value objective says start being willing to commit after just a handful. Be pickier less. Commit sooner.
The intuition in one line: when you only get credit for perfection, it pays to wait for perfection; when you get credit for excellence, waiting for perfection means turning down excellence, which is a bad trade.
This is the centrepiece — go play with it. Start on Maximise expected value and read off where the optimal look sits. Then flip to Pick the single best and watch the sweet spot leap back out to ~37%.
Optimal-stopping lab
Chasing the best vs chasing high expected value
Options arrive one at a time; accept or reject each on the spot, no going back. The rule: look at a fraction without committing, remember the best, then leap at the first that beats it. Drag the look fraction and watch how often the rule wins.
Objective
- Optimal look
- 13%
- Score here
- 85 pctl
- 1/e ≈ 37%
- 0.368
With 30 candidates and the "Maximise expected value" goal, looking at the first 20% then leaping scores 85 pctl. The optimal look is about 13% — close to the theoretical 1/e ≈ 37%.
Watch one sequence
✗ missed the best — the true best was #6
You switch your goal from 'get the single best' to 'maximise the expected quality of my pick.' What happens to how long you should keep looking before you start accepting — and why?
The threshold FALLS as options run out
Now the deep part, and the piece almost everyone forgets. The truly optimal expected-value policy is not “one fixed bar.” It’s a reservation value — a minimum acceptable quality — that decreases with every option that goes by. Early, with a long runway ahead, you can afford to be demanding: only a spectacular option is worth stopping for, because if you pass, you still have loads of chances. Late, with the exit door closing, you must lower your standards: a merely-good option now beats gambling on the two coin-flips you have left. On the very last option, your reservation value hits bottom — you accept anything, because the alternative is walking away empty-handed.
This is the same shape as the house-selling problem (or the job-search / secretary problem with cardinal payoffs): offers arrive one at a time, and your walk-away price should drift down as the selling season runs out. A homeowner who demands their dream price in December, with the market closing, is making the classic error.
The reason the bar moves is pure expected-value bookkeeping. Your reservation value at any moment equals what you expect to get if you decline and play optimally from here. With many options left, that continuation value is high (you’ll probably see something great), so your bar is high. With few left, the continuation value is low (slim pickings ahead), so your bar drops to match. The bar is literally “the value of your remaining options” — and that shrinks toward zero as the options run out.
Here is a concrete descending-threshold table for a pool of 10 options whose values are drawn uniformly at random as percentiles (0–100). “Options remaining” counts the one in front of you plus those still to come. The rule: accept the current option only if its percentile clears the bar.
| Options remaining (incl. current) | Minimum acceptable percentile | Behaviour in plain English |
|---|---|---|
| 10 | 88 | Only a top-tier option is worth stopping — tons of runway left |
| 9 | 86 | Still very demanding |
| 8 | 84 | Demanding |
| 7 | 82 | Slightly relaxing |
| 6 | 78 | ”Very good” now clears the bar |
| 5 | 74 | Willing to take a strong-but-not-elite option |
| 4 | 68 | Getting nervous; a good option looks fine |
| 3 | 61 | Above-average is now acceptable |
| 2 | 50 | Take anything better than a coin flip |
| 1 | 0 | Last option — accept it no matter what |
Read it top to bottom: the bar starts high and marches down. (The exact numbers depend on the value distribution; the shape — monotonically decreasing, hitting rock bottom on the last option — is universal for expected-value optimal stopping.)
The one-sentence version of the whole policy
Set your bar to the value of your remaining options — and since that value falls every time an option passes, your bar must fall with it. Demand brilliance early, accept ‘good enough’ late, and take literally anything on the last move.
Worked example: the same stream, two objectives
Let’s make it painfully concrete. Ten candidates arrive one at a time. Nobody can be recalled. Here are their true quality percentiles in arrival order (you don’t know future ones, only the current and past):
| Position | Arriving percentile | Options remaining | Bar (from table) | Accept? |
|---|---|---|---|---|
| 1 | 71 | 10 | 88 | No — 71 < 88 |
| 2 | 55 | 9 | 86 | No |
| 3 | 90 | 8 | 84 | Yes — 90 ≥ 84 → STOP, expected-value pick = 90 |
| 4 | 42 | 7 | 82 | (never reached) |
| 5 | 96 | 6 | 78 | (never reached) |
| 6 | 63 | 5 | 74 | (never reached) |
| 7 | 30 | 4 | 68 | (never reached) |
| 8 | 88 | 3 | 61 | (never reached) |
| 9 | 12 | 2 | 50 | (never reached) |
| 10 | 47 | 1 | 0 | (never reached) |
The expected-value rule stops at position 3 and banks a 90th-percentile pick. Excellent. Not the true maximum (that was the 96 at position 5), but the policy doesn’t know the future and the 90 was well above its bar. A very happy outcome.
Now run the best-only 37% rule on the same stream. Look at the first 37% — positions 1–3 (round 3.7 down to 3) — reject them all, but remember the best seen: the 90 at position 3. Then accept the first later option that beats 90.
- Position 4: 42 — below 90, reject.
- Position 5: 96 — beats 90! Accept → best-only pick = 96.
Interesting. On this particular stream, the best-only rule actually did better (96 vs 90), because it happened to catch the true max. But notice the cost structure: the best-only rule was willing to blow past the excellent 90 and risk running to the end on the hope of something better. Reshuffle the deck — put the 96 earlier, or leave nothing good after position 3 — and the best-only rule frequently walks away with scraps (forced to accept whatever’s last), because it never lowers its bar. Across many random streams, the expected-value rule wins on average quality and rarely gets skunked; the best-only rule wins on frequency of catching the exact max and occasionally faceplants into the last option.
That trade — reliable excellence vs. occasional perfection at the risk of disaster — is the whole lesson in one worked example.
On the stream above, the expected-value rule took the 90 at position 3 and the best-only rule held out and caught the 96 at position 5. A colleague concludes: 'See? Holding out for the max is just better.' What's the trap in that reasoning?
You're using the descending-threshold rule on 10 options (the table above). You've reached the 9th option — 2 options remaining, bar = 50th percentile. The current option is a 48th-percentile 'meh.' What does the rule say, and what's the risk if you reject it?
Which objective is yours?
Everything above hinges on one question you must answer before you pick a rule: is second place worthless, or is a near-miss almost as good? Get this wrong and no amount of clever stopping math will save you.
Use the strict best-only rule (look ~37%, then leap, never lower the bar) only when the prize is genuinely winner-take-all and near-misses are worth nothing:
- A one-of-a-kind auction lot or a unique collectible — you either own the item or you don’t.
- A sole grand prize in a contest with no consolation value.
- A single irreplaceable slot where the runner-up literally cannot be used.
Use the expected-value / descending-threshold rule for almost everything else, because in almost every real decision a strong runner-up is a great outcome:
- Hiring — a top-3% engineer is a fantastic hire, not a failure.
- Housing / apartments — the second-best flat you can actually get beats the dream flat you can’t.
- Dating / partners — “not the single most compatible human alive” is a wildly wrong bar for a happy relationship.
- Selling a house or a car — a very good offer today beats gambling the season on a slightly better one that may never come.
Sort each decision by which objective actually fits it — is a near-miss almost as good (expected value), or is only #1 worth anything (best-only)?
Place each item in the right group.
- Accepting offers on the house you're selling before the season ends
- Securing the last remaining unit of an irreplaceable, sold-out item
- Bidding to win a single unique painting at auction — you own it or you don't
- Landing the one grand-prize slot in a winner-take-all contest with no runner-up reward
- Choosing an apartment to rent this month
- Finding a life partner
- Hiring a software engineer from a pool of applicants
Match each term from the expected-value view of optimal stopping to its meaning.
Pick a term, then click its definition.
Pulling it together
The 37% rule isn’t wrong — it’s a precise answer to a question most people aren’t asking. It optimises the probability of the single best, a brutal all-or-nothing target that treats a 98th-percentile pick as a failure identical to the worst. Swap in the objective real life almost always has — the best expected value of your choice — and two things move: the optimal look shrinks below 37% (you calibrate fast and commit sooner because holding out for perfection means passing up excellence), and your acceptance bar falls as options run out (demand brilliance with a long runway, accept “good enough” as the exit nears, take anything on the last move).
Three failure modes to burn in: (1) assuming “37%, then leap” is universal — it’s optimal only for the single-best objective; (2) being too picky when near-misses are fine — perfectionism is expensive once excellent nearly equals perfect; and (3) the classic tragedy — forgetting to lower your bar as the runway shortens, holding out for the dream, and ending up with scraps.
Next lesson: what happens when the assumptions crack — when you can recall a rejected option, when options can say no back to you, and when you don’t know how many there are.
The upgrade to remember
Before reaching for any stopping rule, ask one question: is a near-miss almost as good, or is only #1 worth anything? For the rare winner-take-all case, use strict 37%-then-leap. For hiring, housing, dating, and selling — basically everything — chase expected value with a bar that starts high and descends: demand brilliance early, settle for good late, and never, ever hold out so long that you’re forced to take the very last thing in the pile.