Skip to content
Mental Models

Optimal Stopping & the Secretary Problem

Explore, Then Commit

The 37% rule is one instance of a deeper pattern — explore-then-commit — where you spend an opening phase gathering information only to calibrate a bar, then leap at the first option that clears it.

12 min Updated Jul 12, 2026

Last lesson we nailed a specific number: reject the first 37%, remember the best you saw, then take the first later option that beats it. Clean, exact, a little magical. But if you walk away with only the number, you’ve pocketed the answer and thrown away the idea. The number is a special case. The idea is a shape that fits an enormous range of decisions, most of which have nothing to do with 37%.

That shape is explore-then-commit (also called look-then-leap): split a sequential, no-going-back decision into two phases — an explore phase where you gather information and commit to nothing, and a commit phase where you take the first option that clears a bar. The whole trick, the thing that feels backwards until it clicks, is what the explore phase is for. You are not shopping during it. You are building a ruler.

This lesson pulls that ruler out into the open. We’ll see why every rule in this family has exactly two phases with one bar between them, why “explore vs. commit” is a genuine trade-off sharpened by irreversibility, why the bar has to come from the sample rather than from your imagination, why (preview of Lesson 3) the bar should actually move as your runway shrinks, and how to read messy real decisions as explore-then-commit on sight.

Before you read — take a guess

You're going to explore-then-commit on a decision — say, hiring. What is the true job of the 'explore' phase (the opening candidates you deliberately reject)?

Two phases, one bar

Here’s the reframe, and it’s the whole lesson in one breath. Every optimal-stopping rule in this family is the same machine: an explore phase, then a commit phase, with a single bar sitting between them. The bar is the best option you saw while exploring. The commit phase does exactly one thing: take the first option that beats the bar.

The analogy that makes it stick: imagine you’re a wine judge at a competition, tasting one glass at a time, and you must crown a winner the instant you taste it — no going back for a re-taste. If you crown glass one, you’re crowning it against nothing. So the first stretch of glasses, you swirl, sip, spit, and rank in your head but award nothing. You’re calibrating your palate to this competition’s range. Then, the moment a glass beats every one you’ve tasted so far, you slam down the gavel. The tasting was never about those early glasses. It was about learning what “excellent here” tastes like so you’d know it on arrival.

The counterintuitive part is worth saying slowly: the sample is not a shortlist. Options in the explore phase are not candidates to be chosen from — even if one is obviously brilliant. Their entire purpose is to set the number that the commit phase measures against. An option in the explore phase is a ruler tick, not a prize.

Worked example: the two phases in numbers

You’ll interview 12 applicants, one at a time, decide each on the spot, no recall. Using a look fraction near 1/e, your explore phase is the first 4 applicants (about a third of 12). Say their true quality scores turn out to be:

PhaseApplicantScoreWhat you do
Explore16.1Reject — recording only
Explore28.3Reject — but this is now the bar
Explore35.0Reject — recording only
Explore47.2Reject — recording only
Commit57.9Reject — below the 8.3 bar
Commit68.0Reject — still below 8.3
Commit78.6HIRE — first to clear the bar

Applicant 2 scored 8.3 and you rejected them — that stings, but it’s the point. Their 8.3 became your bar. You then held firm through two decent-but-not-better people (7.9, 8.0) and leapt on applicant 7 the instant they cleared 8.3. The explore phase built the ruler; applicant 7 was the first thing tall enough to trip it.

Warning:

The rigidity trap

Notice you rejected an 8.3 during exploration and it turned out nobody in the commit phase was much better. In the pure secretary problem that’s correct — you can’t recall applicant 2. But real life is often kinder: sometimes you can go back. Blindly refusing to ever pick from your sample, when recall is actually available, is a rigidity trap. Lesson 4 handles “what if you can recall?” head-on. For now, hold the pure rule — but don’t mistake the model for a law of physics.

Feel the two phases in the lab

Set the lab to 25 options with a 37% look fraction. Deal a sequence and watch the structure the color-coding makes obvious: the grey cards are the explore phase — thrown away on purpose to set the bar — and the rule sits on its hands, committing to nothing, until a later card beats the best grey one. Then it leaps. Deal a dozen fresh sequences in a row and you’ll see the pattern: a long, disciplined look, then a sudden pounce.

Optimal-stopping lab

Explore to set the bar, then leap

Options arrive one at a time; accept or reject each on the spot, no going back. The rule: look at a fraction without committing, remember the best, then leap at the first that beats it. Drag the look fraction and watch how often the rule wins.

Objective

37%P(get the best)Look fraction →0%
Optimal look
36%
Score here
38%
1/e ≈ 37%
0.368

With 25 candidates and the "Pick the single best" goal, looking at the first 37% then leaping scores 38%. The optimal look is about 36% — close to the theoretical 1/e ≈ 37%.

Watch one sequence

lookleapthe true best was #3

✗ missed the bestthe true best was #3

Deal sequences and watch the two phases: the grey (explore) cards are sacrificed to calibrate the bar, then the rule leaps at the first later card that clears it. Drag the look fraction to 0% (leap blind, no ruler) and to 90% (explore so long you've burned every option) to feel why the sweet spot lives in the middle.

Now abuse the slider to feel the trade-off. Drag the look fraction near 0% — the rule has no ruler, so it grabs the first card and usually gets a mediocre one. Drag it near 90% — the rule explores almost the whole deck, sets a sky-high bar, then discovers nothing is left to clear it and gets stuck with the last card. Neither extreme works. The middle wins because that’s where you’ve looked enough to know the scale but not so long that you’ve burned your options. That tension has a name.

The explore/exploit trade-off under irreversibility

Every sequential decision under uncertainty pits two moves against each other, and it’s worth naming them like the old friends they are. Explore means keep looking — buy more information, learn the landscape, and accept the risk that a great option you passed is now gone. Exploit means commit to what you have — cash in now, and accept the risk that something better was still coming. This is the explore/exploit trade-off, and it shows up everywhere from clinical trials to which restaurant to revisit.

What makes optimal stopping its own beast is the setting: explore/exploit meets no-recall irreversibility. Every option you pass is burned. So the cost of exploring isn’t merely “time” — it’s that the very thing you were sizing up might be the one you can never get back. Exploration here is expensive in a way it usually isn’t, which is exactly why the math cares so much about when to stop.

Contrast this with explore/exploit when you can revisit — the world of the multi-armed bandit (imagine a row of slot machines; each pull tells you a little more about which arm pays best, and you can return to any arm any time). There, exploring is cheap: pull a lever, learn something, and if a different lever looks better you just switch back — nothing is destroyed. Because nothing is burned, the stakes of a “wasted” explore are gentle; you pay a small opportunity cost and move on. Optimal stopping strips that safety net away. You get one pass, one shot per option, and a closing door behind you.

Why is the explore/exploit trade-off HARSHER in an optimal-stopping problem than in a multi-armed bandit?

Why the bar comes from the sample, not thin air

Suppose you announced your bar before looking: “I’ll hire the first person who’s an 8 out of 10.” Sounds decisive. It’s also nonsense — an 8 out of 10 on what scale? You don’t know if this applicant pool tops out at 6 (in which case you’ll wait forever and hire nobody) or routinely hits 9 (in which case your 8 is mediocre and you’ll settle too soon). A bar plucked from imagination is unanchored. You’re measuring with a ruler whose tick marks you invented.

This is the deep reason the explore phase exists: it is literally the act of learning the top of the distribution so your bar sits on real ground. “Best of the first third” is not a guess — it’s an estimate of what excellent looks like here, extracted from actual data. The sample’s whole gift is a grounded threshold. That’s why you throw the sample away: its value was never the options in it; it was the number they taught you.

Worked example: two candidates, two calibrations

Two hiring managers each interview the same 15 people in the same order, both using explore-then-commit with a 5-person sample.

  • Manager A’s first 5 happen to include a genuine star (a 9.1). Her bar is set high at 9.1. She holds out — and because that was a fluke-strong sample, nobody later clears it, so she’s forced onto the last candidate. Calibrated a touch too high by luck.
  • Manager B’s first 5 are all middling (top score 6.8). Her bar is a modest 6.8. She leaps on candidate 8, a solid 7.4 — a good, early, confident hire.

Same rule, same people, different samples — and that’s fine. The point isn’t that the sample is always perfectly calibrated (it isn’t; it’s a random draw). The point is that a bar from the sample is anchored to reality and gets the best option surprisingly often, whereas a bar from your imagination is anchored to nothing and fails predictably. A slightly noisy real ruler beats a confident fake one.

Tip:

Say it back

The explore phase answers a question you literally cannot answer any other way: how good is good, here? You can’t demand an 8 if you’ve never seen the scale. Look first — the top of your sample is your 8.

The bar isn’t fixed — it depends on how much looking remains

Here’s a subtlety that we fully unpack next lesson but that you should meet now, because it corrects the biggest misread of explore-then-commit: the bar is not a single fixed number you hold for the entire commit phase. Or rather — in the pure secretary problem it happens to look fixed (beat the best-of-sample), but the real optimal rule is a threshold that depends on time remaining.

The intuition is pure common sense once stated. Early in your commit phase, with lots of runway left, you can afford to be picky — if you pass a good option, plenty more chances are coming. Late in the commit phase, with only a couple of options left, you must relax the bar — holding out for excellence when the deck is nearly empty is how you get stranded with the dregs. So the optimal policy is really a falling bar: demand a lot when you have time, demand less as the door closes.

Warning:

The 'fixed bar forever' trap

The most common way people break explore-then-commit in real life is treating the bar as sacred and permanent. You calibrated “great” in month one of a six-month house hunt — and by month five, with two weekends left before you must move, you’re still rejecting perfectly good flats for failing a bar you set when you had infinite runway. A bar that ignores your shrinking time is a bar that will strand you. Optimal stopping is a time-dependent threshold rule, and Lesson 3 makes the falling bar precise.

Frame it this way and the whole family snaps into focus: explore-then-commit isn’t “look at 37%, then apply a frozen standard.” It’s “spend an opening phase learning the scale, then apply a standard that starts high and eases down as your options run out.” The 37% rule is the clean, best-only special case of that far more general idea.

You're five months into a six-month apartment search with a hard move-out date. A very good flat appears — not quite the 'best' bar you calibrated back in month one. Optimal-stopping logic says:

Reading a decision as explore-then-commit

The real skill this lesson buys you is a lens: meet a messy sequential decision and immediately split it into its explore phase (gather info, commit to nothing, set the bar) and its commit phase (take the first option over the bar). Here are four, decomposed.

DecisionExplore phase (calibrate the bar)Commit phase (leap over the bar)
Hiring from sequential interviewsInterview an opening batch you won’t hire from; the best of them sets “strong candidate”Offer the job to the first later candidate who beats that best-of-sample
Choosing a contractor from rolling quotesCollect several quotes as they come in, awarding none; note the best value-for-scope so farSign with the first new quote that clearly beats your best-so-far on price and scope
Picking a grad school from rolling offersField the early offers with hard deadlines, declining them, to learn the tier of program that says yes to youAccept the first later offer that tops the best of the ones you already turned down
Buying a used car as listings appearWatch listings for a set window without buying; the best price-condition combo you see sets the barBuy the first new listing that beats that bar — before it, inevitably, sells to someone else

Every row is the same machine. An opening stretch spent learning the scale and committing to nothing, then a disciplined first-past-the-bar leap. What changes is only the units of the bar (skill, dollars, program tier, price-condition) and the length of the explore phase. The one thing they share — and the thing your gut fights — is that you deliberately walk away from good options early so that “good” means something later.

Sort each action into the phase it belongs to in an explore-then-commit strategy.

Place each item in the right group.

  • Interview the first four candidates with no intention of hiring any of them
  • Watch used-car listings for two weeks without buying, tracking the best price-condition combo
  • Note that the strongest of the opening candidates scored an 8.3 and make that your bar
  • Buy the first new listing that clears the bar you set during your watching window
  • Accept the first later grad-school offer that tops the ones you already declined
  • Sign with the first fresh quote that clearly beats your best-so-far on price and scope
  • Offer the job to the first later candidate who beats the best of your opening batch
  • Collect several contractor quotes as they arrive, awarding the job to none of them yet

Explore is not procrastinate

One more distinction, because it’s the difference between using this model and abusing it. Exploring is not procrastinating. They can look identical from the outside — both involve “not deciding yet” — but they are opposites in structure. Procrastination has no defined end and no purpose; it’s avoidance dressed as diligence. Exploration has a fixed end (the sample size — say, the first third) and a specific job (set the bar). The moment the explore phase ends, exploring becomes forbidden: you switch to commit mode and leap at the first option over the bar. If your “research phase” has no scheduled finish line and no threshold waiting on the other side, you’re not exploring — you’re stalling, and the model gives you zero cover for it.

Match each explore-then-commit idea to what it actually means.

Pick a term, then click its definition.

Putting it together

Strip away the specific 37% and here is the reusable machine you now own. A sequential, no-going-back decision splits cleanly into two phases joined by one bar. The explore phase gathers information and commits to nothing — its only product is a bar, calibrated from the top of what you saw, because “good” is meaningless until you’ve seen the scale. The commit phase takes the first option that clears the bar. The trade-off underneath is explore/exploit, made unusually sharp by irreversibility — every look risks burning a great option, unlike a bandit where you can always return. And the bar itself isn’t frozen: the truly optimal rule is a threshold that falls as your runway shrinks, which is exactly where Lesson 3 takes us.

Success:

The pattern to carry forward

Explore-then-commit: spend a defined opening phase gathering information and committing to nothing — its only job is to calibrate a bar from the top of what you see — then take the first option that clears it. The sample is a ruler, not a shortlist. Watch for the three traps: refusing to ever pick from the sample even when recall is genuinely available (rigidity), mistaking aimless stalling for exploration (procrastination has no end and no bar), and clutching a bar you set with infinite runway when the door is nearly shut (a fixed bar strands you). Next up: why the smartest bar falls as your options run out.

Mark lesson as complete