Skip to content
Mental Models

Bottlenecks & the Theory of Constraints

Don't Polish the Wrong Stone

Speeding up any step that isn't the bottleneck moves a system's output by exactly zero — and a faster upstream step is worse than useless, burying the constraint in a growing pile of work that just waits.

12 min Updated Jun 28, 2026

You’ve found the constraint (lesson 1) and you’ve learned Goldratt’s procedure for governing it (lesson 2). Now comes the lesson everyone agrees with in theory and violates in practice — the counter-intuitive heart of the whole course. It is this: work on the wrong step and you accomplish nothing, no matter how hard you work, how much you spend, or how proud you are of the result. Worse, on a bad day, you accomplish something negative.

Picture a jeweller with a necklace of ten stones. Nine are already polished to a mirror shine. One is dull. A visitor will judge the whole necklace by its worst stone — that’s just how eyes work. So where does our jeweller spend the afternoon? On the nine that already gleam, of course, because polishing them feels like progress and the dull one is annoying to deal with. Hours later the nine sparkle a fraction brighter, the dull stone is exactly as dull, and the necklace looks no better than when he started. He polished the wrong stones.

That’s the mistake this lesson is built to inoculate you against. The dull stone is your bottleneck. Everything else is a stone that’s already shiny enough — and every minute spent buffing it is a minute the necklace doesn’t improve.

Before you read — take a guess

A four-stage pipeline runs at 30, 50, 30, 40 units/hr (raw material always available). You spend a fortnight and a big budget upgrading the second stage from 50 to 70 units/hr. What's the line's new throughput?

Local vs. global optimization

The deep error has a name, and naming it is half the cure. Local optimization is improving one part of a system — a single step, a department, a machine, a metric — in isolation, judging success by how good that part now looks. Global optimization is improving the output of the whole system — what actually comes out the far end. The trap is that these two feel like the same thing, and they are not. A system can be a showroom of locally-optimized parts and still be globally terrible.

The analogy that sticks: an orchestra. Local optimization is tuning the loudest trumpet a touch louder, drilling the violinist on her solo, polishing each instrument in isolation until it’s individually flawless. Global optimization is whether the whole orchestra sounds good together. And here’s the rub — you can have eighty individually excellent musicians and a catastrophe of a concert, because one section is playing half a beat behind. Making the trumpet louder doesn’t fix that; it makes the mess louder. The performance is set by how the parts combine, not by how good the best part is in isolation.

A chain is the most ruthless version of this. The output of a chain is its throughput, and you proved last lesson that throughput equals the capacity of the slowest step — the minimum, never the average, never the sum. So “optimize the part” and “optimize the whole” come apart completely: you can drive every non-bottleneck stage to perfection and the whole still produces exactly what the one slow stage allows.

Tip:

The one-sentence version

Local optimization makes one part look great; global optimization makes the whole system deliver more — and in a chain, only one part (the bottleneck) connects the two. Polish any other and the part shines while the system doesn’t budge.

When to use it

Reach for the local-vs-global distinction the instant someone reports a “win” — a department hit its target, a machine ran at record efficiency, a stage got 40% faster. Before you celebrate, ask the only question that matters: did the output of the whole system go up? If the answer is “well, our part improved,” you may be looking at a beautifully polished wrong stone.

Why a faster non-bottleneck adds zero throughput

Let’s make the “exactly zero” claim airtight, because it really is exactly zero — not “a little,” not “diminishing returns,” but none. Recall the rule:

Throughput = the capacity of the bottleneck = the minimum capacity across all stages.

Now take any stage that isn’t the minimum. By definition it has slack: capacity it isn’t using, because work only arrives at the rate the bottleneck lets it flow. A stage running at 50 in a line capped at 30 is already idle for a chunk of every hour — it does its 30 units’ worth of work and then twiddles its thumbs waiting for the next batch to clear the constraint. Raise that stage to 70 and what have you changed? Its slack grew from 20 to 40. It now twiddles its thumbs longer. The minimum is still 30. Throughput is still 30. You bought idle time and called it an upgrade.

Walk the numbers on the pretest line — stages at 30, 50, 30, 40:

Stage 1Stage 2Stage 3Stage 4Throughput
Before305030 (tied bottleneck)4030
Upgrade stage 2 → 703070304030 (unchanged)
Upgrade stage 4 → 603050306030 (unchanged)
Upgrade a bottleneck (stage 1 → 45)4550304030 (other bottleneck still caps it!)
Upgrade both bottlenecks → 454550454040 (constraint moved to stage 4)

Notice three things. Pumping stage 2 or stage 4 does nothing. Lifting just one of the two tied 30-stages also does nothing, because the other 30-stage still caps the line — when a constraint is shared, you have to relieve all of it. And only when you finally exceed every slow stage does throughput climb — and immediately a new slowest stage (the 40) takes over as the constraint. The output never moves until the bottleneck moves.

Don’t take it on faith — drag it yourself. Push the tall bars up and watch the dashed throughput line refuse to react:

Find the bottleneck

Raise the non-bottleneck — watch nothing happen

Drag each stage’s capacity. The whole line can only run as fast as its slowest stage — the bottleneck (red). Speed up a faster stage and nothing happens; speed up the bottleneck and the line finally moves — until a different stage becomes the new slowest one.

Line throughput25units/hr
A40B55BottleneckC25D4525 units/hr
BottleneckLine throughputStage capacityWork piling upStarved / idle

The line runs at 25 units/hr — the capacity of “C”, the slowest stage. Every other stage is faster, so 65 units/hr of capacity sits idle. Speeding up anything but “C” changes nothing.

Stage C is the constraint (red). Drag B or D as high as you like: the dashed throughput line doesn't move a pixel, and you're just growing the faded 'wasted slack' on top of those bars. Only when you drag C upward does the line climb — until A or D becomes the new bottleneck and the red badge jumps. Every drag that isn't on the red bar is a wrong stone, polished.

For a line currently bottlenecked at one slow stage, sort each action by whether it actually raises the system's throughput.

Place each item in the right group.

  • Speed up the bottleneck stage itself
  • Speed up the single fastest stage in the line
  • Add a second machine in parallel at the bottleneck
  • Make the upstream prep cook twice as fast (cook is not the constraint)
  • Cut wasted setup/changeover time at the bottleneck
  • Buy a faster machine for a stage that already has slack

Worse than useless: the WIP pile-up

Up to now the worst case was “wasted effort.” But speed up a non-bottleneck that sits upstream of the constraint — before it in the flow — and you cross from useless into actively harmful. Here’s why.

A faster upstream stage doesn’t push more work out of the system (the bottleneck still caps that). It pushes more work at the bottleneck, faster than the bottleneck can swallow it. The excess has to go somewhere, and it goes into a pile: a growing heap of half-finished work waiting its turn at the constraint. That heap has a name — work-in-progress, or WIP: any unit that’s been started but not yet finished, stuck somewhere in the middle of the pipeline.

And it’s not a one-time heap. The upstream stage produces 50/hr, the bottleneck clears 30/hr, so the pile grows by 20 units every hour, forever. This is exactly the reinforcing loop you met in feedback loops: an unchecked build-up that compounds because nothing balances it. The faster you run the upstream stage, the faster the pile climbs.

Make it concrete with a kitchen. One grill is the bottleneck — it can cook 30 plates an hour, and that’s the restaurant’s real throughput. The prep cook, eager to look productive, doubles her speed and fires tickets at the grill twice as fast. Does the restaurant serve more dinners? Not one extra plate — the grill still does 30. What does happen: a wall of prepped tickets stacks up beside the grill, growing all night. The cook feels like a hero. The kitchen is now a disaster.

That pile isn’t free clutter — it costs, in ways that hide until they hurt:

  • Cash frozen in inventory. Every half-finished unit is money you’ve spent (materials, labour) that you can’t collect on until it’s finished and sold. A growing WIP pile is a growing pile of frozen cash.
  • Lead times explode. A unit entering a 200-deep queue waits for 200 ahead of it to clear the bottleneck before its turn. The same line, same throughput, now takes far longer per item from start to finish — customers wait longer for nothing in return.
  • Defects discovered late. If the prep cook is making a systematic error, it’s now baked into 200 buried tickets before anyone catches it at the grill. Big WIP piles hide problems and turn small mistakes into huge rework.
  • Clutter, confusion, and stress. Floor space, mental load, “which of these 200 do I do first?”, the constant low-grade panic of a visibly-overflowing queue. None of it produces a single extra finished plate.

So the faster prep cook didn’t just waste her own effort — she actively degraded the whole operation: more frozen cash, longer waits, later-caught defects, more chaos. Same throughput, strictly worse system. That is why a non-bottleneck improvement can be worse than useless.

A code team has one reviewer who can approve 8 pull requests a day — that's the bottleneck. To 'go faster,' management tells the eight developers to merge-ready their PRs more aggressively, and now they finish 20 PRs a day between them. Three weeks later, what does the team most likely see?

The efficiency trap: the cult of “keep everyone busy”

So why does this mistake get made over and over by smart, well-meaning people? Because of how we measure work. The instinct — and the management tradition behind it — is to judge each worker, machine, or department by its own utilization: the fraction of time it’s actually busy producing, rather than idle. An idle machine looks like waste. An idle worker looks like a problem. So we push every part to run flat out, all the time, at 100%.

That instinct is poison to a system with a bottleneck. Driving every non-bottleneck to 100% utilization is precisely the recipe for the pile-up you just saw. A non-bottleneck running flat out overproduces by definition — it makes more than the constraint can absorb — and the surplus becomes WIP. The local metric (utilization) goes green; the global metric (throughput) doesn’t move; and a pile of frozen cash grows quietly in the corner. Maximizing the utilization of a non-bottleneck doesn’t speed the system up — it just inflates inventory.

This is the manager’s hardest pill, so let’s state it plainly:

An idle non-bottleneck is not a problem to be fixed.

It feels wrong. The prep cook standing there waiting looks like waste. But her idleness is the system working correctly — she’s pacing herself to the grill, exactly the subordinate step from lesson 2: non-bottlenecks should match the constraint’s rhythm, and their idle time is free, because they had spare capacity to burn. Forcing her to “be productive” by working flat out is what breaks the system. The free thing and the productive-looking thing are opposites here.

There’s a Goodhart’s-Law sting baked into this. Goodhart’s Law: when a measure becomes a target, it stops being a good measure. Make “keep your utilization high” the target for each part, and people optimize for that — running flat out, hoarding work, overproducing — at the direct expense of the thing you actually wanted, which was the system’s throughput. The local efficiency metric becomes the enemy of global output. The metric meant to make the system productive is the very thing strangling it.

A plant manager proudly reports that every machine on the floor is running at 95–100% utilization, with almost no idle time anywhere. The plant's actual finished-goods output, however, is flat, and the warehouse of half-finished parts keeps growing. What's the most likely diagnosis?

Match each term to its precise definition.

Pick a term, then click its definition.

Fill in the core lesson.

Pick the right option for each blank, then check.

Speeding up a step that isn't the constraint changes the system's throughput by , because that step already has . If the step is upstream of the bottleneck, it's actually worse: it floods the constraint with that just waits in a growing pile. And measuring each part by its own pushes everyone to run flat out, which overproduces at non-bottlenecks and builds exactly that pile — so an idle non-bottleneck is to fix.

When to use it

This is less a “sometimes” lens and more a permanent reflex to install. Before you optimize anything — buy the faster machine, hire into the busy team, drill the loud trumpet, push a department to run harder — ask one question:

Is this the constraint?

If yes: push with everything you’ve got; this is the highest-leverage spot in the whole system (it’s the most concrete kind of leverage point from the last course). If no: stop. Improving it will move the system’s output by zero, and if it sits upstream of the real bottleneck, you’ll make things actively worse. You’d be polishing a stone nobody will ever notice, while the dull one — the only one that matters — stays exactly as dull.

The discipline is uncomfortable precisely because the wrong stones are the fun ones to polish: the part you understand best, the metric that’s easy to move, the team that’s eager to look busy. Resist all of it. Find the dull stone first.

Warning:

The pitfall in one sentence

Before optimizing any step, prove it’s the constraint. If it isn’t, every hour and dollar you spend improving it returns zero throughput — and if it’s upstream of the real bottleneck, it returns negative, burying the constraint in a growing pile of frozen, waiting work while you congratulate yourself on the polish.

Check yourself: optimizing the right step

Question 1 of 50 correct

A line runs at 20, 35, 50, 35 units/hr. A manager triples the third stage to 150 units/hr. What happens to the line's throughput?

Check your answer to continue.

Recap

Big picture

Don't Polish the Wrong Stone

  • Optimize the constraint — nothing else
    • Local vs. global
      • Local = one part looks great (a step, dept, metric)
      • Global = the WHOLE system delivers more
      • In a chain they meet only at the bottleneck
    • Faster non-bottleneck = zero
      • Throughput = min capacity; non-bottleneck has slack
      • More capacity → just more idle time, not output
      • Tied bottlenecks: must relieve ALL of them to gain
    • Upstream = worse than zero
      • Floods the constraint → WIP pile grows every hour
      • A reinforcing build-up (feedback loops)
      • Costs: frozen cash, long lead times, late defects, chaos
      • Kitchen: fast prep cook buries the one grill
    • The efficiency trap
      • Measuring each part by its own utilization → run flat out
      • Non-bottleneck at 100% → overproduction → WIP
      • An idle non-bottleneck is NOT a problem (subordinate it)
      • Goodhart: local metric becomes enemy of throughput
    • The reflex
      • Before optimizing anything: "Is this the constraint?"
      • No → zero gain (or negative). Stop.
      • Yes → push hard; it's the real leverage point

Where this goes next

You now know the direction of the rule: a non-bottleneck improvement gives you zero throughput, and an upstream one gives you less than zero. But there’s a sharper, quantitative question lurking underneath. When you do speed up part of a system — say a part that’s genuinely a portion of the work — exactly how much faster does the whole thing get? Is there a hard ceiling on the payoff, no matter how dramatically you accelerate that one part?

There is, and it has a precise formula. The next lesson, Amdahl’s Law, is the math of this whole lesson: it shows that speeding up one fraction of a process has a strict, calculable limit on the total improvement — and that even an infinite speedup of a non-dominant part can only help so much. You’ll be able to do the arithmetic of “how much will this actually help?” before you spend the fortnight and the budget. The intuition you built here — most effort is wasted because it’s aimed at the wrong stone — is about to get its equation.

Mark lesson as complete