Was It Good Judgment, or Did It Just Work? [On Judgment - Part 3]

If judgment is an assessment, the harder question is how to tell whether it was a good one. Outcomes help, but they can also mislead.

Was It Good Judgment, or Did It Just Work? [On Judgment - Part 3]

In Part 1, I started questioning what I actually meant by judgment. In Part 2, I narrowed it down: judgment is an assessment, separate from reasoning, preference, decision and authority.

That definition is cleaner. But it creates a harder question.

If judgment is an assessment, how do we know whether an assessment was actually good?

There is a product postmortem I have seen in many forms.

A team launches something new.

Before launch, the team believes demand is strong.

Early research looks promising. Sales is asking for it. A handful of customers have described the problem with unusual consistency.

The product ships.

Revenue disappoints.

The conclusion arrives quickly:

“Customers did not want it.”

Maybe.

But suppose the onboarding flow was broken for the first six weeks, the pricing page buried the new offer, sales compensation still favored the old product, and performance problems affected the exact customer segment the team was trying to reach.

Which hypothesis did the outcome actually test?

That question has been bothering me because it exposes a mistake I have made in my own thinking about judgment.

In the previous essay, I separated assessment from decision.

That creates another separation we need to make:

assessment quality is not the same thing as outcome quality.

We know this in theory.

We violate it constantly in practice.


The outcome gets into the room too early

In a classic set of experiments, Jonathan Baron and John Hershey showed people decisions made under uncertainty while holding the information available to the original decision-maker constant and varying the outcome. Favorable outcomes led participants to rate the same decision process more positively; unfavorable outcomes led them to rate it more negatively.

Most of us recognize the pattern. It is outcome bias: the result changes how we evaluate the reasoning that came before it.

Organizations add extra layers because we do not just judge decisions privately. We reward and promote people, turn successful launches into internal mythology, and failed bets into evidence that someone lacked judgment.

Outcome bias can therefore distort careers, incentives and the stories an organization uses to teach itself.

But there is a trap on the other side too.


“Good process” can become unfalsifiable

Once leaders learn about outcome bias, a comfortable defense appears:

“The outcome was bad, but we made the best decision with the information available.”

Sometimes that is exactly right. A good assessment can precede a bad result: a well-grounded demand assessment can be followed by weak execution, a reasonable market assessment can be overtaken by an external shock, and a carefully calibrated risk estimate can still be followed by a rare loss.

But “good process” can also become a shield that prevents learning.

If every bad result is explained away by uncertainty, then the process can never lose.

And if the process can never lose, it cannot improve.

So we need a way to do two things at once:

  1. resist treating outcomes as verdicts;
  2. still let outcomes change what we believe.

The distinction I find useful is this:

An outcome is evidence about a prior assessment. It is not the verdict on it.

With one boundary: the outcome has to be relevant to the assessment. Later revenue can tell us something about a demand forecast. It does not, by itself, tell us whether a policy was fair or a product experience coherent.

That sounds small, but it changes the postmortem.


Five things we collapse into one story

Take a product launch that underperforms.

A company enters a new segment. Revenue disappoints two years later. Was the segment thesis wrong, the proposition wrong, the sales motion wrong, the investment level insufficient, or did the market change?

There are at least five different things we might evaluate.

These can diverge. A strong result at one layer does not guarantee a strong result at the next. A team can make a strong assessment, choose well, execute badly and get a poor outcome. A weak assessment and bad choice can still be rescued by brilliant execution or a favorable market. And a team can do everything well and still encounter a shock.

When we compress all of that into “it worked” or “it failed,” we lose the ability to learn what actually happened.

For some assessments, the new evidence can directly show that the assessment was wrong. Probabilistic assessments are different. If I said something had a 70% chance of happening and it did not happen, that outcome should affect how we evaluate the forecast, but it does not by itself prove the 70% estimate was bad. Calibration appears across repeated assessments, not in one outcome.


Outcomes test propositions, not stories

This is where I think product postmortems often go wrong.

We tell the story at the level of the initiative.

“The redesign worked.”

“The new pricing failed.”

“The market entry was successful.”

But initiatives contain multiple claims.

Suppose a redesign increases conversion by 12%. What did that establish?

It supports the claim that more people completed the flow. It may suggest customers preferred the design, but it does not establish that long-term value or trust improved, or that the same design principle should be applied across the rest of the product.

The more general the story, the less likely one outcome has actually tested it.

This is especially important in experimentation cultures because statistical rigor can create a false sense of conceptual rigor.

An experiment can be perfectly well executed against the wrong question.

A metric can move significantly while the strategic belief remains mostly untested.


How much should this outcome update us?

How much an outcome should update us depends on what it actually tells us about the belief we held.

Did we measure the thing we actually care about? Can we reasonably attribute the observed result to what we are evaluating? And how strongly does the evidence distinguish between the explanations that remain?

If we predict that 30% of a clearly defined cohort will activate within seven days and observe 8% under a clean measurement setup, the result should move our belief substantially.

If we predict that a new enterprise offer will create a durable strategic position and then look at three months of bookings during a sales incentive campaign, it should move us much less.

A few things determine how much an outcome should move us.

Attribution. Can we reasonably connect the result to the thing we are evaluating? If adoption falls after a redesign while performance also deteriorates, the outcome is not a clean test of the design judgment.

Confounding. What else changed at the same time? Pricing, distribution, seasonality, competitor action, customer mix and sales behavior can all produce the same visible outcome.

Measurement validity. Does the metric actually represent the thing we care about? “Resolved” is not “customer problem solved”; “engaged” is not “received value”; “clicked” is not “trusted.”

Selection. Who generated the data? A satisfaction result from the happiest customers, or a human-versus-automation comparison built from only the most complex reviewed cases, tells a different story from a representative sample.

Time horizon. Some beliefs resolve quickly; others do not. Short-term conversion can improve while long-term retention falls, just as a pricing change can lift revenue this quarter and damage expansion next year.

Reflexivity. Some decisions change the environment that later evaluates them. Rankings, recommendations, fraud rules and sales routing can partly create the evidence that comes back.

Competing explanations. The more plausible causal explanations remain, the less the result should move our confidence.

None of this is exotic. Product teams already use these disciplines in experiments, analytics and causal inference. We often stop applying them once the conversation becomes “strategic.”


Metrics can validate the wrong judgment

There is a deeper issue here.

Metrics do not simply report outcomes. They define what the organization counts as evidence of an outcome.

Imagine an AI support system evaluated primarily on resolution rate. If “resolved” means a conversation that does not reopen within a fixed period, the metric can improve because the system genuinely solved more problems. It can also improve because customers gave up, difficult cases were routed elsewhere, the time window is too short, or the system became more aggressive about closing interactions.

The number is not lying. The operational definition may be too narrow.

We thought we were testing:

“The system solves customer problems better.”

We actually tested:

“More conversations satisfy our operational definition of resolution.”

Those are related, not identical. A metric is not neutral infrastructure; it carries a claim about what should count. When we forget that, an apparently successful result can validate the wrong assessment and make the feedback loop self-confirming.


Keep a small record before the result arrives

One simple practice helps with reconstruction: record the assessment before the result arrives, and bring that record back into the review.

Before a consequential decision, write down:

  • What do we believe?
  • Why do we believe it?
  • What would we expect to observe if it is true?
  • What would change our view?
  • When will we review it?

This is not a new framework. It is a defense against memory. Fischhoff’s classic work on hindsight bias showed how knowing an outcome changes our reconstruction of what seemed predictable beforehand. An ex-ante record preserves the uncertainty that existed before the outcome cleaned up the story.

Davies later found that bringing back a written record of people’s original foresight thoughts reduced hindsight bias, which is why preserving the assessment before the result matters here.

I am not claiming these five questions automatically produce better judgments. The narrower point is that if we want to learn from an outcome, we need an honest record of what the outcome was supposed to update.


A better postmortem

The postmortem question I want to stop asking is:

“Did this decision work?”

It compresses too much.

Instead, for a consequential product decision, I would ask five questions.

What did we believe before the outcome?

Not what we now say we believed.

What was the actual assessment?

Which part of that belief did this evidence actually test?

Be specific.

Did the result test our belief about demand, usability, willingness to pay, distribution, execution or something else?

What other causes could explain the result?

This is not an invitation to explain away failure.

It is an attribution check.

What should update?

Which belief became stronger, weaker or obsolete?

What should not update?

This is the part I rarely see.

A failed launch should not automatically destroy every belief attached to the initiative.

A successful experiment should not automatically validate the strategy around it.

Learning requires restraint in both directions.


What this changes about “good judgment”

I started this article asking whether a successful outcome means the judgment was good.

The answer is no.

Successful outcomes do not prove that a judgment was well formed, and bad outcomes do not by themselves prove that it was poorly formed.

Later evidence should not cause us to rewrite the historical quality of the reasoning using information that was unavailable at the time. But it can change what we believe about the reliability of the evidence, methods and judgment process that produced it.

But across repeated comparable assessments, ex-post performance becomes evidence about something else: whether the person, method or system producing them is reliably good at that class of judgment.

So outcomes are not verdicts on a single judgment. But for judgments that outcomes can meaningfully test, repeated performance is evidence we need if we want to know whether the person, method or system producing them is reliably good.

Correctability is a different property. It belongs to the system around the judgment, not to the judgment itself.

By correctable, I mean that the system preserves enough of the original assessment to recognize relevant contrary evidence, identify what should change, and update future assessments or decisions without rewriting what was believed before.

A serious system for forming and using consequential judgments must therefore be correctable.

Can the organization tell what the new evidence actually tested? Separate execution failure from assessment failure? Resist turning one result into a universal lesson? Update the belief without rewriting the past?

That is harder than celebrating wins and writing postmortems after losses.

It is also much more useful.


Then another assumption breaks

Suppose we get the feedback loop right: we record the assessment, observe the result, understand what it actually tested and update appropriately.

Does experience now guarantee better judgment?

No.

A person can spend ten years in a role and receive terrible feedback.

They can encounter the same narrow class of cases repeatedly.

They can operate in an environment where results take years to appear.

They can mistake confidence for learning.

They can become expert in a world that no longer exists.

Even perfect feedback is useless if the experience does not contain the right signal, or the person cannot detect it, or the signal does not transfer to the next environment.

So the next question becomes:

Where does good judgment actually come from?


Sources