Judgment Is Smaller Than I Thought [On Judgment - Part 2]
Judgment is not one thing. It is an assessment, separate from reasoning, preference, creation, decision and authority. Making those distinctions visible makes consequential decisions easier to inspect.
A few weeks ago I started looking more carefully at a word I use far too often.
Judgment.
I had used it to describe product leadership, decision-making, AI, strategy, taste, curation, governance and expertise.
I had used it that way in earlier essays too, sometimes as shorthand for several different kinds of valuable human work. I would not use the word that broadly now.
That should have been a warning:
When one word explains everything, there is a decent chance it is hiding something.
So I went looking for a cleaner definition, and I expected the answer to make judgment richer.
It did the opposite.
The useful definition got thinner.
In Part 1, I stopped with a working definition and a list of doubts about it. This time I wanted to see what survived the counterexamples.
Start with a normal product question
Imagine a leadership team asking:
“Should we launch an enterprise tier?”
We might call this a strategic judgment.
But pause for a moment and look inside the question.
To answer it, the team may need to ask questions like:
- What customer problem exists?
- How large is the opportunity?
- How likely are customers to pay?
- Which capabilities would make the offer credible?
- What new complexity would enterprise customers introduce?
- Is that complexity acceptable to us?
- Which option should we choose?
- Who has the authority to commit budget and people?
That single executive “judgment call” contains several different operations.
Some ask for assessments: how large, how likely, how damaging.
Some ask what we value or are willing to accept.
Some ask us to create possibilities.
Some ask us to choose.
And one asks who is allowed to commit the organization.
Calling all of them “judgment” makes the conversation feel simpler than it is.
The research literature on judgment and decision-making makes a distinction that helped me here. A common framing separates judgment, preference and choice:
- Judgment is about assessing what is likely to happen.
- Preference is about how much we value different outcomes.
- Choice combines those elements into a decision.
That definition is too narrow for product work, because not every assessment is a prediction. We also judge whether something is coherent, risky, fair, ready or strategically attractive.
There is no single standard definition of judgment across the fields that study it. Decision research, medicine, auditing and philosophy use the word differently. So for this series, I need a definition narrow enough to be useful without pretending the different traditions all mean exactly the same thing.
And I am going to use judgment in a deliberately operational sense for this series:
A judgment is an assessment about something.
Judging is the act of forming that assessment.
This is a functional convention for system design, not a claim about the essence of judgment.
I know “assessment” is broad too. I am using it because it gives me something more concrete to inspect: a position about a target, including an evaluation of that target against a standard.
I use preference separately for how an actor values or ranks possible outcomes.
In this convention, the assessment does not by itself express that preference, select a course of action or commit to one.
That lets me separate the assessment from how we reached it, how we value the possible outcomes, what we decide to do, and who has the authority to act.
I am looking for the smallest definition that survives the cases I care about without quietly swallowing reasoning, preference, decision or authority.
That is a much smaller claim than the one I started with.
Saying someone “has good judgment” is a different claim: that they reliably produce good assessments of some kind.
Judgment is not reasoning
In Part 1, I left this as an open question. It was the first distinction I needed to test.
I used to treat reasoning as if it were almost constitutive of judgment.
You gather evidence. You reason through it. You reach a judgment.
That sequence often happens.
But it is not the only way assessments form.
Reasoning is a process that connects premises, observations or evidence.
A judgment is an assessment.
Those can come apart.
A product leader can reason for an hour without closing on an assessment.
“I understand the trade-offs, but I still do not know whether this market is attractive.”
That is reasoning without a settled judgment.
The reverse can also happen.
An experienced designer can look at a flow and say:
“This is going to confuse people.”
They may be unable to reconstruct the full chain immediately.
That does not automatically make the assessment mystical or irrational.
Research on expert intuition is useful here. Daniel Kahneman and Gary Klein, coming from traditions that often disagreed about intuition, found common ground around a simple idea: intuition can become skilled when the environment contains learnable regularities and experience provides enough opportunity to learn them through feedback.
In other words, explicit reasoning is not required every time a judgment is formed.
But that does not mean intuition is automatically trustworthy either.
The question shifts from:
“Can you show me every reasoning step?”
to:
“What makes this assessment worth trusting?”
Those are different tests.
Judgment is not decision
Part 1 left another boundary unresolved: if judgment produces an assessment, where does the decision actually begin? Ordinary language often collapses the two. For this series, I want to keep them separate.
Consider these two sentences:
“I think the market is moving toward usage-based pricing.”
and:
“We are moving to usage-based pricing.”
The first is an assessment.
The second is a decision.
Between them there may also be a recommendation:
“Given those assessments, I recommend that we move to usage-based pricing.”
A recommendation goes further than an assessment because it selects or advocates a course of action. But it still need not commit the organization.
That introduces another distinction: authority. A person or system may assess the situation and recommend an option without having the authority to bind the organization to it.
Maybe the company is not operationally ready.
Maybe customers would resist the transition.
Or maybe the economics do not work yet.
Maybe the strategic cost is too high.
A decision can depend on several judgments plus preferences, constraints, alternatives and timing. A recommendation may propose that decision. Authority determines who is entitled to make it binding.
That difference is not semantic housekeeping.
It becomes structurally important once AI enters the workflow.
A model may be better than a human at one assessment inside a decision without having any legitimate authority to make the decision itself.
If we collapse assessment and decision into “judgment,” we lose that design option.
The same is true of authority. The person with the best assessment may not be the person allowed to commit the organization. And the person with authority may not have the best assessment.
“The leader makes the judgment call” often collapses assessment, choice and authority into one act.
Judgment is not generation
This distinction became more important than I expected.
Product work is not only about evaluating options.
It is also about creating them.
Generating an option and assessing an option are different operations, even when product work moves rapidly between the two.
Generative AI has blurred the language further.
A system can generate ten onboarding concepts.
A system can also score those concepts against explicit criteria.
Those are different capabilities.
We use one umbrella word, “intelligence,” for all kinds of machine capability.
We increasingly use “judgment” for whatever we still think a human should do. But that creates a moving boundary. As machines become good at something we had placed inside judgment, we either have to admit that machines can perform some judgments or move the definition again.
That is not enough precision for serious system design.
Product taste is a good boundary case.
When a strong product leader says, “This does not feel like us,” what is happening?
Part of it may be evaluative. The design is being assessed against standards such as coherence, simplicity, distinctiveness or brand principles.
Part of it may be creative. The leader sees an alternative the current design has not expressed.
Part may be preference, part may be authority.
The mistake is not that the word taste is useless.
The mistake is assuming it names one operation.
Different assessments are good for different reasons
This is the distinction I found most useful.
If judgment is an assessment, the next question is what would make that assessment a good one.
I will use “standard” loosely for whatever makes an answer better or worse in that particular case.
Sometimes that is truth. Sometimes calibration. Sometimes causal fit, conformity with professional rules, fit with product principles or performance against criteria learned through experience.
So a practical question appears:
What would make this a good assessment?
Take four assessments:
“This customer will churn.”
This is mainly a predictive assessment. Over time, we can test it against what actually happens and look at calibration, accuracy and error patterns.
“This payment is fraudulent.”
This is also empirically testable, but the standard depends on the form of the answer. If the system produces a fraud probability, we can examine calibration and discrimination. If it produces a classification, we can examine precision, recall and error patterns.
But what the organization does with that assessment is a different question. The threshold for blocking a payment depends on the relative cost of false positives and false negatives, as well as risk appetite and policy.
“This product experience is coherent.”
Now the standard is harder to reduce to ground truth. We may rely on product principles, usability evidence, consistency, customer expectations and expert critique.
“This policy is fair.”
Here, predictive accuracy is not enough. The standard is partly normative, and reasonable people may disagree about what fairness requires.
Once a fairness standard is specified, however, we can still ask how well a policy satisfies it. Choosing the standard and assessing performance against it are different operations.
—
All four are assessments.
But what makes them good differs.
Some can be tested against future outcomes. Some depend on explicit criteria or professional standards. Some require normative standards that have to be chosen or justified before the assessment can even be evaluated.
And this helps clarify where values enter.
Estimating whether a customer will churn does not necessarily express a preference about whether churn is acceptable or how much preventing it is worth. But values can shape how the problem is defined: what counts as churn, which customers matter, which time horizon matters.
In other cases, values help define the standard itself. Whether a policy is fair or a risk is acceptable depends partly on what we think should matter.
So I no longer think values belong inside judgment as a universal component. They can shape the target, the framing or the standard against which an assessment is made without being a necessary ingredient of every assessment.
Standards may come from preferences, principles, law, professional norms or prior judgments. Once a standard is sufficiently specified, assessing how well something satisfies it is a different operation from selecting or justifying that standard.
In practice, those operations can be compressed into the same sentence. When I say, “this risk is acceptable,” I may be both assessing the risk against a standard and expressing or selecting the standard I think should apply.
This is why I no longer think judgment names one cognitive process.
For this series, it is more useful to treat judgment as a family of assessments: different questions, different standards and potentially different ways of arriving at the answer.
That also means “better judgment” is not one generic capability. Before asking who has good judgment, we should ask what kind of assessment they are making and what makes that assessment good.
What I changed my mind about
I tried to make the definition richer. The counterexamples kept making it thinner.
The definition above was not where I started. It was what survived.
First, judgment does not begin only when the rules run out.
In Part 1, I proposed that judgment begins where facts, rules or models no longer settle the answer.
I now think that was too strong.
The boundary is not clean enough to make underdetermination part of the definition. Under the operational convention above, an assessment can still be formed when a task is tightly specified or strongly constrained. I therefore no longer need underdetermination to be part of what makes an assessment a judgment.
Underdetermination is not sufficient either.
Sometimes the evidence simply does not support an assessment yet:
“We cannot assess this.”
There may be ambiguity without a judgment actually being formed.
So I would change the role of underdetermination rather than discard it:
Judgment becomes more important to design and govern when available evidence, rules, models or standards constrain an assessment without uniquely determining it.
Ambiguity amplifies the role of judgment. It does not create it.
Second, machines can perform the assessment function in bounded classes.
That sentence still feels uncomfortable because our everyday language gives judgment a distinctly human tone.
But a system can estimate whether a customer will churn, classify whether an image contains a defect, rank alternatives against a specified criterion or evaluate an output against a rubric. In a decision system, those outputs can occupy the same functional role as assessments produced by people.
Saying that a machine “has judgment” is too underspecified to help me design anything. Does it mean the system can predict churn, diagnose a defect, rank alternatives, evaluate an answer, choose an action or own the consequences? Those are different capabilities. I do not need to decide whether a machine literally “judges” in the human sense. I need to know which assessment it can perform, how well, and under which conditions.
The useful product question is:
For this assessment class, who or what produces the better assessment, under which conditions?
Third, I would now place option generation next to judgment rather than inside it. Generating a possibility and assessing it are different operations.
Where judgment becomes a design problem
That thinner definition creates an obvious edge case.
If judgment is just forming an assessment, does that mean 2 + 2 = 4 is judgment?
Suppose a system applies a perfectly specified eligibility rule and classifies an account as enterprise.
Has it made a judgment?
Under the operational definition I am using here, you could call that a very thin assessment. That is a consequence of choosing a functional boundary rather than trying to define judgment philosophically.
But it is not where judgment becomes an interesting product problem.
If the inputs and rule fully determine the output, there may be almost no discretion in applying the rule. But that does not mean assessment quality disappears. The rule may be wrong, the inputs may be poor, the threshold may be badly chosen, or the rule may be inappropriate for the purpose.
In those cases, the judgment problem has moved upstream: from applying the rule to designing, validating or governing it.
As uncertainty, interpretation, tacit standards, expertise or contested values enter the application itself, assessment quality becomes a more visible product problem.
You can stretch the definition until every classification becomes judgment. I do not need to. For system design, the more useful question is where assessment quality and discretion actually enter.
So this is a practical boundary, not a philosophical one. I am using judgment to name the assessment function inside a decision system, then asking where assessment quality, discretion and governance actually become consequential.
I care most about the places where assessment quality is meaningfully in play.
One question I now want in consequential decisions
I do not think teams need a new framework for this.
I certainly do not think we need a seven-box “Judgment Canvas.”
One question is enough to start:
“What exactly is the assessment inside this decision?”
Take a roadmap decision.
The decision may be “invest in enterprise administration this quarter.”
Inside it may sit several assessments:
- enterprise demand is likely to remain durable over the next two years;
- the addressable segment is large enough to support the revenue target;
- missing administration capabilities explain more lost enterprise deals than any other identified product gap;
- we can deliver the capability within the quarter with the current team;
- doing so would reduce self-serve delivery capacity.
Then comes a different question:
“How much reduction in self-serve delivery capacity are we willing to accept?”
Once those assessments are visible, the conversation improves.
We can ask who has relevant evidence.
We can ask what standard makes each answer good.
We can ask where we are expressing a preference rather than making an empirical claim.
We can ask what is actually a question of authority.
We can ask whether the team is arguing about the same thing at all.
And that matters because different disagreements require different interventions.
More data does not solve a values disagreement.
A better forecast does not create a missing option.
And an executive decision does not resolve a disagreement about what the evidence supports.
That is the practical value of the thinner definition.
It does not make judgment more impressive, but it makes the work around it more inspectable.
And then the next problem appears
Separating assessment from decision creates an immediate problem.
Suppose our assessment was that enterprise demand would be strong.
We launch.
Revenue exceeds plan.
Was that good judgment?
Now reverse it.
Same evidence.
Same reasoning.
Same assessment.
A competitor makes an unexpected move, the market contracts, and the launch fails.
Was that bad judgment?
We all know, intellectually, that outcomes and judgment quality are different.
Organizations still rewrite the past around the result.
And if different kinds of assessment are good for different reasons, there may not be one generic test of “good judgment” either.
So the next question is harder than the definition:
“If the outcome alone does not determine whether a judgment was good, what does?”
Thanks for reading The Thinking Lens. Subscribe for free to receive the next part of On Judgment.
Selected research
- Baruch Fischhoff & Stephen B. Broomell (2020). Judgment and Decision Making. Annual Review of Psychology, 71, 331–355.
Used here for the distinction between judgment, preference and choice. - Robert S. Billings & Lisa L. Scherer (1988). The effects of response mode and importance on decision-making strategies: Judgment versus choice. Organizational Behavior and Human Decision Processes, 41(1), 1–19.
Used here for the distinction between judgment and choice as different decision processes. - Daniel Kahneman & Gary Klein (2009). Conditions for intuitive expertise: A failure to disagree. American Psychologist, 64(6), 515–526.
Used here for the conditions under which expert intuition can become reliable. - Oded M. Kleinmintz, Tal Ivancovsky & Simone G. Shamay-Tsoory (2019). The two-fold model of creativity: The neural underpinnings of the generation and evaluation of creative ideas. Current Opinion in Behavioral Sciences, 27, 131–138.
Used here for the distinction between generating and evaluating creative ideas.