On Judgment, Part 1: I Keep Talking About Judgment. What Do I Actually Mean?
We keep saying judgment will matter more. I have been saying it too. But before we build careers, organizations and AI systems around that assumption, I think we need to understand what we're actually talking about.
August gave me enough space to think about what I wanted to keep exploring.
I asked myself: what do I actually care about enough to keep thinking and writing about for another year?
Looking back across more than a year of The Thinking Lens, across product, data, AI, organizations, governance and leadership, I noticed I kept reaching for the same word: judgment.
I use it when I write about data and the limits of metrics; I use it when I talk about frameworks and decision-making.
It shows up in my thinking about AI, organizations and increasingly about agents and autonomy.
The more I looked, though, the less obvious the word became.
Take a few examples.
A product team has ten metrics available. Which one actually matters?
Five customers ask for the same feature. Is that an early market signal, or just five unusually vocal customers?
An AI agent produces an answer based on three sources. Which of those sources deserves trust?
A company makes a strategic bet with incomplete information.
An agent has performed the same task correctly hundreds of times. At what point do we stop asking a human to approve every action?
A product looks polished, the numbers are fine, but something still feels wrong.
I have called all of these judgment.
But they are clearly not identical.
Some are about interpreting evidence.
Some are about taste, which seems to be a surprisingly controversial word these days.
Some are about making a prediction when you simply do not know enough.
Some are about deciding how much risk to accept.
And somewhere along the way I also mixed judgment with responsibility: who owns the decision when things go wrong?
They are clearly related.
I am just no longer convinced they fit under one definition.
Which is slightly inconvenient, because I use the word all the time.
More importantly, I have repeatedly argued that companies need to preserve, improve or scale judgment.
That is a much bigger claim.
So before I keep arguing that judgment matters, I want to understand what I actually mean by it.
I did not start with judgment
The Thinking Lens did not begin as an exploration of judgment.
I was trying to answer a much more practical question:
Where did the thinking go?
I kept seeing product organizations with all the things we normally expect to see.
Roadmaps. Discovery. OKRs. Dashboards. Prioritization frameworks. Rituals. Backlogs.
There was no obvious lack of process.
And yet, when you looked at some decisions closely, it was surprisingly difficult to reconstruct how the team had arrived there.
Sometimes a framework had stopped being a tool for thinking and had quietly become the answer.
Teams would come out of discovery with useful learning, but a few weeks later much of that context was hard to find in the backlog or in the decision itself.
And teams could usually explain what they were building.
The harder question was “Why this, rather than something else?”
That was the part that kept bothering me.
The process was still there.
The artifacts were still there.
But the reasoning connecting them could disappear.
Over time, though, I realized I kept coming back to roughly the same question:
How do we preserve the thinking that connects evidence, context, trade-offs, choices and learning?
That is where judgment entered the picture.
I did not start with the word and then go looking for a problem that fitted it.
I started with the problem.
Judgment became useful later, because it gave me a way to describe something I had already been trying to understand.
Maybe it became too useful.
When the machinery stops deciding
Let's say enterprise customers use one capability three times more than smaller customers.
I have looked at signals like that and thought:
“Clearly, we should invest here.”
But that conclusion does not actually come from the data.
You could just as reasonably reposition the product, change the target segment, or decide the signal is strategically irrelevant.
This was where I started using a line I still like:
Data informs. Judgment makes the call.
I still think there is something important in that.
What I am less sure about now is what I meant by “the call.”
One thing is working out what the evidence means.
Another is deciding whether that evidence matters.
Another is deciding what to do about it.
I called all of them judgment.
At the time, that distinction felt academic.
Today, I think it matters a lot.
Then judgment started expanding
The word became broader as I kept writing.
When I wrote From Content Is King to Judgment Is the Crown, I was thinking about what happens when AI makes competent production abundant.
A decent first draft, clean slide, research summary, prototype or strategy memo is no longer scarce.
The blank page is becoming less and less the problem.
What becomes interesting is everything around the artifact: the context you have, what you have seen before, what you notice that someone else might miss, and the ability to know that something apparently good is actually generic.
I called that judgment too.
Then came the Judgment Economy.
I had used judgment to describe filtering signal from noise, contextualizing information, deciding what to trust and synthesizing conclusions.
They seem to answer different questions.
What deserves my attention?
What does this mean here?
What should I believe?
And once I have all of that, what can I reasonably conclude?
I originally put all of this under the label of judgment.
Some of these feel like reasoning.
Some feel more like evaluation.
Others may be steps that happen before judgment is even possible.
And once a concept starts combining several different abilities, it becomes much harder to talk about clearly.
But I think the bigger change happened later.
Because I gradually stopped talking only about something a person does.
The Semantic Supply Chain made the ambiguity operational
The ambiguity around judgment became much harder for me to ignore while I was working on the full, unpublished version of the Semantic Supply Chain.
What surprised me was how often questions of judgment appeared between the layers.
Retrieval is a good example.
A semantic system can tell you that one document is more similar to a question than another.
But similarity is not truth. And similarity does not tell you which source deserves trust.
You still have to decide which sources should be trusted, how much evidence is enough, how relevant something needs to be, and when the correct answer from the system is simply “I don't know.”
Then there is evaluation.
A Golden Set takes examples of what we consider “good” and “bad” and turns them into something repeatable.
We can go further and have one LLM evaluate another against a rubric.
When I first looked at this, I mostly thought of it as evaluation infrastructure.
Looking at it through the lens of judgment, something else is happening: we have taken some of the criteria by which someone would judge an output and encoded them into a test.
Then agents made the distinction even more obvious.
A system may be capable of recommending an action.
That still does not tell us whether it should be allowed to execute it.
How large is the blast radius? Can we reverse the action?
How certain do we need to be?
And who owns the outcome when it gets the decision wrong?
Those questions eventually led me to the Autonomy Ladder.
But they also exposed a more fundamental distinction:
Capability is not authority. And authority is not accountability.
A system may be better than any individual person at assessing a situation and still not have the authority to act.
A senior executive may have the authority to make a decision even when someone else has the better assessment.
And the company remains accountable for the outcome even if the recommendation came entirely from a machine.
I used to put all of these ideas quite close together under the word judgment.
I think they need to come apart.
And this is where I think my use of the word changed more fundamentally.
When I write that companies should scale judgment, something else is happening.
Judgment is no longer just something an individual has or does.
It becomes something closer to an organizational capability, spread across people, evidence, context, standards, memory, dissent and decision rights.
The Semantic Supply Chain pushed this even further.
Looking back, I think I had moved from describing a capability toward describing much of the system through which information becomes consequential action.
And I was still calling it judgment.
That is probably a sign that I have been asking one word to do too much work.
One useful correction from decision science
Looking outside product management gave me a distinction that I found surprisingly useful.
One influential decision-science framing separates judgment, preference and choice quite cleanly.
Fischhoff and Broomell use judgment for predicting what outcomes might follow, preference for how we value those outcomes, and choice for combining the two into a decision. But the word is used much more broadly elsewhere.
Tichy and Bennis, writing about leadership, describe judgment as a process that includes preparing for a call, making it, executing it, and learning from what follows. So apparently I am not the only one asking one word to cover different territory.
I had mostly been collapsing these into the same thing.
A simplified example helped me see the difference.
Imagine we are considering entering a new market: “There is a 70% probability that this market will become strategically important in the next three years.”
That is an assessment about what we think will happen.
Then: “Entering early matters more to us than protecting near-term margin.”
Now we are no longer assessing reality.
We are saying what we value.
And finally: “We will invest €5 million and enter now.”
That is the decision.
I suspect product people compress these things quite naturally, because real product decisions rarely arrive neatly separated into assessment, preference and choice.
A serious product call mixes customer understanding, experience, prediction, business context, trade-offs and sometimes taste.
There are good reasons why we compress these things.
But I think something gets lost when we compress all of them into one capability called judgment.
For a long time, this could remain mostly a conceptual problem.
AI makes it much harder to ignore.
AI makes the distinction urgent
For a while, I used a simple line in my writing:
AI generates. Humans judge.
I liked it because it made the division of labor easy to understand.
I no longer think it holds.
AI systems are already doing things that look uncomfortably close to what we normally call judgment.
They rank alternatives. Assess evidence. Critique arguments. Forecast outcomes. Evaluate outputs. Recommend actions.
That does not mean AI has “better judgment.”
I am increasingly unsure what that sentence would even mean.
But there are already bounded situations where machines can make assessments that we would normally associate with experienced people.
One recent 2026 working paper made this particularly concrete for me.
Researchers gave frontier models and experienced managers standardized summaries of 30 live Kickstarter technology projects to evaluate while their crowdfunding campaigns were still running.
Several frontier models predicted the projects' eventual fundraising rankings more accurately than the human benchmarks.
I do not want to overstate what that tells us.
It is one study and one bounded task.
And ranking ventures is obviously not the same thing as being a good CPO.
But it still creates a problem for the convenient definition I had been using.
If assessing uncertain future outcomes in an ambiguous business situation is something we call judgment, and a machine can sometimes do that better than an experienced human, then I cannot define judgment as:
the thing humans do after AI has finished.
That category would just keep shrinking as the technology improves.
Worse, the definition itself would change every six months.
And this also creates a problem for another line I have used repeatedly:
Generation becomes cheap. Judgment becomes expensive.
I still think there is something important behind that idea.
I am just less confident that judgment is the right container for it.
Because some of the things I had placed inside judgment are already getting cheaper.
AI can forecast, compare, critique, evaluate and synthesize.
In bounded situations, it can sometimes do these things extremely well.
So maybe judgment remains scarce.
Maybe only some kinds of judgment do.
Maybe the scarce capability moves somewhere else: setting the right objectives, recognizing when an old pattern no longer applies, deciding which evidence deserves trust, or understanding which trade-offs are actually acceptable.
Or perhaps the scarce capability is designing the environment in which human and machine judgments can challenge each other without converging too quickly on a false consensus.
I am also not sure scarcity is the right lens at all.
That is where I am now.
So what is judgment?
I am not ready to define it properly yet.
But after going back through my own writing and reading outside product management, I keep seeing roughly the same pattern.
Judgment seems to appear when the facts, rules or models do not completely settle the answer.
And when the answer is not fully determined in advance, the eventual outcome cannot by itself tell us whether the judgment was good.
Context changes how we interpret the evidence.
And usually something is at stake: a belief, a priority, money, a hiring decision, a product, a risk.
That gives me a working definition, at least for now:
Judgment is the act of forming an assessment when facts, rules or models do not by themselves settle the answer.
I can already see problems with it.
Does judgment only appear when the answer is not fully determined?
What about an expert who recognizes the answer almost immediately?
Is that judgment, intuition, pattern recognition, or all three?
Is judging whether something is beautiful or good the same kind of act as judging whether something is likely to happen?
And where do values enter?
If I say a risk is “too high”, am I assessing the risk itself, expressing my appetite for risk, or quietly doing both?
Then there is the boundary with decision-making.
If judgment produces an assessment, at what point does the decision actually begin?
I do not have good answers yet.
That feels more useful than pretending the boundaries are already clear.
Separating things for now
For now I am trying to separate evidence, reasoning, evaluation, values, decision, authority, action and accountability, while recognizing that real decisions loop between them rather than moving neatly through them.
There may be several judgments inside a single decision.
Intuition can skip over reasoning that would be difficult to articulate.
Values shape what we notice before we consciously evaluate anything.
And inside organizations, none of this happens in one head anyway.
It is distributed across people, processes, incentives, information and, increasingly, machines.
Still, separating the concepts has already helped me notice a few mistakes in my own thinking.
Reasoning is not necessarily judgment.
Judgment is not the same thing as a decision.
The person who makes the best assessment may not have the authority to decide.
The person with authority may not be the person best equipped to judge the situation.
And accountability tells us who owns the outcome, not necessarily who made the best judgment along the way.
These distinctions matter much more once humans and machines participate in the same decision system.
On Judgment
This is why I want to spend more time on the subject.
Not because I have another framework ready to publish.
Actually, I already have one in my notes.
That is probably a good reason not to publish it yet.
If I build a taxonomy before I understand what belongs inside judgment and what merely sits next to it, I will just give my confusion cleaner boxes.
So I want to take the word apart.
I want to understand what is actually judgment and what is reasoning, preference, decision, authority or accountability.
I want to understand what makes one judgment better than another, and whether good judgment can actually be learned.
I want to understand whether an organization can become better at judgment, rather than simply relying on individuals who seem to have it.
And I want to understand what happens as machines become capable of more of the assessments we have traditionally put inside the same word.
Because the question is no longer simply whether judgment matters.
It is what we actually mean when we say someone has it, how it becomes a decision, how organizations preserve and improve it, and which parts humans and machines should each perform.
And eventually this also leads to a question I suspect will become important: if a machine becomes the better judge in a particular domain, does that mean it should also get to decide?
Capability, authority and accountability are clearly not the same thing. I have put them too close together in some of my previous thinking.
But I am getting ahead of myself.
For now, though, I have reached a simpler conclusion.
I keep writing about judgment.
I am no longer sure I know exactly what I mean by it.
So I am going to spend some time finding out.
Part of why I write publicly is to expose my thinking before it becomes too comfortable.
Think of a decision you still believe was well judged even though the outcome was bad. What made you believe the judgment was sound before you knew how it turned out?
And the reverse interests me too: have you seen a poor judgment rescued by a good outcome?
If you have a real example, I would genuinely like to hear it. Anonymized is fine. I am trying to understand how we tell the difference.