What Remains of Judgment? [On Judgment - Part 7]

Organizations scale decisions by distributing assessments, standards, authority and feedback across people and systems. The challenge is ensuring those systems remain trustworthy, contestable and correctable.

What Remains of Judgment? [On Judgment - Part 7]

What Remains of Judgment?

I started this series because I kept using one word as if I knew what I meant.

Judgment.

Behind it was a more practical question: Where did the thinking go?

But the idea I expected to find did not survive intact.

The research began with a broad intuition:

Organizations scale judgment by externalizing it into products, metrics, processes, rules, software and eventually AI.

Organizations can scale particular assessment capabilities. They can also make standards, preferences and prior decisions durable. But that is not the same as scaling one general faculty called judgment.

What they build are decision systems that connect assessments and standards to authority, action and feedback.

That is the smaller claim I would now defend.


What I now mean by judgment

The definition I would keep is deliberately thin:

A judgment is an assessment about an object or question. Judging is the act of forming it.

I am using this as an operational convention for examining decision systems, not a universal definition across every field that studies judgment.

Saying someone has “good judgment” is a different claim: they reliably produce good assessments of some relevant kind under relevant conditions.

An assessment can be predictive: “This customer is likely to churn.” Evaluative: “This product experience is coherent.” Normative: “This policy is fair.” Practical: “This system is ready to launch.”

But that assessment is not the entire decision system around it.

Values can still shape what we assess and the standard we use. Choosing or justifying that standard and assessing something against it are distinguishable functions, even when they happen together.


What survived, and what did not

Five distinctions survived the earlier essays.

1) Different assessments need different standards. (Part 2)
2) Outcomes can be evidence about assessments without being verdicts on their quality. (Part 3)
3) Skill depends on informative practice and feedback, not tenure alone. (Part 4)
4) Rules and metrics connect assessments or standards to action in different ways. I called that connection operational force. (Part 5)
5) And a machine being the better assessor does not settle who should make the decision. (Part 6)

What did not survive was my larger explanation of scale.

I thought organizations encoded human judgment in rules, metrics and software. Sometimes they do. But durable standards can also come from regulation, negotiation, copied practice, optimization or years of incremental change.

Nor is compression always a loss. It can remove useful distinctions. It can also remove noise, bias and inconsistency.

I would not keep organizational judgment as one general scalable faculty, universal encoding as the central mechanism, or translation loss as the central problem.

That leaves a more useful question: what do these distinctions change in a consequential decision system?

I also need to revisit something I wrote before this series.

In Turn AI into a Judgment Multiplier, I argued that AI should amplify human judgment and that people should remain in the decision loop.

Part of that still stands. AI can help us frame questions, test assumptions and examine evidence.

But it can also form a better assessment than a human in a defined class of cases. And human approval is not meaningful control simply because a person clicks a button.

I would now ask what the human is actually contributing, and what the assessment can cause.


Is judgment still the right umbrella?

Why keep the word at all?

Conceptually, yes. Practically, only sometimes.

Judgment remains useful because an assessment can stay implicit even when the surrounding decision process is well specified.

A strategic decision may have excellent role clarity, clear alternatives and a strong process while still relying on a weak assessment nobody has made explicit.

An AI governance program may have permissions, approvals and monitoring while still failing to ask whether the model is actually the best assessor for the cases it is handling.

A metric can be precisely defined while carrying the wrong evaluative standard.

The judgment lens helps expose that layer.

But once the assessment enters a real organizational system, judgment is no longer the complete unit of design.

The more practical unit is the decision class.

A recurring type of consequential decision or action.

This is a unit for examining recurring operational work, not a universal approach to every one-off strategic decision.

Fraud review. Code merge. Ranking. Pricing exception.

The class needs a defined population and relevant operating conditions. It may contain several assessments, each with its own standard and validity boundary.

It lets us ask who or what should assess, what follows, where challenge exists and how the system learns.

That is more useful than trying to give the organization “better judgment” in the abstract.


Where the lens adds real value

I think the lens is strongest when an assessment becomes consequential beyond the moment in which it was formed.

It gets reused. Automated. Turned into a default. Used by a system with authority to block or execute.

Scale changes more than volume. A shared model or rule can make related errors across many cases.

The question is no longer only whether an assessment is good. It is what standard it serves, what it can cause and whether the system can recognize when it needs to change.


The smallest practical diagnostic I would keep

After all of this, I would publish four questions. They consolidate the durability questions from Part 5 and the human-AI allocation questions from Part 6. They do not replace the more specific checks in either essay.

1. What assessment are we relying on, for which cases, and what makes a good answer?

Make the assessment explicit.

Name the standard.

Specify the population, relevant conditions and uncertainty.

2. Who or what should form it, and what authority or operational force can follow?

Separate capability from the right to act.

Be explicit about what the output can cause.

Where performance can be measured, compare the actual human, machine and combined workflow, including relevant error patterns and rare costly failures.

Identify who has authority to commit to an action. Consider stakes, reach and reversibility, as well as direct actions, defaults and incentives.

3. Who can challenge, defer, appeal, change or stop it, and how?

Do not confuse formal human presence with meaningful control.

Specify the owner and usable route. Human roles need the information, time, competence and authority to perform their function.

4. What evidence should change the assessment, standard or action rule, and who will act on it?

Design the feedback loop before the story hardens around the outcome.

Preserve enough of the assumptions to recognize contrary evidence, staleness or cases outside the valid scope.

Probabilistic forecasts need evaluation across relevant cases. Contested normative standards need justification and challenge as well as feedback.

Then one conditional question:

If automation changes the cases humans will later need to handle or supervise, how will the required capability be formed, measured or sourced?

Use existing records where they capture the assumptions, ownership and change history. The questions do not require a new artifact.


What the diagnostic changes

Imagine a B2B software company deciding which customers can enable a new autonomous AI feature.

It has a rollout checklist, named approvers and a release date. A model provides a readiness assessment.

The process can look controlled while leaving three design choices unresolved.

First, separate the assessments from the rollout decision.

“Ready” may mean technical compatibility, reliable performance under tested conditions and sufficient customer controls.

Those are different assessments, potentially with different standards and evidence. The model might be reliable at checking compatibility but not at evaluating the customer's operational safeguards.

The team can allow automatic enablement for a defined set of tested configurations, defer other cases, and keep authority over that policy separate from whoever or whatever forms the assessment.

Second, make exclusions testable.

If the company observes only customers it enabled, it cannot tell whether excluded accounts would have performed safely.

Sampling the rejected accounts is not enough if the relevant outcome remains unobserved.

The company could independently review a sample of rejected configurations, run technical or shadow tests where these are informative, and consider tightly controlled trials only where the risks allow.

In other cases, it should record the uncertainty rather than treat an untested exclusion as a confirmed correct decision.

Third, design the human capability the control path depends on.

If AI takes over routine readiness reviews, humans may eventually see only exceptions and appeals.

A formal appeal route is weak if the reviewer no longer has the competence to assess the case.

Independent case samples before seeing the model recommendation, exercises using difficult cases and periodic proficiency checks could preserve and test that competence without requiring human approval for every account.

None of those choices follows automatically from having a checklist and named approvers.

The three choices are connected.

What the readiness assessment is allowed to cause affects which outcomes the company can observe.

Automating the routine reviews changes which cases human reviewers get to practice on.

Both shape whether the company can recognize and correct a mistake later.

The lens earns its place only if it reveals a consequential assumption or leads to a better system design that the normal process would have missed.


When existing frameworks are enough

None of these questions is unique to this lens.

Decision Quality already addresses framing, alternatives, information, values, reasoning and commitment. RAPID clarifies decision roles. NIST AI RMF addresses the broader system of AI risk management.

I am not claiming those approaches overlook assessment quality.

The narrower contribution is where the analysis starts: name the assessment inside a recurring decision class, then follow its standard, validity boundary, operational force, challenge route and update path.

Those concerns may already be covered, but not always connected to the same assessment.

For a low-frequency strategic decision, Decision Quality and decision-rights tools may be enough. For a mature AI risk program, the relevant controls may already exist.

If they make the consequential assessment and its effects inspectable, I would use them.

A lens should be able to tell you when it adds nothing.


What remains unvalidated

Across this series, I have examined the ideas against research literature and used public cases to illustrate the mechanisms.

I have not demonstrated that these four questions improve business outcomes.

The next test would compare the diagnostic with normal practice and the nearest appropriate existing framework on comparable consequential decision classes.

Define quality criteria beforehand, and measure the time and operating costs of each approach.

Then examine whether it exposes consequential assumptions or leads to better-justified decisions and system designs relative to the alternatives.

A lens should survive because it provides useful information or improves design relative to its cost, not because it has a good name.


So what happens to judgment when it scales?

I no longer think that is quite the right question.

Organizations can scale particular assessment capabilities. They can make standards, preferences and prior decisions durable. They can connect those things to action through rules, software, people and incentives.

But those are different operations. Calling all of them judgment hides the design choices.

The more useful question is whether the system can tell what it is relying on, when that reliance is justified and how to change it when it is not.

That brings me back to where I started.

Where did the thinking go?

Some of it did not disappear. It became distributed across assessments, standards, metrics, rules and systems.

Not all of those things began as someone's judgment.

I used to say, "Data informs. Judgment makes the call." I still think there was something useful in that line. But I was asking judgment to cover too much. Assessing something, choosing what to do and having the authority to act are different functions.

I also thought the task was to preserve judgment as organizations scale.

I now see a more specific problem.

A consequential assessment or standard can keep working exactly as designed after the conditions that justified it have changed.

A human approval step does not solve that by itself.

The task is not to preserve judgment as one general human faculty. It is to make the assessments that shape consequential action visible, challengeable and correctable.

That is the smaller claim I can defend now.


Sources