> ## Content Index
> Fetch the complete content index at: https://www.the-thinking-lens.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# If AI Can Judge Better Than Us, What Should Humans Still Do? [On Judgment - Part 6]
- URL: https://www.the-thinking-lens.com/if-ai-can-judge-better-than-us-what-should-humans-still-do-on-judgment-part-6/
- Published: 2026-10-03T13:24:30.000Z
- Updated: 2026-10-06T05:27:42.000Z
- Description: If AI makes better assessments, what should humans still do? Human presence is not control. The real task is to design useful oversight and preserve the capability people need to challenge the system.
- Author: Amodiovalerio Verde
- Tags: #series-on-judgment, #lens-judgment-decisions

Suppose an AI system becomes better than your human reviewers at a bounded task.

Not better at everything.

Better at one defined assessment.

Fraud risk.  
Support intent classification.  
Code defect detection.  
Document triage.  
Default probability.

The obvious governance answer is familiar:

> Let the AI recommend. Keep a human as the final decision maker.

It sounds safe.

It also skips the most important design question.

**What exactly is the human adding?**

Imagine a high-volume review workflow.

The model sees more relevant evidence than the reviewer. Its error rate is lower on the defined case class.

Human reviewers see the recommendation first. They rarely override it.

When they do, the override is, on average, worse.

Policy still requires a final human click.

What function is that click performing?

Accountability?  
Legitimacy?  
Exception handling?  
Quality control?

Or ceremony?

The deeper I went into this part of the research, the less useful “human in the loop” became as a design principle.

**“Human in the loop” is a location. It does not tell you the job.**

So, I want to go one level deeper than the position of the human in the loop.

I want to know the function.

---

## Start by separating the functions

A human can play many distinct roles in an AI-enabled decision system.

They can:

- *define* the standard;
- *assess* the individual case;
- *approve* the action;
- *handle* exceptions;
- *hear* appeals;
- *audit* samples;
- *monitor* system performance;
- *stop* the system;
- *remediate* harm;
- *own* the policy;
- *remain accountable* for the organizational outcome.

Those are different functions. A system can need some of them without needing all of them at case level.

When AI is the better assessor, the human role can move from checking every case to owning standards and policy, handling exceptions and appeals, monitoring performance and retaining meaningful intervention rights.

That sounds obvious.

Yet many governance conversations still compress them into one question:

> “Is there a human in the loop?”

The result is often a human reviewer placed at the end of a workflow with little time, weak information, and a strong automation anchor.

Formal authority remains human.

**Practical control may already be elsewhere.**

---

## First question: who is better at the assessment?

If a judgment is an assessment, then the first allocation question should be empirical where the task allows it.

For a defined class of cases:

> Who or what produces the better assessment?

Not the best human in theory.

Not the model on a benchmark disconnected from the workflow.

The *actual* comparison.

Human baseline.  
AI baseline.  
Combined system.

Relevant population.  
Relevant environment.  
Relevant errors.  
Relevant time period.

And a clear validity boundary.

Lower average error is not enough to settle the allocation.

Humans may still add value by catching rare, costly failures or recognizing cases outside that validity boundary. The comparison needs to test those functions too.

This matters because “human + AI” does not automatically beat both.

A [2024 systematic review and meta-analysis](https://doi.org/10.1038/s41562-024-02024-1?ref=the-thinking-lens.com) in *Nature Human Behaviour* examined 106 experiments that compared humans alone, AI alone and human-AI combinations.

On average, the combinations improved performance relative to humans alone.

But they performed worse than the better of human or AI alone.

The effect also varied by task. Decision tasks showed losses from combination on average, while creation tasks looked more promising.

That does not mean “do not combine humans and AI.”

It means **combination itself is not a safety guarantee or a performance strategy**.

Sometimes the human corrects the model. Sometimes the model corrects the human.

Sometimes they anchor each other into a worse result.

So before adding a mandatory human approval step, we should be able to say what failure mode that step is intended to catch.

---

## The better assessor does not automatically own the decision

This is where capability and authority separate.

Suppose a model predicts default risk more accurately than a human credit analyst.

That tells us something important. It tells us the model deserves substantial evidential weight for that defined prediction task, assuming the evaluation is valid.

It does not tell us:

- what level of default risk is acceptable;
- which false positives the institution is willing to tolerate;
- what fairness constraints apply;
- which customers deserve recourse;
- which products should be offered;
- who is legally permitted to make the decision;
- who is accountable when the policy causes harm.

**Prediction quality does not create policy legitimacy.**

This is a distinction I think AI product teams need to make much more explicitly.

The model may answer:

> “How likely is this event?”

The organization still has to answer:

> “What should follow if that likelihood is high?”

Those are not the same question.

And the second question is often not a prediction at all. It contains values, rights, risk appetite, and institutional authority.

---

## The same assessment can have very different force

Return to [operational force](https://www.the-thinking-lens.com/what-happens-when-a-company-tries-to-scale-judgment-on-judgment-part-5/).

A fraud risk score displayed to an analyst is one product.

The same score automatically blocking a payment is another.

The model may be identical. The consequence is not.

That means governance should attach not only to the model but to the **model-action connection**.

What can this output cause?  
Under which confidence or risk conditions?  
Who can change that mapping?  
Where does recourse exist?  
Who can stop it?

### Case 1: Stripe Radar

Stripe Radar is a useful example because it makes several of these layers visible.

Radar uses machine-learning fraud detection and exposes risk information. Businesses can also configure rules that allow, block, review or request 3D Secure authentication based on transaction attributes and risk levels.

Notice the decomposition.

The model can *assess* risk.

The business *sets rules* around what should happen at different levels of risk.

Some cases can be blocked automatically.  
Some can trigger additional authentication.  
Some can be sent to manual review.

The same underlying risk assessment can therefore receive different operational force. That is much more useful than saying “AI decides fraud.”

The real system asks several questions:

- How good is the risk assessment for this transaction class?
- Which risk level should trigger which intervention?
- What is the cost of a false positive?
- How can a legitimate customer recover from a wrong block?

A manual-review queue is one part of the design.

It is not the definition of responsible control.

### Case 2: John Deere's See & Spray

John Deere's See & Spray Gen 2 makes the model-action connection physical.

Cameras and machine learning distinguish weeds from crops as the sprayer moves through a field. When the system identifies a weed, it activates individual nozzles to apply herbicide.

The operator does not approve each detection before the nozzle opens.

The human function sits at a different level.

The operator configures the spraying operation. In See & Spray's separate variable-rate capability, the operator can set rate and biomass thresholds, and the machine adjusts application rates within those settings.

The distinction is concrete.

Identifying a weed is an assessment.

Opening a nozzle gives that assessment immediate physical force.

Setting application parameters and deciding when to use the system are different functions.

This example does not establish that the machine identifies weeds better than a person. It shows what must be designed once an assessment is connected to action:

- In which crops and field conditions is detection reliable?
- How are missed weeds and mistaken sprays noticed?
- Which application settings can the operator change?
- Who can pause or stop the application?
- What evidence should change how the system is used?

A human click on every detected plant would not answer those questions.

Across the Stripe and Deere cases, control can live in rules, operating settings, monitoring, recourse and stop rights.

Sometimes a case-level human review is the right control.

Sometimes a machine-enforced boundary is stronger.

Often you want both at different layers.

---

## The human override can make things worse

This is the part many teams find uncomfortable.

If the AI assessment is better on a bounded task, forcing a human to review every case can degrade performance.

Humans can overweight the recommendation.  
Or distrust it inconsistently.  
Or override it based on weak intuition.  
Or spend so little time on each case that the review adds latency without adding signal.

This is not an argument for removing humans. It is an argument for specifying the human function.

If human reviewers are demonstrably valuable on ambiguous edge cases, route those cases.

If humans are better at appeals because they can consider new evidence or contested standards, design that function.

If humans are needed to own the policy, keep policy authority human.

If humans need stop rights, make them real and usable.

But do not make someone click Approve on 20,000 routine cases so the architecture can claim to be human-controlled.

That is not respect for human judgment.

**It is administrative theatre.**

---

## Then the second problem arrives: human capability

Suppose we do the allocation well.

AI handles 95% of routine cases because it is demonstrably better and faster there.

Humans handle the remaining 5%.

The system may perform better immediately.

Now ask what happens to the humans over time.

Their case distribution has changed. They no longer see the normal pattern.

They see edge cases.  
Ambiguous cases.  
Failures.  
Novel conditions.

The exact cases where strong judgment may be hardest to maintain.

**Automation allocates work. Work allocates practice. Practice shapes capability.**

How does someone become good at the exception if they no longer experience the normal case?

This is the strongest version of the deskilling concern.

But again, I do not think the answer is to preserve routine work automatically.

### Automation can also improve formation

The customer-support evidence from [earlier in this series](https://www.the-thinking-lens.com/where-does-good-judgment-come-from-on-judgment-part-4/) matters again.

In the [published field study](https://doi.org/10.1093/qje/qjae044?ref=the-thinking-lens.com) of thousands of support agents, generative AI assistance produced larger gains for less-experienced workers and evidence consistent with learning.

That is a direct challenge to the idea that AI involvement necessarily destroys apprenticeship.

**Routine work is not automatically formative work.**

Some repetition teaches.

Some repetition is just repetition.

AI can remove low-value effort while increasing access to expert patterns.

It can also remove exactly the cases humans need to learn.

The effect is conditional.

So the formation question should not be:

> “How much work are humans still doing?”

It should be:

> **Which experiences contain the signal the future human role still needs?**

That is much more actionable.

### Selective automation is selective practice

Every automation choice that changes which cases, evidence, feedback or decisions humans encounter changes their learning distribution.

If AI handles easy cases, humans increasingly encounter harder ones.  
If AI handles common cases, humans see rarer ones.

If AI proposes the answer before humans think, independent assessment practice decreases.  
If AI gives feedback after a human attempt, practice may improve.

If AI removes tedious information gathering, humans may spend more time on interpretation.  
If AI removes interpretation, humans may become supervisors of a capability they rarely exercise.

**That is why workforce design cannot be separated from workflow design.**

---

## Design formation for the role that remains

If the future human role is exception handling, supervision or policy ownership, train for that role directly.

Possible mechanisms include:

### Independent shadow samples

Give humans a sample of cases to assess before they see the model recommendation. This preserves independent assessment practice, and supports calibration and helps detect model drift.

### Simulation

Rare failures may be better practiced in simulation than by waiting for production. Incident response already uses this logic.

### Edge-case drills

Curate representative difficult cases and compare assessments, not only final answers.

### Contrastive review

Show why two similar cases received different outcomes. That often teaches the boundary better than another generic policy document.

### Coaching

Use expert feedback where it produces learning rather than requiring experts to perform every production case.

### Proficiency testing

If a role exists to supervise a consequential system, test whether the people in that role can still perform the relevant assessment.

### External expertise

Not every low-frequency capability must be developed internally. Sometimes reliable access to specialist expertise is better than pretending every organization can maintain it through occasional production exposure.

The principle is simple:

**Do not preserve bad production design in the name of apprenticeship if you can create better practice deliberately.**

---

## A decision-class allocation review

For one AI-enabled workflow, I would ask:

1. **What is the assessment?**
2. **Who or what is empirically better at it in this defined class?**
3. **Who owns the standard?**
4. **What operational force can the output receive?**
5. **What gets deferred, challenged or appealed?**
6. **Who monitors the system and who can stop it?**
7. **What human capability remains necessary?**
8. **How will that capability be formed, measured or sourced?**

You may not need all eight questions for every workflow.

But if the only design answer is “a human approves it,” I would assume the system is under-specified until proven otherwise.

---

## What I no longer believe

I no longer believe “human final say” is a safe universal rule.

Sometimes it is necessary.  
Sometimes it adds value.  
Sometimes it is required by law or policy.

Sometimes it creates the illusion of control while making performance worse.

I also do not believe the opposite claim.

AI outperforming humans on a benchmark does not create legitimate authority over the workflow.

Capability matters. So do standards, rights, accountability, recourse and control.

And I no longer think automation inevitably deskills people.

It changes the learning environment. We have to inspect how.

At this point, the series has accumulated a suspicious number of distinctions.

Assessment versus decision.  
Outcome versus assessment quality.  
Experience versus informative experience.  
What a standard requires versus what it can cause.  
Capability versus authority.  
Human presence versus human control.  
Formation versus workload.

This is the moment where a framework diagram starts looking very tempting.

And that is exactly what worries me.

The final question is not whether I can turn all of this into a comprehensive model.

It is whether a smaller answer survives.

---

## Sources

- Michelle Vaccaro, Abdullah Almaatouq & Thomas Malone (2024). [When combinations of humans and AI are useful: A systematic review and meta-analysis](https://doi.org/10.1038/s41562-024-02024-1?ref=the-thinking-lens.com). *Nature Human Behaviour*, 8, 2293–2303.
- Stripe Docs. [Radar](https://docs.stripe.com/radar?ref=the-thinking-lens.com) and [Fraud prevention rules](https://docs.stripe.com/radar/rules?ref=the-thinking-lens.com).
- John Deere. [See & Spray Gen 2](https://www.deere.com/en-us/products-solutions/sprayers-and-applicators/see-spray?ref=the-thinking-lens.com).
- Erik Brynjolfsson, Danielle Li & Lindsey Raymond (2025). [Generative AI at work](https://doi.org/10.1093/qje/qjae044?ref=the-thinking-lens.com). *The Quarterly Journal of Economics*, 140(2), 889–942.