Weekend Reflections #18 | The Expected Answer

A certification needs one correct answer. Real work has an annoying habit of adding an asterisk.

Weekend Reflections #18 | The Expected Answer

[Views are my own]

What do you do on a Sunday morning?

Apparently, I get certified.

I took the Claude Certified Associate exam this weekend.

Not really for the badge.

Although I do have an unhealthy affection for collecting those.

I wanted to assess the assessment.

Does an AI certification actually test something useful?

Or can someone who has spent enough time prompting, understands the basic vocabulary, and is reasonably good at multiple-choice exams get through it with common sense and educated guessing?

There was more substance than I thought.

A lot of the material was transferable beyond Claude: how you frame a task, evaluate an output, decide what should be delegated, structure a workflow, and recognize when the human still needs to do the work.

But one particular kind caught my attention.

The questions where I understood the concept, understood the options, and still did not entirely agree with the expected answer.


That sounds like a criticism of the exam, but it is not.

It made me notice something about assessments more generally.

An assessment does not only test your model of good practice. It also exposes the assessor's model of good practice.

Every answer key contains assumptions.

About what "good" looks like, which risk matters most, what should happen first, what is acceptable to delegate, and where a boundary should sit.

Most of the time, you do not see those assumptions because you agree with them.

They become visible when you do not.

I knew what answer the question was looking for.

I could explain why that answer made sense.

And in a real situation, I might still have made a different choice.

Not because the expected answer was obviously wrong.

Because I could imagine the missing context.


A rule is easy to state.

Experience can teach you to notice the conditions around it.

"When X happens, do Y."

Fine.

Until you have seen the case where Y creates another problem.

Or where two principles conflict.

Or where the answer changes because the stakes, user, data, organization, or failure cost are different.

Experience adds asterisks. Standardized assessments have to remove most of them.

An exam cannot give every question fourteen paragraphs of organizational context and then accept three defensible answers depending on how you interpreted paragraph eleven.

So messy practice has to be compressed into something testable.

Usually one best answer.

If I select something else, the scoring system knows only that I did not select the expected answer.

It cannot know why.

Maybe I did not understand the concept.
Maybe I guessed.
Maybe I misunderstood a word.
Or maybe I understood the model perfectly and disagreed with where it drew the boundary.

Those are very different things.

To the answer key, they look identical.


That does not make certification useless.

Actually, I came away thinking almost the opposite.

A good assessment can make a professional model unusually visible.

Reading documentation tells you what a product can do.

Training tells you what someone wants you to learn.

An assessment forces someone to decide what distinctions matter enough to test.

In that sense, an assessment is a model made executable. It has to turn principles into choices about what matters more.

The plausible wrong answers were particularly useful for me.

They showed where two approaches that sound almost equivalent are considered meaningfully different.

In a way, they map the boundaries the assessor cares about.

When two answers both sound reasonable, the reason one is still considered wrong tells you which distinction the assessment considers decisive.

And occasionally I found the edge of the model and thought:

I understand why that is the expected answer.

I am still keeping my asterisk.

A credential is evidence that I can recognize and apply the model being assessed.

It cannot show you every place where I would challenge it.

And perhaps it should not try.

The exam needs an answer key. Practice needs judgment about when the answer key stops being enough.