Claudefishing Is a Real Problem. The Detector Cannot Solve It.
Substack’s new detector can estimate whether text looks AI-generated. It cannot determine whether a human thought, understood, and owned the argument.
[Views are my own.]
On 21 July 2026, Substack introduced an AI-detection feature powered by Pangram. Readers can scan posts, Notes, comments and replies to see an estimate of how much was written by a person and how much involved AI.
Substack calls the underlying problem “Claudefishing”: publishing machine-generated writing in a way that suggests a real person is thinking and speaking to the reader. It is a new term for a legitimate concern.
But I do not think the detector measures the problem Substack has described. Claudefishing is about whether a person genuinely did the thinking. Pangram can only analyze the resulting text.
The real problem
Chris Best was precise in his announcement. The problem he named is a specific one: a reader invests attention in a piece of writing expecting a human mind on the other end, and there is no human mind there. The writing is automated. The connection is simulated.
That is a real and meaningful concern. Readers have a right to know whether someone is thinking in public or running a production line. It affects trust, attention, and the economics of independent writing.
But notice what the problem actually is: the absence of human thought, not sentence patterns, lexical variation, or any other stylistic score.
Whether a person did the thinking is not directly observable from the final text alone. Text may provide evidence of thought, but it cannot reveal the complete process that produced it. A detector cannot observe that process.
What Pangram can see
Pangram is a neural classifier. It classifies textual patterns using representations learned from labeled training data, samples marked as human-written or AI-generated. It estimates how much of the text appears consistent with machine-generated or AI-assisted writing versus unassisted human writing.
The estimate is inferred from textual patterns; it does not observe the process that produced the text.
Take a writer who develops the argument independently and then uses AI to challenge the wording. They reject many of its suggestions and publish a position they genuinely hold. Pangram may still report substantial AI assistance, even though the thinking was theirs.
Now reverse the situation. Someone generates an article, skims it, paraphrases a few sections and publishes it. The score might be low, but the intellectual engagement is also low.
The detector classifies patterns in text. It cannot measure a person behind the argument.
This gap is not a technical limitation that better training data will close. It is structural. Stylistic patterns and cognitive processes are different kinds of thing. Assume for a moment that Pangram is highly accurate at identifying AI-generated or substantially AI-edited text. It still cannot distinguish thoughtful AI assistance from handing the thinking to a model.
From the final paragraph alone, the detector cannot know what happened during drafting. I might have spent three days developing the argument and used a model to shorten two sentences. Someone else might have generated the whole piece and changed a few phrases.
The category error is not in estimating whether AI was involved. It occurs when that estimate is treated as a proxy for, or verdict on, whether a human engaged with, directed, and intellectually owned the work. It is like penalizing an engineer for using a calculator.
Then I tested the distinction
I had already finished The Detector Doesn’t Know Cicero before running this experiment.
It argued that detectors classify surface patterns in final text, and that those patterns cannot establish who originated an argument, whether the writer understood it, or whether they intellectually owned it. I wrote it before uploading anything to Substack.
When I uploaded the finished draft, Pangram returned a verdict: 100% AI-generated or AI-assisted.

I was still making the same argument, using the same sources and evidence. The framework and conclusion had not changed. The changes were mostly stylistic: sentence rhythm, punctuation, contractions and a reduction in compressed parallel constructions. Several qualifications and explanatory passages also changed, but the underlying argument and intellectual position did not.
The second version received a verdict: 100% human.
The two versions were not identical. They did not need to be. The question was whether the thinking behind them had changed.
It had not. I was still the author, using the same sources to reach the same conclusion. What changed was how that thinking was expressed.
What remained stable, and what changed

The comparison makes the distinction visible. The intellectual content remained materially stable. The surface style changed: sentence rhythm, punctuation, compression, structural regularity and the prominence of parallel rhetorical constructions.
What this case establishes is narrow:
The detector’s absolute-looking verdict was not stable under changes to the linguistic envelope that preserved the article’s underlying intellectual position.
The detector’s output could not establish whether the thesis was mine, whether I understood the evidence, whether I held the conclusion, or whether the argument represented human conviction. Those dimensions remained constant while the verdict changed.
This does not establish that Pangram is generally inaccurate, that every detector is useless, or that the two texts are linguistically identical. Strong benchmark performance and this result can both be true. In this case, the score changed even though the thinking did not.
Both complete versions, the original detector screenshots, the comparison methodology, the rhetorical annotation and the full line-by-line diff are available in the evidence pack.
¹ A conservative manual count of passages using adjacent parallel clauses or fragments, repeated openings, or compressed triadic structures. Overlapping devices within the same passage were counted once. The count is descriptive and does not identify which features Pangram used internally.
² Calculated using lowercase word-frequency vectors after removing Markdown formatting.
³ Calculated using Python’s difflib.SequenceMatcher on lowercase ordered word tokens.
Why universities stepped back
Several universities have restricted or abandoned AI detection in formal academic-integrity processes since 2023.
Oxford’s AI Competency Centre stated in February 2026 that it does not endorse any digital AI detector for academic decision-making. Curtin University disabled Turnitin’s AI detection from January 2026. Waterloo followed in September 2025 after internal tests classified human-written text as 100% AI-generated “in more than one instance.” Washington State University cancelled its AI-detection contract in February 2026: 33% of Review Board cases involving alleged AI use ended in a finding of not responsible because detector output had been submitted without independent supporting evidence.
These decisions concerned Turnitin in formal academic misconduct processes, not Pangram in a reader interface. The stakes, controls, and consequences are different, and the analogy should not be overstated.
The relevant comparison is not whether Pangram and Turnitin make the same errors. It is whether a probability based on writing style can bear the weight of judgments about who wrote something and whether it is genuine in any context, institutional or reader-facing.
A probabilistic estimate of textual AI-likeness has become, in reader experience, a social verdict on authenticity.
Substack imposes no formal penalty. Writers can report an error, remove a disputed scan, or disable detection on individual posts. Those controls matter. They cannot prevent a reader from screenshotting a result before it is removed and circulating it without context. The score is private to the requester in principle. In practice, it is whatever the reader chooses to do with it.
Who is most exposed
Recent Pangram-specific evaluations report far fewer false positives than older detectors. That evidence matters.
What it does not establish is how the newly deployed Substack feature will perform across the range of professional second-language writing, translated prose, or mixed human-AI editing workflows.
Substack describes the result as estimating how much text was written “by hand or with AI assistance.” That is broader than detecting AI-generated writing. Translation, grammar editing, vocabulary polishing and structural rewriting could all contribute to an AI-assisted classification, even when the argument, research and intellectual position are the writer’s own. A non-native English writer using AI to correct grammar while their ideas, research, and positions are their own may score as AI-assisted. The percentage does not tell the reader whether the assistance concerned grammar, expression, argument formation or intellectual substitution. The detector cannot distinguish that.
The 2023 Stanford study demonstrated severe false-positive rates against non-native writers in earlier systems. That study did not test Pangram. Whether Pangram reproduces or avoids that failure in real-world Substack use, across professional writers from different rhetorical traditions, writing about complex subjects, editing with AI tools, has not yet been independently established.
That is not a claim that Pangram is broken. It is an unresolved risk for a reader-facing feature with reputational consequences.
Detection is only one layer
The problem Substack named is real. The detector is insufficient on its own. That does not mean self-description alone is the answer.
A writer willing to publish machine-generated text as personal thought may also be willing to submit a false process description. Self-attestation selects for honest disclosure. It is least effective against the exact behavior called Claudefishing.
Medium combines disclosure requirements, distribution restrictions on predominantly AI-generated content, automated signals, and human review. The Authors Guild’s Human Authored certification requires a signed license agreement and issues a numbered mark listed in a public database; non-members also undergo identity verification.
A better system is layered: clear disclosure standards, author attestation that has consequences when someone lies, human review for disputed cases and any action affecting the writer rather than treating the automated score as the decision, and automated signals that trigger inquiry rather than deliver verdicts. No single component solves the problem. The combination changes the incentives.
Transparency built on self-description creates accountability only when the description carries consequences if it is knowingly false.
What the score cannot tell readers
Substack is right that readers deserve to know whether a human mind is behind what they are reading.
Substack explicitly acknowledges that the estimate cannot establish whether human care went into a text. The product risk is that readers may nevertheless interpret the percentage as a verdict that the text either contains, or lacks, real human judgment.
It establishes how the text looks relative to a training set. That is a different question. And treating it as the same question, as an answer to whether human thought was present, is the central risk in how the Claudefishing feature may be interpreted.
A detector can be a useful signal that AI was involved in producing a text. It cannot determine whether that involvement was deceptive, whether a human directed and owned the argument, or whether what looks like AI assistance is actually an ESL writer editing for fluency.
Claudefishing is real. The detector cannot solve it on its own, not because it is necessarily inaccurate, but because the problem is about presence of thought, and the tool measures patterns in text.
Better training data may improve estimates of AI involvement. It cannot turn a text-only classification into proof of who did the thinking, what the writer intended, or whether deception occurred. Those are not the same thing.
The research behind this argument, on detector calibration, intercultural rhetoric, and the formation penalty, is in The Detector Doesn’t Know Cicero.