The Detector Doesn't Know Cicero

The Detector Doesn't Know Cicero

[Views are my own.]

Editorial note, 26 Jul 2026: I completed this article before testing Substack's Pangram-powered AI scanner. When I uploaded the finished draft, it was classified as 100% AI-assisted. I then made a set of mostly stylistic edits (sentence rhythm, punctuation, some contractions) while preserving the argument, evidence, structure, and conclusion. The revised version was classified as 100% human. I have preserved both versions. The next article documents what changed, what the comparison can establish, and what it cannot.

AI detectors are appearing everywhere.

On academic platforms. In publishing workflows. And now on platforms where readers can scan an author's work themselves.

The premise behind these tools is straightforward: machines write differently from humans, and the difference is measurable.

The premise becomes unreliable when a population-level classification is treated as proof about an individual author without appropriate calibration, context and independent evidence.


What detectors actually measure

AI text detectors do not identify an author. They classify textual patterns learned from samples labeled "human" and "AI."

Early heuristic detectors assumed that human writing would show greater variation in predictability, sentence construction and diction, while generated writing would be more statistically uniform. GPTZero's original public version used perplexity and burstiness to operationalize this assumption. Perplexity measures how predictable a piece of text is to a language model. Burstiness measures variation in sentence length and structure. In detectors using these signals, structurally regular prose could receive a more machine-like classification.

GPTZero says it moved away from perplexity and burstiness, migrating to a supervised deep-learning architecture. Other current commercial detectors describe transformer-based classifiers that operate on learned representations rather than explicit statistical measures. The internal features of these systems are not fully observable from the outside.

Research has identified two related failure modes. Some detectors have shown substantial disparities when evaluating particular non-native-English writing populations. Separately, feature-based detectors have lost performance under domain and generator shift. The available evidence does not establish that every current detector fails uniformly across all author populations or rhetorical traditions.

That baseline, wherever it sits, is not universal. It reflects which writers, genres and linguistic traditions were represented as human.

What looks statistically human depends partly on which writers, genres, and linguistic traditions are represented in a system's training and evaluation data.


What the research has established

Stanford researchers tested seven AI detectors on 91 TOEFL essays written by non-native English speakers (drawn from a Chinese educational forum) and 88 essays written by US eighth-grade students. On average, the detectors falsely classified 61.3% of the non-native essays as AI-generated, compared with 5.1% of the US essays. At least one detector falsely flagged 97.8% of the non-native essays. All seven unanimously flagged 19.8%.

The detector was not finding AI. It was finding a particular kind of human.

The researchers then tested one likely mechanism. They used ChatGPT to enhance the word choices in the TOEFL essays to sound more like those of native speakers. Detector classifications were highly sensitive to lexical predictability rather than authorship alone. Human-written essays with more constrained word choice were frequently classified as machine-generated. After machine-assisted vocabulary enhancement, the same essays were much more likely to be classified as human.

The two groups were not closely matched: they differed in age, task context and educational setting. The result demonstrates a strong and systematic disparity. It does not isolate native-language status as the only possible contributing cause.

A 2026 preprint found that feature-based detectors with high in-domain performance deteriorated substantially under domain and generator shift. Explainability analysis indicated the models frequently relied on dataset-specific stylistic features rather than stable markers of machine authorship.

OpenAI reached a similar practical conclusion in 2023. Its public classifier identified only 26% of AI-written text as likely AI-generated while falsely flagging 9% of human text. OpenAI withdrew it on 20 July 2023, citing low accuracy.


The calibration problem

Since Kaplan's 1966 paper, contrastive and intercultural rhetoric research has documented differences in textual organization, authorial stance and structural patterning across languages, genres, disciplines and educational settings. Contemporary research treats these patterns as dynamic and context-dependent rather than as fixed national styles. The field documents tendencies, not rules.

Studies of English and Italian academic writing have found measurable differences in authorial positioning and rhetorical stance. Studies of Arabic-English transfer have documented how additive organization and lexical repetition, patterns central to Arabic rhetorical tradition, can persist in English writing. These modern studies do not prove direct inheritance from classical Latin or Arabic rhetoric. They establish that formation shapes prose in ways that can remain observable in advanced academic second-language writing.

A writer formed in one of these traditions may produce formal English that differs systematically from the writing represented as human in a detector's training data. Different in ways that may overlap with the statistical patterns a detector has learned to associate with AI.

Not worse. Not machine-generated. Just different in a way the training data may not represent adequately.

Classical rhetoric rewards deliberate structural regularity: parallelism, antithesis, tricolon, anaphora. These devices could overlap with features a detector has learned to classify as machine-like. That possibility has not been tested directly.


The evidence ladder

What detector research has established: some AI detectors systematically misclassify non-native English writing at significantly higher rates than native English writing. Part of that disparity has been linked to statistical features such as lower perplexity and constrained lexical variability.

What contrastive and intercultural rhetoric research has separately established: language, education and professional formation shape how writers organize arguments in a second language, in ways that can remain observable in advanced or professional writing. The extent varies with genre, task and conscious adaptation.

What has not yet been tested: whether formal rhetorical devices (parallelism, antithesis, tricolon, anaphora) contribute directly and independently to detector false positive rates. No study identified in this area has run that experiment.

The hypothesis that classical rhetorical formation could contribute to some false positives is plausible and consistent with two separate bodies of research. No study has tested the connection directly.


What the detector cannot see

A text-only detector applied to the final prose has access only to that prose. It cannot observe the document's production history, the origin of the argument, the rejected alternatives, or the writer's decision-making process.

It cannot measure whether the argument is original. Whether the observation is specific to an experience no model shares. Whether the position is held against resistance, against the simpler, more plausible version the AI would have produced if left alone. Whether the framework names something no one had named before. Whether the conviction behind the claim is earned or borrowed.

Those are the dimensions that matter when judging intellectual contribution. They do not, by themselves, establish provenance.

A writer who uses parallelism because their education taught them to value structural balance, and a model that produces parallel structures because they are statistically probable, may arrive at similar output through entirely different processes. The detector sees the output. It has no mechanism to see the process, the conviction, or the intellectual territory the writer has spent years accumulating.

A meaningful assessment of intellectual contribution would require something harder: the novelty of the argument, the robustness of the reasoning, the specificity of the observation, the value the piece adds that was not there before. These are the dimensions that require judgment to evaluate. That is precisely why they have been replaced by a classifier.

I call this possible failure mode the formation penalty: the risk that writers from highly patterned rhetorical traditions produce prose that overlaps with a detector's machine-like category. In such cases, a writer could be penalized not for using AI, but for writing in the ways their education and professional formation taught them to value. It is a hypothesis consistent with two separate bodies of research. No study has isolated the mechanism directly.


The organizational failure

Organizations find classifiers difficult to resist.

A classifier offers scalable certainty. It converts a complex judgment about authorship, intent and process into a percentage that can be recorded, compared and acted upon. The certainty is administrative, not epistemic.

The score allows the organization to avoid the harder work: reviewing drafts, examining sources, discussing the argument with the author, deciding what forms of AI assistance are acceptable and at what threshold they become a problem. Judgment is not improved. It is displaced.

That displacement changes behavior. Writers learn that formal structure, polished language and rhetorical consistency may be treated as suspicious. Non-native writers face pressure to make their English less predictable, less precise, less like the traditions that formed them. Employees optimize for what the classifier recognizes as human rather than for clarity, accuracy or intellectual value.

The organization then creates the behavior its detector expects. Human writing becomes deliberately irregular. AI writing is edited to look more human. The signal degrades while confidence in the score remains.

The detection problem is a symptom. The governance failure is the cause.

A responsible organization can use automated signals to trigger inquiry. It cannot use them to replace inquiry. The moment a probabilistic classification becomes a verdict, the organization has outsourced judgment without outsourcing accountability.


The institutional verdict

Several universities have reached a related conclusion in practice.

Vanderbilt disabled Turnitin's AI detector in 2023, saying it did not believe "AI detection software is an effective tool" and that the decision was made "in pursuit of the best interests of our students and faculty." Waterloo followed in 2025 after internal tests classified human-written text as 100% AI-generated "in more than one instance." Washington State University cancelled the AI-detection component of its Turnitin contract in 2026: among all Review Board cases involving alleged AI use from 2023 to 2025, 33% ended in a finding of not responsible because detector output had been submitted without independent supporting evidence.

These decisions did not make prohibited AI use acceptable. They reflected a judgment that detector scores were not sufficiently reliable, transparent or independently probative for high-stakes individual decisions.

Turnitin itself suppresses scores in the 1–19% range because false positives are more common there. That is a vendor acknowledging that its own output should not be treated as equally reliable across the full scale.


What AI assistance actually requires

I use AI. I have no discomfort saying so.

Every piece I produce starts the same way: I write the full draft. The argument, the structure, the examples, the position. Then the iterations begin: ten to twenty of them on average.

Not because the first version is wrong. Because the AI, left without resistance, drifts toward the most plausible version of what I was trying to say. Plausible is not the same as true. Plausible is statistically probable. It is the smooth output that arrives when no one is pushing back.

Pushing back is the work.

The AI does not have the specific conviction behind the argument. It does not know which positions I am unwilling to soften, which word choices carry weight I am not prepared to trade away, which observations came from watching the same dysfunction appear in different organizations under different names. What it has is probability. Given my draft, it proposes statistically likely improvements. When those improvements diverge from my position, I reject them and redirect. The iterations are not refinements toward something better. They are repeated corrections of an instrument that, without steering, would produce something more polished and less mine.

A detector sees the final text. It cannot see the iterations that preceded it, the suggestions rejected, or the convictions held. A score that cannot see the process cannot, on its own, establish who owns the output.


What my case suggests

I am Italian. I studied Latin. I studied marketing and communication theory. I spent years as an advertising strategist before I ever wrote for a general audience.

My writing scores close to 100% AI, including articles I published before large-language-model writing tools became part of ordinary publishing workflows. I now think in English after years abroad. The structural fingerprint of my formation remains.

I do not present this as proof of the causal mechanism. My combination of Italian prose formation, Latin rhetorical training and advertising theory offers a credible explanation for why features of my English may overlap with detector signals. It is my current working hypothesis, not an established explanation.

I am not asking to be exempted from scrutiny. I am asking for scrutiny that can distinguish a machine from a writer whose formation, intentions and writing process the detector cannot observe.

The detector doesn't know Cicero. That is the detector's limitation, not mine.