Letting AI Models Judge AI Models!
Substack’s Pangram, AI slop, and why detectors can’t measure human care.
Substack said their readers do not want to invest attention in content that feels empty, automated, or mass-produced. so they brought in Pangram.
The company has framed the in-built AI detection tool feature as a response to “Claudefishing”, ( This is actually the first time I heard this word 🧐😶 🙈) with the goal of helping readers judge whether text was likely written by a human, written with AI assistance, or generated more heavily by machines.
Definitely nothing wrong with the concern, its understandable even if the proposed solution isn’t.
The first problem is that this is plainly an AI-for-AI solution.
One model is being asked to identify and score the outputs of another class of models, then present that score as a meaningful signal to human readers.
That may sound practical at first, but it immediately creates an arms race dynamic:
generators improve,
detectors adapt,
“humanizers” appear,
detectors respond again, and
everyone spends more time interpreting scores than discussing whether the writing is actually useful, accurate, or thoughtful.
Everyone writes!
This is not only a creator issue. Writing is now part of almost every serious job, especially in technical, strategic, and leadership roles. Founders write applications, technical updates, product notes, investor messages, partner emails, research summaries, and public positioning pieces. Engineers, scientists, and operators do the same in different forms. The platform context may be Substack, but the real issue is broader:
how people write in an AI-assisted world.
And that wider view changes the frame.
The debate should not be reduced to whether “creators” are cheating. It should ask what counts as legitimate assistance in modern knowledge work, what still requires human judgment, and whether platforms are measuring the right thing when they flag text.
No one ever disclosed ( or had to) that they used Google?
One basic argument still stands: no one ever had to add a disclaimer saying “I used Google here.” Search engines changed writing permanently, but they were treated as infrastructure rather than authorship.
People searched for facts, checked context, located sources, and then wrote. No one demanded a running label every time a search engine shaped the path to an idea.
AI is treated differently because it can generate direct text rather than only retrieving information.
That distinction is real, but it is often overstated.
In practice, many professionals use AI as a tool inside the same workflow that already includes search, notes, references, redrafting, and editing. The relevant question is not whether a tool was present, but whether the human remained responsible for the argument, the verification, the nuance, and the final judgment.
AI is not good enough to clone thoughtful writing, yet!
The strongest practical point is that AI is still not efficient enough to clone a person doing serious, thoughtful, high-context writing. It can help produce direct content quickly, and it is often useful for structure, summarization, rewriting, copy-editing, and routine explanatory text.
But thoughtful writing is not just sentence production. It requires layered research, selective reading, domain judgment, strategic positioning, iterative drafting, and often multiple rounds of self-critique.
That matters even more in technical and founder-led contexts. High-quality writing in those settings usually depends on choosing which evidence matters, deciding what to omit, aligning the message with commercial or scientific stakes, and understanding second-order consequences.
Those are not merely stylistic tasks.
They are acts of judgment.
So the claim that AI can simply “replace the writer” is not convincing in this category of work.
AI can accelerate parts of the process, but acceleration is not authorship.
Assistance is not equivalence. A model can produce prose in a familiar tone, but producing credible, deeply reasoned writing with real stakes still depends on a human carrying the intellectual load from beginning to end.
Not all AI use is the same
Another confusion in the current discourse is the habit of flattening all AI use into one category.
Not all AI use is slop, and not all slop is made with AI.
A person can use AI to brainstorm headings, tighten sentences, summarize source material, or remove repetition while still doing all the substantive thinking themselves. Another person can press a button, publish whatever comes out, and contribute noise at scale. Those are not ethically or intellectually the same thing.
This is why percentage labels are misleading.
A detector may try to estimate how much text “looks AI-generated,” but that does not tell a reader whether the piece reflects care, expertise, originality, or accountability.
It does not capture whether sources were checked, whether claims were challenged, or whether the author has real skin in the game. In serious writing, those are the variables that matter.
SO why detectors like Pangram exist?
There are understandable reasons these tools exist.
Platforms, universities, and publishers want a shortcut for handling a trust problem that has arrived faster than social norms and institutional policies can adapt. They want a signal that says - this might be machine-written, take a closer look. They also want some kind of visible response to fears that AI-generated volume will degrade quality and erode confidence.
So detectors exist partly because institutions want enforcement, partly because platforms want to show they are acting, and partly because readers are anxious about being manipulated or wasting attention.
In that sense, Pangram is not random. It is a product of institutional demand for a simple metric in a messy transition period.
But then why the detector logic breaks down?
The trouble is that the measurement does not match the thing people actually care about.
Even Substack’s own framing acknowledges that Pangram cannot detect whether “great human care” went into a piece. it can only estimate whether AI was used in making the text. and that is a major limitation because -
most readers do not ultimately care about tool purity.
They care about whether the writing is worth their time and whether someone thoughtful stands behind it.
Research and criticism around AI detection also show recurring accuracy problems. Evaluations report variable reliability across tools, along with false positives and false negatives, especially when human-edited AI text or atypical human writing is involved.
Once that is true, a detector becomes less an instrument of trust and more a mechanism for suspicion.
It starts training people to doubt the text, the writer, and sometimes their own judgment.
This is where the phrase “AI stopping AI slop” becomes too neat. If the system cannot distinguish between low-effort machine churn and high-effort human-guided drafting, then it is not actually measuring slop. It is measuring statistical traces in language and presenting them as if they map cleanly onto authenticity. They do not.
The AI-for-AI trap
Even if forget about substack for a moment and just look out, its a familiar pattern in the AI era.
use AI to scale production,
then use more AI to monitor, rate, rank, and filter the outputs of the first layer.
That may be operationally convenient, but it shifts attention away from substance and toward machine legibility.
The model is no longer helping the work directly, it is just helping institutions police the side effects of other models.
That is why this feels circular.
The internet is flooded with low-effort generated material, so the answer becomes another layer of automated judgment. Then people begin optimizing for the detector, arguing with the score, or buying tools that evade the detector. This does not restore trust.
It industrializes distrust.
What would make more sense?
A better frame would start from responsibility rather than purity.
Who stands behind the argument?
Were the claims checked?
Is there real domain knowledge present?
Does the piece show evidence of thought, synthesis, and accountability?
Those questions are harder than asking a detector for a percentage, but they are closer to what readers actually need.
That does not mean transparency is useless. It can be reasonable for platforms or authors to describe how AI was used, especially in cases where the distinction matters. But disclosure should clarify process, not replace judgment, and detectors should not be mistaken for a measurement of human value.
The fact that a tool touched the workflow is not the same as saying the work itself is empty.
The position underneath all of this.
The most defensible position is not anti-AI. It is anti-laziness, anti-deception, and anti-false certainty.
AI can be a meaningful support tool for research, drafting, editing, and communication. It can also enable high-volume, low-accountability publishing. The line between those outcomes is not whether AI exists in the workflow, but whether a human has done the thinking and is willing to own the result.
That is why the detector debate can feel so unsatisfying.
It asks the wrong question.
Instead of asking whether a model touched the prose, the more serious question is whether the final work carries evidence of thought, care, verification, and responsibility. Those qualities are still human, and no percentage score has solved for them.





