The Controls Don’t Scale Down. Only the Effort Does.
Four checks for anyone publishing with AI help
Not all slop is the same
“AI slop” has turned into a catch-all, and the catch-all is costing us something useful. Three different problems are hiding under one label.
First, some posts are accurate and completely forgettable. No voice, no expertise, nothing that only this author could have written. Call that voiceless slop.
Second, some are just bot output at industrial scale. That is a platform problem, and the platforms are working on it. This one is bot slop.
Third, some posts are confident, nicely formatted, and simply not true. They contain a statistic with no source. A quote that traces back to nothing. A citation to a book that does not exist. This last one is evidentiary slop, and it is the only one I have a fix for.
The third is also the only one that costs the author anything. Nobody’s career suffers because a post was bland. Careers suffer when somebody in the audience actually knows the subject, checks the claim, and finds it isn’t real.
Why it happens
The usual explanation is hallucination. The AI made something up, created something out of whole cloth. That explanation isn’t wrong, but it makes the problem sound rare and exotic, when it is neither. There are two ordinary mechanisms underneath, and they fail in different ways.
Common is not the same as true
A language model reproduces what shows up most often in its training data. Frequency is not accuracy.
Here is a claim you have probably seen on LinkedIn this month. 93% of communication is nonverbal. 55% body language, 38% tone, 7% words. It gets stated as fact, attributed to research, and used to sell everything from interview coaching to AI avatars.
It does come from two small lab studies published in 1967.[1] In the first, seventy-five undergraduates judged how much a speaker liked someone from nine single words spoken in different tones. In the second, sixty-two women judged the same thing from the word “maybe” paired with photographs of faces. Neither study looked at conversation, presentations, negotiation, or job interviews. Both measured whether one person seemed to like another, under artificial conditions, when the signals did not match.
Now here is the part almost nobody who quotes the formula has checked.
Those numbers do not appear in either study’s results. They show up once, on page 252, in the discussion section of the second paper, introduced with the words “it is suggested.” A proposed way of combining two experiments. No derivation. No account anywhere of how the two sets of results were combined to produce those three figures.
The study’s author, Albert Mehrabian, has objected to the generalization for years. The caveat sits on his own website, directly under the formula: unless someone is talking about feelings or attitudes, “these equations are not applicable.”[2] David Lapakko wrote a journal article about the claim in 2007 and titled it, wearily, “Communication is 93% Nonverbal: An Urban Legend Proliferates.”[3]
One of the most quoted statistics in professional communication was a closing suggestion in a discussion section, and the actual author objects to using the data that way.
One thing worth being careful about here. AI did not hallucinate this claim. It has been circulating for decades, and humans propagated it perfectly well long before any of us had a chatbot. So the claim I am making is a narrower one: the AI training data was already contaminated before the models arrived, and models inherit contamination.
Which has an uncomfortable implication I will come back to. Two AI platforms agreeing does not confirm anything. They may be agreeing because they read the same wrong thing.
AI models want you to be happy, and they work hard to make it happen
The second mechanism has a name in the research literature. Sycophancy. It comes from how AI systems are trained: human raters prefer agreeable answers, so AI models learn to be agreeable, and they will go along with a user’s stated or implied position even when that position is wrong.[4]
Two things to watch for if you draft a post in a chat window.
It gets worse the longer you talk. A model’s first answer is the one closest to its actual assessment. Push back, restate, keep building in one direction, and later answers drift toward whatever you appear to want. Long working sessions feel productive from the inside. That feeling is not evidence that they are getting more reliable.
It affects what gets found, not just what gets said. A model can quietly favor search results that match the angle of your question. You do not get a fabricated source. You get a real one, selected because it agrees with you. That is an insidious problem for research, because the model is quietly steering where the research goes.
It is worth putting this next to the failure everybody already knows about. Straight hallucination is the model producing a quotation nobody said, a citation to a paper that was never written, a study with a real journal name and an author who does not exist. That one has a clean fix. A fabricated citation is an artifact of one model’s particular training run, so a differently trained model asked to find that same source simply comes back with nothing. Ask a second platform for the source rather than for the answer, and invented references fall over almost immediately.
Sycophancy is the harder problem. Hallucination produces claims that fall apart the moment you check the source. Sycophancy produces claims that survive source-checking and fail against the question. The citation is real, the quote is accurate, the study exists, and the emphasis has quietly slid toward the answer you signaled three prompts ago.
Sourcing will not catch that. Only your process will.
Where I’m coming from
I’m an independent researcher working across multiple domains including economic and social history, genealogy, and constitutional history, with articles on track for publication in peer-reviewed journals. Working outside an institution means nobody down the hall reads a draft before it goes out and no research assistant checks my footnotes. So over the past two years I built an AI-assisted research and verification methodology that puts several AI platforms into defined, separated roles to do part of what a research group would otherwise do.
The full version is documented elsewhere.[5] It is much heavier than anything a “simple” blog post needs. But it scales down, and the part that scales is what follows.
What the heavy version catches
My example is a current research project on evidential claims in nineteenth-century religious literature.[6] It is a useful example because verification there is genuinely hard. The topic is contested, so the secondary sources lean in predictable directions. The primary sources I am working with are in Arabic, Urdu, and Greek, none of which I read. And it is a field where a fabricated citation looks exactly like a real one. Plausible author, plausible title, plausible page range, plausible argument.
I ran the full methodology. Four different platforms in defined roles, audit prompts written to provoke disagreement rather than agreement, and one rule I never break: no platform assesses its own work. External auditors get the question and the sources. They never get the conclusion I’m testing.
In the first audit session, six errors surfaced. None of them reached publication, which is the entire point of running the checks. The distribution turned out to matter more than the count.
• Perplexity Pro Academic found four, and each one was a citation or source error. That is what a search-first tool with scholarly connectors is built to catch. It caught nothing else.
• SuperGrok found one, and it was a reasoning error rather than a factual one. However, in the same round it caught that, SuperGrok’s own most confident assertion was unsourced and wrong. This is why I run multiple rounds of external audits.
• A parallel Claude chat session found the sixth: invented first names attached to two real authors, in a citation that was otherwise perfect. Right work, right year, right journal, wrong human beings.
Two further findings were worth reporting.
Neither Perplexity nor SuperGrok found the other’s errors. The split looks structural rather than lucky. Perplexity is search-first, and everything it caught was a source problem. SuperGrok is adversarial, and it caught a reasoning problem. The one they both missed took a third independent reading. One external pass is not a quality gate. It is one opinion.
Cross-checking did nothing for a failure I hadn’t anticipated. It works beautifully on fabrication, because an invented citation fails differently on a differently-trained model. It failed completely when both platforms answered questions about a primary text by reading summaries of it, while the text itself sat one click away. What fixed that wasn’t a better prompt. It was handing over the document instead of a description of the document. The moment I put page images of the original in front of a model, a question that had eaten three rounds of searching resolved in one exchange.
That was the first pass of the heavy version. More than eighty documented sessions across four platforms, structured audit prompts, and a written log of every correction.
A LinkedIn post about communication skills does not need all of that. But look at what happens when a post gets none of it.
Even a “small” LinkedIn post can end badly
In early January 2025, attorney Carolyn Elefant published a LinkedIn post asking an interesting question. Could a twenty-dollar-a-month ChatGPT subscription have improved the persuasiveness of the reply briefs in the Supreme Court’s TikTok case? She gave the model the actual briefs, set the original opening of each one beside ChatGPT’s redraft, and invited readers to judge.
A commenter, Adam S. Sieff, spotted the problem. ChatGPT had not pulled those originals out of the briefs she had supplied. It had made them up. The post had been praising the model’s improvements to passages the lawyers never wrote.
To her considerable credit, she appended an update on January 6 saying exactly that, posted about it a second time, and put the real openings in the comments with a note that they needed no edits at all.[7]
A few things about this one are worth thinking about.
Her post was not AI-written. The prose was hers. What came from the AI was the evidence, and the evidence was fabricated. That is the version of this problem most likely to catch a careful writer, because nothing about the writing process feels automated.
The invented text wasn’t decoration in paragraph six. It was the entire demonstration. Take it out and there’s no post left.
She supplied the source documents. That’s the assumption most of us are working from, that if you hand the AI the real material it will use the real material. She did. It didn’t.
And a reader caught it, not the author. If you don’t check, somebody in your audience will, and they’ll let you know about it in the public comments.
Then there is the last thing, which is the one I keep coming back to and the point of this post. Her own explanation was that she would have verified everything had this been a filing for a court rather than a post for LinkedIn. That is an honest account of how nearly all of us operate. It is also the diagnosis rather than an excuse. She had a verification standard, she knew exactly what it was, and she applied a lower one because of where the writing was headed.
Verification has to happen before anything leaves draft, because the draft doesn’t know where it’s going.
On the plus side, this “small” post is an example of more than AI hallucination. It is also an example of an author’s exceptional response to an error. She fully owned it, posted the updates transparently, and turned a possibly embarrassing situation about an interesting question into a much deeper learning opportunity for all of us. Contrast that with those who quietly stealth-correct, or simply delete the post.
A side note about this post
Honestly, this was meant to be a quick sixty-minute effort over coffee. I had an idea about scaling my research methodology down for people writing social media posts. The plan was to have Claude help size the problem, run a novelty check to see whether somebody had already done it better, rough out an outline, a couple of rounds of edits, one external validation pass, and publish by lunch. Man plans, God laughs.
While sizing and scoping the problem, Claude, Perplexity, and Grok all handed me good figures. How much LinkedIn content is actually AI-generated. What share of it contains unverifiable claims. How often AI-sourced citation lists turn out to be fabricated. They arrived in confident prose and would have made a persuasive paragraph.
Then I ran Controls 1 and 2 on the figures. The writer is not the checker. Source it or cut it.
The most quotable item Claude returned was a specific percentage of AI-generated LinkedIn posts said to contain unverifiable claims, based on an analysis of five hundred posts. It traced back to a commercial content-marketing blog. No named researchers, no dataset, no sampling method, no definition of the key term, no publication anywhere. It exists as an assertion on a page selling a product.
A second figure, on hallucination rates in an AI-sourced citation list, turned out to be one copyeditor’s account of one manuscript, posted on LinkedIn. A genuinely useful anecdote. Not a study, but it was handed to me as one.
A third appears in no source I could locate, from any organization, at all. Complete hallucination.
And a data range I was offered, a low end and a high end presented as one span, combines two different companies, two different detectors, two different word counts, and two different definitions of “AI-generated.” The range is an artifact of putting them in the same sentence.
Four claims. All eminently quotable. One traceable to a marketing page, one to an anecdote, one to nothing at all, and one to a category error. Without the controls, I’d have published every one of them, in an article about controlling AI-generated evidentiary slop.
The four controls
1. The writer is never the checker, and the checker shouldn’t know what you want.
Never ask the AI model, in the same chat, to verify a claim it drafted. At a minimum, open a new session in the same model, or even better, a different platform. Models miss most of the errors in their own output while catching those same errors easily when they arrive as somebody else’s work.[8] The second half of this is the part people skip. When you hand over the claim to be checked, hand over the claim and nothing else. No preferred answer, no framing, no “am I right that…”. An AI checker who can work out what you’re hoping for isn’t an independent check. It’s the sycophancy problem with an extra step.
2. Source it or cut it.
Every specific factual claim needs a real, checkable source before it publishes. A number, a quotation, a study finding, a company detail. If you can’t find one in a few minutes, the claim comes out. And a citation existing isn’t the same as the claim being verified. The source can be real and still not say what you’re citing it for.
3. Agreement is not proof.
When you cross-check, ask the second platform for the source, not for the answer. Two models producing the same claim may just be reflecting the same widely-repeated error. Convergence counts only when each one independently points you to primary evidence. And if two platforms disagree about whether something even exists, treat that as disqualifying rather than a tie to be broken.
4. Say what you didn’t verify.
If a claim matters but you couldn’t confirm it, say so in the post, or drop it. A stated limitation reads as credibility. Silent confidence is what got us here.
Scaling it down
| Control | Research-grade | Publishing a post |
| Writer isn’t checker | Separate platform, blind to prior conclusions; audit prompt gives the question and sources only | New chat window; paste the claim alone; “what’s the source for this?” and never “am I right that…” |
| Source it or cut it | Primary sources checked against the original document or image | One clickable link per factual claim; no link, no claim |
| Agreement isn’t proof | Cross-platform search diversity; disagreements logged as findings | Ask the second platform for the source, not the answer |
| Say what you didn’t verify | Written log of corrections and unresolved items | One line: “I couldn’t verify X” |
Notice what didn’t change between those columns. All four controls survive. What shrank is scope. How many claims, how many platforms, how much documentation.
The author of a LinkedIn post doesn’t need a documented audit trail. They just need to check the sources and then refuse to publish an unsourced claim.
This failure is not limited to “small” social media posts, it runs all the way up the scale. The past year has produced three examples from firms that sell governance for a living.
· Deloitte’s Australian arm repaid the final installment on a A$440,000 government contract after its 237-page report turned out to contain fabricated academic references and an invented quotation from a federal court judgment. Roughly twenty errors, cataloged by one outside researcher who sat down and read it.[9]
· EY Canada withdrew a cybersecurity report in May 2026, the same day an AI-detection firm published its finding that most of the report’s citations were fabricated, misattributed, or led nowhere, including a McKinsey report that does not exist.[10]
· A month later KPMG pulled a report on agentic AI after the same detection firm and the Financial Times identified false claims in it, though KPMG hasn’t publicly confirmed the cause.[11]
The EY case carries an extra sting worth noticing. A fabricated McKinsey citation published under a Big Four masthead is exactly how a false source enters the record and gets picked up by the next writer, and eventually by the next model. That is the contamination which leads to evidentiary slop, happening in real time.
One last thing
Verifying the claims in this post took three AI platforms working from a structured prompt. They disagreed with each other on several items, including one where two of them gave flatly incompatible accounts of whether a source existed at all. Those claims aren’t in this post.
That isn’t a failure of the tools. That’s the process working. The step we skip isn’t the AI. It’s the part afterward where somebody goes and looks.
AI disclosure
This post was written with AI assistance, using the methodology it describes. Claude Chat (Opus 5, high reasoning) handled outlining, drafting, and editing. ChatGPT, Grok, and Perplexity Pro Academic handled research and independent verification. No platform verified its own output. Every factual claim originating in the drafting session was checked on a platform that hadn’t produced it, and the verification prompt supplied the claims and the standard of evidence without indicating which answer was wanted.
About the author
Richard E. Rudd is an independent researcher working across economic and social history, genealogy, and constitutional history. He spent more than twenty years as an IT Product Owner and Program Manager at large regulated multinational financial services and insurance firms before applying those governance disciplines to AI-assisted research.
Related writing on AI-assisted research methodology:
• “Four Platforms, One Standard: Why Serious AI-Assisted Research Needs More Than a Better Prompt,” 3 April 2026. https://ruddresearch.com/2026/04/03/four-platforms-one-standard-why-serious-ai-assisted-research-needs-more-than-a-better-prompt/
• “Governance by Design: What Twenty Years in Regulated IT Taught Me About AI-Assisted Research,” 16 July 2026. https://ruddresearch.com/2026/07/16/governance-by-design-what-twenty-years-in-regulated-it-taught-me-about-ai-assisted-research/
• “What My Research Framework Caught: Seven Errors, None Published,” 5 August 2026. https://ruddresearch.com/2026/08/05/what-my-research-framework-caught-seven-errors-none-published/
The full methodology is set out in the preprint article: Rudd, R. E. (2026). Four Platforms, One Standard: A Multi-AI Verification Methodology for GPS-Compliant Genealogical Research. Zenodo. https://doi.org/10.5281/zenodo.21829158
More research topics at ruddresearch.com
[1]Albert Mehrabian and Morton Wiener, “Decoding of Inconsistent Communications,” Journal of Personality and Social Psychology 6, no. 1 (1967): 109–114, doi:10.1037/h0024532; Albert Mehrabian and Susan R. Ferris, “Inference of Attitudes from Nonverbal Communication in Two Channels,” Journal of Consulting Psychology 31, no. 3 (1967): 248–252, doi:10.1037/h0024648. The coefficients appear at Ferris, 252, in the discussion. Some secondary sources cite the Ferris paper as 31, no. 6; the article’s own header and PubMed both give 31, no. 3.
[2]Albert Mehrabian, kaaj.com, accessed 20 August 2026, archived 8 March 2026 at https://web.archive.org/web/20260308150406/http://www.kaaj.com/psych/smorder.html.
[3]David Lapakko, “Communication is 93% Nonverbal: An Urban Legend Proliferates,” Communication and Theater Association of Minnesota Journal 34, no. 1 (2007): 7–19, doi:10.56816/2471-0032.1000. The Cornerstone repository record dates the article to 2015, reflecting digitization; volume 34 is Summer 2007.
[4] Mrinank Sharma, Meg Tong, et al., “Towards Understanding Sycophancy in Language Models,” ICLR 2024; arXiv:2310.13548 (v4, 10 May 2025), CC BY 4.0. Ethan Perez, Sam Ringer, et al., “Discovering Language Model Behaviors with Model-Written Evaluations,” arXiv:2212.09251 (19 December 2022). Both find sycophancy present in pretrained models before human-feedback training, indicating that RLHF amplifies rather than originates the behavior.
[5]Richard E. Rudd, “Four Platforms, One Standard: A Multi-AI Verification Methodology for GPS-Compliant Genealogical Research,” Zenodo, 2026, doi:10.5281/zenodo.21829158 (concept DOI, all versions). First set out in “Four Platforms, One Standard: Why Serious AI-Assisted Research Needs More Than a Better Prompt,” ruddresearch.com, 3 April 2026. See also “Governance by Design: What Twenty Years in Regulated IT Taught Me About AI-Assisted Research,” ruddresearch.com, 16 July 2026.
[6]Richard E. Rudd, “What My Research Framework Caught: Seven Errors, None Published,” ruddresearch.com, 5 August 2026, https://ruddresearch.com/2026/08/05/what-my-research-framework-caught-seven-errors-none-published/, where the project is described at greater length.
[7]Carolyn Elefant, LinkedIn post, January 2025, updated 6 January 2025, linkedin.com/posts/carolynelefant_chatgpt-activity-7281437579143462912-68ak. Details confirmed against the live post, viewed 21 August 2026; see also her follow-up post of the same date. The update is dated in the post’s own text. LinkedIn displays only a relative age for the original, and the post is marked as edited, so the exact original posting date remains unconfirmed. No public archival capture is available, LinkedIn being login-walled.
[8]”Self-Correction Bench: Revealing and Addressing the Self-Correction Blind Spot in LLMs,” arXiv:2507.02778 (2025). See also P. Verga et al., “Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models,” arXiv:2404.18796 (2024).
[9]Report for the Australian Department of Employment and Workplace Relations, July 2025; 237 pages, roughly twenty errors identified by Christopher Rudge, University of Sydney. The refund covered the final installment, not the full contract value. The revised report disclosed use of Azure OpenAI. https://apnews.com/article/australia-ai-errors-deloitte-ab54858680ffc4ae6555b31c8fb987f3
[10]Points of Attack: Uncovering Cyber Threats and Fraud in Loyalty Systems (EY Canada), a 44-page report withdrawn on 14 May 2026, the same day GPTZero published its investigation. GPTZero reported that 16 of the 27 cited sources were fabricated, misattributed, or led to broken pages, including a McKinsey report that does not exist. EY Canada said it was reviewing how the report came to be published. https://www.computing.co.uk/news/2026/ai/ey-cybersecurity-report-withdrawn-ai-hallucinations
[11]Total Experience: Redefining Excellence in the Age of Agentic AI (KPMG, October 2025), withdrawn June 2026. GPTZero reported that only five of the report’s 45 citations pointed to real, intact sources; the findings were verified by the Financial Times, and UBS, the NHS, Swiss Federal Railways, and Transport for London each disputed claims made about their AI use. KPMG removed the report while investigating and has not publicly attributed the errors to generative AI. https://techcrunch.com/2026/06/13/kpmg-pulls-report-on-ai-usage-due-to-apparent-hallucinations/
Corrections, Clarifications, and Revisions:
Correction, 25 August 2026. Note 4, supporting the sycophancy discussion in “AI models want you to be happy, and they work hard to make it happen,” has been replaced. The version published on 22 August carried an unresolved internal marker instructing that the two citations be verified against the primary papers before publication. The marker was live on the site until 25 August. The citations have since been checked against the papers themselves. Both are real, and both support the claim they were attached to. Two errors were corrected: Sharma et al. was cited as a preprint only, omitting that it was published as a conference paper at ICLR 2024; and a journal reference and page range for Perez et al., supplied by a second platform rather than read from the paper, could not be confirmed and has been removed. No claim or conclusion in the post changes.