Why should anyone care which AI tools I used, as long as the work meets the standard?
I think the honest answer is: they shouldn’t. Nobody asks an auditor what calculator they used. Nobody asks a developer which IDE they coded in. The deliverable either satisfies the standard or it doesn’t. The tools are implementation details, disclosed for transparency — not because they require justification.
But that answer only holds if something guarantees the standard was met. In regulated software development, that something has a name. It’s called governance, and it is the most underrated idea in the current conversation about AI and research.
The Question Nobody Is Asking
The AI-and-research conversation is currently stuck on two questions. Which tool is best? And is it ethical to use one at all?
Both skip past the question that actually determines whether the work is trustworthy: what structure ensures the output meets the standard?
That is a governance question. And governance is not an unsolved problem. It was solved decades ago, in industries where getting it wrong is expensive.
I spent more than twenty years as an IT Product Owner and Program Manager at large multinational financial services and insurance firms, building business-critical systems under regulatory scrutiny. In that world, you cannot ship code because the developer says it works. The developer’s confidence is not evidence. What makes a deliverable trustworthy is never the skill of any individual — it is the structure built around them.
Separation of duties. Independent QA. Code review by someone who did not write the code. Audit trails. Change control. Defect logging. Testing rigor scaled to risk.
None of that assumes the developer is careless. It assumes the developer is human — subject to blind spots, overconfidence, and investment in their own work.
AI has precisely the same failure modes, in sharper form. It is fluent, confident, eager to please, and — this is the part that should worry us — its errors are indistinguishable in appearance from its successes. A fabricated citation is formatted exactly like a real one.
Which means the governance pattern transfers. Not perfectly — and the imperfections matter, so I’ll be specific about them. But the core structure transfers more directly than most analogies in this conversation, and it transfers because the underlying problem is the same: how do you trust output from a fluent, confident, self-interested producer?
The Load-Bearing Principle
Of everything I carried across from regulated development, one rule does most of the work:
The platform that produced a finding cannot be the sole platform that verifies it.
This is separation of duties. In financial auditing, the team that prepares the statements cannot be the team that audits them — not because they are dishonest, but because they are invested. Sarbanes-Oxley §404 requires management to assess its internal controls and an external auditor to attest to that assessment independently; ISA 610 governs the conditions under which an external auditor may lean on internal audit work at all, and how much of it must still be done afresh.[8] In software, the developer who wrote the code does not sign off as the sole reviewer of their own work.
Applied to AI-assisted research, the rule becomes: any platform can do QA. No platform does QA on its own work.
That single constraint dissolves most of what people are actually afraid of when they worry about AI in research. Sycophancy is structurally defeated, because the auditing platform has no stake in the analyzing platform’s conclusion. Hallucination is substantially caught, because the same fabrication rarely surfaces on two independently trained models.[1] Confirmation bias is blunted, because the auditor receives the research question and the sources — not the conclusion it is meant to confirm.
The rule is structural, not preferential. It does not depend on any model being good.
Where the transfer is not clean
I want to be precise about the disanalogies, because a careful reader will find them and I would rather name them first.
In regulated software, the independent reviewer is a credentialed professional, legally accountable, operating under compliance regimes with real enforcement. A second AI platform has none of that. It is a different commercial product, not an independent professional bound by duty. So the transfer is structural, not institutional — and the difference is worth stating plainly.
What survives the difference is the mechanism. A human auditor’s credential guarantees competence. An AI auditor’s independence guarantees something narrower but still valuable: a different error surface. The second platform was trained on different data with a different architecture, so it fails in different places. That is not the full institutional protection of regulated auditing. It is a real, structural reduction in correlated error — and for research quality control, that is most of what you need. The rest, as I’ll argue at the end, stays human.
One further limit belongs in the same honest accounting: today’s frontier models share substantial training-corpus overlap, so the diversity of error surfaces is real but not total. Where a mistake is baked into the common substrate of the web these models learned from, two platforms can still agree and still be wrong. Cross-platform verification reduces correlated error; it does not eliminate it. That residual is one more reason the final judgment stays with the researcher, not the machines.
The Field Is Converging on This — From Several Directions
The interesting thing about multi-model verification is that it is being independently rediscovered right now, from several directions at once.
Andrej Karpathy built an “LLM Council” that routes a question to several frontier models, has them review one another’s answers anonymously, and appoints a chairman model to synthesize (he called it a weekend project, but the architecture makes a serious point).[2] The platform vendors have begun to build the pattern in — some agentic modes now spawn parallel sub-agents that check a result before returning it, though most current implementations are orchestration for speed rather than structured cross-verification.[3] And Steve Little, the National Genealogical Society’s AI Program Director, has published a clean demonstration of it: put two different reasoning engines at the same workbench, give them the same evidence and the same method, and have them read independently before either sees the other’s answer. When they disagree, the disagreement is the signal — a question surfaced that one confident answer would have hidden.[4]
Three different arrivals at the same underlying instinct: one model is not enough, and the way to trust the output is to make independent readers check each other. That convergence is a good sign the instinct is sound.
It also lets me draw a distinction the convergence tends to blur — one that matters for how you actually build the thing.
Start with the multi-agent modes. A council of agents inside a single model is still one training corpus, one set of corporate response policies, one architecture. It catches inconsistency — places where the model contradicts itself — but it cannot catch a blind spot the whole model shares. If the training data carries a systematic error, every agent inherits it, and unanimous agreement among them means only that they are all wrong together.[5] Cross-agent verification within one model reduces variance. Cross-platform verification across independently trained models is what actually catches shared blind spots. Karpathy’s own design implies this — he built a council of genuinely different models, not one model talking to itself.
Now the sharper design question, the one Steve Little raises directly. If you do use different platforms, should they hold fixed roles? His worry is vendor mythology — the comfortable story that this model is the careful analyst and that one is the creative writer, as though these were personalities rather than products. His remedy is role reversal: today’s extractor becomes tomorrow’s skeptic, so no model hardens into an unquestionable authority. As he puts it, “a model personality is merely a story we may be tempted to tell.”[6]
He is right, and the mechanism is a good one. Role reversal does something a fixed audit function doesn’t: it keeps the researcher fluent across every model and prevents any one of them from becoming a black box you stop checking. For the task his article addresses — two engines reading the same evidence to answer the same question — it’s the right call. But it isn’t the whole picture.
Where I’d extend the picture is one layer up. The danger in vendor mythology is not specialization itself. It is deference — trusting the designated expert without checking the work. Those two things are separable, and regulated software separates them every day. You do not cure developer overconfidence by rotating developers into QA on a schedule; you cure it with a QA function that reviews everything, permanently. Fixed roles, zero self-review.
That distinction matters because the differences between platforms are real — and some of them are not stories at all.
Some differences are behavioural: a model’s characteristic verbosity, its hedging, its appetite for speculation. These emerge from training data and corporate response policy, which is why the quirks are consistent rather than random. They are also largely mitigable through prompting — a well-built prompt flattens them, and Steve’s method depends on exactly that flattening. On this point the vendor-mythology worry is fair.
But other differences are structural. A search-first architecture with mandatory citations wired into a scholarly index will surface and verify literature that a general-purpose model simply cannot reach. Access to data corpora a competitor cannot license at any price. Agentic tooling that operates directly on files and archives. No prompt closes those gaps, because they are not behavioural — they are plumbing. Assigning a scholarly-literature task to the platform actually engineered for scholarly literature is not mythology. It is matching the work to the tool.
So the full picture has two layers. When platforms read the same evidence to answer the same question — adjudication — roles should be symmetric and reversible, precisely as Little argues. When platforms occupy different stages of a pipeline — discovery, retrieval, deep analysis, literature verification, drafting — differentiated assignment is justified by real structural advantages. Both layers live under one rule: no platform grades its own homework.
I expect the structural gaps to narrow as the platforms advance, and when they do, the pipeline layer will look more like the adjudication layer — more interchangeable, more reversible. We are not there yet. And a framework organized by task rather than by model absorbs that convergence without breaking: as the differences shrink, you simply reverse and rotate more freely, under the same governance. The one refinement I’d add even now is that structural advantages shift over twelve to eighteen months, so even fixed assignments deserve periodic re-benchmarking. Fixed where the structural edge is real and current; reversible where it isn’t; never self-reviewing anywhere.
What Else Transfers
Errors are the audit trail, not an embarrassment. In regulated development, defects are logged, not buried. Every error caught gets documented: what was claimed, by which platform, what flagged it, what the primary source actually showed. A methodology that reports no errors is not rigorous. It is untested. My own errors — including one in a draft arguing for citation verification, which itself contained a misattributed citation, caught by the audit step — are the strongest evidence I have that the process works.
Locked decisions. Once a conclusion clears full verification, it is locked. A later session cannot quietly revert it. That is change control. You do not roll back a production release because someone had second thoughts on a Tuesday.
Governance scales to risk. Nobody applies the same rigour to a two-week bug fix and a two-year platform migration. Verifying a single reported figure needs a quick cross-check between two models. A contested conclusion bound for peer review needs the full apparatus.
The human role is non-delegable. In SDLC terms, the researcher is the Product Owner. They define the question, accept or reject the deliverable, and remain accountable for whether it meets the standard. The AI platforms are specialized developers. They do not write their own requirements, and they do not sign their own acceptance. AI enables, the methodology governs, the researcher decides — and the last of those three can never be delegated.
Why This Isn’t About Any One Field
I am an independent researcher, and I work across several fields. The framework I am describing did not come out of any of them. It came out of regulated IT, and I carried it in.
I have developed and tested it most fully in genealogy, for a specific reason: genealogy has an unusually clean published evidentiary standard. The Genealogical Proof Standard defines what a trustworthy conclusion looks like independently of who reached it or what tools they used. That makes it an ideal proving ground — there is an external yardstick, and the yardstick does not care about my methods. It is one of my current research areas, not my field, and that distinction is the whole point: nothing in the governance framework is genealogical.
It is:
• Tool-agnostic in principle, specialized in practice. The framework does not require any particular platforms — only that you use at least two, and that none of them grades its own homework. But the best way to run it today exploits the specific structural advantages different platforms actually have. The framework will outlive the current lineup; the current lineup is still the sharpest way to implement it.
• Topic-agnostic. I am currently running it against medical research as well, where the producers are just as fluent and the failure modes just as costly.
• Standards-agnostic. Any field with a defined evidentiary standard can be plugged in: legal, clinical, archival, journalistic.
• Scalable. Two free-tier models for a simple question. The full apparatus for publication-track scholarship.
The claim I am making is not “here is my workflow.” Workflows are personal, and they are obsolete within the year. The claim is that regulated industries already solved the problem of trusting output from fallible, confident, self-interested producers — and that the solution transfers to AI-assisted research more or less intact. Notably, the major AI-governance frameworks don’t yet make this transfer: they govern how organizations build and deploy AI systems, not how an individual researcher should structure AI-assisted inquiry against an evidentiary standard.[7]
What Comes Next
Two related papers are in preparation. The first develops the governance paradigm and its transfer into the digital humanities. The second applies it to a specific evidentiary standard, with worked case studies — including the failures, which are the interesting part. Preprint articles are expected to follow shortly.
The tools will keep changing. The governance does not have to.
Notes
Richard E. Rudd is an independent researcher working across multiple fields. He spent more than twenty years as an IT Product Owner and Program Manager at large multinational financial services and insurance firms before applying those governance disciplines to AI-assisted research.
[1] The empirical basis for this is developing quickly. See P. Verga et al., “Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models” (arXiv:2404.18796, 2024), which finds a panel of diverse models reduces intra-model bias relative to any single-model judge. Enterprise-scale studies in regulated sectors report substantial hallucination reduction from cross-platform verification.
[2] Andrej Karpathy, “LLM Council,” github.com/karpathy/llm-council (2025). Three-stage architecture: parallel individual responses, anonymized peer review and ranking, and a chairman model that synthesizes. Karpathy described it as a weekend project; the design nonetheless illustrates the cross-model principle precisely.
[3] As of mid-2026, OpenAI, Anthropic, and Google have shipped multi-agent SDKs and orchestration frameworks, and some “deep research” products deploy parallel sub-agents. Most current implementations orchestrate for parallelism rather than structured cross-verification; the verification pattern is more fully realized in open-source projects and dedicated verification layers.
[4] Steve Little, “Trifecta: One Genealogist, Two AI Assistants, One Folder,” Vibe Genealogy (vibegenealogy.ai), 10 July 2026. The “vendor mythology” and role-reversal argument, and the workbench method of independent reads adjudicated against the record, are set out there in full. Little’s role as the National Genealogical Society’s AI Program Director is per the NGS announcement of the position (ngsgenealogy.org).
[5] On the self-review limit, see the “self-correction blind spot” literature (e.g., “Self-Correction Bench,” arXiv:2507.02778, 2025), which finds that models fail to catch a majority of their own internal errors while reliably catching the identical errors when presented as external input. This is the empirical core of the case for cross-platform rather than same-model checking.
[6] Steve Little, “Trifecta: One Genealogist, Two AI Assistants, One Folder,” Vibe Genealogy (vibegenealogy.ai), 10 July 2026. The “vendor mythology” and role-reversal argument, and the workbench method of independent reads adjudicated against the record, are set out there in full. Little’s role as the National Genealogical Society’s AI Program Director is per the NGS announcement of the position (ngsgenealogy.org).
[7] The NIST AI Risk Management Framework and ISO/IEC 42001 both address separation of duties and independent oversight — but at the level of organizational governance of AI systems, not as a methodology for structuring an individual researcher’s AI-assisted work against an evidentiary standard. That specific application is the gap this work addresses.
[8] Sarbanes-Oxley Act of 2002, §404 (15 U.S.C. §7262); International Standard on Auditing 610 (Revised 2013), Using the Work of Internal Auditors, IAASB. These are the specific instruments the separation-of-duties principle in this article is carried across from.
Corrections, Clarifications, and Revisions:
Revision, 18 August 2026. Added the specific instruments behind the separation-of-duties principle in “The Load-Bearing Principle” — Sarbanes-Oxley §404 and ISA 610 — with a supporting note. The original text described the practice without naming its sources. No argument or conclusion in the post changes. The omission became apparent while reading Zhou and Yu (arXiv:2608.10858), which cites the adjacent §302 for a similar transfer.




Leave a comment