This site hosts research updates, pre-publication summaries, and reflections on methodology. Specific evidence, proof arguments, and detailed analysis are reserved for peer-reviewed publication.

Tag: genealogy

  • Chasing one word across seven languages for three weeks.

    Arabic, Urdu, German, Turkish, French, Greek, English. I read English, but high school French and Duolingo Spanish weren’t going to help. A multi-AI approach with OCR, vision, translation, and transcription capabilities was the toolset I needed.

    The side quest

    I am working on a historical evidentiary research project across four religious traditions. Not to validate or invalidate a tradition, or even to question those traditions, but to review evidentiary claims to a historical standard. In the middle of the project a specific question needed answering, and finding the answer turned into a three-week side quest of its own.

    The situation: in 1864 Rahmatullah Kairanawi, a scholar in Istanbul, was working with a pamphlet from Calcutta, written in Urdu and printed in 1268 AH, which is roughly 1851 to 1852 AD. The pamphlet argued that Muslims had misunderstood a particular Greek word in John’s gospel when they translated it into Arabic. The word is παράκλητος, paraklētos, and the argument in dispute is whether it was originally περικλυτός, periklytos, meaning renowned, which is close enough to the name Ahmad to be interesting. The two differ by their vowels and very little else.

    The question: did Kairanawi accept the pamphlet’s argument that paraklētos is the correct reading, or refute it and hold that the original was periklytos? The question seems straightforward… little did I know what I was in for.

    Kairanawi had debated the missionary Karl Gottlieb Pfander at Agra over two days in April 1854 and wrote a very long reply over the following decade. His discussion of the word runs to a few pages of Arabic in a Saudi critical edition of 1,552 printed pages, and I wanted to know what those pages said.

    A rule when researching anything, especially a historical event: go to the original, and treat every copy, abstract and translation as a step away from it where something can go missing. You do not cite the calendared abstract when the will exists. What follows is that rule running across seven languages.

    A brief note on what this post is and is not: The details around the question, responses, and scholarly treatment are being held for a peer-reviewed article, so I am not publishing the answer here, but I will link to it when it appears. Also, while the context of the post touches upon religious traditions, the post is emphatically not about that. This post is about AI assisted research during the side quest, not the destination.

    Arabic

    The quest started with the Kairanawi Arabic edition which runs to 1,552 printed pages. The scan I worked from is a separate digitization of 1,416 images, one of several floating around, and reconciling the two took longer than it should have.

    Tesseract has Arabic language data. Claude has access and installed it. Once it was up and running, the editor’s 136-page introductory study came out at 34,306 words in about ten minutes, and was searchable for the first time.

    Then it produced a number. A word count on the OCR output said 4,080, roughly one word in eight. I assumed the counting tool did not handle Arabic, which is probably wrong: Arabic is space-delimited much like English and the tool handles it fine on clean text. The cause is more likely the extraction. What matters is that a plausible number arrived, was wrong by a factor of eight, and would have told me the corpus was too thin to bother with.

    Searching it had a second problem. A term I was looking for appears seven times spelled one way and once spelled another, the difference being a single diacritic. Search the first form alone and you have found seven-eighths of the evidence, with nothing to tell you that.

    The passage itself, when I reached it, could not be searched at all. What mattered was not the text but the editor’s footnotes, and specifically which footnote attached to which paragraph. That is layout. OCR flattens layout by design. Those two pages had to be looked at.

    So I had the passage, and it was where the secondary literature said it would be. What I did not have was any reader other than me. I cannot read Arabic, and a load-bearing claim resting on my own machine-assisted reading of a language I do not know is not the time to stop and declare success. The obvious check is obtain a second quality translation, ideally in a language I do read.

    English

    A quick Perplexity and Grok search later found an English translation of the work, published in London in four parts between 1989 and 1990. Those four parts render the first five of the Arabic work’s six books. The sixth book, which is where the material I wanted lives, is not in it.

    I established that by comparing two tables of contents, a job needing nothing but Claude’s ability to diff the two contents pages and four minutes.

    Getting to those four minutes took rather longer, because I first worked out the mapping by reasoning from subject matter, and got it backwards. The Arabic contents page would have settled it in ninety seconds. I spent three exchanges building on the inference before I looked.

    The scans, when I did search them, are two pages to an image with no text layer. OCR interleaves the facing pages, so the output finds a term perfectly well and cannot be quoted from.

    That was not the answer I needed. It was another question. A translation that renders five of six books and stops is either an unfinished job or a decision by an editor, and either way something has gone missing at some step. The English edition told me where to look next, because it says on its own title pages that it was translated from Urdu.

    Urdu

    The Urdu edition, in three volumes, contained what the English translation does not.

    However, the Urdu here is written in nastaliq, a script that slopes, overlaps and joins in ways that defeated every attempt at OCR I made. Internet Archive holds all three volumes with text layers attached, and those text layers are noise. Not poor quality, not partially garbled: noise. Current work on nastaliq recognition reports usable results from newer multimodal models, so this is a statement about what I tried rather than about the state of the field.

    The route into the Urdu was a publisher’s contents listing and screenshots of pages, uploaded and read as images. Which is what one does with bad OCR, and it is slower than it sounds.

    The Urdu has all six Kairanawi books. So the missing sixth book from the English translation did not start with the Urdu, it happened on the English side and why it happened is still a question. That settles part of what the English raised and leaves the original question exactly where it was. I still had no second reader of the passage.

    French

    At this point I was hunting specifically for a translation that contains the sixth book, and staying as close to 1864 as I could, on the same principle that fewer steps mean less loss. Back to Perplexity and Grok.

    The nearest found is French. Two volumes, published in Paris in 1880, sixteen years after Kairanawi finished writing, and it contains the sixth book the English transcription omits. It has been sitting there for nearly a century and a half.

    Claude gave its translator’s name as Cadetti, having copied it from a bibliography without confirmation. The name is Carletti. The name was wrong for three drafts until an external verification from Perplexity, specifically validating citations, caught it.

    The correction produced a second correction. The title page credits the translation to an unnamed young Tunisian, with Carletti as reviser and author of the preface and notes. So I had the name wrong and the role wrong, and was not confident in any answer.

    German

    A translation is a second reader, but it is not scholarship. So the next question was whether another researcher since had studied the passage and could shed light on it.

    Perplexity found a monograph on the controversy and on the question I wanted answered: Christine Schirrmacher, 1992. It is a substantial German-language study of the Pfander and Kairanawi dispute, and it is not available in any digitized form I could reach without paying for it.

    Luckily the pages I needed were available on Google Books preview. I obtained screenshots of twelve pages, loaded into Claude and read as images. That worked well, and better than OCR would have, because footnote numbers matter here and OCR routinely drops superscripts.

    Searching for the keyword Kairanawi returned a sentence that appeared to answer my question outright:

    So argumentiert auch al-Kairânawî selbst im ersten Kapitel des sechsten Buches.

    “Al-Kairânawî himself argues this way too, in the first chapter of the sixth book.”

    So argumentiert auch means argues this way too. This way is anaphoric: it points back at whatever was described immediately before, and what was described immediately before was the periklytos claim. I read it as an answer. Kairanawi refuted the pamphlet and held that periklytos was the original. That reading is mine, not Schirrmacher’s, and the whole of it rests on where a single backward-pointing word lands.

    I had my answer.

    A bonus was that Schirrmacher’s scholarship also gave me the citations of the information leveraged in the monograph, so I could get another viewpoint.

    Turkish

    Schirrmacher’s footnotes led to a Turkish thesis.

    A 2017 master’s thesis from Ankara Üniversitesi runs to 152 pages, is fully digitized, and has a clean text layer. Searchable in every sense. The passage I needed is on page 130, and it matters because it summarizes Kairanawi’s treatment of the word from the Urdu rather than the Arabic. A different reader, working from a different version, in a different language, arriving by a different route. I found it by having Claude read the thesis through from the beginning rather than by searching for keywords, which allowed Claude to ingest the full argument in context.

    The German pointed to periklytos. The Turkish thesis, working from the Urdu, pointed the same way. A second account, in a second language, from a different version of the text. Belt and suspenders, I thought, though I had reached the Turkish through the German’s own footnotes, which is not the same as finding it independently.

    Side Quest Complete?

    I had my sentence, translated accurately, and it answered the question. Why didn’t I feel I was done?

    What made me uneasy was that a load-bearing claim was resting on a partial transcription, in a language I cannot check, and from a book I had only seen twelve pages of. And the Turkish thesis was a translation of a translation more than a century and a half later.

    Back to German

    So I went back to Claude and asked for the full chapter rather than the sentence. I wanted context and not just a translation.

    Here is what the original quotation had cut off. First, the sentence immediately before it on the page:

    …daß anstelle des im Neuen Testament angekündigten ‘parakletos’ früher einmal ‘periklytos’ (der ‘Gesegnete’) zu lesen gewesen sei, womit der Prophet Aḥmad bzw. Muḥammad im Neuen Testament prophezeit worden sei.

    “…that in place of the parakletos announced in the New Testament, periklytos, the Blessed One, was once to be read, whereby the Prophet Ahmad or Muhammad was prophesied in the New Testament.”

    Schirrmacher’s parenthetical gloss there, der ‘Gesegnete’, is looser than the lexicon, which gives much-heard-of or renowned. Nothing in the sentence turns on it, but in a piece about one word it is worth saying that the secondary source and I are not using quite the same meaning of it.

    That is the disputed claim, stated in full. And here is where the next sentence actually ends:

    So argumentiert auch al-Kairânawî selbst im ersten Kapitel des sechsten Buches, wenn er Muḥammad für den in den Büchern der Schriftbesitzer angekündigten Propheten Gottes hält.

    “Al-Kairânawî himself argues this way too, in the first chapter of the sixth book, when he holds Muhammad to be the prophet of God announced in the books of the People of the Book.”

    Same words, in the same order, up to the comma. What changed is the context: the sentence before it, and the half of its own sentence that had been cut off. The wenn clause names the broad claim, that Muhammad is the prophet foretold in earlier scripture, rather than the narrow one about whether the Greek read paraklētos or periklytos. Whether that redirects the so at the head of the sentence or merely specifies it is a question about the page, and German readers can disagree about it.

    What is not arguable is what I was given. A main clause, stopping one clause short of where the sentence ends, arriving punctuated as though it were complete. My reading had rested on where I took that backward-pointing word to land. The clause that would have put the landing in doubt was the part I never saw.

    I have described that as a translation problem, and it is not quite. The words in the clause were rendered accurately. What went wrong is upstream of translation: the extraction stopped at a comma and reported it as a full stop, and the model had the page image in front of it when it did. That is not a request for insufficient context. That is a transcription error, arriving in exactly the register it would have used had the quotation been complete.

    What the tooling could not do was tell me the unit was wrong. Asking for the context was a decision, not a capability.

    The German section also produced something quieter that I nearly missed. Schirrmacher’s discussion runs across two pages. The sentence I found by searching his name is on the first page. The sentence that actually sets out the vowel-change argument is on the second. A search for his name near the disputed word needs both in the same window, and they are not. The concept was there. The search terms never met.

    Where three weeks got me

    Six sources around one passage. The Arabic original. An English translation that stops one book short, and still no explanation. The Urdu it was translated from, which has all six. A French rendering from 1880 that has been quietly complete this whole time. A German study that reaches the section. And a Turkish thesis working from the Urdu.

    None of them is the answer to my original question. The answer is in the differences between them, which is why it takes a journal article rather than a blog post, and why I am not going to pretend I can land it in a paragraph here.

    The journey is the focus today. Every step away from the Arabic cost something. The English lost a book. A quoted sentence lost a clause. A sampled introduction lost the paragraph that mattered. None of those losses announced itself, and every one of them arrived looking fully complete.

    Three things I did not fully appreciate at the start

    The calendar. The pamphlet is dated 1268 AH. In addition to translating languages I was also translating Hijri and Gregorian dates, which do not line up, and which made searching by date its own small exercise in not being confidently wrong.

    In Kairanawi’s text, the Greek does not appear in Greek. Both forms reach him transliterated into Arabic script, بيركلوطوس for periklytos and باراكلي طوس for paraklētos, which is how they sit on the page. You cannot search the Arabic for either Greek word. You can only search it for a Greek word wearing Arabic clothes, and you have to know in advance which spelling the editor settled on.

    And the work divided across three platforms. I turned to Claude for OCR, translation, transcription and context; Perplexity Pro Academic for literature research; and SuperGrok for broad searches. Each was good at something the others were not, which I would like to claim I designed.

    What I would tell you about the tools

    OCR to find. Vision to see. Where I needed to know whether a term occurs anywhere in a corpus, OCR was the practical way to do it. Where I needed to know which footnote attached to which paragraph, or whether a contents page lists six items or four, OCR turned the page into a stream of words with the arrangement thrown away.

    Language decides which. Printed Arabic OCRs well. Nastaliq did not, for me. English print in two-up scans OCRs adequately for search and unusably for quotation. German I never attempted, because twelve pages is not worth a pipeline.

    Every tool returns something. The word count returned a number. The nastaliq text layer returned text. The name-near-term search returned a clean result. The inference about which book returned a conclusion. All four were wrong and none of them looked it. Two of those four are not AI at all, which is worth saying, because the failure mode is not specific to AI. It is specific to instruments that answer.

    I could not have done this work without multiple AI models. Individual steps were ordinary tooling. But the chain, across six languages I do not read, with translation, page reading, and context at every joint, does not exist for me without it.

    AI disclosure

    Written with AI assistance. Claude Chat (Opus 5) handled drafting and editing, ran the Arabic and English OCR, and read the Arabic, German and Urdu page images. Perplexity Pro Academic handled literature searching and independent fact-checking. SuperGrok handled broad searching and adversarial review. ChatGPT provided editorial review. No platform verified its own output. The Schirrmacher pages were read in Google Books preview; I have not purchased the book, and I cannot know whether pages outside the preview qualify what is on page 186.

    The errors described above were caught by a second method, a second AI platform, and going back to the object. That is the detection mechanism, and it has a blind spot: it cannot find a fluent, complete, wrong answer that is consistent across platforms and never checked against the source. These are the ones that were caught. I have no way to count the ones that were not.

    About the author

    Richard E. Rudd is an independent researcher working across religious history, constitutional history, genealogy and medical research. He spent more than twenty years as an IT Product Owner and Program Manager at large multinational financial services and insurance firms before applying those governance disciplines to AI-assisted research. The reasoning is set out in Governance by Design at https://ruddresearch.com/2026/07/16/governance-by-design-what-twenty-years-in-regulated-it-taught-me-about-ai-assisted-research/, and at Four Platforms, One Standard at https://ruddresearch.com/2026/04/03/four-platforms-one-standard-why-serious-ai-assisted-research-needs-more-than-a-better-prompt/

    Sources

    Raḥmatullāh al-Kairānawī, Iẓhār al-Ḥaqq, ed. Muhammad Ahmad Malkāwī (Riyadh: General Presidency for Scholarly Research and Iftāʾ, 1410/1989), 4 vols., 1,552 pp. Scan: Internet Archive `WAQ32899WAQ`.

    Izhar-ul-Haq: The Truth Revealed, trans. Muhammad Wali Raazi, notes by Muhammad Taqi Usmani (London: Ta-Ha Publishers, 1989–90), 4 parts. Part 1 ISBN 0-907461-68-9; Part 4 ISBN 0-907461-79-4. Scan: Internet Archive `IZHARULHAQ_ENGLISH`.

    Akbar ʿAlī Khān, trans., Bāʾibil se Qurʾān Tak, commentary by Muḥammad Taqī ʿUsmānī, 3 vols. (Karachi: Maktaba Dār al-ʿUlūm).

    Idh-har-ul-Haqq ou Manifestation de la Vérité, translated from the Arabic by an unnamed Tunisian, revised with preface and notes by P. V. Carletti, 2 vols. (Paris: Ernest Leroux, 1880).

    Christine Schirrmacher, Mit den Waffen des Gegners: Christlich-muslimische Kontroversen im 19. und 20. Jahrhundert, Islamkundliche Untersuchungen 162 (Berlin: Klaus Schwarz, 1992), 186. Digital reissue: De Gruyter, ISBN 9783112401088.

    Rizwanullah, “19. yüzyılda Hindistan’da Müslümanlar ve Hıristiyanlar arasındaki dini tartışmalar (İzharü’l-hak örneği)” (MA thesis, Ankara Üniversitesi, 2017), v + 152 pp., at 130. İSAM catalogue no. 26169; YÖK Açık Bilim handle 20.500.12812/68301.

    On nastaliq recognition: “From Press to Pixels: Evolving Urdu Text Recognition,” arXiv:2505.13943.

    Notes: what I could not reach

    Worth listing, since a search that returns nothing and a source you never opened are not the same thing, and only one of them tells you anything.

    – Brill’s Christian-Muslim Relations: A Bibliographical History, the standard bibliographic reference for this literature. Paywalled.

    – Gordon Nickel’s book-length English response to Kairanawi. In print; I did not obtain a copy.

    – Avril Powell’s 1976 article, behind an academic paywall.

    – Schirrmacher’s monograph beyond the pages available in preview.

    None of those is a null result. They are places I did not get to. Any one of them could contain something that changes the picture, and I would rather say so than round them off.

  • Claude Chat proposed that the nineteenth-century scholar Rahmatullah Kairanawi had originated a particular argument. The argument is that the Greek paraklētos in John’s gospel was originally periklytos, which would make it a rough match for the name Ahmad. It sounded right. Correct period, correct figure, correct kind of claim. It shaped three rounds of research prompts before I confirmed it.

    When I downloaded Kairanawi’s argument in the original Arabic and had my Claude OCR worker transcribe it, the argument had actually been proposed by the English orientalist George Sale in 1734, and Kairanawi had examined it and refused it.

    Claude had proposed the attribution as a hypothesis, with a hedge. It wasn’t a fabricated citation. It was a plausible guess. Claude stopped flagging it as a guess in later turns, and it hardened into a premise because nobody tested it. What let it sit unchallenged is that Kairanawi’s refusal lives in the sixth book of a work whose standard English translation stops at the fifth. Nothing in English contradicts the claim, because nothing in English contains the passage.

    Anyone who works with primary sources knows this shape of problem. A will summarised in a calendar entry rather than transcribed. A clinical finding that reaches you through a review article instead of the trial. The error is small, it sits in the part nobody translated, and everything downstream inherits it.

    A wrong footnote costs you a correction. A wrong hypothesis costs you weeks.

    Seven errors turned up during this project. None of them reached publication, because all seven got caught at the gate the process specifies for catching them.

    The test case

    I have been having some genuinely interesting conversations with a small group of very diverse college-age friends about the historical facts and other proofs that major religions use. So I decided to apply my current AI assisted research methodology to it.

    Not a comparison of religions, and not a defence or attack on any of them. Just the same discipline I would apply to any other research question. A claim is made, check it against the evidence, and see whether it survives its own stated terms.

    I am an independent researcher working across several domains. The framework is domain-agnostic, and it’s currently running in genealogy, healthcare and elsewhere.

    What made this a useful test wasn’t that the subject was unfamiliar. It was three properties of the material. The topic is contested, with adversarial literatures on both sides, so secondary sources lean in predictable directions. The primary sources are in a script I cannot read. And it is a field where a fabricated citation looks exactly like a real one: plausible author, plausible title, plausible page range, plausible argument.

    If the framework was going to fall over, it had good conditions to fall over in.

    What transferred, and what didn’t

    The governance transferred cleanly. Staged development, internal audit at release candidate, external verification on platforms that didn’t produce the work, primary-source checking, and honest disclosure of any gate left open.

    The research instrument did not transfer at all. In genealogy I build transmission graphs and test directional sequence. Here I built a claim-testing rubric with rated domains: attestation, dating security, what premises a reader has to grant, how strong the best surviving counter-argument is. Completely different tools, same workflow.

    That distinction is the useful part. The governance is domain-agnostic. The instruments it governs are domain-specific. Anyone hoping for a single method that answers questions in genealogy, healthcare and comparative religion is hoping for the wrong thing. What travels is the discipline around the method, not the method itself.

    The staging comes from twenty-odd years of software delivery in a highly regulated industry, which I’ve written about in Governance by Design and am developing into a peer-reviewed article. Early drafts are development and unit test: internal iteration, adversarial review inside the same session, no external audit. External QA starts when the argument stops moving between drafts, which on this project was version six. Sending an unstable draft out for verification wastes audit cycles on text that will change structurally before anyone reads it again.

    The errors, by type and by cost

    Fabricated attribution. A claim attached to a real author and a real book, except the source doesn’t make it. This is the one unit testing cannot catch, because the reviewing model has no independent access to the artifact. It sits there among the correct citations looking exactly like them.

    Fabricated bibliographic detail. Invented first names for two real authors, in a citation that was otherwise fine. Neither external platform caught it. A parallel session did. Slightly uncomfortable lesson: the QA gate isn’t one external pass, it’s however many independent readings you can afford.

    Truncated quotation. An interpretive option presented as an author’s settled position. I’ll come back to this one, because it turned into something more interesting than a correction.

    An unchecked hypothesis. The error from the opening, and the expensive one by a distance. Everybody verifies citations, because everybody knows citations can be wrong. A hypothesis that a platform floats mid-conversation just gets absorbed into the direction of the work, and a bad one takes everything downstream with it. Same gate, applied earlier.

    On the tools, for the record: seven errors in total. Perplexity Pro Academic found four. SuperGrok found one, and it was methodological rather than factual, spotting that I had retrofitted a claim about predictive success onto research that had merely noticed a pattern after the fact. A parallel session found the invented first names. The Kairanawi hypothesis fell to all three searches converging on the same null. In the same round it caught the retrofit, SuperGrok’s most confident assertion was unsourced and wrong.

    I’m reporting the asymmetry rather than tidying it up, because “the external audit caught the errors” is less useful than what actually happened. I suspect the split is structural rather than luck. Perplexity Pro Academic is search-first with scholarly connectors wired in, and every error it caught was a citation or source error. SuperGrok caught a reasoning error, which is what an adversarial engine is for. Neither found the other’s, and the one both missed took a third reader entirely.

    The control that wasn’t in my framework

    Cross-platform verification works beautifully on fabrication. A fabricated citation fails differently on a differently-trained model, so the second platform catches what the first invented.

    It did nothing at all for summary-level drift, at least here, and this project made that painfully clear.

    Perplexity Pro Academic and SuperGrok both answered questions about a primary text by reading summaries of it, while the text itself sat one click away. One couldn’t extract a PDF that opened cleanly for another tool. A search tool told me a document appeared in only two places online while I was looking at a third instance sitting in my own files.

    What worked, over and over, was supplying the document. Not a report of the source. The source. The moment I put page images of the Arabic critical edition in front of a model, a question that had eaten three rounds of searching resolved in one exchange.

    So the framework formalized an intuitive rule. Any load-bearing claim about what a text says requires someone to open the text. Cross-platform verification is necessary, and it has to validate against the original document wherever possible.

    How hard that is varies enormously. Colonial probate records are straightforward. Esoteric nineteenth-century Arabic is not, and neither is Latin in secretary hand or a Norwegian township register. The answer isn’t a specialist on call for every language you meet. It’s the same chain you’d run anywhere else: OCR with confidence scoring, independent transcriptions compared against each other, published translations used as a check, and the original text quoted in full so a reader can verify it themselves. That’s more auditable than a private consultation, because it leaves a trail anyone can follow.

    Two findings, not corrections

    The standing complaint about AI-assisted research is that it just rearranges what already exists. Two results from this project push back on that, and both came from chasing errors instead of quietly patching them.

    The truncation that wasn’t the AI’s fault. The system quoted the medieval commentator Fakhr al-Din al-Razi asserting that certain biblical narratives matched the Qur’an exactly, without any difference at all. Verification against the Arabic showed the clause was the second of four interpretive options al-Razi lists. Completely standard exegetical practice. Not a settled claim at all.

    But the AI hadn’t invented the truncation. That is how the passage circulates in the polemical literature, and it had been picked up faithfully. Worth noting that this is the second time a draft of mine arguing for citation discipline has turned out to contain the exact fault it was arguing against. I am starting to think the failure mode clusters around whatever you happen to be writing about.

    Then, while checking something unrelated, the same structure showed up running the other way. Raymond Brown, the Catholic biblical scholar, opens an appendix by noting that various scholars have doubted a particular identification, and then spends nine pages testing that view and rejecting it. The version circulating in Muslim apologetics stops at the doubt. It deletes his next sentence, the one announcing that he is about to test the claim. Some versions add a sentence he never wrote, promising that proof is coming.

    Two cases, opposite confessional directions, one technique. Take an expert from the other tradition, strip off the frame that makes the statement provisional, present the result as testimony against his own side.

    Two cases are not a pattern, and I’m not going to pretend otherwise. What they are is a hypothesis worth testing against a corpus, and it exists only because somebody followed an error back to its source instead of fixing it and moving on.

    The passage nobody had read. Verifying that opening hypothesis meant getting page images from the Arabic critical edition. What they showed was better than the correction. Kairanawi had received the argument from a named missionary tract, printed in Calcutta in a specified year, published by the very people the argument was going to be used against. He credited it as plausible. Then he set it aside as not good enough for his purposes and built his case on different ground entirely.

    That discussion sits in the sixth book of his work. The standard English translation contains the first five. Four independent checks, three searches across six languages plus a recent full-length study of the book by a specialist, found nobody engaging that passage. One reference volume remains unchecked behind a paywall, which is the caveat that belongs with the finding rather than after it.

    I’m not claiming a major discovery here. I’m claiming a verification chain can produce findings and not just catch mistakes, which is a different thing from rearrangement.

    What “it worked” actually means

    Seven errors. All caught before publication, at the gate the process specifies for catching them.

    The value was never error prevention. No matter how good your prompt, it doesn’t stop a system from generating a plausible fabrication, which is the argument I made at more length in Four Platforms, One Standard. Prompting is the first discipline, not the last defence. The value is that errors surface while they are still cheap. In a draft, not in print. In a footnote, not in a conclusion resting on one.

    One gate is still open. That paywalled reference volume is the last place a contrary finding could be hiding, and it’s disclosed in the article rather than quietly closed, which is the other half of the discipline.

    The article is unpublished anyway, since I have two remaining religions to finish plus a comparative summary. But the seven errors are already dealt with, and that is what a working method looks like. Not one that prevents mistakes. One that makes them cheap.

    Richard E. Rudd is an independent researcher working across multiple fields. He spent more than twenty years as an IT Product Owner and Program Manager at large multinational financial services and insurance firms before applying those governance disciplines to AI-assisted research. This post describes one project run under that framework; the article it discusses is in preparation.

  • Every AI model gets things wrong. That’s not controversial — it’s well-documented. Chat GPT hallucinates confident citations to papers that don’t exist.[1] Claude can construct logically elegant arguments from flawed premises. Gemini sometimes conflates similar but distinct concepts. Perplexity occasionally returns outdated information as current fact. And the problem goes deeper than factual errors: researchers have documented that even semantically similar prompts can produce drastically different outputs from the same model, a fragility known as “prompt brittleness.”[2]

    In my own research — tracing nine armigerous ancestral lines through the clothier and gentry families of Tudor Kent — I’ve watched three different AI platforms confidently produce three different, mutually exclusive interpretations of the same sixteenth-century will. Each interpretation was internally consistent, well-reasoned, and wrong.

    The standard response to these problems is to write better prompts. Be more specific. Add constraints. Use chain-of-thought reasoning. Provide few-shot examples. And all of that genuinely helps — these are real advances, and researchers teaching prompt engineering at conferences and in workshops are doing valuable work. A well-structured prompt with clear context, role assignment, and explicit formatting constraints will produce measurably better output than a vague one.[3]

    But you’re still asking a single model to check its own work. You’ve improved the employee, but you haven’t addressed the fundamental problem: one perspective, one training dataset, one set of blind spots.

    The problem is worse than most users realize. Researchers at MIT CSAIL recently formalized what they call “delusional spiraling” — situations where extended conversations with a single AI chatbot lead users to high confidence in demonstrably false beliefs.[4] The Human Line Project has documented nearly 300 such cases, including an accountant who, after weeks of single-chatbot conversation, came to believe he was trapped in a false universe.[5] The MIT team’s formal model demonstrates something unsettling: even an idealized, perfectly rational user is vulnerable to this effect when interacting with a sycophantic model. And two commonly proposed mitigations — preventing the chatbot from hallucinating, and informing users about model sycophancy — do not eliminate it.[6]

    This is the single-model problem stated in its most extreme form. But milder versions of the same dynamic affect every serious research project conducted with a single AI platform: the model’s agreeable, confident, internally consistent output gradually shapes the researcher’s thinking, and there’s no independent check in the system.

    Beyond Single Models

    Andrej Karpathy proposed an elegant next step — the LLM council, where you query multiple models and synthesize their responses.[7] It’s a real improvement over single-model work, and for many use cases it’s sufficient. But there’s a limitation that matters for research: large language models are trained on overlapping internet-scale datasets. Because they share training data, they are likely to share blind spots — a form of correlated error that an LLM council may not catch, even when the individual models would each flag a different mistake in isolation.

    This limitation is real but relative, not absolute. Current frontier models do use different architectures, different fine-tuning approaches, and different post-training processes, which produces genuinely different error profiles on many tasks. The methodology treats multi-platform verification as an additional layer of defense, not a final one — just as the Genealogical Proof Standard never permits reliance on a single source, this framework never treats cross-platform agreement as a substitute for the primary archival record.

    None of this means that better prompting or LLM councils are wasted effort. Prompt engineering is the first discipline of serious AI-assisted research, and this methodology assumes that competence. Chain-of-thought reasoning, few-shot examples, role-assignment techniques — these are meaningful advances that make AI more useful for everyone. An LLM council that queries multiple models is better than querying one. The methodology I’m describing doesn’t replace these approaches. It builds on them. It starts from the premise that you’ve already written a good prompt for a capable model, and asks: what structural verification do you still need?

    The chess world offers a useful parallel for thinking about this progression. As Dr. Dominic Ng recently observed, chess is “30 years ahead of every other profession in dealing with AI.”[8] In 2005, two amateurs with laptops famously beat both grandmasters and supercomputers by combining human and machine input more effectively than either could operate alone.[9] It was the triumph of “human in the loop.” But by 2026, the dynamic has reversed: adding a human to a chess engine actually makes it play worse.[10] “Human in the loop” wasn’t a permanent solution — it was a transitional phase.

    But genealogical research is not chess. In chess, the outcome is computationally verifiable — there is a correct move, and a sufficiently powerful engine will find it. Genealogical conclusions require judgment about ambiguous, incomplete, and contradictory evidence that no AI can resolve for itself. The Genealogical Proof Standard codifies that requirement: its five elements demand not just evidence but reasoned human evaluation of evidence, including the resolution of conflicts that may have no algorithmic solution.[11] The lesson from chess isn’t that human judgment is obsolete. It’s that how humans and AI interact matters more than whether they interact. The value isn’t in being “in the loop” as a vague reassurance — it’s in designing the loop itself.

    When Standards Are at Stake

    In fields where research must meet a defined evidentiary standard — where conclusions aren’t just opinions but claims that can be independently verified — single-model AI assistance creates a specific and dangerous problem: it can produce work that looks rigorous while containing errors that the model itself cannot detect.

    The separation-of-duties principle in regulated industries offers a conceptual parallel. In financial auditing, the people who prepare the accounts are never the same people who audit them — a principle codified in standards like ISA 610 and SOX Section 404.[12] Pharmaceutical companies apply similar logic to drug discovery validation. The architectural principle is the same: no single system, human or artificial, should be trusted to validate its own output. Some regulated industries are actively developing independent AI validation workflows; the principle is well established even where the specific multi-platform practice is still emerging.

    Genealogy has its own published evidentiary standard — the Genealogical Proof Standard (GPS), a five-element framework that governs how conclusions are reached and documented in peer-reviewed genealogical scholarship.[13] The GPS requires reasonably exhaustive searches, complete and accurate source citations, skilled analysis of evidence, resolution of conflicting evidence, and soundly reasoned written conclusions. It doesn’t care whether your tools are digital or analog. It cares whether your process is defensible.

    The same structural question faces historians working with digitized archives, anthropologists analyzing field data, and social scientists conducting qualitative research: how do you integrate AI into a research process that must meet a defined evidentiary or methodological standard?

    A Field Ready for Governance

    The research community is already moving toward an answer. In genealogy, the Coalition for Responsible AI in Genealogy (CRAIGEN) has published guiding principles for ethical AI use, emphasizing accuracy, disclosure, and compliance with existing standards.[14] Steve Little, the National Genealogical Society’s AI Program Director, has led workshops and webinars on responsible AI adoption, including GPS-compliant narrative writing with AI assistance — work that has helped establish the vocabulary the community uses to discuss AI and evidentiary rigor. In a recent cross-platform test, Little fed the same 1909 newspaper article to both Claude and ChatGPT and got meaningfully different results: one platform extracted 34 individuals connected by family relationships, the other extracted 55 names including unlinked mentions — same source, same prompt, different architectural decisions about what constitutes genealogically relevant data.[15] Brian’s Ancestors and Algorithms podcast has recently explored GPS-mapped multi-platform workflows, assigning specific AI tools to specific GPS elements based on functional strengths — a practical demonstration that researchers are independently arriving at multi-platform architectures.[16]

    In adjacent fields, multi-agent AI pipelines for fact verification and bias reduction are emerging in clinical research and journalism, and university programs are beginning to teach cross-platform AI output verification as a core research skill.[17] The tools are available. The community recognizes the need for governance. What’s missing is a formalized framework — one that goes beyond best practices and workflow tips to define roles, verification protocols, and documentation requirements that make AI-assisted research as defensible as traditional scholarship.

    A Governance Methodology, Not a Tool Review

    Over the past year, I’ve been developing and field-testing exactly that: a structured multi-platform AI verification methodology for standards-governed research, demonstrated first in genealogy because that’s my current research domain, but designed to be portable to any field with a published evidentiary standard. The framework assigns different AI platforms to specialized roles — deep analysis, project orchestration, adversarial auditing, independent cross-verification — and governs the entire process through the five elements of the GPS.

    This is not “ask another chatbot.” It is a role-based verification system with locked decisions, confidence labels, and an audit trail. The methodology defines specific platform roles, establishes locked analytical decisions that prevent later sessions from reverting confirmed conclusions, enforces temporal rules about what kinds of evidence can upgrade what kinds of claims, and documents errors and corrections as rigorously as successes. It’s the difference between “I got a second opinion” and a structured audit process with defined responsibilities and a paper trail.

    The methodology rests on two pillars.

    Platform Specialization. Each AI platform has measurable strengths. Independent benchmarks — from Stanford’s HELM evaluation to LMSYS’s Chatbot Arena — confirm that no single model leads across all task categories.[18] Rather than using one platform for everything, the methodology assigns roles based on demonstrated capability — the same way a research team assigns tasks based on expertise. If you were hiring researchers, you wouldn’t hire five people with the same degree from the same university. You’d want a data analyst, a subject-matter expert, a skeptic whose job is to find holes, and a project manager. The methodology applies the same logic to AI platforms.

    Multi-platform architecture also unlocks capabilities that single-platform work cannot access. Research on automated prompt optimization — including Google DeepMind’s OPRO framework and Zhou et al.’s work demonstrating that LLMs are “human-level prompt engineers” — has shown that AI models can generate prompts that match or exceed human-crafted ones.[19] In a multi-platform methodology, this means one platform can generate queries optimized for another platform’s specific strengths — a form of inter-platform collaboration that builds on human prompt engineering skill rather than replacing it.[20]

    Cross-Platform Verification. No conclusion stands unless it has been independently examined by at least two platforms with different training data and different architectures. This is the digital equivalent of what the GPS already requires: no responsible genealogist trusts a single source for a critical conclusion.[21] The methodology extends that principle to the AI tools themselves. Crucially, this is the structural defense against the “delusional spiraling” that the MIT team documented — by breaking the single-model feedback loop, cross-platform verification introduces the independent check that single-model interaction inherently lacks.

    The framework is meant to be adopted in stages: start with two platforms and a no-single-source rule, then add roles and controls as the project grows. Many AI platforms offer free tiers; a practical starting configuration — one analytical hub and one adversarial research engine, both at paid tiers runs about $40 per month and delivers a meaningful improvement in both capability and verification rigor over any single-platform workflow. It maps to all five GPS elements, documenting not just what the AI contributed but what it got wrong and how the errors were caught.

    The key distinction is important: AI enables, the methodology governs, the researcher decides. In this framework, AI-generated suggestions are treated as research leads, never as evidence. No citation, transcription, abstraction, or proof statement enters the final argument until the human researcher has verified it against the underlying records. The tools expand what a single researcher can accomplish. The methodology ensures that expansion doesn’t come at the cost of evidentiary rigor. And the researcher — with their subject-matter expertise and professional judgment — remains the final authority on every conclusion.

    What It Produces

    I’ve applied this methodology across a demanding research project: tracing the English Tudor ancestry of a colonial American woman through nine armigerous lines converging in the clothier and gentry families of Cranbrook, Kent, c.1460–1620. The project has run across eighty-plus research sessions spanning multiple AI platforms, producing six publication-track articles across three peer-reviewed disciplines over the course of a year.

    The Courthope correction is perhaps the most telling example of the methodology in action. One AI platform generated a plausible claim about a generational relationship in a prominent Kent family pedigree. The adversarial auditor — a different platform with different training data — flagged it as suspicious. Investigation across platforms confirmed it was a false positive. But the investigation also uncovered something real: a generational error that had propagated through published sources for nearly two centuries, since William Courthope, Somerset Herald, first identified it — but the correction never reached the genealogical literature. The cross-platform workflow surfaced an error that had remained unresolved in the published line of transmission across multiple generations of scholarship. A correction article is in preparation for a leading county history journal.

    The methodology also produced a quantitative analysis of fiduciary trust networks among eleven Tudor clothier families — measuring whether formalized financial trust preceded or followed marriage alliances. Across twenty-one documented family pairings, formalized trust preceded marriage in every testable case, with zero counter-examples. An article is in preparation for a peer-reviewed social history journal.

    It identified a pedigree collapse — the same individual appearing twice in the ancestral tree through two different descent pathways — that connects nine documented armigerous lines through two intermarried gentry networks. A monograph documenting these lines with full GPS-compliant proof arguments is in preparation for a genealogical register.

    And it developed a diagnostic framework for assessing armigerous claims through extinct male lines — a methodological problem that traditional heraldic scholarship doesn’t address, but which cognatic-descent hereditary societies routinely encounter. An article is in preparation for a heraldic journal.

    The sycophancy researchers at MIT proposed mitigations that include “informing users of the possibility of model sycophancy.”[22] This methodology goes further. It doesn’t just warn researchers that AI might agree with them too readily — it builds structural disagreement into the workflow. The adversarial auditor’s job is literally to find reasons the analysis is wrong. When it can’t, that’s evidence of robustness. When it does, that’s the methodology working as designed.

    What’s Next

    A full article documenting this methodology — with detailed case studies, the complete GPS mapping, and the cross-platform verification protocol — is in preparation for submission to a peer-reviewed genealogical journal. This post describes the methodology and its general results; the specific evidence, proof arguments, and detailed analysis are reserved for peer-reviewed publication. The framework is portable: any field with an evidentiary standard — from legal scholarship to historical research to clinical case reporting — can use it to govern AI without tying the method to any one platform. The governance layer doesn’t depend on any specific AI platform — it governs whatever tools are in use, and it will continue to govern whatever tools replace them.

    I’ll be sharing more about specific findings — including the error that persisted across centuries of scholarship, the trust networks that preceded marriage, and the diagnostic framework for extinct male lines — in upcoming posts.

    Richard E. Rudd is an independent researcher. One of his current projects traces nine armigerous ancestral lines through two Tudor gentry networks in the Kentish Weald, c.1460–1620. This research was conducted using a multi-platform AI verification methodology; a full disclosure of AI tools and methods will appear in the published articles.


    Notes

    1. Salvagno, M., Taccone, F.S., & Gerli, A.G. (2023). “Artificial intelligence hallucinations.” Critical Care, 27(1):180. doi:10.1186/s13054-023-04473-y.
    2. Lee, J.H. & Shin, J. (2024). “How to Optimize Prompting for Large Language Models in Clinical Research.” Korean Journal of Radiology, 25(10):869–873. doi:10.3348/kjr.2024.0695. The authors document “prompt brittleness” — minor prompt variations producing substantially different outputs.
    3. Google (2025). Prompt Engineering White Paper. Available at kaggle.com/whitepaper-prompt-engineering. See also Anthropic’s published prompting documentation and OpenAI’s GPT Best Practices guide.
    4. Chandra, K., Kleiman-Weiner, M., Ragan-Kelley, J., & Tenenbaum, J.B. (2026). “Sycophantic Chatbots Cause Delusional Spiraling, Even in Ideal Bayesians.” (Preprint, arXiv 2602.19141.) MIT CSAIL, University of Washington, MIT Department of Brain & Cognitive Sciences.
    5. Hill, K. (2025). “They Asked ChatGPT Questions. The Answers Sent Them Spiraling.” The New York Times, 13 June 2025. The Human Line Project has documented nearly 300 cases of “AI psychosis” or “delusional spiraling.” Cited in Chandra et al. (2026).
    6. Chandra et al. (2026): “This effect persists in the face of two candidate mitigations: preventing chatbots from hallucinating false claims, and informing users of the possibility of model sycophancy.”
    7. Karpathy, A. LLM council concept. See e.g. https://x.com/karpathy — widely discussed in AI research community since 2023.
    8. Dr. Dominic Ng (@DrDominicNg), post on X, 17 March 2026, 1.8M views. https://x.com/DrDominicNg/status/2034252746996785213
    9. The 2005 PAL/CSS Freestyle Chess Tournament, in which amateurs using commodity hardware and chess engines defeated both grandmasters and dedicated supercomputers. Widely documented in chess history.
    10. Ng (2026). Current top engines (Stockfish, Elo ~3,653) exceed the highest human rating by approximately 800 points, making human intervention computationally harmful. See also Rao, V. (@VivekVRao1), X post, 18 March 2026.
    11. Board for Certification of Genealogists, Genealogy Standards (2d ed., 2019; rev. ed. 2021). The GPS comprises five interdependent elements. Element 4 (resolution of conflicts in evidence) and Element 5 (soundly reasoned conclusions) inherently require human judgment about ambiguous and contradictory evidence.
    12. ISA 610 governs the use of internal auditors’ work; SOX Section 404 requires management assessment of internal controls over financial reporting. The principle — that those who produce work should not be the sole evaluators of that work — provides the conceptual parallel for multi-platform AI verification.
    13. Board for Certification of Genealogists, Genealogy Standards (2d ed., 2019; rev. ed. 2021).
    14. Coalition for Responsible AI in Genealogy (CRAIGEN). Guiding principles available at craigen.org. CRAIGEN’s five principles address accuracy, disclosure, privacy, education, and compliance with existing genealogical standards.
    15. Little, S. NGS AI Program Director. Presentations include “Uses of AI in Genealogy” (RootsTech, NGS conferences, 2024–2026) and the “Genealogy Narrative Assistant” project for GPS-compliant AI-assisted writing. See aigenealogyinsights.com. Cross-platform GEDCOM comparison: Little, S. “Turn Anything into a GEDCOM File — With Any AI Tool.” Vibe Genealogy (Substack), 31 March 2026. vibegenealogy.ai/p/turn-anything-into-a-gedcom-file. Both platforms independently caught contradictory evidence in the source and documented the discrepancy rather than silently resolving it.
    16. Ancestors and Algorithms: AI for Genealogy podcast, hosted by Brian. Episode 30 (March 2026) maps GPS elements to specific AI platforms by functional strength. Available at ancestorsandai.com.
    17. See e.g. Shan et al., “Community-Driven AI Support for Genealogy Research” (Virginia Tech, CSCW 2023 workshop); Rice University Fondren Fellows project (2026), teaching cross-platform AI output verification. Multi-agent AI pipelines for fact verification: arXiv 2510.22751 (2025).
    18. Liang, P., et al. (2022). “Holistic Evaluation of Language Models (HELM).” Stanford CRFM, arXiv 2211.09110. LMSYS Chatbot Arena: Chiang, W.-L., Zheng, L., et al. (2024), arXiv 2403.04132. As of early 2026, no single model dominates across all task categories.
    19. Zhou, Y., et al. (2022). “Large Language Models Are Human-Level Prompt Engineers.” arXiv 2211.01910. Google DeepMind’s OPRO: Yang, C., et al. (2023). “Large Language Models as Optimizers.” arXiv 2309.03409. OPRO’s optimized instructions exceeded zero-shot human-crafted prompts by up to 8% on GSM8K, with substantially larger gains on Big-Bench Hard tasks.
    20. The MAPS framework (arXiv 2501.01329, 2025) demonstrates, in the context of software test generation, that prompts optimized for one LLM perform measurably differently on others — confirming that platform-specific optimization yields real advantages.
    21. The GPS requirement for “resolution of conflicts in evidence” (Element 4) inherently demands consultation of multiple independent sources. The methodology extends this principle from sources to tools.
    22. Chandra et al. (2026), abstract.