This site hosts research updates, pre-publication summaries, and reflections on methodology. Specific evidence, proof arguments, and detailed analysis are reserved for peer-reviewed publication.

Tag: language

  • Chasing one word across seven languages for three weeks.

    Arabic, Urdu, German, Turkish, French, Greek, English. I read English, but high school French and Duolingo Spanish weren’t going to help. A multi-AI approach with OCR, vision, translation, and transcription capabilities was the toolset I needed.

    The side quest

    I am working on a historical evidentiary research project across four religious traditions. Not to validate or invalidate a tradition, or even to question those traditions, but to review evidentiary claims to a historical standard. In the middle of the project a specific question needed answering, and finding the answer turned into a three-week side quest of its own.

    The situation: in 1864 Rahmatullah Kairanawi, a scholar in Istanbul, was working with a pamphlet from Calcutta, written in Urdu and printed in 1268 AH, which is roughly 1851 to 1852 AD. The pamphlet argued that Muslims had misunderstood a particular Greek word in John’s gospel when they translated it into Arabic. The word is παράκλητος, paraklētos, and the argument in dispute is whether it was originally περικλυτός, periklytos, meaning renowned, which is close enough to the name Ahmad to be interesting. The two differ by their vowels and very little else.

    The question: did Kairanawi accept the pamphlet’s argument that paraklētos is the correct reading, or refute it and hold that the original was periklytos? The question seems straightforward… little did I know what I was in for.

    Kairanawi had debated the missionary Karl Gottlieb Pfander at Agra over two days in April 1854 and wrote a very long reply over the following decade. His discussion of the word runs to a few pages of Arabic in a Saudi critical edition of 1,552 printed pages, and I wanted to know what those pages said.

    A rule when researching anything, especially a historical event: go to the original, and treat every copy, abstract and translation as a step away from it where something can go missing. You do not cite the calendared abstract when the will exists. What follows is that rule running across seven languages.

    A brief note on what this post is and is not: The details around the question, responses, and scholarly treatment are being held for a peer-reviewed article, so I am not publishing the answer here, but I will link to it when it appears. Also, while the context of the post touches upon religious traditions, the post is emphatically not about that. This post is about AI assisted research during the side quest, not the destination.

    Arabic

    The quest started with the Kairanawi Arabic edition which runs to 1,552 printed pages. The scan I worked from is a separate digitization of 1,416 images, one of several floating around, and reconciling the two took longer than it should have.

    Tesseract has Arabic language data. Claude has access and installed it. Once it was up and running, the editor’s 136-page introductory study came out at 34,306 words in about ten minutes, and was searchable for the first time.

    Then it produced a number. A word count on the OCR output said 4,080, roughly one word in eight. I assumed the counting tool did not handle Arabic, which is probably wrong: Arabic is space-delimited much like English and the tool handles it fine on clean text. The cause is more likely the extraction. What matters is that a plausible number arrived, was wrong by a factor of eight, and would have told me the corpus was too thin to bother with.

    Searching it had a second problem. A term I was looking for appears seven times spelled one way and once spelled another, the difference being a single diacritic. Search the first form alone and you have found seven-eighths of the evidence, with nothing to tell you that.

    The passage itself, when I reached it, could not be searched at all. What mattered was not the text but the editor’s footnotes, and specifically which footnote attached to which paragraph. That is layout. OCR flattens layout by design. Those two pages had to be looked at.

    So I had the passage, and it was where the secondary literature said it would be. What I did not have was any reader other than me. I cannot read Arabic, and a load-bearing claim resting on my own machine-assisted reading of a language I do not know is not the time to stop and declare success. The obvious check is obtain a second quality translation, ideally in a language I do read.

    English

    A quick Perplexity and Grok search later found an English translation of the work, published in London in four parts between 1989 and 1990. Those four parts render the first five of the Arabic work’s six books. The sixth book, which is where the material I wanted lives, is not in it.

    I established that by comparing two tables of contents, a job needing nothing but Claude’s ability to diff the two contents pages and four minutes.

    Getting to those four minutes took rather longer, because I first worked out the mapping by reasoning from subject matter, and got it backwards. The Arabic contents page would have settled it in ninety seconds. I spent three exchanges building on the inference before I looked.

    The scans, when I did search them, are two pages to an image with no text layer. OCR interleaves the facing pages, so the output finds a term perfectly well and cannot be quoted from.

    That was not the answer I needed. It was another question. A translation that renders five of six books and stops is either an unfinished job or a decision by an editor, and either way something has gone missing at some step. The English edition told me where to look next, because it says on its own title pages that it was translated from Urdu.

    Urdu

    The Urdu edition, in three volumes, contained what the English translation does not.

    However, the Urdu here is written in nastaliq, a script that slopes, overlaps and joins in ways that defeated every attempt at OCR I made. Internet Archive holds all three volumes with text layers attached, and those text layers are noise. Not poor quality, not partially garbled: noise. Current work on nastaliq recognition reports usable results from newer multimodal models, so this is a statement about what I tried rather than about the state of the field.

    The route into the Urdu was a publisher’s contents listing and screenshots of pages, uploaded and read as images. Which is what one does with bad OCR, and it is slower than it sounds.

    The Urdu has all six Kairanawi books. So the missing sixth book from the English translation did not start with the Urdu, it happened on the English side and why it happened is still a question. That settles part of what the English raised and leaves the original question exactly where it was. I still had no second reader of the passage.

    French

    At this point I was hunting specifically for a translation that contains the sixth book, and staying as close to 1864 as I could, on the same principle that fewer steps mean less loss. Back to Perplexity and Grok.

    The nearest found is French. Two volumes, published in Paris in 1880, sixteen years after Kairanawi finished writing, and it contains the sixth book the English transcription omits. It has been sitting there for nearly a century and a half.

    Claude gave its translator’s name as Cadetti, having copied it from a bibliography without confirmation. The name is Carletti. The name was wrong for three drafts until an external verification from Perplexity, specifically validating citations, caught it.

    The correction produced a second correction. The title page credits the translation to an unnamed young Tunisian, with Carletti as reviser and author of the preface and notes. So I had the name wrong and the role wrong, and was not confident in any answer.

    German

    A translation is a second reader, but it is not scholarship. So the next question was whether another researcher since had studied the passage and could shed light on it.

    Perplexity found a monograph on the controversy and on the question I wanted answered: Christine Schirrmacher, 1992. It is a substantial German-language study of the Pfander and Kairanawi dispute, and it is not available in any digitized form I could reach without paying for it.

    Luckily the pages I needed were available on Google Books preview. I obtained screenshots of twelve pages, loaded into Claude and read as images. That worked well, and better than OCR would have, because footnote numbers matter here and OCR routinely drops superscripts.

    Searching for the keyword Kairanawi returned a sentence that appeared to answer my question outright:

    So argumentiert auch al-Kairânawî selbst im ersten Kapitel des sechsten Buches.

    “Al-Kairânawî himself argues this way too, in the first chapter of the sixth book.”

    So argumentiert auch means argues this way too. This way is anaphoric: it points back at whatever was described immediately before, and what was described immediately before was the periklytos claim. I read it as an answer. Kairanawi refuted the pamphlet and held that periklytos was the original. That reading is mine, not Schirrmacher’s, and the whole of it rests on where a single backward-pointing word lands.

    I had my answer.

    A bonus was that Schirrmacher’s scholarship also gave me the citations of the information leveraged in the monograph, so I could get another viewpoint.

    Turkish

    Schirrmacher’s footnotes led to a Turkish thesis.

    A 2017 master’s thesis from Ankara Üniversitesi runs to 152 pages, is fully digitized, and has a clean text layer. Searchable in every sense. The passage I needed is on page 130, and it matters because it summarizes Kairanawi’s treatment of the word from the Urdu rather than the Arabic. A different reader, working from a different version, in a different language, arriving by a different route. I found it by having Claude read the thesis through from the beginning rather than by searching for keywords, which allowed Claude to ingest the full argument in context.

    The German pointed to periklytos. The Turkish thesis, working from the Urdu, pointed the same way. A second account, in a second language, from a different version of the text. Belt and suspenders, I thought, though I had reached the Turkish through the German’s own footnotes, which is not the same as finding it independently.

    Side Quest Complete?

    I had my sentence, translated accurately, and it answered the question. Why didn’t I feel I was done?

    What made me uneasy was that a load-bearing claim was resting on a partial transcription, in a language I cannot check, and from a book I had only seen twelve pages of. And the Turkish thesis was a translation of a translation more than a century and a half later.

    Back to German

    So I went back to Claude and asked for the full chapter rather than the sentence. I wanted context and not just a translation.

    Here is what the original quotation had cut off. First, the sentence immediately before it on the page:

    …daß anstelle des im Neuen Testament angekündigten ‘parakletos’ früher einmal ‘periklytos’ (der ‘Gesegnete’) zu lesen gewesen sei, womit der Prophet Aḥmad bzw. Muḥammad im Neuen Testament prophezeit worden sei.

    “…that in place of the parakletos announced in the New Testament, periklytos, the Blessed One, was once to be read, whereby the Prophet Ahmad or Muhammad was prophesied in the New Testament.”

    Schirrmacher’s parenthetical gloss there, der ‘Gesegnete’, is looser than the lexicon, which gives much-heard-of or renowned. Nothing in the sentence turns on it, but in a piece about one word it is worth saying that the secondary source and I are not using quite the same meaning of it.

    That is the disputed claim, stated in full. And here is where the next sentence actually ends:

    So argumentiert auch al-Kairânawî selbst im ersten Kapitel des sechsten Buches, wenn er Muḥammad für den in den Büchern der Schriftbesitzer angekündigten Propheten Gottes hält.

    “Al-Kairânawî himself argues this way too, in the first chapter of the sixth book, when he holds Muhammad to be the prophet of God announced in the books of the People of the Book.”

    Same words, in the same order, up to the comma. What changed is the context: the sentence before it, and the half of its own sentence that had been cut off. The wenn clause names the broad claim, that Muhammad is the prophet foretold in earlier scripture, rather than the narrow one about whether the Greek read paraklētos or periklytos. Whether that redirects the so at the head of the sentence or merely specifies it is a question about the page, and German readers can disagree about it.

    What is not arguable is what I was given. A main clause, stopping one clause short of where the sentence ends, arriving punctuated as though it were complete. My reading had rested on where I took that backward-pointing word to land. The clause that would have put the landing in doubt was the part I never saw.

    I have described that as a translation problem, and it is not quite. The words in the clause were rendered accurately. What went wrong is upstream of translation: the extraction stopped at a comma and reported it as a full stop, and the model had the page image in front of it when it did. That is not a request for insufficient context. That is a transcription error, arriving in exactly the register it would have used had the quotation been complete.

    What the tooling could not do was tell me the unit was wrong. Asking for the context was a decision, not a capability.

    The German section also produced something quieter that I nearly missed. Schirrmacher’s discussion runs across two pages. The sentence I found by searching his name is on the first page. The sentence that actually sets out the vowel-change argument is on the second. A search for his name near the disputed word needs both in the same window, and they are not. The concept was there. The search terms never met.

    Where three weeks got me

    Six sources around one passage. The Arabic original. An English translation that stops one book short, and still no explanation. The Urdu it was translated from, which has all six. A French rendering from 1880 that has been quietly complete this whole time. A German study that reaches the section. And a Turkish thesis working from the Urdu.

    None of them is the answer to my original question. The answer is in the differences between them, which is why it takes a journal article rather than a blog post, and why I am not going to pretend I can land it in a paragraph here.

    The journey is the focus today. Every step away from the Arabic cost something. The English lost a book. A quoted sentence lost a clause. A sampled introduction lost the paragraph that mattered. None of those losses announced itself, and every one of them arrived looking fully complete.

    Three things I did not fully appreciate at the start

    The calendar. The pamphlet is dated 1268 AH. In addition to translating languages I was also translating Hijri and Gregorian dates, which do not line up, and which made searching by date its own small exercise in not being confidently wrong.

    In Kairanawi’s text, the Greek does not appear in Greek. Both forms reach him transliterated into Arabic script, بيركلوطوس for periklytos and باراكلي طوس for paraklētos, which is how they sit on the page. You cannot search the Arabic for either Greek word. You can only search it for a Greek word wearing Arabic clothes, and you have to know in advance which spelling the editor settled on.

    And the work divided across three platforms. I turned to Claude for OCR, translation, transcription and context; Perplexity Pro Academic for literature research; and SuperGrok for broad searches. Each was good at something the others were not, which I would like to claim I designed.

    What I would tell you about the tools

    OCR to find. Vision to see. Where I needed to know whether a term occurs anywhere in a corpus, OCR was the practical way to do it. Where I needed to know which footnote attached to which paragraph, or whether a contents page lists six items or four, OCR turned the page into a stream of words with the arrangement thrown away.

    Language decides which. Printed Arabic OCRs well. Nastaliq did not, for me. English print in two-up scans OCRs adequately for search and unusably for quotation. German I never attempted, because twelve pages is not worth a pipeline.

    Every tool returns something. The word count returned a number. The nastaliq text layer returned text. The name-near-term search returned a clean result. The inference about which book returned a conclusion. All four were wrong and none of them looked it. Two of those four are not AI at all, which is worth saying, because the failure mode is not specific to AI. It is specific to instruments that answer.

    I could not have done this work without multiple AI models. Individual steps were ordinary tooling. But the chain, across six languages I do not read, with translation, page reading, and context at every joint, does not exist for me without it.

    AI disclosure

    Written with AI assistance. Claude Chat (Opus 5) handled drafting and editing, ran the Arabic and English OCR, and read the Arabic, German and Urdu page images. Perplexity Pro Academic handled literature searching and independent fact-checking. SuperGrok handled broad searching and adversarial review. ChatGPT provided editorial review. No platform verified its own output. The Schirrmacher pages were read in Google Books preview; I have not purchased the book, and I cannot know whether pages outside the preview qualify what is on page 186.

    The errors described above were caught by a second method, a second AI platform, and going back to the object. That is the detection mechanism, and it has a blind spot: it cannot find a fluent, complete, wrong answer that is consistent across platforms and never checked against the source. These are the ones that were caught. I have no way to count the ones that were not.

    About the author

    Richard E. Rudd is an independent researcher working across religious history, constitutional history, genealogy and medical research. He spent more than twenty years as an IT Product Owner and Program Manager at large multinational financial services and insurance firms before applying those governance disciplines to AI-assisted research. The reasoning is set out in Governance by Design at https://ruddresearch.com/2026/07/16/governance-by-design-what-twenty-years-in-regulated-it-taught-me-about-ai-assisted-research/, and at Four Platforms, One Standard at https://ruddresearch.com/2026/04/03/four-platforms-one-standard-why-serious-ai-assisted-research-needs-more-than-a-better-prompt/

    Sources

    Raḥmatullāh al-Kairānawī, Iẓhār al-Ḥaqq, ed. Muhammad Ahmad Malkāwī (Riyadh: General Presidency for Scholarly Research and Iftāʾ, 1410/1989), 4 vols., 1,552 pp. Scan: Internet Archive `WAQ32899WAQ`.

    Izhar-ul-Haq: The Truth Revealed, trans. Muhammad Wali Raazi, notes by Muhammad Taqi Usmani (London: Ta-Ha Publishers, 1989–90), 4 parts. Part 1 ISBN 0-907461-68-9; Part 4 ISBN 0-907461-79-4. Scan: Internet Archive `IZHARULHAQ_ENGLISH`.

    Akbar ʿAlī Khān, trans., Bāʾibil se Qurʾān Tak, commentary by Muḥammad Taqī ʿUsmānī, 3 vols. (Karachi: Maktaba Dār al-ʿUlūm).

    Idh-har-ul-Haqq ou Manifestation de la Vérité, translated from the Arabic by an unnamed Tunisian, revised with preface and notes by P. V. Carletti, 2 vols. (Paris: Ernest Leroux, 1880).

    Christine Schirrmacher, Mit den Waffen des Gegners: Christlich-muslimische Kontroversen im 19. und 20. Jahrhundert, Islamkundliche Untersuchungen 162 (Berlin: Klaus Schwarz, 1992), 186. Digital reissue: De Gruyter, ISBN 9783112401088.

    Rizwanullah, “19. yüzyılda Hindistan’da Müslümanlar ve Hıristiyanlar arasındaki dini tartışmalar (İzharü’l-hak örneği)” (MA thesis, Ankara Üniversitesi, 2017), v + 152 pp., at 130. İSAM catalogue no. 26169; YÖK Açık Bilim handle 20.500.12812/68301.

    On nastaliq recognition: “From Press to Pixels: Evolving Urdu Text Recognition,” arXiv:2505.13943.

    Notes: what I could not reach

    Worth listing, since a search that returns nothing and a source you never opened are not the same thing, and only one of them tells you anything.

    – Brill’s Christian-Muslim Relations: A Bibliographical History, the standard bibliographic reference for this literature. Paywalled.

    – Gordon Nickel’s book-length English response to Kairanawi. In print; I did not obtain a copy.

    – Avril Powell’s 1976 article, behind an academic paywall.

    – Schirrmacher’s monograph beyond the pages available in preview.

    None of those is a null result. They are places I did not get to. Any one of them could contain something that changes the picture, and I would rather say so than round them off.