New Ostrog All articles
Medieval History

Light Through Damaged Pages: How Spectral Imaging and Artificial Intelligence Are Recovering Lost Medieval Slavic Manuscripts

New Ostrog
Light Through Damaged Pages: How Spectral Imaging and Artificial Intelligence Are Recovering Lost Medieval Slavic Manuscripts

Photo: NASA / Johns Hopkins University Applied Physics Laboratory / Southwest Research Institute, Public domain, via Wikimedia Commons

The manuscript had been considered illegible for more than two centuries. Held in the collections of a monastic library in western Ukraine, the codex had sustained severe water damage sometime in the eighteenth century, and subsequent attempts to stabilize the parchment had, inadvertently, made the ink even harder to read. By the time a team of digital humanities researchers from the University of Vienna and the National University of Kyiv-Mohyla Academy examined it in 2019, the text appeared to most observers as little more than a brownish smear across degraded vellum. Six months later, after processing through multispectral imaging and a neural network trained on comparable Cyrillic scripts, the manuscript yielded forty-seven previously unknown pages of a thirteenth-century ecclesiastical commentary — including passages that appear to have no parallel in any other surviving source.

This is no longer an isolated achievement. Across Eastern Europe and in American research universities with significant Slavic studies programs, a convergence of technologies developed largely for other purposes is being redirected toward one of the most pressing challenges in medieval scholarship: recovering the written record of a civilization that suffered centuries of conquest, deliberate destruction, and institutional neglect.

The Technology Explained

Multispectral imaging works on a principle that is, in retrospect, almost obvious: different materials absorb and reflect light differently at different wavelengths, and the human eye captures only a narrow band of that spectrum. Iron gall ink — the dominant writing medium in medieval European manuscripts — retains chemical properties that distinguish it from the surrounding parchment even when it has faded to invisibility in ordinary light. By illuminating a manuscript with controlled light sources across wavelengths ranging from ultraviolet through near-infrared and capturing the results with specialized sensors, imaging systems can reconstruct the contrast between ink and substrate that the naked eye can no longer detect.

The results can be striking. At wavelengths around 850 nanometers, text that appears completely absent in visible light frequently re-emerges with startling clarity. The technique is not new — it has been applied to Western European manuscripts, including the celebrated Archimedes Palimpsest, for decades. What has changed is the cost and portability of the equipment, which has made it practical to deploy in archives and monastic libraries that lack the infrastructure of major research institutions, and the availability of computational tools capable of processing the resulting images at scale.

Artificial intelligence enters the workflow at the point where human transcription becomes impractical. A manuscript that has been successfully imaged may still present thousands of pages of text in archaic script variants that only a handful of specialists can read fluently. Machine-learning models trained on digitized examples of historical Cyrillic, Glagolitic, and related scripts can now produce draft transcriptions of sufficient quality to be usable — not as final scholarly editions, but as working texts that human experts can verify and correct far more rapidly than they could transcribe from scratch. The combination of imaging and automated transcription is compressing timelines that once spanned decades into periods measured in months.

What Is Being Found

The scholarly significance of the recovered materials extends well beyond the satisfaction of recovering previously inaccessible texts. In several cases, the newly readable documents are providing evidence for historical claims that had rested on fragmentary or indirect grounds.

A project centered at Harvard's Houghton Library, in collaboration with partners at the Jagiellonian University in Kraków, has been working with a collection of damaged commercial documents from a medieval Ruthenian trading center. The documents — a mixture of contracts, inventories, and correspondence — had been partially legible since their acquisition in the nineteenth century, but significant portions were obscured by mold damage. Multispectral processing has now recovered enough additional text to confirm the existence of trade relationships between this community and merchants in the Baltic and Black Sea regions that historians had hypothesized but could not document. The commercial vocabulary preserved in the texts is also providing new data for historical linguists studying the development of early East Slavic business terminology.

Elsewhere, the discoveries have been more unexpected. A set of marginalia recovered from a damaged liturgical manuscript at a Belarusian archive turned out to contain what appears to be a fragment of vernacular poetry — a genre extremely poorly attested in this region for the period in question. The fragment is short, and its interpretation remains contested, but its existence alone has prompted a reassessment of assumptions about the relationship between learned and popular literary culture in medieval Belarusian communities.

American Institutions and the Collaborative Framework

The United States has become an important node in this international scholarly network, for reasons that are partly historical and partly practical. American research universities acquired significant Eastern European manuscript collections during the Cold War, when émigré scholars and cultural organizations sought to preserve materials outside the reach of Soviet institutional control. Those collections now serve as training data for AI models and as comparative resources for interpreting newly recovered texts.

The Digital Slavonic Manuscripts Initiative, a consortium that includes Columbia University, the University of Illinois at Urbana-Champaign, and several European partners, has been developing shared protocols for imaging, transcription, and metadata standards that allow recovered texts to be integrated into searchable databases accessible to researchers worldwide. The initiative's publicly available corpus has already been used in more than two hundred published studies since its launch, and its training datasets have been adapted by projects working on Armenian, Georgian, and other non-Slavic scripts from the same region.

Funding has come from a mixture of sources: the National Endowment for the Humanities has supported several American-led components of these projects, while European Research Council grants have underwritten the fieldwork in Eastern Europe itself. Private foundations with interests in Eastern European cultural heritage have contributed to equipment acquisition and digitization costs.

The Limits of the Technology

Scholars working in this field are careful to articulate what these methods cannot accomplish. Multispectral imaging cannot recover text that has been physically removed from a manuscript — only text whose chemical signature remains present even when visually obscured. AI transcription models perform unevenly across script variants and deterioration types, and their error rates, while declining, remain high enough to require systematic human review. The recovery of a text, furthermore, is only the beginning of the interpretive work; establishing what a newly readable document means within its historical context requires the full range of conventional scholarly methods.

Nor can technology address the prior question of which manuscripts survive to be imaged. The destruction of Eastern European cultural heritage across the twentieth century was systematic and vast. What digital humanities can recover is what remains — a subset, of uncertain representativeness, of what once existed. The recovered texts are genuinely illuminating. They are also, inevitably, a fraction of what was lost.

That caveat registered, the pace of recovery is accelerating in ways that would have seemed implausible to the previous generation of medievalists. Manuscripts that were catalogued as damaged and set aside are being reexamined. Archives that lacked the resources to process their holdings are gaining access to tools that make processing feasible. And a body of evidence about medieval Eastern European intellectual and commercial life is emerging from the obscurity into which centuries of conflict had consigned it — legible, at last, to those with the instruments to read it.

All Articles

Related Articles

Recovered Voices: How Ideas Suppressed Under Soviet Rule Are Transforming American Intellectual Life

Recovered Voices: How Ideas Suppressed Under Soviet Rule Are Transforming American Intellectual Life

Type as Defiance: The Underground Printers Who Kept Eastern European Civilizations Alive Under Imperial Rule

Type as Defiance: The Underground Printers Who Kept Eastern European Civilizations Alive Under Imperial Rule

Guardians of the Written Word: How Eastern Europe's Monastic Scriptoriums Kept Learning Alive Through Centuries of War