New Ostrog All articles
Medieval History

Invisible in the Archive: How Digital Platforms Are Failing Eastern European Cultural Heritage

New Ostrog
Invisible in the Archive: How Digital Platforms Are Failing Eastern European Cultural Heritage

Photo: Julian Kücklich, CC0, via Wikimedia Commons

In the spring of 2023, a research team at a major American university attempted to use a leading AI language model to analyze a corpus of medieval Ukrainian chronicles. The model's performance was, by the researchers' own description, embarrassingly poor—not because the technology was incapable of textual analysis, but because the training data underpinning it contained almost no material in Church Slavonic, Old Ruthenian, or the hybrid literary languages common to medieval Eastern European manuscripts. The model had simply never encountered these texts in any meaningful quantity. It had been trained, like most of its peers, on a digital world in which Eastern European heritage is conspicuously underrepresented.

This is not an isolated technical complaint. It is a symptom of a structural problem that runs through the entire architecture of digital cultural preservation: the platforms, institutions, and corporations that now control access to the global cultural record have, through a combination of market logic, linguistic bias, and institutional inertia, replicated and amplified historical patterns of marginalization. Eastern European heritage—medieval, early modern, and modern alike—is being left out of the digital age in ways that will have consequences for generations of scholars, students, and communities.

The Uneven Geography of Digitization

The digitization of cultural materials did not happen uniformly. It happened according to funding priorities, institutional capacity, and the interests of the organizations—primarily American and Western European—that led the effort. Google Books, launched in 2004, was a transformative project, but its partnerships were overwhelmingly concentrated in Anglophone and Western European library collections. The Bodleian, the Library of Congress, Harvard, and the New York Public Library contributed millions of volumes. Libraries in Kyiv, Kraków, Vilnius, and Sofia contributed a fraction of that.

The Internet Archive, for all its admirable mission, reflects similar imbalances. Its Wayback Machine has preserved vast quantities of English-language web content from the early internet era. Eastern European web content from the 1990s and early 2000s—a period of extraordinary cultural and political ferment following the collapse of communism—exists in the archive only patchily, if at all. Entire digital communities, early blogging ecosystems, and online discussions that documented the post-Soviet transition have vanished without a trace.

Europeana, the European Union's digital cultural heritage platform, has made genuine efforts to incorporate Eastern European collections. Yet even there, the contributions from Western European institutions dwarf those from Eastern partners, partly because the latter often lack the technical infrastructure, staff capacity, and funding to digitize and format materials according to the platform's metadata standards.

When Algorithms Inherit Bias

The consequences of this uneven digitization landscape become most acute when the resulting datasets are used to train artificial intelligence systems. Large language models, optical character recognition systems, and machine translation engines are only as capable as the data they have processed. When that data systematically excludes a region's textual heritage, the resulting tools perform poorly on materials from that region—and that poor performance becomes self-reinforcing.

Consider the practical implications for a scholar attempting to work with digitized Eastern European archival materials. OCR software trained primarily on Latin-script Western European texts struggles with Cyrillic, Glagolitic, or the hybrid scripts common in medieval Slavic manuscripts. Machine translation tools for languages such as Belarusian, Rusyn, or Sorbian are, where they exist at all, markedly inferior to their counterparts for French, German, or Spanish. Search algorithms optimized for English-language content return Eastern European materials inconsistently, burying them beneath layers of more heavily indexed Western sources.

For American researchers, this creates a practical barrier to scholarship. For communities whose cultural identity is bound up in those materials—Ukrainian Americans, Polish Americans, Czech Americans, and others—it represents something more troubling: a technological infrastructure that implicitly signals their heritage as peripheral.

The Commercial Logic of Exclusion

It would be convenient to attribute this situation entirely to malice or deliberate policy. The reality is more mundane and, in some ways, more difficult to address. Digital platforms and AI companies operate according to commercial logic that rewards scale and penalizes complexity. English-language content is abundant, relatively standardized, and commercially valuable. Medieval Church Slavonic manuscripts are rare, require specialized expertise to process, and represent a tiny potential market.

This logic is not unique to technology companies. It mirrors the economic pressures that have long disadvantaged Eastern European studies within American universities—where Slavic departments have faced cuts, consolidations, and enrollment pressures that Western European language programs have often escaped. The digital world did not invent the marginalization of Eastern European heritage; it inherited and automated it.

But that inheritance is not inevitable, and it is not irreversible. Several American institutions have demonstrated that targeted investment in Eastern European digital heritage can yield significant results. The Slavic Reference Service at the University of Illinois at Urbana-Champaign has long provided specialized support for researchers working with Eastern European materials. The Digital Humanities program at the University of Pittsburgh has supported projects digitizing Ukrainian and Polish historical records. These efforts demonstrate what is possible when institutions commit resources to the problem.

What American Institutions Must Do

The argument here is not that American universities and technology companies bear sole responsibility for preserving Eastern European cultural heritage. That responsibility rests primarily with Eastern European institutions themselves, and many are doing remarkable work under severe resource constraints. The argument is that American institutions, which have benefited enormously from Eastern European intellectual contributions and which house significant Eastern European diaspora communities, have both the capacity and the obligation to do more.

Concretely, this means several things. Major digitization initiatives should establish explicit targets for Eastern European collection inclusion, with dedicated funding streams rather than reliance on ad hoc grants. AI training datasets used for cultural heritage applications should be audited for regional and linguistic representation, with deficiencies treated as technical problems requiring systematic correction rather than acceptable limitations.

Technology companies with significant resources—and significant influence over the shape of the digital cultural record—should establish partnerships with Eastern European libraries and archives analogous to the partnerships they have built with Western European institutions. The technical challenges of working with Cyrillic scripts, non-standard historical orthographies, and damaged physical materials are real but not insurmountable. They require investment, not innovation.

American universities should resist the institutional pressures that have led to the contraction of Slavic and Eastern European studies programs, recognizing that these programs provide the human expertise without which no digitization initiative can succeed. Machines cannot contextualize a fourteenth-century Ruthenian chronicle; scholars trained in the relevant languages, paleographies, and historical contexts can.

The Stakes of Inaction

The digital archive is not a neutral record. It is a constructed artifact, shaped by the choices—commercial, institutional, and intellectual—of the people and organizations that built it. When those choices systematically exclude a region's cultural heritage, the resulting archive does not merely fail to represent that heritage. It actively shapes future perceptions of what is important, what is worth preserving, and whose past deserves to be remembered.

Eastern European civilizations have survived centuries of imperial conquest, deliberate cultural suppression, and physical destruction of their written records. They should not now be rendered invisible by the indifference of digital infrastructure. The technology to prevent that invisibility exists. The question is whether American institutions have the will to deploy it.

All Articles

Related Articles

The Cookbook as Chronicle: How Eastern European Immigrant Kitchens Preserved What History Tried to Erase

The Cookbook as Chronicle: How Eastern European Immigrant Kitchens Preserved What History Tried to Erase

Light Through Damaged Pages: How Spectral Imaging and Artificial Intelligence Are Recovering Lost Medieval Slavic Manuscripts

Light Through Damaged Pages: How Spectral Imaging and Artificial Intelligence Are Recovering Lost Medieval Slavic Manuscripts

Recovered Voices: How Ideas Suppressed Under Soviet Rule Are Transforming American Intellectual Life

Recovered Voices: How Ideas Suppressed Under Soviet Rule Are Transforming American Intellectual Life