Preserving What We May Not Yet Know We Need – A Conversation with Prof. David Zeitlyn

Digital preservation is often described as a technical challenge: keeping files secure, maintaining formats, checking data integrity, and making sure information remains accessible as technologies change. But a conversation with Professor David Zeitlyn of the University of Oxford suggests that the hardest questions surrounding preservation may have little to do with technology and more with ethical, social, institutional, and legal questions about what we preserve, who has the right to access it, who decides what should remain private, and what future researchers will need when the people making those decisions today are no longer around. As Prof. Zeitlyn points out, preservation is ultimately about creating a bridge between the present and people we cannot yet know, including researchers who may ask questions we cannot currently imagine.

Zeitlyn's own experience provides an excellent example of how easily important digital history can disappear. He recalls being involved in some of the earliest web experiments at Oxford, including helping create what he describes as the Institute of Social and Cultural Anthropology's first website, at a time when the web itself was still in its infancy and before there was a website for the University of Oxford itself. Yet today, there is no surviving record of that website. There is documentary evidence that came later, but the original digital artifact itself is gone, leaving behind little more than a personal recollection and an unverified claim. It is a small example of a much larger problem: digital information can feel permanent precisely because it is digital, when in reality it can be remarkably fragile. A website that once seemed like a milestone can disappear completely, leaving future historians unable to reconstruct what happened.

This raises a fundamental question about what we mean when we say that something is "open" or "accessible." In scholarly communication, open access has become an essential principle, and rightly so, but openness does not automatically mean meaningful accessibility, particularly when the materials involved contain information about people. Zeitlyn's work as a social and cultural anthropologist brings this tension into focus because research involving human subjects often carries obligations around privacy, confidentiality, consent, and the responsible use of information. The challenge becomes even more complicated when we consider that the people represented in research materials may themselves have an interest in how that information is preserved and eventually accessed.

One of the most provocative issues raised in the conversation is the practice of default anonymization. Anonymization is often treated as the safest and most responsible solution for sensitive research materials, particularly when personal information is involved, but Zeitlyn argues that applying it automatically can have consequences that are rarely considered. Removing names may protect privacy in the short term, but it can also erase relationships, identities, histories, and forms of evidence that future researchers or communities may consider essential. In some cases, descendants may want to identify people in historical photographs or records because they are trying to understand their own family histories, while communities may need historical records as evidence in questions relating to land, identity, or collective memory.

The distinction between anonymization and pseudonymization therefore becomes particularly important. Rather than permanently destroying identifying information, one possible approach is to preserve the original material in a secure, dark environment while creating a public version in which names or other identifying information have been removed or protected. The original could remain inaccessible for a defined period, potentially decades, while the anonymized or pseudonymized version could support appropriate research in the present. In this model, preservation does not mean immediate access, and restricting access does not mean destroying information. Instead, preservation gives future generations choices that would otherwise be lost forever.

That idea becomes particularly powerful when we think about the timescale of scholarship. A researcher may be working today with material that seems too sensitive to release, but what about 50, 75, or 100 years from now? The ethical circumstances may have changed completely. The people involved may no longer be alive, the social context may be different, and new generations of researchers may be asking questions that today's researchers cannot anticipate. If the original material has been permanently anonymized or destroyed, those future researchers will have no opportunity to reconsider it. If it has been preserved securely, however, they may eventually have access to a richer historical record.

The conversation also highlights an important tension between preservation and the increasing pressure for immediate openness. Funders increasingly require researchers to make publications, data, and analytical materials openly available, and from a research and public-interest perspective, this can be enormously valuable. But there are circumstances in which a five-year embargo may be considered long, while researchers working with sensitive qualitative or historical materials may believe that even 20 years is too short. The problem is not simply deciding whether information should be open or closed; it is determining what kind of access is appropriate, to whom, under what conditions, and for how long.

This is where the concept of the dark archive becomes particularly relevant. A dark archive is not a place where information is forgotten; it is a place where information can be deliberately protected until the circumstances for access are appropriate. In the context of sensitive research materials, this model offers an important middle ground between unrestricted access and permanent destruction. As Zeitlyn suggests, there may even be possibilities for creating controlled environments in which researchers can work with sensitive data without being able to remove the underlying material. Such models already exist in other areas of research, where scholars can access detailed datasets within secure environments and leave only approved analyses behind.

For preservation organizations, this changes the conversation considerably. The role is no longer simply to ensure that a PDF, dataset, image, or website survives a technological transition. It is also to create the conditions under which that material can remain trustworthy and usable over time, even when access needs to change. At CLOCKSS, this long-term perspective is central to the preservation model, and the conversation illustrates why maintaining multiple geographically distributed copies matters beyond the technical question of redundancy. When archives exist across different countries and continents, they are also better positioned to withstand political upheaval, conflict, institutional change, or other circumstances that could threaten access in a particular location.

Geopolitics is rarely the first thing people think about when they think about digital preservation, but it is increasingly difficult to separate the two. Wars, political conflicts, institutional instability, and changes in national policy can all affect who controls information and who can access it. A preservation strategy that depends entirely on a single institution, country, technology, or political environment creates vulnerabilities that may not be visible until it is too late. Geographic distribution therefore becomes more than a technical backup strategy; it becomes part of a broader commitment to ensuring that knowledge survives changes in the world around it.

Trust sits at the center of all of these questions. Researchers need to trust that sensitive information will be handled responsibly. Research participants need to trust that their contributions will not be misused. Institutions and funders need to trust preservation organizations to protect materials over very long periods. And future researchers need to be able to trust the records that survive. Building that trust requires more than policies and technical infrastructure. It requires institutions to make clear commitments about what they will preserve, how they will protect it, when access may be granted, and what safeguards will remain in place as technologies, institutions, and laws change.

Libraries have an important role to play in this ecosystem. Zeitlyn describes librarians at Oxford as active participants in preservation and points to the role of data librarians in helping researchers think about how their materials should be managed. This role becomes particularly visible when researchers are documenting events that are happening now, such as the work of archivists in Ukraine who are recording experiences from the war, including firsthand accounts from people on the front lines. The question is not simply how to store those recordings today, but how to ensure that they remain meaningful, trustworthy, and accessible to future generations while respecting the people whose experiences they document.

There is also a deeply personal dimension to preservation. Zeitlyn reflects on how much his own research has benefited from archival materials created by people who lived generations before him, and he has begun thinking about what will happen to his own substantial collection of analog and digital materials. This is perhaps one of the most powerful reminders that preservation is not only about institutions preserving "important" materials. It is about creating a chain between generations in which today's researchers become part of the historical record that tomorrow's researchers will rely upon.

Ultimately, the conversation challenges us to think beyond the idea that preservation is simply about keeping information safe. Preservation is about keeping possibilities alive. A document that survives gives a future researcher the possibility of asking a new question. A photograph that retains its original context gives a family the possibility of understanding its history. A dataset preserved securely gives future scholars the possibility of revisiting conclusions with new methods. A website that survives gives historians the possibility of understanding how scholarship and technology evolved. Once information has been destroyed or permanently stripped of its context, those possibilities disappear with it.

Perhaps the most important message from the conversation is therefore a simple one: we should question our assumptions about what future generations will need. Default anonymization may seem like the safest answer today, just as immediate open access may seem like the most responsible approach, but neither should become an automatic response to every kind of research material. Preservation gives us another option: protect information carefully, preserve its context and integrity, and allow future generations to make informed decisions about how it should be used.

Because we cannot know what questions future researchers will ask, we cannot know today which details they will need. What we can do is make sure those details are not lost before they have the chance to ask. That is the deeper promise of digital preservation: not simply to protect the record of what we know today, but to preserve the evidence that will allow others to discover what we do not yet know.

You can re-watch the webinar, below:

David Zeitlyn ‘For Augustinian archival openness and laggardly sharing: trustworthy archiving and sharing of social science data from identifiable human subjects’ 2021 Frontiers in Research Metrics and Analytics 6 (63). DOI: 10.3389/frma.2021.736568. http://journal.frontiersin.org/article/10.3389/frma.2021.736568

David Zeitlyn  ‘Anthropology In and of the Archives: Possible Futures and Contingent Pasts. Archives as Anthropological Surrogates’, 2012. Annual Review of Anthropology 41, 461–80. https://doi.org/10.1146/annurev-anthro-092611-145721.

Scroll to Top