{"id":922567,"date":"2026-07-26T07:14:24","date_gmt":"2026-07-26T12:14:24","guid":{"rendered":"https:\/\/newsycanuse.com\/index.php\/2026\/07\/26\/why-digital-government-records-are-so-hard-to-preserve\/"},"modified":"2026-07-26T07:14:24","modified_gmt":"2026-07-26T12:14:24","slug":"why-digital-government-records-are-so-hard-to-preserve","status":"publish","type":"post","link":"https:\/\/newsycanuse.com\/index.php\/2026\/07\/26\/why-digital-government-records-are-so-hard-to-preserve\/","title":{"rendered":"Why digital government records are so hard to preserve"},"content":{"rendered":"<div data-click-position=\"inline-link\">\n<p class data-block=\"sciam\/paragraph\">In May <a href=\"https:\/\/www.courtlistener.com\/docket\/73154145\/24\/american-historical-association-v-trump\/\">a federal judge ordered<\/a> White House staff to comply with the Presidential Records Act, the 1978 law that <a href=\"https:\/\/www.congress.gov\/crs-product\/R46129\">makes a president\u2019s official records public property<\/a> and governs their preservation and eventual release.<\/p>\n<p class data-block=\"sciam\/paragraph\">A month earlier the Justice Department <a href=\"https:\/\/www.justice.gov\/olc\/media\/1434131\/dl?inline\">had argued<\/a> that the law exceeds Congress\u2019s constitutional authority. The American Historical Association and the watchdog group American Oversight sued, warning that the opinion could let the White House abandon policies meant to restrict officials from conducting government business through personal e-mail or encrypted messages. The risk, they argued, was a current loss of accountability and a permanent gap in the historical record.<\/p>\n<p class data-block=\"sciam\/paragraph\">Judge John D. Bates has so far found the law \u201clikely constitutional.\u201d But the court fight is just one part of a much broader challenge. The records that reveal how governments and public figures make decisions are now born in email, chat apps and cloud documents, often inside proprietary systems whose lifespans are measured in product cycles. Preserving them long enough for the public to see them has become a technical problem in its own right, one that grows harder as the volume climbs. The National Archives <a href=\"https:\/\/www.archives.gov\/about\/info\/national-archives-by-the-numbers\">added 463 terabytes of electronic records<\/a> to its permanent collection in 2024 alone.<\/p>\n<hr>\n<h2>On supporting science journalism<\/h2>\n<p>If you&#8217;re enjoying this article, consider supporting our award-winning journalism by <a href=\"http:\/\/www.scientificamerican.com\/getsciam\/\">subscribing<\/a>. By purchasing a subscription you are helping to ensure the future of impactful stories about the discoveries and ideas shaping our world today.<\/p>\n<hr>\n<p class data-block=\"sciam\/paragraph\">\u201cThe world is creating digital records at a pace no organization anticipated,\u201d says Mike Quinn, CEO of digital preservation company <a href=\"https:\/\/preservica.com\/about\">Preservica<\/a>.<\/p>\n<p class data-block=\"sciam\/paragraph\">Before archivists can preserve a record, the record must survive long enough to make it into their hands. Public-records laws can require preservation, and the technology exists to capture and store messages even from some encrypted platforms when accounts or devices are configured to retain them. The digital preservation company <a href=\"https:\/\/www.smarsh.com\/capture\/channels\/\">Smarsh<\/a>, for instance, advertises it can capture data from more than 100 communications channels. But recent incidents, from <a href=\"https:\/\/www.scientificamerican.com\/article\/psychologys-groupthink-helps-explain-the-signal-chat-fiasco\/\">U.S. Cabinet officials discussing military plans<\/a> via the encrypted app Signal to U.K. Prime Minister Keir Starmer\u2019s <a href=\"https:\/\/www.bbc.com\/news\/articles\/cp8pgg0108qo\">reported use<\/a> of disappearing WhatsApp messages, suggest how easily significant records can still vanish.<\/p>\n<p class data-block=\"sciam\/paragraph\">The same fragility follows private archives, too. Even when individuals such as politicians or artists\u2014or their estates\u2014donate physical papers to a university library, the digital material that once sat alongside them can be overlooked and lost, says <a href=\"https:\/\/liberalarts.utexas.edu\/germanic\/faculty\/tr24969\">Thorsten Ries<\/a>, an assistant professor at the University of Texas at Austin who <a href=\"https:\/\/biblio.ugent.be\/publication\/01JZRH3046JKSV2VMXA9CQ7GFN\">applies digital-forensics techniques to archival work<\/a>.<\/p>\n<p class data-block=\"sciam\/paragraph\">Pulling the data off a hard drive or USB drive without altering files or metadata such as time stamps also takes skill, Ries says. Different software versions, and even different storage media, can preserve different file fragments and automatic backups. Those offer valuable clues to how a document was drafted and how its creators thought, but recovering and interpreting them is painstaking, specialized work. \u201cThis kind of knowledge and expertise is actually still very sparse,\u201d he says.<\/p>\n<p class data-block=\"sciam\/paragraph\">Cloud-based systems such as Google Docs can hold the most detailed file histories of all, but extracting files from them without the original passwords and two-factor authentication is its own challenge, he adds.<\/p>\n<p class data-block=\"sciam\/paragraph\">Survival is just the first step; the material also must remain readable as software changes. \u201cAll these types of digital content don\u2019t age like paper,\u201d Quinn says. \u201cThey become unreadable when formats become obsolete.\u201d<\/p>\n<p class data-block=\"sciam\/paragraph\">That often requires regularly migrating material like word processing documents, spreadsheets and computer-aided design files to current file formats while keeping a careful log of exactly what\u2019s been done. If handled carelessly, those conversions can misrepresent the original, says Christopher J. Prom of the University of Illinois Urbana-Champaign library. That appears to be what happened when<a href=\"https:\/\/www.yahoo.com\/news\/articles\/fact-check-doj-epstein-librarys-095825612.html\"> the Justice Department released e-mails<\/a> tied to the late financier and sex offender Jeffrey Epstein that were marred by rendering errors.<\/p>\n<p class data-block=\"sciam\/paragraph\">A preserved file can still be hard to use. Digital archives can contain copyrighted material alongside sensitive correspondence, including personal messages and medical bills, sitting in the same inboxes and folders as the files a researcher wants. That makes institutions cautious about opening collections broadly. And though a digital file could in theory be opened from anywhere with an internet connection, archives still routinely require an onsite visit, if they grant access at all, says <a href=\"https:\/\/www.lboro.ac.uk\/subjects\/communication-media\/staff\/lise-jaillant\/\">Lise Jaillant,<\/a> a professor of digital cultural heritage at Loughborough University in England. Researchers must schedule and pay for travel, then comb through enormous collections on potentially unfamiliar systems in whatever time they have.<\/p>\n<p class data-block=\"sciam\/paragraph\">The \u201cstaggering volumes\u201d of digital material produced by U.S. government agencies have likewise slowed the handling of <a href=\"https:\/\/www.scientificamerican.com\/article\/fauci-faces-congressional-committee-over-covid-e-mails\/\">Freedom of Information Act requests<\/a>, says <a href=\"https:\/\/ischool.umd.edu\/directory\/jason-r-baron\/\">Jason R. Baron<\/a>, a professor at the University of Maryland\u2019s College of Information and former director of litigation at the National Archives and Records Administration. Agencies must first try to locate potentially relevant files, often by keyword search, then remove or redact anything <a href=\"https:\/\/www.scientificamerican.com\/article\/chatbot-honeypot-how-ai-companions-could-weaken-national-security\/\">classified, sensitive<\/a> or otherwise exempt from disclosure.<\/p>\n<p class data-block=\"sciam\/paragraph\">\u201cIt is not unusual for a requester to wait years or even in some cases over a decade to receive complete responses,\u201d Baron says.<\/p>\n<p class data-block=\"sciam\/paragraph\">Automation may help, with substantial human oversight. In a 2025 <a href=\"https:\/\/link.springer.com\/article\/10.1007\/s00146-025-02256-3\">paper<\/a>, Baron explored using artificial intelligence and machine-learning techniques to flag paragraphs likely to be exempt under the FOIA provision that shields an agency\u2019s \u201cdeliberative process.\u201d Software can also help spot pieces of sensitive information, such as Social Security numbers, and extract text from scanned documents or archived video through optical character recognition and automated transcription.<\/p>\n<p class data-block=\"sciam\/paragraph\">AI can also surface files relevant to a particular question in a sprawling archive, including documents a simple keyword search would miss. As Baron points out, the same techniques are already used for electronic discovery in litigation, when vast sets of corporate files, e-mails and other records often must be searched for material bearing on a lawsuit.<\/p>\n<p class data-block=\"sciam\/paragraph\">Still, challenges remain, says Jaillant, who is <a href=\"https:\/\/lustre-network.net\/about\/\">leading an international project <\/a>on <a href=\"https:\/\/www.scientificamerican.com\/article\/ai-slop-is-spurring-record-requests-for-imaginary-journals\/\">AI\u2019s applications to government records<\/a>. One is a shortage of publicly available e-mail data to <a href=\"https:\/\/www.scientificamerican.com\/article\/your-personal-information-is-probably-being-used-to-train-generative-ai-models\/\">train AI<\/a> to handle messages of various types and origins. Partly because of privacy concerns, researchers still often lean on a now decades-old set of messages that government investigators <a href=\"https:\/\/en.wikipedia.org\/wiki\/Enron_Corpus\">obtained from Enron<\/a>, Jaillant says.<\/p>\n<p class data-block=\"sciam\/paragraph\">And even as AI gets better at parsing archival material, it is unlikely to relieve human researchers of the need to read the relevant documents themselves. \u201cIt\u2019s still important for a human user to go back to the documents and be able to read individual e-mails just to understand the context,\u201d she says.<\/p>\n<p class data-block=\"sciam\/paragraph\">All of that assumes the records survive long enough to be read\u2014which is precisely what the fight in Washington has put in doubt. Archivists, and the software they depend on, are working to make sure they do, before the records of today\u2019s decisions become trapped in dead formats or erased from message threads without the public ever getting the chance to see them.<\/p>\n<\/div>\n<p> Lawanda Wiers<br \/><a href=\"https:\/\/www.scientificamerican.com\/article\/why-digital-government-records-are-so-hard-to-preserve\/\" class=\"button purchase\" rel=\"nofollow noopener\" target=\"_blank\">Read More<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>In May a federal judge ordered White House staff to comply with the Presidential Records Act, the 1978 law that makes a president\u2019s official records public property and governs their preservation and eventual release. A month earlier the Justice Department had argued that the law exceeds Congress\u2019s constitutional authority. The American Historical Association and the<\/p>\n","protected":false},"author":1,"featured_media":922568,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[22067,284],"tags":[6541,5170],"class_list":["post-922567","post","type-post","status-publish","format-standard","has-post-thumbnail","category-digital","category-government","tag-digital","tag-government"],"aioseo_notices":[],"_links":{"self":[{"href":"https:\/\/newsycanuse.com\/index.php\/wp-json\/wp\/v2\/posts\/922567","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/newsycanuse.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/newsycanuse.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/newsycanuse.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/newsycanuse.com\/index.php\/wp-json\/wp\/v2\/comments?post=922567"}],"version-history":[{"count":0,"href":"https:\/\/newsycanuse.com\/index.php\/wp-json\/wp\/v2\/posts\/922567\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/newsycanuse.com\/index.php\/wp-json\/wp\/v2\/media\/922568"}],"wp:attachment":[{"href":"https:\/\/newsycanuse.com\/index.php\/wp-json\/wp\/v2\/media?parent=922567"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/newsycanuse.com\/index.php\/wp-json\/wp\/v2\/categories?post=922567"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/newsycanuse.com\/index.php\/wp-json\/wp\/v2\/tags?post=922567"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}