Source of Truth in a Digital Age: Link Rot Explained (2026)

Quick summary: Roughly half of the URLs cited in published U.S. Supreme Court opinions are dead. A 2014 Harvard Law Review study (Zittrain, Albert, Lessig) found 49.9% of SCOTUS-opinion links and 70%+ of links in Harvard law journals had suffered link rot. A 2024 Pew Research analysis found ~25% of all webpages that existed between 2013 and 2023 are gone. The average webpage is deleted or significantly changed in about 100 days. Wikipedia gets silently re-edited. AI models hallucinate citations. The Internet Archive — the closest thing the web has to a permanent record — was DDoS-attacked offline for weeks in late 2024 and had 31 million records breached. This post is about why some things still belong on paper, what to physically archive, and why source-of-truth is now an active responsibility, not a passive default. Updated 2026-05-15.

A historian writing a book in 2025 about an event that happened in 2010 sat down to gather her sources. The blog post a key witness had written had moved domains; the original URL gave a 404. The local newspaper that had run the most thorough early coverage had been bought by a hedge fund, paywalled, then quietly purged all articles older than five years. The witness’s Facebook post — the one that had circulated widely at the time — was now visible only to the original author’s friends because Facebook had silently changed default sharing settings. The Wikipedia article had been edited 87 times in the intervening years and now told a meaningfully different version of the story. The historian still had the book her aunt had given her as a child: a published bound volume that had documented the same event, immutable on her shelf since 2012. That book ended up being her most reliable source.

This is not an abstract problem. The web is decaying faster than most people realize. The institutions that traditionally kept the record — newspapers, libraries, courts, government agencies — have largely outsourced their archives to digital systems that are far more fragile than the paper they replaced. AI tools, which are now most people’s first interface to the world’s information, are trained on this decaying material, cite it confidently, and sometimes hallucinate citations that never existed at all. If a family or institution wants to know what happened, what was true, or what was decided — in many cases, the only durable answer is a physical copy. This post is about why, how much, and what to do.

Get Smarter About AI Every Morning

Free daily newsletter — one story, one tool, one tip. Plain English, no jargon.

Free forever. Unsubscribe anytime.

How fast does information actually disappear from the web?

Faster than is comfortable. The single most-cited study on this question is Jonathan Zittrain, Kendra Albert, and Lawrence Lessig’s 2014 Harvard Law Review paper, “Perma: Scoping and Addressing the Problem of Link and Reference Rot in Legal Citations.” The authors examined the URLs cited in published opinions of the United States Supreme Court and in three Harvard legal journals. The findings:

  • 49.9% of the URLs cited in Supreme Court opinions no longer pointed to the originally-cited content.
  • More than 70% of the URLs cited in the Harvard Law Review, the Harvard Journal of Law and Technology, and the Harvard Human Rights Journal had suffered reference rot — either the link was dead or the content at the link had changed substantially.

These are not obscure citations. They are the published opinions of the highest court in the country, citing sources the justices considered authoritative. Half of them have effectively disappeared. The response was Perma.cc, a Harvard-Law-Library project that captures snapshots of cited URLs at publication time so that the citation remains stable even if the original page disappears. Many legal journals now require Perma links. The underlying problem — that the web is not built for permanence — hasn’t been solved; it’s been quarantined to specific corners that are willing to do the archival work.

A 2024 Pew Research Center analysis updated the picture for the general web. Pew sampled webpages that had existed at any point between 2013 and 2023 and checked which were still accessible. Roughly 25% of all webpages from that decade are now gone. The decay is not evenly distributed — older pages and pages on shuttered platforms (Geocities, MySpace, Vine, Google+, large numbers of personal blogs) account for disproportionate losses — but the overall trend is one in four pages lost per decade.

The Internet Archive’s founder Brewster Kahle has cited a more sobering figure for individual pages: the average webpage is deleted or substantially changed within about 100 days of being published. Most of the web you read today won’t exist in the same form in three months.

What about silent edits — pages that still exist but say different things now?

This is the harder problem and the less-discussed one. Link rot at least announces itself with a 404. Silent edits don’t.

Wikipedia is the textbook example. The English Wikipedia article on any actively-contested topic gets edited dozens of times per year. Most edits are improvements; some are reversals. Without consulting the article history (which most readers never do), there is no way to tell whether the version of the article you’re reading is the same as the version your friend cited last year, or whether a contested fact has been swapped out, softened, or sharpened. The Wayback Machine captures snapshots, but only some are preserved, and finding the relevant version requires effort.

Beyond Wikipedia, the silent-edit problem extends to: news articles that get quietly corrected without retraction notices, corporate websites that delete inconvenient old positions, government pages that update without retaining the historical version, product documentation that drifts away from older but still-deployed software, and academic web resources whose authors changed their minds. None of these typically issue a public notice when they change. The version your AI was trained on, the version you saw last year, and the version on the live page today may all differ.

For students and researchers, the practical implication: a citation to a live URL is no longer self-contained evidence. You need either a Wayback Machine snapshot, a Perma link, a downloaded PDF, or — best — a physical printout of the relevant page at the moment you used it as a source.

Where does AI fit into the source-of-truth problem?

Two ways, in opposite directions. The compounding-the-problem direction is that AI models trained on the open web inherit the same decay problem: if their training data included pages that are now gone, the model can cite or paraphrase content that’s effectively unverifiable. Worse, modern LLMs sometimes hallucinate citations — fabricating plausible-looking URLs, paper titles, or court cases that don’t exist at all. Lawyers have been sanctioned by judges for filing briefs containing AI-hallucinated case citations. Students have submitted papers citing made-up academic articles. AI tools that present themselves as research aids can produce dangerously confident misattribution.

The mitigating direction is that AI tools can also speed up verification. Asking an LLM to fetch a URL and tell you what’s there (which the better tools can now do live), asking it to cross-reference a claim across multiple sources, or asking it to find the original primary source rather than secondary commentary, is faster than the same work done manually. The user has to actually do this — and most users don’t. The default behavior is to take what the AI says at face value, which is exactly the wrong response given the citation-quality problem.

This isn’t an AI-only issue. The mainstream press has always had retraction rates and fact-checking failures. The novel feature in 2026 is the volume and confident tone of AI output combined with the fragility of the underlying record. A high-confidence claim from an AI tool, citing a URL, on a topic where the URL no longer exists or has been silently edited, is an information-integrity hazard that did not exist at this scale ten years ago.

Is the Internet Archive going to save us?

Partially. The Internet Archive, founded by Brewster Kahle in 1996, runs the Wayback Machine, which as of October 2025 had archived more than 1 trillion webpages across 99+ petabytes of storage. It is the closest thing the web has to a comprehensive historical record. It is also a single private 501(c)(3) nonprofit with finite resources, ongoing legal pressure from publishers (the 2023 ruling against Internet Archive’s National Emergency Library digital-lending program is still being adjudicated), and increasingly hostile actors.

In late 2024 the Internet Archive endured what was probably the most serious sustained attack on a major web-preservation institution in history. In September 2024 a data breach exposed approximately 31 million records containing user email addresses and hashed passwords. In October 2024 the organization was hit by a sustained DDoS attack that took the Wayback Machine offline for several days; the site returned in read-only mode and remained partially operational until November 2024. The attacks did not destroy the underlying archive, but they demonstrated that even the world’s most important web-history institution is one bad week away from going dark.

Other preservation efforts exist — the Library of Congress’s Web Archive, Common Crawl, perma.cc, the Wikipedia Foundation’s archival commitments, ArchiveTeam’s community-driven efforts — but none of them comes close to the Internet Archive’s coverage. The realistic picture: digital preservation depends on a small number of fragile institutions, any of which could be lost. The institutions deserve support. They also should not be the only safeguard.

What should you actually keep on paper or in physical form?

CategoryWhy physicalWhat specifically
Vital recordsGovernment-issued and irreplaceable if the digital copy failsBirth certificates, marriage certificates, passports, Social Security cards, naturalization papers, military records (DD-214), property deeds, mortgage docs
Financial recordsTax-audit window is up to 7 years; IRS prefers paper proof when challengedLast 7 years of tax returns + supporting docs, major-investment cost-basis records, retirement-account statements, large-purchase receipts
Health recordsProvider portals shut down; insurance records get archived inconsistentlyImmunization records, major-surgery summaries, allergy and medication history, advance directives, organ-donation forms
Family archivesPhoto cloud services have changed pricing or shut down repeatedlyPrinted photo albums of important years, family genealogy, letters of significance, baby books, school records
Foundational referenceReference works that get silently edited online; you want a frozen versionOne good dictionary, one good thesaurus, an atlas, a one-volume world history, classic encyclopedia volumes if you have a child homeschooling
Children’s reading libraryPer the Evans 2010 research, physical books at home predict educational attainment50-500 children’s books at the child’s reading level + reach level
Foundational books for an educationThe canonical “great books” lists; these stay relevant for decadesSt. John’s College reading list, Susan Wise Bauer’s lists, Mortimer Adler’s How to Read a Book, your child’s curriculum’s recommended-reading lists
Legal and estate documentsSame as vital records, but with attorney-prepared specificsWills, trust documents, power of attorney, healthcare proxy, beneficiary designations
Insurance policiesInsurer portals change; printed policy text is the authoritative version in disputesCurrent policy declarations, the actual policy text, claim records

The list is not “everything.” The list is “things you would deeply regret losing, would not be able to reconstruct, or would need to authenticate years from now in front of a skeptical official.” That’s a much narrower scope than “back up the cloud,” and it’s manageable for a typical family in a single banker’s box or fireproof safe.

How should you back up your digital life — practically?

Two principles, both unglamorous.

First, the 3-2-1 rule, adopted from the IT industry. Three copies of important data. Two different media types. One copy off-site. In practice: the live copy on your computer, a local external hard drive that’s plugged in periodically, and a cloud backup or off-site hard drive at a trusted family member’s house. For most households, a $100 external hard drive plus a Backblaze or iDrive subscription ($60-$100/year) covers this entirely.

Second, the annual print ritual. Once a year — pick a date you’ll remember — print physical copies of the year’s most important documents and add them to a binder or filing cabinet. The list above is the starting set. Add to it items specific to your family’s year (school records, medical visits, major correspondence). It takes an afternoon. It is more durable than any cloud service has ever been.

For families with children, add an annual photo book printed via a service like Artifact Uprising, Chatbooks, or a local printer. Photos in the cloud are easy to lose to platform decisions; photos in a bound book on a shelf will be in your grandchildren’s hands.

How does this connect to learning and education?

Three direct connections.

Textbooks remain a legitimate source of truth. A printed textbook from a reputable publisher has been through editorial review, was a snapshot of consensus knowledge at publication, and does not silently change. Web-based curricula are not bad — they’re often better in many ways — but they should not be the only source. Every child should learn what it’s like to look something up in a real textbook or a real dictionary.

Cite primary sources, not aggregators. The web is good at telling you about something; primary sources are what allow you to verify. For an academic question, that means the published paper, not the news article summarizing the paper. For a historical question, that means the actual document, not the Wikipedia version of it. AI tools can help you find primary sources faster than you used to be able to. Teach children to ask AI for primary sources rather than secondary commentary.

Build a small reference library. A family doesn’t need a hundred reference works — they need three. A good dictionary, a good atlas, and a good single-volume encyclopedia. Add one good world history. Your home library replaces approximately zero things you’d otherwise look up online; it replaces something more important — the felt presence of authoritative reference in your child’s environment.

For more on the cognitive case for physical books, read the dedicated post. For the broader argument about AI in learning, read the pillar guide.

Learn Our Proven AI Frameworks

Beginners in AI created 6 branded frameworks to help you master AI: STACK for prompting, BUILD for business, ADAPT for learning, THINK for decisions, CRAFT for content, and CRON for automation.

Frequently asked questions

Is this a doomer take? Should I panic about digital decay?

No. The web is mostly fine for most things. The point is not that digital is bad; the point is that physical and digital have different properties, and a thoughtful approach uses both deliberately. Don’t panic. Don’t stop using the cloud. Do print the things you’d be devastated to lose, and do keep at least a small physical reference library.

What’s the easiest first step for a family that’s done none of this?

One afternoon, one banker’s box. Print or gather: vital records, last 7 years of tax returns, immunization records, current insurance policies, will and estate documents, one printed photo book of the past year. Done. That covers ~80% of the catastrophe scenarios with one afternoon of work.

Are e-books a long-term archive option?

Less than buyers usually think. When you “buy” a Kindle book, Amazon retains the right to remove access to it under various circumstances (this has happened). DRM-locked e-books are accessible only as long as the platform supports the format and your account remains in good standing. For long-term reference, EPUB files stored locally on hardware you own are better; physical books are better still.

Should I trust the Wayback Machine for citation purposes?

Mostly yes, with caveats. The Wayback Machine is the best free option for capturing a moment-in-time snapshot of a webpage. Snapshots are stamped with the date captured. For high-stakes citations (legal, academic), use Perma.cc instead — it’s purpose-built for this and has institutional commitments to long-term preservation. For everyday research, “save to Wayback Machine” before citing is a low-effort high-value habit.

What about cryptocurrency / blockchain-based archives?

Promising in theory, unconvincing in practice. Several blockchain-based archive projects exist (Arweave is the most prominent); they offer technical immutability for stored data. The weaknesses are economic (will the protocol still be funded and running in 30 years?), practical (the data is preserved but the indexes that let you find it may not be), and accessibility (most users don’t know how to use these systems). For now, the Internet Archive, Perma.cc, and physical copies remain the more reliable approaches.

How does this affect homeschoolers specifically?

More than it affects most families. Homeschool transcripts, samples of work, standardized test scores, curriculum records, and any portfolio assessments may need to be produced for college admissions, military enlistment, or scholarships years in the future. The college that asks for a transcript in 2034 will not accept a screenshot of a defunct online platform. Keep printed records. Keep them organized. The full guide is in the homeschool hub.

Sources

You may also like

Two ways to go further

The AI Prompt Library

1,000+ ready-to-use prompts for Claude, ChatGPT, and Gemini. Stop staring at a blank box.

Get it for $39 →

2-Hour Live AI Crash Course

A private, beginner-friendly session across Claude, ChatGPT, Gemini, and the wider landscape.

Book for $125 →

Discover more from Beginners in AI

Subscribe now to keep reading and get access to the full archive.

Continue reading