Domain name evidence, forensics and litigation support
Abstract horizontal band illustration representing Website Archive Analysis

EvidenceDocumentary

Website Archive Analysis

Produces
Preserved, timestamped captures of what a site displayed
Sources
Internet Archive Wayback Machine and other web archives
How it is obtained
Publicly retrievable; site operator records need legal process
Authority
Internet Archive published policies; FRE 901, 902(13), 902(14)

What an archived capture proves, what it cannot, and how to preserve it before it disappears

What an archived capture proves

An archived capture proves that a crawler retrieved specified content from a URL at a stated time. That is a narrower proposition than it first appears, and it is the whole of what the artifact supports. It does not establish when the content was first published, how long it stayed up, who authored it, or that any human visitor ever saw it.

Understood that way, archive evidence carries real weight. In domain matters it is often the only surviving record of what a site displayed during the relevant period: the parked page carrying competitor advertising, the offer to sell, the mark used in a heading, the abrupt content change in the week a demand letter arrived.

Its custody sits entirely outside the domain name system. The Internet Archive is a US non-profit, not a registry or registrar, so a capture's provenance has to be established on its own terms and cannot be inferred from anything about the domain itself. That separation is a feature: the archive had no stake in the dispute when it made the capture.

How captures come to exist

Captures come from crawls. The Internet Archive states that the Wayback Machine gathers content from its own crawls and from crawls contributed by others, and that its crawlers tend to find sites that are well linked from other sites (Internet Archive help). There is a stated processing delay of three to ten hours between the time a site is crawled and the time the capture appears.

Each capture is addressed by a fourteen-digit timestamp in YYYYMMDDhhmmss form, recorded in GMT — so 20080414172354 is April 14, 2008 at 17:23:54 GMT. That timestamp is part of the capture's URL, which is why the full URL belongs in the exhibit rather than a bare screenshot. Converting it silently into a local time zone in a report is a small error that produces a large argument.

The Internet Archive began archiving the web in 1996, which sets the outer boundary of what any Wayback analysis can reach.

A rendered capture is assembled from parts

What a browser displays at a Wayback URL is not a single stored document. The HTML is one capture; every embedded image, stylesheet and script is a separate capture, potentially made on a different date. The archive assembles them at view time.

The consequence is measurable. Research presented at Hypertext 2015 by Ainsworth, Nelson and Van de Sompel found that only about 17.9 percent — roughly one in five — of composite archived pages were both temporally coherent and completely intact, because embedded resources are frequently archived at different datetimes than the main page. A rendered Wayback page can therefore show a combination of elements that never coexisted at any single moment.

In practice this means checking the capture dates of the embedded resources before asserting that a page "looked like this" on a date. Where the visual appearance of the page is the point at issue — as it often is in confusion and diversion questions — that check is not optional.

Pull the whole index, not the favorable captures

The Archive publishes a CDX index: the list of every capture it holds for a URL, with dates and HTTP status codes. Retrieving it costs nothing and changes the character of the exhibit.

Put the entire index on the record, including the dates that do not help. It shows the court or panel the full sampling pattern rather than a curated selection, it makes the gaps explicit instead of leaving them to be discovered, and it forecloses the suggestion that captures were cherry-picked. It also frequently reveals the more interesting fact: not what a page said, but that crawling stopped, or started, or suddenly intensified around a particular date.

The status codes matter too. A capture recording a 404 or a redirect is evidence about the state of the site at that time, and those entries are routinely omitted from analyses that look only at pages that rendered.

What archived captures cannot establish

Three limitations decide how much weight a capture carries, and stating them yourself is cheaper than having them drawn out of you.

Crawl frequency is irregular. Crawlers tend to find well-linked sites; orphan pages with no inbound links, password-protected areas, JavaScript-driven content and anything excluded by robots.txt may never be captured. A domain with two captures a year apart supports nothing about the eleven intervening months, and absence of captures is not evidence that a site did not exist.

Captures can disappear retroactively. The Internet Archive publicly described the mechanism in April 2017: when a domain changes hands through expiration, squatting or hacking, a new owner's restrictive robots.txt could cause previously accessible historical snapshots to be withheld, and the Archive reported receiving complaints about such "disappeared" sites almost daily (Internet Archive, April 17, 2017). The evidentiary consequence is blunt: a capture visible today may not be visible later. Preserve it now rather than putting a live link in an exhibit and hoping.

A capture is what the crawler received. Content served conditionally by geography, login state, split testing or user-agent detection was never captured as a visitor would have seen it. And a broken image in a rendering may mean the image is unavailable on the Archive's servers, not that the live page lacked it.

Corroborating the archive

An archive capture is stronger in company. Where the Wayback record is thin, other web archives, national library collections and subscription archiving services may hold captures of the same URL. Contemporaneous press coverage, search-engine cached snapshots and the parties' own marketing materials can independently place content at a date.

Two other classes of evidence corroborate that a site was live when a capture says it was: DNS resolution history for the same window, and certificate records showing that a certificate covering the name existed then. None of these prove content, but together they make a capture much harder to dismiss as an artifact.

The strongest corroboration is held by the site operator or its host: the original files, the server access logs and the content management system's revision history. That is the only source showing what was served to real visitors rather than to a crawler, and it is legal-process territory with a short retention window. If what a page showed to actual users is genuinely at issue, those records are the ones to preserve first.

Availability is not guaranteed

Evidence that exists only as a live link to a third party's website is exposed to that third party's continuity. In October 2024 the Internet Archive suffered cyberattacks that took services offline: the Wayback Machine resumed on October 13, Archive-It was restored on October 17, and archive.org returned in provisional read-only availability on October 21, with uploading and other services unavailable.

The practical implications are two. A capture window that depends on on-demand saving is not guaranteed to be open when it is needed, so live captures of currently important pages should be made early rather than at filing. And every archive exhibit should exist as a preserved local copy with its own hash, not as a URL that a reader is expected to visit. An exhibit that cannot be opened during a hearing is not an exhibit.

Authenticating archived captures (US federal rules)

This describes United States federal rules only; other jurisdictions and US state courts apply their own, and application is a matter for counsel.

Under Rule 901(a) the proponent must produce evidence sufficient to support a finding that the item is what the proponent claims it is. The relevant illustrations are 901(b)(1), testimony of a witness with knowledge — the analyst who retrieved and preserved the capture — and 901(b)(9), evidence describing a process or system and showing that it produces an accurate result, which is where the archive's own published description of its crawling and timestamping process fits.

Rules 902(13) and 902(14) provide for self-authentication of records generated by an electronic process and of data copied from a file and authenticated by a process of digital identification (FRE 902). Hash verification is the ordinary form of that digital identification, which is the practical reason to hash every preserved file at the moment of collection. Both routes require a qualified person's certification and the advance notice Rule 902(11) requires.

What goes in the exhibit

For each capture relied on, I preserve and produce: the complete Wayback URL including its fourteen-digit GMT timestamp; the archived HTTP response, including status code and content type; the rendered page as both PDF and screenshot; the capture dates of the embedded resources where appearance is at issue; and a cryptographic hash of every preserved file, recorded at the time of collection.

Alongside the individual exhibits sits the CDX index for the URL, so the full set of known capture dates — and the gaps — is on the record. The report states the archive's documented limitations in its own section rather than waiting to be asked about them, because an expert who names the limits of his evidence before opposing counsel does is more useful to the retaining attorney, not less.

Frequently Asked Questions

Is a Wayback Machine capture admissible evidence?

Admissibility is for the court, and it turns on authentication, hearsay and relevance analysis that counsel conducts. What an expert can do is put the capture in the best possible position: preserve the full capture URL with its fourteen-digit GMT timestamp, the archived HTTP response, rendered copies and a hash of each file, and be able to describe the retrieval process and the archive's own published account of how captures are made. In US federal practice the authentication question is whether there is evidence sufficient to support a finding that the item is what it is claimed to be.

What if there are no captures for the dates I need?

Gaps are common and are not evidence that a site was inactive. Crawlers tend to find well-linked pages, so lightly linked sites, orphan pages, password-protected sections and JavaScript-driven content are systematically under-captured. The response is to check other web archives, look for search-engine cached copies and contemporaneous press, and corroborate the domain's activity from DNS resolution and certificate records for the same period. Where the content itself is essential, the site operator's own files, server logs and content management system revision history are the records to pursue through counsel.

Can archived pages be removed after the fact?

Yes, and this is the practical reason to preserve rather than link. The Internet Archive has publicly described how a new domain owner's restrictive robots.txt could cause previously accessible historical snapshots to be withheld, a problem it said generated complaints almost daily, particularly after domains changed hands through expiration or hijacking. Exclusion requests are also possible. A capture visible today may not be visible when an exhibit is filed or when a hearing occurs, so every capture relied on should be preserved locally with a hash at the moment it is found.

Does a capture show what visitors actually saw?

No. It shows what the crawler received. Sites that serve content conditionally — by geography, by logged-in state, by split testing, or by detecting the requesting software — may have shown something different to a human visitor. A rendered capture can also assemble elements archived on different dates, so the page as displayed may be a combination that never existed at one moment. Where what real users saw is the issue, the operator's server logs and content management records are the only sources that answer it, and they are usually short-lived.

Why does the capture timestamp matter so much?

Because it is the only thing tying the content to a date, and because it is recorded in GMT while reports, filings and testimony are usually framed in local time. The fourteen-digit timestamp forms part of the capture URL, so quoting the URL preserves the date in its original form and lets anyone reproduce the retrieval. Silently converting it into a local time zone introduces an unexplained discrepancy between the exhibit and the narrative, which is exactly the sort of small inconsistency that consumes an afternoon of cross-examination.

Should captures of the current site be made now?

Yes, and early. Present content is observable at no cost and stops being observable the moment it is edited or the domain moves. On-demand archiving depends on a third-party service that has been unavailable for days at a time, so relying on it at the moment of filing is a risk with no upside. Making dated, hashed captures of the pages at issue at the outset of a matter costs very little and creates a record that does not depend on anyone else's crawl schedule or continued operation.
Keep reading

The guides put the pieces in order

An entry covers one kind of work and the record it produces. A guide runs the sequence: when an expert is retained, what is preserved first, what has to be authenticated, and what the report has to carry.

A reference, not an intake page. This site describes what a domain name expert witness does and what the domain record can be made to show. It is not legal advice, nothing on it creates any relationship, and no engagement is taken through this website. The current record of credentials is at hartzer.com.

Top