Attribution is an inference, not a lookup
Attribution is the most consequential thing an expert does in a domain matter, and the honest description of it is this: it is an inference built from converging records, not a fact read off a register. No public database will tell you who stood behind a domain name. Since 2018 most public registration records do not even purport to.
What is available instead is a set of artifacts, each consistent with a particular operator and none of which identifies anybody on its own: an organization field nobody remembered to clear, a stable contact address that recurs across a portfolio, an archived page carrying an analytics identifier, a name server shared with other domains, a certificate naming several hostnames at once. The opinion is about how tightly those artifacts converge, and how plausible it is that they converge for a dull reason instead.
I say that in the report, in those terms, because the alternative is worse. An expert who presents attribution as a retrieved fact has written opposing counsel's cross-examination for him: one innocent explanation for one link, and the whole opinion reads as overstated. An expert who has already graded each link, named the competing explanations, and stated what would change the conclusion is describing his reasoning rather than defending a claim, which is far harder to take apart.
What the public record still shows after redaction
A registrant is the person or organization a domain name is registered to. Redaction is the removal of a field's value from the published registration data, with a marker left behind showing that something was removed. Under ICANN's Registration Data Policy, effective August 21, 2025, where redaction applies a provider "MUST NOT include the value of the data element in the RDDS output and MUST indicate that the value is redacted" (ICANN Registration Data Policy).
Three fields survive that treatment. Registrant organization remains publishable with the registrant's consent, and country and state or province remain published. In a live lookup, country and state are frequently the only identity signal left. An organization field that was never cleared is, in my experience, one of the most productive artifacts in this work, because it was populated once by someone who was not thinking about litigation.
The contact channel is the other survivor. RDAP — the Registration Data Access Protocol, the structured JSON successor to WHOIS — handles the registrant email by replacement: either a syntactically valid substitute address or a contact-uri pointing at a web form, not both. Neither identifies anyone. But a replacement value that stays stable across a portfolio is a correlation key, and correlation is what a portfolio is built from.
Why the pre-2018 record is often the whole exercise
Redaction was prospective. It changed what registries and registrars publish going forward; it did not reach back and scrub observations that third parties had already captured. Records collected before May 25, 2018 — the effective date of ICANN's Temporary Specification for gTLD Registration Data — routinely carry a full name, a postal address, a telephone number and an email address, in plain text.
That date is the practical dividing line in attribution work, and the first question I ask about any disputed name is which side of it the domain's history sits on. A name registered in 2011 and quietly renewed may have a decade of open records behind it. A name registered in 2021 may never have had a public identity record at all.
The second question is about siblings. Attribution very often turns not on the disputed domain but on another name in the same portfolio — one that was registered earlier, or was allowed to lapse, or was simply never tidied up, and whose pre-2018 record names somebody. Linking the disputed name to that sibling is then the analytical work, and the sibling supplies the identity. This is why the portfolio, and not the single domain, is the unit of analysis.
Linking a portfolio through shared technical identifiers
Websites carry identifiers so that analytics and advertising platforms can attribute activity to an account. Legacy Universal Analytics tags took the form UA-123456789-1, where the middle digits identify the account and the trailing digit identifies a property within it. Two sites carrying the same account number were reporting to the same analytics account. That is a link between operators, not between registrants, but it is a link with a documented basis.
The method is measurable rather than folkloric. In a peer-reviewed study of roughly 145,000 malicious URLs per day over two weeks, researchers extracted 9,395 unique analytics identifiers and used them to surface about 11,000 additional live domains; 8,182 domains shared 7,945 Google Analytics identifiers, with a mean campaign size of 7.6 domains and a largest campaign of 480 (Starov et al., WWW 2018). Comparable open-source research published in 2015 and again in 2024 established the technique for investigative use.
Two practical points decide whether it works. First, operators strip these tags once a dispute is live, so the extraction is run against archived copies of the page source rather than the live site. Second, the same researchers had to discard any identifier appearing on more than 500 domains, because at that scale the tag belongs to a hosting provider's error pages or a template vendor rather than to a campaign.
Independence is what makes a matrix mean anything
The output of this work is a matrix. Each candidate link is a row: shared analytics account, shared name server, common registrant email replacement value, common postal address in a pre-2018 record, common payment identifier disclosed by a registrar. Each row carries its source, the observation date, the hash of the captured artifact, and a stated strength.
The column that does the real work is independence. Five links that all derive from a single hosting account are one link, not five, and a matrix presenting them as five is inflated in a way a competent opposing expert will find. So each row is assessed for whether it could have arisen from the same underlying fact as another. Shared hosting, shared name servers and a shared certificate are frequently one observation wearing three hats, because one control panel produced all three.
Conversely, genuinely independent links multiply. A pre-2018 postal address, an analytics account in an archived capture from another year, and a payment instrument produced by a registrar are three kinds of record, created by three systems, for three purposes. When those converge the inference is strong — and I still call it an inference.
What attribution evidence cannot establish
It cannot establish identity from public records alone. For a post-2018 registration behind a privacy or proxy service — one that appears in the public record in place of the actual registrant — the customer identity is held by the provider, and there may be no historical public record to recover because none ever existed.
It cannot treat a shared identifier as identification. Analytics linkage produces false positives at a rate high enough that the researchers who quantified the technique had to exclude any tag appearing on more than 500 domains to strip out benign shared services. A single IP address can host thousands of unrelated sites. Neither is evidence of a common operator standing alone.
The technique is also decaying. Google's move from Universal Analytics to GA4 replaced the old codes with less uniform identifiers that yield far less, so the method is strongest for the historical period and weakest for recently built sites.
Archive coverage is incomplete: a page can be absent because crawlers never knew it existed, because robots.txt blocked collection, or because the site owner requested exclusion. An archive gap is not evidence that a page did not exist.
Finally, registration data was never identity-verified. Registrar accuracy obligations test whether a contact point works — whether an email or a phone number reaches somebody — not whether the named person exists or is the person named. A complete, unredacted, pre-2018 record can be entirely fictitious, and sometimes is.
Reaching the record that actually names someone
The record that names a human is the registrar's account file: the account holder's name and address, the payment instrument, the correspondence, and where retained, the login and session history. That file is not public and is not reachable by analysis. It is reached by a channel, and there are three worth knowing.
- Legal process served on the registrar. Every accredited registrar publishes its own requirements, the courts whose process it accepts, and its practice on notifying the customer. Which process is available, and whether it reaches a particular registrar, is a question for counsel.
- ICANN's Registration Data Request Service. A centralized channel for requesting nonpublic gTLD registration data from participating registrars. Participation is voluntary and most requests are refused.
- Registrar verification in a filed UDRP. Once a complaint is filed, the provider requests verification, and the registrar supplies the full registration data — which is how privacy and proxy customer details commonly surface.
One channel is not available at all: registrar data escrow. Registrars deposit registration data, including the identity behind privacy and proxy services, under their ICANN agreements — but escrow is an ICANN-facing contingency mechanism, not a discovery route.
How the opinion is written, and how it holds up
An attribution section of a report has four parts. The matrix, with every row sourced, dated and hashed. A narrative explaining what each link is and how it was obtained. An assessment of strength, link by link and then overall, with the independence analysis shown rather than asserted. And an explicit statement of what the evidence does not establish, written before anyone asks.
The underlying records — a registrar's certified account file, an archived capture, a certificate log entry — are documents, and in United States federal practice their authentication is generally approached through the certification routes for records of a regularly conducted activity and for records generated by an electronic process (FRE 902). That is US-specific, and whether and how any rule applies in a given matter is for counsel. The attribution opinion built on top of those records is something different in kind: it is expert opinion, not a self-authenticating record, and it is examined as opinion.
Which is the point of stating confidence honestly. I have testified in domain-related legal cases and provided expert witness reports in others, and the reports that hold up best are the ones that already contain the concession. The current record of engagements and credentials is maintained at hartzer.com.