Domain name evidence, forensics and litigation support
Abstract scattered dot illustration representing Typosquatting Evidence

EvidenceAnalytical

Typosquatting Evidence

Produces
An enumerated variant table and a linkage exhibit
Sources
Zone data, registration data, passive DNS, CT logs, archives
How it is obtained
Public retrieval; registrant linkage needs legal process
Authority
UDRP 4(b)(ii); 15 U.S.C. 1125(d)(1)(B)(i)(VIII); FRE 702

Showing that a name is a deliberate variant rather than an accident, by enumerating the variant space and locating it inside

The question is deliberateness, not resemblance

Resemblance is easy to assert and nearly worthless on its own. The evidentiary question in a typosquatting matter is narrower and harder: is this name a deliberate variant of a target string, or a coincidence?

Typosquatting is the registration of names a user reaches through an input error. What makes it tractable as evidence is that the error classes are finite, enumerable and machine-generable. That lets an expert say something precise — this string sits inside a small, well-defined neighborhood of the mark, generated by a published process, rather than being one of the effectively unlimited strings available to a registrant.

Two forums make the analysis operative rather than academic. The UDRP treats registration to prevent a mark owner from reflecting its mark in a corresponding domain name as evidence of bad faith, expressly conditioned on a pattern of such conduct. And in the United States, the ACPA directs a court to consider a registrant's registration or acquisition of multiple domain names known to be identical or confusingly similar to marks of others. US-specific: the ACPA and the Federal Rules of Evidence referred to on this page are United States law only.

Both provisions operate on the set. That is why the enumeration exists.

The variant space is finite, and it is generated rather than guessed

The measurement literature works in edit distance — the minimum number of insertions, deletions, substitutions or transpositions needed to turn one string into another — supplemented by a keyboard-adjacency measure capturing letters that sit next to each other on a physical keyboard. Those operations map onto observable variant classes:

  • Omission — a dropped character. Transposition — two adjacent characters swapped. Repetition — a character doubled.
  • Substitution, most often by a physically neighboring key, and insertion of an adjacent key struck alongside the intended one.
  • Hyphenation, vowel swap, pluralization, and dictionary or combination terms appended — the combosquatting pattern, where nothing is misspelled at all.
  • Homoglyph substitution — a visually confusable character, within ASCII or across scripts.
  • Bitsquatting — a single-bit difference in the character's binary representation, catching lookups corrupted in memory.
  • Extension swap — the same second-level label in a different extension.
  • Subdomain insertion — the target string relocated into a label position the target does not control.

That list is not an invention for a particular matter. It is the published permutation set of a widely used open-source generator, dnstwist, whose generator methods carry those names.

Name the generator, the version and the parameters

Using a published, independently maintained generator rather than a hand-picked list is the single decision that most affects whether the exhibit survives. It matters for a reason that has nothing to do with convenience: the other side can re-run it.

A hand-curated list of variants is a selection, and a selection made by the party asserting the pattern. Every name on it invites the question of what was left off. A full permutation run over a named tool at a named version, with the wordlist and parameters stated, is a process someone else can reproduce and check — which is the standard the reliability framing of FRE 702 asks about, and which is simply better practice regardless of forum.

The run produces candidates, not findings. Each candidate is then resolved and recorded: does it resolve now, what address and mail records does it carry, which registrar sponsors it, what is its creation date, and is its registrant field redacted. The negative results stay in the table. A variant that does not exist is a fact about the neighborhood, and removing it inflates every ratio in the exhibit.

Homoglyphs and internationalized names: codepoints, not screenshots

Confusable characters have formal machinery behind them, and a report that uses it is on much firmer ground than one that asserts two strings look alike. The Unicode Consortium's security mechanisms specification, UTS #39, defines a skeleton transformation — normalize, strip default-ignorable characters, map each character to its prototype exemplar, normalize again — and defines two strings as confusable precisely when their skeletons match.

It also classifies single-script, mixed-script and whole-script confusables (a whole Latin word reproduced in another script), defines mixed-script detection by resolved script set, and sets restriction levels running from ASCII-only upward. It points in turn to the registry-side label generation rules mechanism, by which a registry defines a permitted character repertoire and blocks variant labels.

So a homoglyph exhibit reports the codepoints, the script of each character, the skeleton result, and the registry's label-rule posture for that extension. It does not report a screenshot and an assertion. Codepoints survive re-typing, copy-paste and font substitution; an image does not, and an image of two visually identical strings makes the opposite of the intended point.

One variant is weak; the patterned set is the evidence

A single close variant, standing alone, is weak evidence of anything. What carries weight is the set, and specifically four properties of it:

  • Multiple variants of the same target held by one registrant or one account.
  • Registered in a burst — same date, same registrar, sequential registry transaction ordering.
  • A shared monetization identifier across the set. Published measurement work found this concentration directly: in one 2010 study of typo domains displaying advertising, a majority of those observed used one of only five advertising partner identifiers, and roughly four in five of the crawled typo domains monetized through pay-per-click advertising (Moore and Edelman).
  • Shared infrastructure — name servers, address ranges, certificates whose subject alternative names cover several of the variants at once.

Those are the signals that survive redaction, and post-2018 they carry more weight than the registrant-name field does, because that field is usually not there. Each remains an inference. A shared name server can mean a shared hosting provider; a shared address can mean shared hosting. The exhibit states the alternative explanation next to each signal, and treats registrar account records as the confirmation.

The prevalence literature cuts both ways, and the page says so

Published measurement studies establish that patterned typo registration is real, measured and industrial in scale. One 2010 study identified on the order of hundreds of candidate typo domains per popular target site. A 2014 USENIX Security study identified millions of candidate typo names in a single extension's zone and estimated that a substantial fraction of registrations in that extension were true typo domains — while also finding that only a small percentage of lexicographically similar registrations targeted popular domains, so the activity sits overwhelmingly in the long tail. Work on combosquatting, drawing on hundreds of billions of DNS records over roughly six years, reported that most abusive combosquatting domains persist for more than a thousand days.

Now the part an expert has to volunteer. The same literature establishes that close variants of any given string are extremely common — which is precisely the base-rate argument a respondent will make. An expert who cites prevalence research to demonstrate intent has cited evidence that also undercuts the inference. The honest use of these figures is to establish that the phenomenon exists and has measurable structure, not to bridge from one similar name to a conclusion about why it was registered.

What variant analysis cannot establish

It cannot establish intent, and it cannot establish common ownership from public data alone.

  • Base rates defeat naive inference. A large share of any big extension is lexicographically similar to something else, so "this string is one edit from the mark" is on its own an extremely weak signal.
  • Enumeration is generator-dependent. A different tool yields a different variant set, so the generator, version, parameters and wordlist are named or the exhibit is not reproducible.
  • A variant may have an independent meaning. Short strings, dictionary words, acronyms and other parties' marks collide innocently. The expert reports the collision; whether it is innocent is not the expert's call.
  • Redaction breaks registrant linkage. "One registrant holds all of these" is usually an inference from infrastructure until disclosure occurs, and privacy services and bulk registrar accounts mask common ownership the same way.
  • Passive DNS coverage is partial and sensor-dependent; absence of a record is not evidence a name never resolved.
  • Deleted variants vanish from live sources, leaving only archive, certificate log and passive DNS traces — and registration records for dropped names age out first.
  • Country-code variants may have no public registration data at all.

What a retaining attorney should expect

The deliverable is an enumerated variant table — every generated permutation, its operation class, whether it resolves, its address and mail records, its registrar, its creation date and its registrant field state — cross-joined against passive DNS and certificate transparency data so that names which once resolved but no longer do are still captured. Alongside it sits a linkage exhibit: shared name servers, addresses, certificate coverage, and advertising or affiliate identifiers recovered from archived page source. The target's own defensive registrations get their own column rather than sitting unlabeled in the total.

US-specific: a generated variant list is not a record. It is the output of a process, and it comes in under the reliability framing of FRE 702 as amended effective 1 December 2023, which is why the tool, version and parameters go in the report. The underlying registration and DNS records are records, authenticated under FRE 901 with the 902(13) and 902(14) certification routes available; a large table is a candidate for summary presentation under FRE 1006 with the underlying data produced.

Bill Hartzer has testified in domain-related legal cases and has provided expert witness reports in others, and has worked in this field since 1996. The current record is at hartzer.com. What the set means legally is for counsel.

Frequently Asked Questions

How is a typo domain distinguished from a coincidence?

Not by looking at it. The method is to enumerate the full variant space around the target with a published permutation generator, locate the disputed name inside that space by operation class, and then examine whether the registrant holds a patterned set rather than a single name. The signals that matter are multiple variants of one target, registrations clustered in time at one registrar, shared advertising or affiliate identifiers across the set, and shared infrastructure. A single close variant, standing alone, remains weak evidence and the report should say so.

Why does the tool used to generate variants matter?

Because a hand-curated list is a selection made by the party asserting the pattern, and every name left off it is a question waiting to be asked. A full permutation run using a named, independently maintained generator at a stated version, with the wordlist and parameters recorded, is a process the other side can reproduce and check. Reproducibility is the basis on which a technical exhibit is accepted, and in United States practice it is what the reliability framing of the expert-testimony rule is asking about.

How should a homoglyph domain be presented in evidence?

With codepoints, not a screenshot. Give each character's Unicode codepoint and the script it belongs to, the result of the standard confusable-skeleton transformation, and the registry's label generation rule posture for that extension, alongside a rendering that states the environment it was produced in. Two visually identical strings printed in a report look like a typographical mistake; the underlying data is what makes the point and what survives re-typing, copy-paste and font substitution when the exhibit is handled by someone else.

Does prevalence research prove that a registration was deliberate?

No, and an expert who uses it that way has handed the other side an argument. Published measurement studies establish that patterned typo registration exists at industrial scale and has measurable structure. The same studies establish that close variants of any given string are extremely common, which is the base-rate objection a respondent will raise. The honest use of the literature is to describe the phenomenon and its structure, and to state the base-rate limitation in the report rather than waiting for it to arrive in cross-examination.

What about variants that have already been dropped?

They still belong in the analysis, and they are the ones most likely to be missed. A live DNS check finds only what exists today, so the enumeration is cross-joined against passive DNS and certificate transparency records to capture names that once resolved and no longer do, plus archived captures of what they served. Registration records for dropped names also age out first under registrar record-keeping obligations, which is a preservation argument for acting early rather than a reason to leave them out of the table.

Can an expert say one person owns the whole set?

Usually not from public data. Registrant name and email are typically withheld from public registration output, so common ownership is an inference built from infrastructure and account-level signals: shared name servers, address ranges, certificates covering several variants, matching template or favicon fingerprints, shared advertising and analytics identifiers, and clustered registration timestamps. Each signal carries an innocent explanation that belongs beside it in the exhibit. The confirmation is registrar account records, which reach the file through a disclosure channel or legal process.
Keep reading

The guides put the pieces in order

An entry covers one kind of work and the record it produces. A guide runs the sequence: when an expert is retained, what is preserved first, what has to be authenticated, and what the report has to carry.

A reference, not an intake page. This site describes what a domain name expert witness does and what the domain record can be made to show. It is not legal advice, nothing on it creates any relationship, and no engagement is taken through this website. The current record of credentials is at hartzer.com.

Top