What an internet investigation is for
Not every online problem starts as a lawsuit. More often somebody needs to know what is actually happening before deciding whether there is anything worth pursuing: who is behind a site attacking a business, whether a group of storefronts selling counterfeits is one operator or twelve, whether a departing employee stood up a competing operation while still employed, whether a domain being acquired carries a history that will cause damage, or whether a review campaign is organic or manufactured.
That is a project with a defined scope and a defined deliverable - a written report with exhibits - rather than an ongoing advisory relationship. The scope should be written down before work starts, because online investigations expand naturally and an unbounded one gets expensive without getting more useful.
The organizing principle is that an investigation which might later become evidence has to be conducted to that standard from the first hour. Collection method, capture quality and documentation cannot be retrofitted. I would rather spend the extra time on preservation at the start than explain later why the most important exhibit is a phone photograph of a monitor.
Preserve before you look
The first mistake in most investigations is looking around before capturing anything. It costs evidence in two ways.
The obvious one is that content disappears. Sites get taken down, posts get deleted, listings are removed, and registration records change - frequently within hours of the subject sensing attention. Anything not captured at the moment it is found may not exist tomorrow.
The less obvious one is that visiting a site leaves a record in the subject's own logs. Address, time, referring page and client string all land in a file the other party controls. An investigation run carelessly from a corporate network announces itself, and an operator who notices tends to respond by cleaning up. So I plan the collection sequence deliberately: capture the widely-available public material first, work outward from records that do not touch the subject's infrastructure, and treat direct interaction as a decision rather than a reflex.
Contact through accounts connected to the client is off the table without explicit instruction from counsel, and even then the approach is counsel's call rather than mine. Investigation is not the same activity as engagement, and mixing them contaminates both.
Mapping the infrastructure behind an operation
Online operations that present as unrelated are usually built by people reusing their own components. The work is finding the reuse. The signals I look at, roughly in descending order of strength:
- Reused analytics and advertising identifiers. The strongest signal available, because a human being had to paste the same identifier into multiple properties. Historical page captures often preserve identifiers that were later removed.
- Registration records. Registrant e-mail addresses and organization names surfaced through reverse lookups, and historical records from before contact data was redacted.
- Certificate records. Public certificate transparency logs record every issued certificate, and a certificate covering several names ties them together at a moment in time. These logs also reveal subdomains that were never linked publicly.
- DNS history. Nameserver and address changes over time, which frequently show a cluster of properties moving together on the same day.
- Content and template fingerprints. Identical boilerplate, matching image assets, shared favicons, the same content management fingerprints, recycled phone numbers and addresses.
- Hosting adjacency. The weakest of these. Large providers place thousands of unrelated sites behind one address, so co-location is a lead to be checked, never a finding to be reported.
Assembled properly, a cluster built from several independent signals is persuasive. Assembled carelessly from shared hosting alone, it is the kind of exhibit that gets an entire report discounted.
Domain history and background checks
A large share of this work concerns domain names, either because a name is the subject of a dispute or because someone is about to buy one and wants to know what comes attached.
A domain history examination reconstructs prior ownership from registration records and lifecycle events, prior content from archived captures, and prior reputation from the traces that earlier use leaves behind - an inherited link profile built for a different purpose, blocklist history, spam or malware association, adult or gambling use, prior brand association that still generates traffic and confusion, and evidence of past search penalties. Names that dropped and were re-registered often changed character completely, and the creation date will not tell you that.
This grew out of an algorithm and scoring process I developed in 2013 for performing a background check on a domain name, which I describe as patent-pending, and which I have run commercially since through several ventures. Stolen domain recovery came out of the same practice - by my own count, which is my figure and not an audited one, I have helped recover more than 500 stolen domain names. That work is where I learned to read registrar and registry records at the level a contested matter requires.
Anonymous actors and the honest limits of attribution
The most common request is also the hardest to satisfy: identify the person behind an anonymous site, review campaign or account.
What is achievable through open sources is usually a cluster - a set of properties, accounts and records that plainly belong together, sometimes with an operational error that names someone. Those errors are real and they happen: a registration made before privacy was enabled, an old cached record, an identifier reused on a personal project, a reused profile photograph, a business filing matching a domain contact, a support address that resolves to a named individual.
What is frequently not achievable is closing the final gap without records held by third parties. When that is where the investigation lands, the most useful thing I can produce is not a guess. It is a precise specification of what to subpoena and from whom: the registrar records that would identify the account holder, the platform records that would tie the account to a device or payment instrument, the hosting provider records that would show who provisioned the server, and the date ranges and identifiers each request must name to be answerable.
Naming somebody because the client expects a name is the fastest way to be wrong in public, and it is not something I will do.
The report and its confidence levels
The deliverable is a written report built the way an expert report is built, because it may become one.
Its structure: the scope as agreed and any limits on it; the method, including what was searched, which sources were consulted and what was deliberately not done; findings; exhibits; and a closing section on what could not be determined and what would be required to determine it.
The part that matters most is that every finding carries a confidence level, and the levels are defined in the report itself:
- Observed - directly recorded in a preserved source, with the exhibit cited.
- Strongly supported - multiple independent signals point the same way and the alternatives were tested.
- Indicated - consistent with the evidence, plausible alternatives remain open.
- Unsupported - asserted by someone, not established by anything I found.
Exhibits are collected with the response data and hashed at collection, with a log recording who collected what, when, in which time zone and with which tool. A report that grades itself is more useful to counsel than one that reads as uniformly confident, because it tells you which findings you can build on and which ones need more work before anyone relies on them in a filing.
What I will not do, and how the work is arranged
Boundaries are part of the method, so they belong in the description of it. I do not access accounts or systems I am not authorized to access. I do not use pretexting or false identities to obtain information from people. I do not purchase breached or stolen data. I do not contact parties I have been told are represented by counsel. And I do not deliver an identification that the evidence does not support, whatever the client is hoping for.
These are not just ethical positions. Evidence gathered improperly is worse than no evidence, because it can taint the material around it and turn the investigation itself into the story - which is a poor position for the party that commissioned it.
Investigations run as fixed-scope projects with a written deliverable, and where they touch domain history and stolen name recovery they connect to work I run through DNAccess. Engagements are arranged through Hartzer Consulting. This site is my professional record and does not take work directly.