This site is a record, not a business. BillHartzer.net publishes Bill Hartzer’s professional history and practice. It sells nothing, quotes nothing and takes no engagements — consulting and expert witness inquiries go to Hartzer Consulting.

BillHartzer.net logo mark — the professional record of Bill HartzerBillHartzer.netThe professional record of Bill Hartzer
Practice area
Engagement: ProjectExpert Witness and InvestigationsLegal and Investigative

Internet Investigations

Defined-scope investigation of sites, domains, networks and online actors, delivered as a documented report with exhibits

Abstract spiral coil illustration representing Internet Investigations

What an internet investigation is for

Not every online problem starts as a lawsuit. More often somebody needs to know what is actually happening before deciding whether there is anything worth pursuing: who is behind a site attacking a business, whether a group of storefronts selling counterfeits is one operator or twelve, whether a departing employee stood up a competing operation while still employed, whether a domain being acquired carries a history that will cause damage, or whether a review campaign is organic or manufactured.

That is a project with a defined scope and a defined deliverable - a written report with exhibits - rather than an ongoing advisory relationship. The scope should be written down before work starts, because online investigations expand naturally and an unbounded one gets expensive without getting more useful.

The organizing principle is that an investigation which might later become evidence has to be conducted to that standard from the first hour. Collection method, capture quality and documentation cannot be retrofitted. I would rather spend the extra time on preservation at the start than explain later why the most important exhibit is a phone photograph of a monitor.

Preserve before you look

The first mistake in most investigations is looking around before capturing anything. It costs evidence in two ways.

The obvious one is that content disappears. Sites get taken down, posts get deleted, listings are removed, and registration records change - frequently within hours of the subject sensing attention. Anything not captured at the moment it is found may not exist tomorrow.

The less obvious one is that visiting a site leaves a record in the subject's own logs. Address, time, referring page and client string all land in a file the other party controls. An investigation run carelessly from a corporate network announces itself, and an operator who notices tends to respond by cleaning up. So I plan the collection sequence deliberately: capture the widely-available public material first, work outward from records that do not touch the subject's infrastructure, and treat direct interaction as a decision rather than a reflex.

Contact through accounts connected to the client is off the table without explicit instruction from counsel, and even then the approach is counsel's call rather than mine. Investigation is not the same activity as engagement, and mixing them contaminates both.

Mapping the infrastructure behind an operation

Online operations that present as unrelated are usually built by people reusing their own components. The work is finding the reuse. The signals I look at, roughly in descending order of strength:

  • Reused analytics and advertising identifiers. The strongest signal available, because a human being had to paste the same identifier into multiple properties. Historical page captures often preserve identifiers that were later removed.
  • Registration records. Registrant e-mail addresses and organization names surfaced through reverse lookups, and historical records from before contact data was redacted.
  • Certificate records. Public certificate transparency logs record every issued certificate, and a certificate covering several names ties them together at a moment in time. These logs also reveal subdomains that were never linked publicly.
  • DNS history. Nameserver and address changes over time, which frequently show a cluster of properties moving together on the same day.
  • Content and template fingerprints. Identical boilerplate, matching image assets, shared favicons, the same content management fingerprints, recycled phone numbers and addresses.
  • Hosting adjacency. The weakest of these. Large providers place thousands of unrelated sites behind one address, so co-location is a lead to be checked, never a finding to be reported.

Assembled properly, a cluster built from several independent signals is persuasive. Assembled carelessly from shared hosting alone, it is the kind of exhibit that gets an entire report discounted.

Domain history and background checks

A large share of this work concerns domain names, either because a name is the subject of a dispute or because someone is about to buy one and wants to know what comes attached.

A domain history examination reconstructs prior ownership from registration records and lifecycle events, prior content from archived captures, and prior reputation from the traces that earlier use leaves behind - an inherited link profile built for a different purpose, blocklist history, spam or malware association, adult or gambling use, prior brand association that still generates traffic and confusion, and evidence of past search penalties. Names that dropped and were re-registered often changed character completely, and the creation date will not tell you that.

This grew out of an algorithm and scoring process I developed in 2013 for performing a background check on a domain name, which I describe as patent-pending, and which I have run commercially since through several ventures. Stolen domain recovery came out of the same practice - by my own count, which is my figure and not an audited one, I have helped recover more than 500 stolen domain names. That work is where I learned to read registrar and registry records at the level a contested matter requires.

Anonymous actors and the honest limits of attribution

The most common request is also the hardest to satisfy: identify the person behind an anonymous site, review campaign or account.

What is achievable through open sources is usually a cluster - a set of properties, accounts and records that plainly belong together, sometimes with an operational error that names someone. Those errors are real and they happen: a registration made before privacy was enabled, an old cached record, an identifier reused on a personal project, a reused profile photograph, a business filing matching a domain contact, a support address that resolves to a named individual.

What is frequently not achievable is closing the final gap without records held by third parties. When that is where the investigation lands, the most useful thing I can produce is not a guess. It is a precise specification of what to subpoena and from whom: the registrar records that would identify the account holder, the platform records that would tie the account to a device or payment instrument, the hosting provider records that would show who provisioned the server, and the date ranges and identifiers each request must name to be answerable.

Naming somebody because the client expects a name is the fastest way to be wrong in public, and it is not something I will do.

The report and its confidence levels

The deliverable is a written report built the way an expert report is built, because it may become one.

Its structure: the scope as agreed and any limits on it; the method, including what was searched, which sources were consulted and what was deliberately not done; findings; exhibits; and a closing section on what could not be determined and what would be required to determine it.

The part that matters most is that every finding carries a confidence level, and the levels are defined in the report itself:

  • Observed - directly recorded in a preserved source, with the exhibit cited.
  • Strongly supported - multiple independent signals point the same way and the alternatives were tested.
  • Indicated - consistent with the evidence, plausible alternatives remain open.
  • Unsupported - asserted by someone, not established by anything I found.

Exhibits are collected with the response data and hashed at collection, with a log recording who collected what, when, in which time zone and with which tool. A report that grades itself is more useful to counsel than one that reads as uniformly confident, because it tells you which findings you can build on and which ones need more work before anyone relies on them in a filing.

What I will not do, and how the work is arranged

Boundaries are part of the method, so they belong in the description of it. I do not access accounts or systems I am not authorized to access. I do not use pretexting or false identities to obtain information from people. I do not purchase breached or stolen data. I do not contact parties I have been told are represented by counsel. And I do not deliver an identification that the evidence does not support, whatever the client is hoping for.

These are not just ethical positions. Evidence gathered improperly is worse than no evidence, because it can taint the material around it and turn the investigation itself into the story - which is a poor position for the party that commissioned it.

Investigations run as fixed-scope projects with a written deliverable, and where they touch domain history and stolen name recovery they connect to work I run through DNAccess. Engagements are arranged through Hartzer Consulting. This site is my professional record and does not take work directly.

Frequently asked questions

Can you find out who is behind an anonymous website attacking my business?

Sometimes, and I will tell you early which situation you are in. Open sources often produce a solid cluster of connected properties and accounts, and operators make mistakes that surface a name - a registration made before privacy was enabled, an identifier reused on a personal project, a business filing matching a contact record. When the last gap requires records held by a registrar, a platform or a host, the deliverable becomes a precise specification of what counsel should subpoena, from whom, naming which identifiers and date ranges. I do not supply a name the evidence does not support.

How do you determine whether several websites are run by the same operator?

By looking for reuse that required a deliberate human action. A shared analytics or advertising identifier across supposedly unrelated properties is the strongest signal, and historical page captures frequently preserve identifiers that were later removed. Registration records surfaced through reverse lookups, certificates covering multiple names, coordinated DNS changes, identical templates and recycled boilerplate all add weight. Shared hosting is the weakest signal, since large providers place thousands of unrelated sites behind one address, and I report it as a lead rather than a conclusion. Findings are graded by confidence, not asserted flatly.

What should we do before starting an investigation?

Preserve first and browse second. Capture the material at issue immediately with proper page capture that records the address, the underlying response and a hash, because content disappears within hours once a subject senses attention. Do not have staff poke around the target site from a company network - visits land in the other party's server logs, complete with address and timing, and an operator who notices usually starts cleaning up. Avoid any contact from accounts connected to your business. Then define the question narrowly, because an unbounded investigation gets expensive without getting more useful.

Is an investigation report usable later in litigation?

It is if it was built for that from the start, which is why I build every one that way. Exhibits are collected with their response data rather than as bare screenshots, hashed at the moment of collection, and logged with the collector, time, time zone and tool version. The method section records what was searched and what was deliberately not done. Findings are graded - observed, strongly supported, indicated, unsupported - so a reader can see which ones will bear weight. Collection quality cannot be added afterward, and an uncorroborated screenshot cannot be repaired.

What can an investigation tell us about a domain name before we buy it?

A good deal, and this is where several purchases get called off. The examination reconstructs prior ownership from registration and lifecycle records, prior content from archived captures, and prior reputation from what earlier use left behind - an inherited link profile built for a different purpose, blocklist and malware history, adult or gambling use, prior brand association that still sends confused traffic, and signs of past search penalties. A name that lapsed and was re-registered may have changed character entirely, and its creation date will not reveal that on its own.
Top