This site is a record, not a business. BillHartzer.net publishes Bill Hartzer’s professional history and practice. It sells nothing, quotes nothing and takes no engagements — consulting and expert witness inquiries go to Hartzer Consulting.

BillHartzer.net logo mark — the professional record of Bill HartzerBillHartzer.netThe professional record of Bill Hartzer
Practice area
Engagement: TestimonyExpert Witness and InvestigationsLegal and Investigative

Web Analytics Expert Witness

Analytics data read for what it actually measures, before anyone builds a damages number on top of it

Abstract vertical bar illustration representing Web Analytics Expert Witness

Why analytics ends up in litigation

Analytics data shows up whenever a dispute requires a number: lost traffic in a business claim, conversion performance in an agency dispute, delivery and quality in an advertising matter, or audience size in a valuation.

It gets used this way because it is the only quantitative record most businesses keep about their website, and because it produces clean-looking charts. That is precisely the risk. Analytics was built to support marketing decisions, not to serve as evidence. It was built to be directionally useful and cheap to collect, with undocumented accuracy tradeoffs at every layer, because nobody expected a lawyer to read it.

My role is to establish what a given dataset actually measured during the period at issue, whether the configuration was consistent across that period, and how far the numbers can be pushed before they stop supporting the conclusion resting on them. Sometimes that strengthens a claim considerably. Often it means telling the retaining attorney that the comparison at the center of the case is not a valid comparison, which is far better heard early.

How the data is actually produced

To evaluate an analytics number you have to know what had to succeed for it to exist. A script loads in the visitor's browser, executes, and sends a message describing the interaction to a collection endpoint, where it is processed and aggregated before anyone sees a report.

Every stage drops data:

  • Blocking. Content blockers, privacy browsers and network-level filtering prevent the script from loading at all. That loss is not random - it correlates with audience technical sophistication, so some sites lose far more than others.
  • Consent. Where a consent banner governs collection, declined or ignored banners mean no data. A change to banner design or default behavior can move reported traffic sharply while actual traffic is unchanged.
  • Implementation faults. Missing tags on some templates, duplicated tags inflating counts, tags firing before content loads, single-page applications that never signal a page change, and redirects that strip campaign parameters.
  • Cross-domain breakage. A visitor moving between domains without proper linkage becomes two users, and the second one arrives as a referral from the first.
  • Cookie lifetime. Browser restrictions on storage mean a returning visitor is frequently counted as new.

The conclusion follows directly: analytics is a sample of behavior collected under conditions that changed over time. Server logs, by contrast, record requests that actually reached the server. Neither is the truth by itself, and where the two disagree, the disagreement is often the most informative thing in the file.

Platform migrations and the comparison that is not a comparison

The most consequential error I encounter in analytics-based damages models is a before-and-after comparison spanning a change in measurement.

The industry-wide move from a session-based measurement model to an event-based one is the clearest example. The two generations count differently at a definitional level - what starts a session, how a user is identified and estimated, how engagement is defined, how conversions are attributed - so a decline across that boundary can be an artifact of counting rather than a change in business. Reported figures can move substantially in either direction with no change in visitor behavior at all.

Related traps in the same family:

  • Retention windows. Detailed event data is commonly retained for a limited period by default, so a matter filed two years after the conduct may find the underlying detail already purged even though summary reports survive.
  • Legacy access. When a measurement platform is retired, historical data can become permanently unavailable unless someone exported it in time.
  • Modeling. Some current figures are modeled estimates rather than counted observations, filling gaps left by consent and blocking. A modeled number is a projection, and a report should say so.
  • Thresholding. Reports suppress rows with small counts for privacy reasons, so totals will not reconcile with the sum of their parts.

Before running any period comparison, I establish that the measurement itself was constant. If it was not, the comparison has to be rebuilt on a source that was - usually server logs or order records.

Attribution, channels and where numbers get moved

Attribution decides which channel gets credit for a conversion, and it is a configuration choice, not a fact about the world. Move from a last-interaction model to a distributed one and the same conversions redistribute across channels, sometimes dramatically. If a dispute concerns whether a particular channel performed, the attribution setting in force during the period is a threshold question, not a footnote.

Other mechanisms that quietly move numbers:

  • Direct traffic as a catch-all. It absorbs untagged campaigns, e-mail clients, messaging apps, document links and stripped referrers. A large direct segment is a measurement artifact as often as it is brand strength.
  • Campaign tagging errors. Inconsistent capitalization or naming splits one campaign into several, and missing tags dump paid traffic into organic or direct - which matters enormously when the dispute is about which party generated the traffic.
  • Bot and internal traffic filters. These are usually applied going forward only. Turning a filter on midway through a disputed period creates a step change in the data that has nothing to do with visitors.

This is why I ask for the account configuration history, not just the reports. The settings tell you what the numbers mean; the numbers alone do not.

Reconciling traffic data with money

When damages depend on revenue, the accounting system is the better source and analytics is the supporting one. Analytics ecommerce figures routinely diverge from the order database, and the reasons are structural rather than suspicious: refunds and cancellations that never flow back, tax and shipping included in one system and excluded in the other, currency handling, transactions completed by phone or through a channel the site never sees, purchase events blocked on the confirmation page, and duplicate events from a page refresh.

A discrepancy of a few percent is ordinary. Larger gaps point at a specific implementation fault that can usually be identified and explained, and the explanation is often worth more to the case than the number.

The practical rule I apply is straightforward. Use order and accounting records for revenue. Use analytics for behavior and proportion - which pages, which sources, which devices, what changed and when. A damages model that takes analytics revenue at face value while an order database sits available in discovery is the easiest thing in the world for an opposing expert to take apart.

Building an analysis that survives scrutiny

Triangulation is the whole discipline. No single dataset here is authoritative, so I look for the same event across independent systems: analytics, raw server logs, search console data, advertising platform reporting, the order database, and the site's own deployment and content history. Signals that appear in several independent records are findings. Signals that appear in one are hypotheses.

The rest is documentation discipline. Every figure in a report names its source system, account and property, the exact report or query, the date range, the segment and filters, and the time zone configured on the property - which is a routine source of one-day discrepancies that opposing counsel enjoys pointing out. Where a raw data export exists, I work from it rather than from the interface, because it can be re-run.

I also verify the tracking configuration for the period at issue instead of assuming it matched the present. Tag manager version history, archived copies of the page source showing which scripts were present, and account change logs will show what was actually measuring the site on the dates in dispute. That verification step has changed my conclusions often enough that I no longer treat it as optional.

Working with counsel

Ask for access, not artifacts. Read-only access to the analytics property, the advertising accounts and the search console property is worth more than any volume of exported PDFs. Also request the account change history, the tag manager container versions, any raw data export destination the business configured, the order or accounting export covering the same period, and the raw server access logs while they still exist.

Watch the clocks. Server logs rotate in weeks, detailed analytics data ages out under retention settings nobody remembers choosing, and account access is commonly held by a departing agency, lost the day that relationship ends.

Where analytics evidence is central, bring the technical analysis in before the damages model is built rather than after. Rebuilding a model whose foundation turns out to be a measurement artifact is expensive; testing the foundation first is not. Engagements are arranged through Hartzer Consulting, and this site remains a record of the practice rather than a place to retain me.

Frequently asked questions

Can analytics data prove how much revenue a business lost?

It can support the analysis, but it should not be the primary source for revenue. Analytics ecommerce figures routinely diverge from order records because refunds and cancellations do not flow back, tax and shipping are handled differently, some transactions never touch the website, confirmation-page tracking gets blocked, and page refreshes create duplicate purchase events. My approach is to take revenue from the accounting or order system and use analytics for behavior and proportion - which pages, which sources, what changed and when. A model resting on analytics revenue when order data exists is easy to attack.

Our traffic dropped after we changed analytics platforms. Is that a real decline?

Frequently it is not, and this is the single most common flaw I see in analytics-based claims. Successive measurement generations count differently at a definitional level - what begins a session, how a user is identified, how engagement and conversions are defined - so figures can shift substantially with no change in visitor behavior. The same applies to consent banner changes, new bot filters, and tag deployments. Before accepting any period comparison I confirm that the measurement was constant across it, and if it was not, I rebuild the comparison on server logs or order records.

Why do analytics and server logs disagree?

Because they count different things. Analytics requires a script to load and execute in the visitor's browser and to transmit successfully, so content blockers, privacy browsers, declined consent, implementation faults and cookie restrictions all remove data. Server logs record requests that actually reached the server, including automated traffic that analytics filters out. Neither is the complete truth. The disagreement is often the most useful item in the file, because its size and its pattern point directly at what was failing, when it started and which segment of the audience it affected.

What analytics evidence should be preserved, and how quickly?

Secure read-only account access first, for analytics, advertising platforms and search console, because access typically lives with an agency or an employee and disappears when that relationship ends. Then capture the account change history, tag manager container versions, and any raw data export destination the business configured. Request raw server access logs immediately, since rotation is often measured in days. Detailed analytics data also expires under retention settings that were never deliberately chosen. Exported PDFs are a poor substitute for access, because a static report cannot be re-segmented when the question changes.

Can you tell whether an agency's reported results were real?

Usually, yes, and the method is reconstruction rather than argument. I rebuild the reported figures from the underlying accounts and check whether the report's numbers can be reproduced from the same property, date range, segment and filters. Common findings include reporting on a filtered view that excluded unflattering traffic, comparisons spanning a tracking change, campaign tagging that reassigned paid traffic to organic, and conversion definitions that counted events no reasonable client would call a conversion. Sometimes the reconstruction confirms the reporting was accurate, and that is a finding worth having early.
Top