A task ticket on a recent engagement read almost exactly like this: “Remove the old pricing-page A/B test, it ended months ago.” The ticket had a date, a name, and a confident tone. It also turned out to be describing a decision that was made and then never carried out. The test was still live, still splitting real visitors, months after everyone involved believed it was over.
The Ticket Said “This Test Ended Months Ago”
A ticket that says a test “ended months ago” is a claim about a document, not a claim about production. Someone decided the test was over. Whether anyone actually went into the testing platform and stopped it is a separate, unverified fact.
This is a B2B SaaS marketing site, anonymized here as with every client pattern on this blog. The task read as routine cleanup: an old pricing-page experiment, superseded by a redesign months earlier, flagged for removal so the codebase and the analytics setup could both be simplified. Nothing about the ticket suggested anything unusual. A reasonable person reading it would assume the only work left was deleting some now-inert code.
The habit that changed the outcome here was simple and almost boring: before touching anything that could affect live traffic or data collection, check the actual running state first, not the description of it. That habit is the entire subject of this post, because the alternative, trusting the ticket, is the default almost everyone reaches for under time pressure.
Why You Check the Live State Before You Touch Anything That Touches Data
A ticket, a wiki page, or someone’s memory can all be wrong about production in the same specific way: they describe a decision, and production reflects whatever was actually implemented, which drifts from the decision the moment nobody circles back to confirm it landed. Checking live state closes that gap. Trusting the description leaves it open.
The gap exists because “the test ended” and “the test was stopped” are two different events, separated by a manual step that is easy to skip. A team runs an experiment, reaches a conclusion, and moves on to the next priority. Stopping the experiment in the testing platform, removing its traffic allocation, and cleaning up the resulting code are three more steps after the decision, and none of them happen automatically just because the team stopped thinking about the test. Industry analysis of this exact pattern in engineering contexts describes abandoned test variants as commonly leaving inactive logic running long after a team has moved on, precisely because most workflows have no forcing function that requires the cleanup step to actually happen[1].
What Checking Live Actually Found
Checking the live state, opening the actual testing platform’s dashboard for that experiment instead of the ticket that described it, found the experiment status set to running, with roughly half of pricing-page visitors still being split into a variant nobody had looked at the results of in months.
Nothing about the ticket was written in bad faith. The person who filed it almost certainly believed the test had ended, because the decision to end it had genuinely been made, discussed, and acted on for every purpose except the one that mattered here: the actual toggle in the testing platform. A calendar invite, a Slack thread, and a project-management ticket can all agree the test is over while the platform itself disagrees, and the platform is the only one of the four that decides what real visitors actually see.
What This Actually Costs When Nobody Checks
A forgotten test that is still splitting traffic does not fail loudly. It fails by quietly polluting every downstream number that assumes all visitors are seeing the same page. Campaign attribution, landing page conversion rate, and any pricing-page benchmark built during the months the split was still running are all measuring two different experiences and reporting the blended result as one number.
This connects directly to a silent default nobody caught for months on a different client’s site: the specific mechanism differs, a lead-classification default versus a leftover traffic split, but the shape is identical. Something kept running past the point everyone believed it had stopped, the dashboard kept producing a plausible-looking number the whole time, and nothing about the reporting layer distinguished a healthy number from a quietly compromised one. The defaults GA4 applies without telling anyone is the same failure at the platform-configuration layer instead of the experiment layer. All three are variations of one underlying rule: a system that changes behavior with no visible signal will eventually be trusted past the point it deserves it.
McKinsey’s 2024 survey of 104 C-suite marketing executives found only 41% consider their organization mature at performance measurement[2]. Silent infrastructure like a forgotten test is exactly the kind of gap that keeps that number low. It is invisible in any single dashboard review and only shows up as a general, hard-to-place sense that the numbers do not quite add up, which is precisely the state most of the other 59% are describing.
The Checklist for Removing Anything That Touches Data Collection
The fix here is not a smarter ticket-writing process. It is a mandatory verification step before any removal work starts, applied to test infrastructure, tracking pixels, feature flags gating an experience, or anything else that can silently change what a visitor sees or what gets recorded about them.
Before removing anything in this category, confirm four things directly in the system of record, not in a ticket, wiki, or chat thread: the actual status field shows stopped or archived inside the tool that controls it, live traffic allocation is at zero or fully weighted to a single variant, no active experiment ID appears in a fresh page load’s network requests, and the change log or audit trail (where the platform has one) shows an actual stop action, not just a decision recorded somewhere else. Four checks, each one independent of what anyone remembers or documented elsewhere, done in the minutes before the removal work starts rather than assumed from the ticket that requested it.
The document that would have caught this on schedule is the same idea applied proactively instead of reactively: a tracking plan with a last-verified date would have flagged this experiment’s continued existence at the next quarterly check, long before a cleanup ticket got filed based on an assumption nobody had re-checked.
Who Should Run This Check
Whoever executes the removal is the person who runs the four-point check, not the person who filed the ticket. The ticket author’s job was to flag the cleanup. The executor’s job includes confirming the premise before acting on it, because the executor is the one who will actually be running code, not just describing a decision.
This is a small discipline with an outsized payoff, because the failure it prevents is invisible until someone notices a number that does not make sense, and by then the cost has already compounded for as many months as the test kept quietly splitting traffic. Verifying live state first costs a few minutes. Not verifying it cost this particular engagement months of quietly blended conversion data on a page every paid campaign pointed at.
Sources
- Harness, Managing Feature Flag Retirement and Technical Debt – Abandoned test/flag infrastructure commonly leaves inactive logic running in production after a team moves on, because most workflows have no forcing function requiring the cleanup step ↩
- McKinsey, Connecting for Growth: A Makeover for Your Marketing Operating Model – 2024 Global Consumer Marketing Leader Survey, n=104 C-suite executives; 41% mature in performance measurement ↩
Seeing these patterns at your company?
Book a free WebOps Diagnostic. I'll review your site before the call and share specific observations.
Book a Free Call →Frequently Asked Questions
Because 'ended' usually describes a decision, not an action. Someone decided the test was over and moved on, but nobody went back into the testing platform and actually turned the experiment off or removed its traffic split. The decision and the live state drift apart, and nothing alerts anyone when they do.
Check the live state in the testing platform itself, not the ticket, the wiki page, or someone's memory. Confirm the experiment status is stopped, confirm traffic allocation is at zero or fully weighted to one variant, and confirm real visitors are no longer being split. A status field that says 'archived' in a project tool is not the same fact as a status field that says 'stopped' inside the testing platform.
It silently splits a portion of visitors into a variant nobody is tracking or acting on anymore, which pollutes every conversion metric downstream: campaign attribution, landing page performance, and any dashboard built on the assumption that all visitors are seeing the same experience.
No. The same failure applies to any infrastructure that touches data collection and gets removed based on assumed state: old tracking pixels, feature flags gating an experience, redirect rules, consent-mode configurations. A ticket title is a claim someone made at some point. Live production is the only source of truth.