Before You Remove That Old A/B Test, Check If It Actually Ended

A ticket describing old test infrastructure as "ended months ago" is a claim about a document, not a claim about production. Before removing anything that still touches live traffic or conversion data, check the actual running state. A test that a ticket calls finished can still be splitting real visitors and quietly skewing the numbers a team trusts.

Yasser Soliman

Yasser Soliman

Fractional Head of WebOps

Published

Updated

10 min read

A task ticket on a recent engagement read almost exactly like this: “Remove the old pricing-page A/B test, it ended months ago.” The ticket had a date, a name, and a confident tone. It also turned out to be describing a decision that was made and then never carried out. The test was still live, still splitting real visitors, months after everyone involved believed it was over.

The Ticket Said “This Test Ended Months Ago”

A ticket that says a test “ended months ago” is a claim about a document, not a claim about production. Someone decided the test was over. Whether anyone actually went into the testing platform and stopped it is a separate, unverified fact.

This is a B2B SaaS marketing site, anonymized here as with every client pattern on this blog. The task read as routine cleanup: an old pricing-page experiment, superseded by a redesign months earlier, flagged for removal so the codebase and the analytics setup could both be simplified. Nothing about the ticket suggested anything unusual. A reasonable person reading it would assume the only work left was deleting some now-inert code.

The habit that changed the outcome here was simple and almost boring: before touching anything that could affect live traffic or data collection, check the actual running state first, not the description of it. That habit is the entire subject of this post, because the alternative, trusting the ticket, is the default almost everyone reaches for under time pressure.

Why You Check the Live State Before You Touch Anything That Touches Data

A ticket, a wiki page, or someone’s memory can all be wrong about production in the same specific way: they describe a decision, and production reflects whatever was actually implemented, which drifts from the decision the moment nobody circles back to confirm it landed. Checking live state closes that gap. Trusting the description leaves it open.

The gap exists because “the test ended” and “the test was stopped” are two different events, separated by a manual step that is easy to skip. A team runs an experiment, reaches a conclusion, and moves on to the next priority. Stopping the experiment in the testing platform, removing its traffic allocation, and cleaning up the resulting code are three more steps after the decision, and none of them happen automatically just because the team stopped thinking about the test. Industry analysis of this exact pattern in engineering contexts describes abandoned test variants as commonly leaving inactive logic running long after a team has moved on, precisely because most workflows have no forcing function that requires the cleanup step to actually happen[1].

The gap between “decided” and “done” A horizontal chain of four boxes. Box one, decision made: the team concludes the test is over. Boxes two through four, each a manual follow-up step: stop the experiment in the platform, remove traffic allocation, clean up the code. A caption states that skipping any one of the three follow-up steps leaves the experiment live in production while everyone believes it has ended. The gap between “decided” and “done” Three manual steps stand between a decision and production actually matching it. DECISION Team concludes the test is over. → STOP IT Stop the experiment in the platform. → REMOVE SPLIT Remove traffic allocation entirely. → CLEAN UP Remove the leftover code. Skip any one step and the test stays live while everyone believes it ended.

What Checking Live Actually Found

Checking the live state, opening the actual testing platform’s dashboard for that experiment instead of the ticket that described it, found the experiment status set to running, with roughly half of pricing-page visitors still being split into a variant nobody had looked at the results of in months.

Nothing about the ticket was written in bad faith. The person who filed it almost certainly believed the test had ended, because the decision to end it had genuinely been made, discussed, and acted on for every purpose except the one that mattered here: the actual toggle in the testing platform. A calendar invite, a Slack thread, and a project-management ticket can all agree the test is over while the platform itself disagrees, and the platform is the only one of the four that decides what real visitors actually see.

Three sources agree. One source controls reality. Three small cards on the left, calendar invite, Slack thread, project ticket, each marked ended in green. One larger card on the right, the testing platform itself, marked running in red. A caption states that only the platform’s own status controls what real visitors actually see. Three sources agree. One source controls reality. Only the testing platform’s own status decides what real visitors see. CALENDAR INVITE Says: ended SLACK THREAD Says: ended PROJECT TICKET Says: ended THE TESTING PLATFORM ITSELF Says: running Still splitting roughly half of pricing-page visitors. Illustrative · yassersoliman.com · live state vs. described state

What This Actually Costs When Nobody Checks

A forgotten test that is still splitting traffic does not fail loudly. It fails by quietly polluting every downstream number that assumes all visitors are seeing the same page. Campaign attribution, landing page conversion rate, and any pricing-page benchmark built during the months the split was still running are all measuring two different experiences and reporting the blended result as one number.

This connects directly to a silent default nobody caught for months on a different client’s site: the specific mechanism differs, a lead-classification default versus a leftover traffic split, but the shape is identical. Something kept running past the point everyone believed it had stopped, the dashboard kept producing a plausible-looking number the whole time, and nothing about the reporting layer distinguished a healthy number from a quietly compromised one. The defaults GA4 applies without telling anyone is the same failure at the platform-configuration layer instead of the experiment layer. All three are variations of one underlying rule: a system that changes behavior with no visible signal will eventually be trusted past the point it deserves it.

McKinsey’s 2024 survey of 104 C-suite marketing executives found only 41% consider their organization mature at performance measurement[2]. Silent infrastructure like a forgotten test is exactly the kind of gap that keeps that number low. It is invisible in any single dashboard review and only shows up as a general, hard-to-place sense that the numbers do not quite add up, which is precisely the state most of the other 59% are describing.

One silent split, three polluted numbers A single source box, forgotten traffic split, feeds three downstream cards: campaign attribution, landing page conversion rate, and pricing benchmarks. Each downstream card is marked as blending two different visitor experiences into one reported number, with no flag distinguishing a healthy figure from a compromised one. One silent split, three polluted numbers Nothing in any single dashboard distinguishes the healthy number from the blended one. FORGOTTEN TRAFFIC SPLIT Campaign attribution Two experiences, one blended figure. Landing page conversion rate Same blend, reported as a single trend. Pricing-page benchmarks Built on months of quietly split data. Illustrative · yassersoliman.com · downstream pollution

The Checklist for Removing Anything That Touches Data Collection

The fix here is not a smarter ticket-writing process. It is a mandatory verification step before any removal work starts, applied to test infrastructure, tracking pixels, feature flags gating an experience, or anything else that can silently change what a visitor sees or what gets recorded about them.

Before removing anything in this category, confirm four things directly in the system of record, not in a ticket, wiki, or chat thread: the actual status field shows stopped or archived inside the tool that controls it, live traffic allocation is at zero or fully weighted to a single variant, no active experiment ID appears in a fresh page load’s network requests, and the change log or audit trail (where the platform has one) shows an actual stop action, not just a decision recorded somewhere else. Four checks, each one independent of what anyone remembers or documented elsewhere, done in the minutes before the removal work starts rather than assumed from the ticket that requested it.

The four-point check before you remove anything Four stacked rows, each a required check. Row one, status field: confirm stopped or archived inside the controlling tool. Row two, traffic allocation: confirm zero or fully weighted to one variant. Row three, network requests: confirm no active experiment ID on a fresh page load. Row four, audit trail: confirm an actual stop action, not just a recorded decision. The four-point check before you remove anything Each check is independent of what anyone remembers or documented elsewhere. 1. STATUS FIELD Stopped or archived, inside the tool that controls it 2. TRAFFIC ALLOCATION Zero, or fully weighted to one variant 3. NETWORK REQUESTS No active experiment ID on a fresh page load 4. AUDIT TRAIL An actual stop action, not just a recorded decision Illustrative · yassersoliman.com · the four-point check

The document that would have caught this on schedule is the same idea applied proactively instead of reactively: a tracking plan with a last-verified date would have flagged this experiment’s continued existence at the next quarterly check, long before a cleanup ticket got filed based on an assumption nobody had re-checked.

Who Should Run This Check

Whoever executes the removal is the person who runs the four-point check, not the person who filed the ticket. The ticket author’s job was to flag the cleanup. The executor’s job includes confirming the premise before acting on it, because the executor is the one who will actually be running code, not just describing a decision.

This is a small discipline with an outsized payoff, because the failure it prevents is invisible until someone notices a number that does not make sense, and by then the cost has already compounded for as many months as the test kept quietly splitting traffic. Verifying live state first costs a few minutes. Not verifying it cost this particular engagement months of quietly blended conversion data on a page every paid campaign pointed at.

Sources

  1. Harness, Managing Feature Flag Retirement and Technical Debt – Abandoned test/flag infrastructure commonly leaves inactive logic running in production after a team moves on, because most workflows have no forcing function requiring the cleanup step ↩
  2. McKinsey, Connecting for Growth: A Makeover for Your Marketing Operating Model – 2024 Global Consumer Marketing Leader Survey, n=104 C-suite executives; 41% mature in performance measurement ↩

Seeing these patterns at your company?

Book a free WebOps Diagnostic. I'll review your site before the call and share specific observations.

Book a Free Call →

Frequently Asked Questions

Yasser Soliman

Written by Yasser Soliman

Fractional Head of WebOps

I've spent 5+ years embedded in marketing teams at B2B SaaS companies. I own the marketing website — performance, analytics, SEO, integrations — so your team ships without bottlenecks.

Let's talk about your site.

Book a free WebOps Diagnostic. Send me your URL and what you'd like me to look at — I'll come prepared with specific observations.

Book a Free Call