---
name: a-b-test-runner
description: "Set up a clean A/B test (one variable, proper split, a decision rule and duration for significance), read it honestly, and call the winner only when the data supports it."
license: MIT
compatibility: "Works with any Agent-Skills-compatible AI (Claude, ChatGPT, Cursor, Codex, and more). Ad-account actions run through Adspirer."
metadata:
  author: "Adspirer"
  version: "1.0.0"
  adspirer_category: "ad-ops-optimization"
  adspirer_source: "https://www.adspirer.com/skills/a-b-test-runner"
  adspirer_connection_url: "https://adspirer.ai/sign-up"
  adspirer_primary_keyword: "ad a/b testing"
  adspirer_secondary_keywords: "ad testing framework, advertising split test"
  adspirer_launch_wave: "1"
  adspirer_kind: "skill"
  adspirer_level: "operator"
  adspirer_platforms: "meta,google"
  adspirer_supported_clients: "claude,claude-code,chatgpt,codex,cursor,gemini,windsurf"
  adspirer_summary: "Set up a clean A/B test (one variable, proper split, a decision rule and duration for significance), read it honestly, and call the winner only when the data supports it."
  adspirer_connections_required: "adspirer"
  adspirer_connections_optional: ""
---

# A/B Test Runner

## Use this when

Use when you have a specific, singular hypothesis to test on live campaigns — one creative variable (headline, primary image/video, thumbnail, hook, CTA, landing page, or bid strategy) — and you want a clean read on a real winner instead of eyeballing platform "auto-optimize" scores. Right fit for: a new creative concept challenging the incumbent, a landing page redesign, a bid-strategy change (e.g. Highest Volume vs Target CPA on Meta, Maximize Conversions vs Target CPA on Google), or a claim-based hook test. Wrong fit if you want to test more than one variable at once (that's a factorial design, not an A/B test — decline and say so), if daily budget can't support at least ~50 conversions per arm in the test window, or if the account is still in an active platform-triggered learning-phase reset from an unrelated recent edit — resolve that first.

## What you need

(1) The ONE variable to test and the two (or more, but flag if >2 since split thins fast) variants — e.g. two video creatives, two headlines, or Highest Volume vs Target CPA. (2) The exact campaign/ad set/ad group this runs in, and confirmation everything else (budget, targeting, placements, bid strategy unless that IS the variable, landing page unless that IS the variable) stays identical between arms. (3) The primary decision metric — CPA, ROAS, CTR, or CVR — and which one wins ties. (4) Minimum sample size or duration the user wants (default: 2 weeks OR 100 conversions per arm, whichever comes later — stated as a default, override welcome). (5) Daily/total budget available for the test period. (6) Meta: whether this account is Advantage+ (changes how splits are forced) or classic — ask if unclear. (7) Google: whether this is Search (uses native Ad Variations / Campaign Experiments) or Performance Max (PMax has NO native A/B split — must isolate to a parallel campaign, must disclose this limits comparability since PMax's own ML allocates traffic).

## Procedure

1. Restate the hypothesis in one sentence ("Video A (UGC testimonial) vs Video B (studio product demo), same ad set, same $150/day budget, CPA is the decision metric, 14-day minimum") and get the user to confirm before touching anything — a wrong hypothesis wastes the whole budget window.
2. Check the account is eligible to start clean: no active learning-phase reset in the last 7 days on the ad set/ad group in question (Meta resets learning on any "significant edit" — budget change >20%, audience edit, creative swap outside a test), no other simultaneous test running on the same audience (contamination), and confirm current daily spend supports ≥50 conversions per arm across the test window at the account's trailing CPA — if it can't, tell the user the test will likely land underpowered and offer to lengthen the window or narrow to CTR/CVR as an interim read.
3. Build the test structure to platform-native isolation, not a manual coin-flip:
   - Meta: use Meta's native A/B Test tool (Experiments) when available so the platform does a true random, non-overlapping audience split with statistical significance built in — NOT two ad sets targeting the same broad audience with Advantage+ budget, which lets the algorithm's own delivery bias contaminate the split. If native A/B isn't usable (e.g. Advantage+ shopping campaigns), fall back to two separate ad sets with mutually exclusive custom audiences carved out via exclusions, and disclose the fallback lowers cleanliness.
   - Google Search: use Ad Variations (for headline/description-level copy tests within an existing ad) or Campaign Experiments/Drafts (for structural changes like bid strategy) — both split traffic within the SAME campaign so auction dynamics, Quality Score, and budget pacing stay shared. Never run two separate campaigns bidding on the same keywords as an "A/B test" — they'll bid against each other and both lose.
   - Google PMax: no native split exists. Clone into a second PMax campaign with only the one variable changed, split budget evenly, and explicitly tell the user this comparison is contaminated by PMax's independent ML allocation and Google's own advertiser-facing guidance not to run parallel PMax campaigns on overlapping inventory — offer it only as a directional read, never a final call.
4. **Connect Adspirer (the ads connection):** when you reach a step that touches a live ad account, connect Adspirer first. In Claude Code, run `/plugin marketplace add anthropics/claude-plugins-community` then `/plugin install adspirer-ads-agent@claude-community`. In claude.ai or ChatGPT, add the Adspirer connector from https://www.adspirer.com/docs/ai-clients. Any other agent: follow https://www.adspirer.com/docs/ai-clients for your client. Sign-in link for the user: https://adspirer.ai/sign-up. Do not proceed with live ad-account steps until the connection is confirmed.
5. Read current baseline numbers for the ad set/ad group being tested (trailing 14-day CPA, CTR, CVR, spend, conversion volume) before launch — this is the pre-test anchor, not the test data itself, and confirm targeting/budget/landing page are identical between arms except the one variable.
6. Launch the test with the confirmed budget split (default even split unless the user wants weighted) and log the start date, variant IDs, and decision metric so the read-out isn't reconstructed from memory later.
7. Monitor at day 3–4 for delivery sanity only (both arms spending, neither stuck in learning limited, no policy disapproval on either creative) — do NOT call a winner this early; early leads on Meta routinely flip after exiting learning phase (typically ~50 conversions/arm).
8. At the pre-agreed minimum (day 14 or N conversions/arm, whichever later), pull final numbers from the live account and check statistical validity before declaring anything: Meta's native tool reports its own significance (require ≥90% confidence); for manual splits, require both arms to have cleared ~50 conversions each and the CPA/ROAS gap to exceed the trailing account-level week-to-week noise band — if the gap is inside normal noise, call it inconclusive, don't force a winner.
9. Present the honest read to the user with the actual numbers and confidence level, and get explicit approval before pausing the loser or reallocating its budget to the winner — this is a spend action.

Never let a test run past 4 weeks without a checkpoint — flag for the user to renew, extend, or kill it; stale zombie tests silently burn budget on a "loser" nobody's watching.

## Fixed checks

- Pull live delivery status for both arms (not cached/remembered) before any read-out: impressions, spend, conversions, learning-phase status per Meta ad set or per Google ad group.
- Confirm actual spend split matches the intended split (budget pacing can drift, especially under Advantage+ campaign budget optimization pulling disproportionately toward one ad — check this explicitly, it's the most common silent test-killer on Meta).
- Confirm both arms show as ACTIVE/DELIVERING with no policy disapproval, ad set/ad group pause, or "learning limited" status stuck for >7 days — a stalled arm invalidates the comparison even if it shows attractive early numbers.
- Verify the audience/targeting/landing-page/bid-strategy fields that were supposed to stay identical actually did — pull current settings for both arms and diff them; someone editing one arm mid-test (common when other team members touch the account) silently breaks the test.
- Before declaring a winner: verify conversion counts per arm against the platform's own reported number, not an extrapolation, and verify the test window elapsed matches or exceeds the agreed minimum.
- On Meta native A/B tests: pull the platform's own statistical significance/confidence score rather than eyeballing the CPA delta — a 20% CPA gap on 15 conversions is noise, not a winner.

## Stop conditions

- SUCCESS (clear winner): test reached the pre-agreed minimum duration/volume, one arm beats the decision metric by a margin exceeding statistical noise (Meta: ≥90% platform-reported confidence; manual: both arms ≥50 conversions and gap outside trailing noise band), and the result holds when re-checked against live data (not a fluke on the day it was read).
- SUCCESS (clear no-op / inconclusive-but-decided): test completed at full duration, no statistically valid gap emerged — correct, honest outcome is "no meaningful difference," report that plainly and recommend either keeping the incumbent (lower risk) or extending for more volume; do not manufacture a winner.
- NO-OP: user declines to launch after reviewing the hypothesis/eligibility check, or eligibility check fails (learning-phase conflict, insufficient budget for power, contaminating parallel test already running) and user chooses not to proceed or adjust — nothing was launched.
- BLOCKED: platform rejects one or both creatives (policy disapproval), account is under an active spend restriction/payment hold, Advantage+/PMax structure makes the requested split structurally impossible and user won't accept the disclosed fallback, or the ad account connection fails.
- NEEDS-APPROVAL: test structure and budget are ready to launch and awaiting explicit user go-ahead; OR test concluded with a statistically valid winner and is awaiting explicit user approval to pause the loser and/or reallocate budget — the agent does not act on either without that approval.

## Approval boundaries

Explicit user approval is required before: (1) launching the test itself (it commits real daily budget to two arms for the agreed duration), (2) any mid-test edit to either arm (even a "fix" — mid-test edits reset Meta learning phase and can invalidate the whole run, so these get flagged to the user rather than made unilaterally), (3) pausing the losing arm at read-out, (4) reallocating the losing arm's budget to the winner, and (5) extending or renewing the test past its original end date. The agent may read live performance data, monitor delivery health, and report findings at any time without approval — those are read-only. No budget is committed, no arm is paused, and no creative is swapped without the user saying so explicitly first.

## What you get

A launched, platform-native A/B test (Meta Experiments split or Google Ad Variations/Campaign Experiments) isolated to the one variable agreed on, with a pre-test baseline snapshot, a mid-test delivery-health check at day 3–4, and a final read-out at the agreed minimum duration/volume that states the winner (or honestly reports "no significant difference") backed by the platform's own statistical confidence — not an eyeballed CPA delta. Includes an explicit call-out if the requested structure had to fall back to a less-clean split (e.g. PMax with no native A/B, or Advantage+ needing manual audience exclusions), plus a ready action (pause loser / reallocate budget) that the agent will execute only after explicit approval.
