Test Plans

Test Plans: The Testing Overview

An AI-planned testing program — pick a north star, and Split Test Pro maps your funnel, ranks test ideas by projected impact, lays them out on a surface-laned roadmap, and keeps the plan updated as results come in.

Intermediate9 min read

The Testing Overview is the home of Test Plans: instead of inventing one experiment at a time, you pick a goal and Split Test Pro maps your funnel, ranks test ideas by projected impact, and lays them out on a timeline — then keeps the plan updated as results come in.

You’ll find it under Experiments → Testing Overview in the app.

Setting Up Your Program

The first visit shows a single call to action: Set up your testing program. It opens a three-step wizard — Confirm, North star, Generate:

Confirm your business

The wizard prefills from your brand profile (the same one in Settings — one profile, reused everywhere) and asks only what it lacks. You also confirm your surface map: the areas of your site where tests can run, like Home, Collections, Product pages, Cart, and Checkout. One test runs per surface at a time; different surfaces test in parallel. HTML sites also pick a site type (online store vs. something else) and can enter a rough monthly-sessions estimate.

Pick your north star

The one number this program should move — for stores: purchase conversion rate (recommended), revenue per session, or average order value. Other site types pick any conversion goal, optionally with a value per conversion (e.g. a booked demo ≈ $300) so the program can report estimated value. A default guardrail on revenue per session is set for you — see Global Goals.

Generate

A data-readiness checklist shows exactly what your plan will be built from — analytics history, past tests, your brand profile, surface map, and site snapshot — and what’s not connected yet. Your plan is built from what you have; it gets smarter as you connect more. Generating a plan uses 5 AI generations from the same monthly meter as the AI assistant.

If a generation fails, nothing is used from your allowance — the generations are refunded automatically.

The Page at a Glance

Once a plan exists, the Testing Overview has four parts:

  • KPI strip — your north star (current value vs. baseline), guardrail status, program impact, and activity (what’s running, what’s next). The first three tiles click through to Global Goals.
  • Roadmap — a surface-laned timeline. Each lane is one surface; bars are tests, positioned by their projected schedule.
  • Queue — the ranked backlog of test ideas, each with a hypothesis, surface, projected impact, and effort.
  • Plan updates — when results land, a proposed update card appears at the top with the reasons for each change. See plan updates below.

Projected, Not Scheduled

All dates on the roadmap are projections, not scheduled starts. They come from a queue simulation: how long each test is expected to need on its surface’s real traffic, and what has to finish before it can start. Nothing launches on a date — tests start when you (or Autopilot) start them, and the projections re-flow around what actually happens.

Two chips flag feasibility on queue items:

  • Low power — this surface’s traffic can only detect fairly large lifts in a reasonable time; smaller effects won’t reach significance.
  • Traffic unknown — no traffic data for this surface yet, so duration is unknown until data arrives. Opting in to site-wide tracking fills these in for surfaces you haven’t tested.

The Queue

Every plan item moves through a simple lifecycle: Idea → Queued → Building → Ready → Running → Done (or Dropped). Projected impact is categorical — High, Medium, or Quick win — never a made-up number. You stay in control of the backlog: drag to reorder, drop what you disagree with.

Build turns a plan item into a real, ready-to-review experiment — variants, targeting, and goals — for 1 AI generation. Built tests go through the same pre-launch QA as any other experiment, and you review before anything starts.

Plan Updates: Suggest and Approve

The plan is not static. When an experiment concludes, an anomaly is flagged, metrics refresh, or you change your north star, Split Test Pro proposes a plan update: a card listing each change (add, drop, move, update, follow-up) with the reason for it — for example, a follow-up test added because a variant won.

Nothing changes until you approve. You can approve all changes or only some, or reject the update (the AI proposes again only on new evidence). Approving uses 2 AI generations; proposals are free to receive. Every decision lands in the revision history, so the plan’s whole evolution is auditable.

When you’ve built a track record, you can hand parts of this loop to Autopilot — it’s off by default and earned per workspace.

Surfaces and the Site Snapshot

The Surfaces manager (page-head button) edits your surface map and shows a per-surface crawl status. Refresh site snapshot re-crawls your pages so generated ideas are grounded in what’s actually on your site — see the crawler for how crawling works and what’s respected (robots.txt, rate limits).

AI Generations

One AI generation meter covers everything: assistant drafts (1), full plan generation (5), plan updates you approve (2), building a planned test (1). The meter and remaining balance are always shown before anything is spent, and failed generations are refunded.

Next Steps

Ready to ship a winner?
Start free today.

Sign up free in under a minute — no credit card, no installation steps until you're ready.

No credit card14-day free trial