Autopilot & the Trust Ladder
Opt-in automation for your testing program — auto-stop, prebuild, auto-start, and auto-rollout, each unlocked by track record, with hard caps, kill switches, and an action log + email digest for every move.
Autopilot lets Split Test Pro run the program you approved. It’s off by default, opt-in per action, and every action is logged and emailed to you. You’ll find it behind the Autopilot button in the Testing Overview page head (workspace owners can change it; members see it read-only).
The Four Actions
Each action is a separate toggle under one master switch:
- Auto-stop — “Conclude tests at ≥95% confidence on two consecutive daily checks, or on futility — never earlier than 7 days and a minimum sample.”
- Prebuild — “Build the next queued tests ahead of time so they’re ready to start. 1 generation each.”
- Auto-start — “Start the next ready test when its surface frees up — only if built, safety-scanned, QA-passed, and within your caps.”
- Auto-rollout — “Serve a clean winner at 100% automatically after it concludes.” (See Winner Rollout.)
The Trust Ladder
You don’t get all four on day one. Each action is earned per workspace, matched to its blast radius — the more a mistake would cost, the more track record it takes to unlock:
| Action | Unlocks when |
|---|---|
| Auto-stop | 1 experiment has completed in this workspace. |
| Prebuild | 2 plan items built with the manual Build button pass Launch QA without manual edits. |
| Auto-start | 3 approved plan updates, an edit rate of at most 20% over the last 5 decided proposals, and 2 plan-created experiments manually started and completed without SRM or anomaly flags. |
| Auto-rollout | 1 planned winner is manually rolled out. |
Locked actions show their progress in the drawer (e.g. “Build 2 plan items with the Build button first — 1 of 2”), so you always know what’s left to earn.
Trust can also be lost. If an autopilot-driven test triggers an anomaly or a rollout goes wrong, the action is demoted — paused with a note explaining the cause and the criteria to re-earn it.
Limits and Kill Switches
Autopilot is bounded in several independent ways:
- Master switch — one toggle turns all automatic actions off. Off means off: “All automatic actions are off.”
- Daily action cap — at most this many automatic actions per day (default 3).
- Monthly autopilot generations budget — autopilot stops (and emails you) instead of spending past this. It never buys add-ons or overages.
- Self-suspension — autopilot suspends itself after 3 consecutive failures; you review the log and re-enable explicitly.
- Anomaly hold — if an anomaly is flagged on an autopilot-started test, auto-start pauses until you acknowledge it.
- Held for review — a test whose generated code was flagged by the safety scan, or that came from an automatically applied plan update, is never auto-started; it waits in a “Held for review” list for your explicit approval.
- Fleet pause — Split Test Pro operations can pause autopilot globally if something is wrong on our side. You’ll see it in the drawer when active.
Plan changes that add or drop tests always require your approval — autopilot can auto-apply only low-risk reorder and follow-up updates, and even that respects everything above.
The Action Log and Emails
Every autopilot action lands in the drawer’s action log with a timestamp, and an email digest goes to the workspace owner after each action tick. Each email includes a one-click pause link, so you can stop autopilot from your inbox without logging in.
Manual Actions Always Win
Autopilot reconciles against reality instead of fighting it. If you manually start, stop, or edit an experiment, autopilot treats your action as the truth and re-plans around it — it never reverts or blocks something you did by hand.
Next Steps
- The plan autopilot executes: Testing Overview
- What auto-rollout actually does: Winner Rollout
- The guardrails that constrain every decision: Global Goals