A/B testing: the complete guide
Compare a control and a variant, measure a primary goal, and decide from evidence—not opinions. Below is the reference section; further down, how RunPivot runs modern tests from a plain-English brief.
A/B testing is a controlled experiment that compares two versions of a web experience (a control and a variant) by showing each to a random subset of visitors and measuring which version performs better on a chosen goal. Businesses use it to replace guesswork with evidence: instead of debating which headline, layout, or offer might work, you let real traffic decide under fair conditions.
How A/B testing works
Every test starts with a hypothesis: a specific, testable belief about what will improve a metric. You keep the current experience as the control and ship one or more variants that change only what your hypothesis targets.
Visitors are assigned through traffic allocation (often 50/50 for a simple A/B test). The platform records exposures and conversions for each arm. When enough data accumulates, you analyse whether the difference is likely real or random noise using statistical significance.
- 1Hypothesis
- 2Variant
- 3Traffic
- 4Measurement
- 5Analysis
- 6Decision
Metrics, sample size, and duration
Choose one primary metric before launch (for example completed signup, purchase, or demo request). Secondary metrics help you interpret side effects but should not override the primary decision without a pre-planned rule.
Sample size depends on baseline conversion rate and the minimum lift you care about. Low-traffic sites need bolder changes or longer runtimes. Use a sample size calculator to plan before you start.
Run tests through at least one full business cycle (often two to four weeks) so weekday, weekend, and campaign traffic are represented. Stopping early because a variant looks ahead is a common source of false winners.
- Statistical significance: confidence that an observed lift is not explained by chance alone.
- Experiment duration: long enough for sample size and seasonality, not arbitrary calendar dates.
- Primary vs secondary metrics: one decision metric; others for learning and guardrails.
When A/B testing helps, and when it does not
A/B testing shines on high-traffic pages where small conversion improvements compound: homepages, pricing, signup flows, and core product landing pages. It is weaker when traffic is tiny, when the change is purely brand opinion with no measurable action, or when technical implementation breaks measurement.
- Useful: clear conversion goal, enough visitors, one focused change, stable implementation.
- Risky: very low traffic, many simultaneous changes, broken tracking, constantly shifting pages.
- Common mistakes: testing without a hypothesis, peeking and stopping early, ignoring mobile segments.
A/B testing vs related disciplines
| Approach | What it optimises | Typical trade-off |
|---|---|---|
| A/B testing | One or few changes between control and variant | Clear causality; needs traffic and time |
| Multivariate testing (MVT) | Combinations of several elements at once | More combinations; needs much more traffic |
| Personalisation | Different experiences per segment | Relevance; harder to isolate causal lift per change |
| Traditional CRO audits | Heuristic review and backlog of ideas | Fast ideation; still needs validation in market |
Web experimentation is the broader program that may include A/B tests, split tests, feature flags, and UX research. A/B testing is the workhorse inside that program.
Practical examples to test
- Homepage headline clarity for first-time visitors.
- Primary CTA label and placement above the fold.
- Pricing page plan order, anchoring, and risk-reversal copy.
- Signup flow field count, social proof, and error messaging.
- Product page benefit order and demo vs trial emphasis.
- Navigation labels that reduce confusion between similar pages.
Traditional vs modern workflows
| Stage | Traditional workflow | AI-assisted workflow |
|---|---|---|
| Idea | Workshop or backlog ticket | Prompt or insight from analytics/replay |
| Build | Designer → developer → QA | AI draft variants for review |
| Launch | Manual configuration | Configure goal, audience, allocation in one workspace |
| Analysis | Export to spreadsheet or BI | Live significance and guardrails in product |
AI does not remove the need for sound hypotheses or valid statistics. It compresses build and operate time. See AI A/B testing and prompt-based A/B testing for how that category works in practice.
Where RunPivot fits
RunPivot is built for teams that want standard A/B discipline without a permanent queue of design and development tasks. You describe the change in plain language, review variants before launch, and measure against a defined primary metric with significance monitoring and optional automatic rollout. Deeper product detail lives on A/B testing in RunPivot and live results.
Continue learning
The full program: lifecycle, ideation, measurement, and learning loops.
AI A/B testingWhat AI actually automates, and what still requires human judgement.
Automated A/B testingStage-by-stage automation from opportunity to rollout.
Prompt-based A/B testingNatural-language briefs that become experiments.
Every A/B testing type
your program needs
RunPivot is not a single-test tool. It is a full experimentation platform: classic A/B, split URL, multivariate, sequential programs, targeting, and the automation that runs them end to end.
| A/B testing type | What you use it for | In RunPivot |
|---|---|---|
| A/B testing | Two variants on the same page. The classic controlled experiment. | Prompt-built or visual editor. Live in minutes. |
| Split URL testing | Compare entirely different page URLs, not just on-page changes. | Set up from a plain-English brief. No redirect gymnastics. |
| Multivariate testing | Test multiple page elements at once to find what combination wins. | RunPivot structures the experiment and handles traffic allocation. |
| Sequential testing | Run tests in sequence, rolling each winner forward before the next starts. | Winners called at 95% confidence. What's next stays queued. |
| Audience targeting | Test for mobile visitors, returning customers, one market, or one campaign. | Describe who in plain language. Targeting configured automatically. |
| Cross-domain testing | Run experiments that span multiple domains in one program. | One script tag, one program view across your properties. |
| Personalization | Show different experiences to different visitor segments based on what won. | Winning variants roll out to the audiences that responded. |
| Guardrail metrics | Protect revenue and core KPIs while a test is running. | Guardrails watched continuously. The line holds. |
Prompt-built variants
Describe the change in plain English. Variants, targeting, and structure come back for your approval.
Visual editor
Hands-on control when you want it. Edit what you can see without touching code.
Auto rollout
Winners ship to 100% of traffic at 95% confidence. Or hold for approval per experiment.
From idea to live
winner in three moves
Prompt it
Type what you want to test. A sharper headline, a new offer, a different layout. RunPivot turns your words into ready-to-run variants.
Test it
Your experiment goes live on your real site in minutes. RunPivot splits traffic, watches results, and holds the line on statistical discipline.
Ship the winner
The moment a result is statistically real, RunPivot calls it and rolls the winner out to everyone. No stale tests running at half traffic.
What changes when
the loop runs itself
Agentic experimentation is the shift from software that helps you run tests to programs that keep running. RunPivot is AI-native, built from the first line for a world where you describe the test and the platform does the rest.
Describe the test in plain English
One sentence is the brief. Variants, targeting, and experiment structure come back in minutes.
Prompt-Built TestsRunPivot watches every experiment
Traffic split, sample size, significance, guardrails. No peeking, no forgotten tests, no false positives.
Agentic ExperimentationWinners roll out automatically
Winning variants ship to 100% at 95% confidence. Set any experiment to wait for approval instead.
See auto rolloutThe next test is already queued
RunPivot tracks what won, what lost, and what hasn't been tried. The gap between experiments closes.
Explore the loopProof arrives ready for the room
Every winner comes with lift, confidence, and revenue impact in a report leadership actually reads.
Revenue-Ready ReportingThe highest-leverage
growth work you can do
A/B testing holds the world constant and isolates your change. When the variant wins, you know the variant caused it. That is the closest thing marketing has to proof.
Doubling conversion doubles revenue from the traffic you already have. The gain is permanent.
Roughly 88% of untested changes do nothing or hurt. A/B testing is about not shipping the losses.
If you are running a program, you are ahead of 99.8% of the web.
Most A/B testing programs die from friction, not lack of ideas. The idea takes five minutes. Everything after used to take five weeks. RunPivot closes that gap so your program compounds instead of stalling at two tests a quarter.
Why agentic
experimentation?
You describe the test. RunPivot is AI-native and closes the loop: launch, monitor, call winners, ship results.
Move faster
Describe the test you want in plain language and it's live in minutes. RunPivot builds the variants and handles the setup, so you go from idea to live experiment the same day.
Test smarter
Run better experiments without the statistics degree. RunPivot picks the right approach, shifts traffic toward what's working, and tells you the moment a result is real.
Scale without compromise
Fast, flicker-free experiences on every page. Run more experiments at once without slowing your site down or losing control of what goes live.
Prove impact
Every experiment ties back to revenue. Clear reporting connects each winner to the numbers your leadership actually cares about, so the value of your program is never in question.
One platform.
Three ways to win.
Prompt-Built Tests
Describe a change in plain language and it becomes a live experiment.
ExploreRun the programAgentic Experimentation
RunPivot runs the program, calls winners, and keeps what comes next moving.
ExploreProve the impactRevenue-Ready Reporting
Every winner comes with lift, confidence, and revenue impact.
ExploreWhat should you
actually test?
When a test costs a sentence, you do not need to guess which of your ten ideas is best. Run them properly, one after another, and let the traffic tell you.
Headlines and value props
Small wording changes here regularly produce the largest lifts because everyone sees the headline.
Calls to action
Copy, placement, and framing. One button quietly costs you leads every day.
Forms and checkout
Every field is friction. Checkout is the most valuable A/B testing ground most businesses own.
Social proof
Reviews, guarantees, and logos near the point of decision, where doubt actually lives.
Pricing presentation
Monthly versus annual emphasis, what is anchored first, how the tiers tell a story.
Mobile experiences
Mobile is most of your traffic and converts at a fraction of desktop. Test it on its own.
Start this week,
not this quarter
You need one page and one idea. The first test proves the workflow. The tenth test is where the compounding starts.
- 1
Pick your highest-traffic page with a conversion problem. Start where the visitors already are.
- 2
Write one hypothesis in one sentence. The change, the expected outcome, the metric.
- 3
Describe it to RunPivot. Variants, targeting, and experiment structure come back in minutes for your approval.
- 4
Let it run to significance. RunPivot holds the line on statistical discipline so you do not have to.
- 5
Ship the winner, read the report, run the next one. The recommendation for test two will be waiting when test one ends.
Comparing A/B testing platforms?
Honest, current comparisons of RunPivot against the major experimentation platforms.
A/B testing FAQ
What is A/B testing?
A/B testing compares a control experience and one or more variants by splitting traffic randomly and measuring which version achieves more of a predefined goal, such as signup or purchase. It is a controlled way to learn from visitors instead of debating opinions.
What is A/B testing in simple terms?
A/B testing shows two versions of a page or experience to different halves of your audience at the same time, then measures which version drives more of the action you want. It replaces opinions with evidence from your real visitors.
What is the difference between A/B testing and multivariate testing?
A/B testing usually changes one focused element or a single variant experience. Multivariate testing combines multiple elements to learn which combinations perform best and typically requires much more traffic.
What is agentic experimentation?
Agentic experimentation describes programs where the platform keeps tests operating: monitoring significance, applying rollout rules, and queueing follow-ups. You still set hypotheses and metrics; less time goes to manual dashboard work.
How is AI changing A/B testing?
AI can draft variants, suggest opportunities, and monitor results continuously. Statistical validity and traffic requirements do not disappear. See the AI A/B testing guide at https://www.runpivot.com/ai-ab-testing for a full breakdown.
How much traffic do I need to A/B test?
Enough to reach statistical significance in a reasonable window, which depends on your conversion rate and how big an effect you're testing for. Lower-traffic sites should test bigger, bolder changes, because large effects need smaller samples to detect. RunPivot calculates what each experiment needs before it launches, so you're never guessing.
What does statistical significance mean in A/B testing?
It's the confidence that your result reflects a real difference rather than random chance. At 95% significance, the standard threshold, there's only a 5% probability the result is noise. RunPivot won't call or roll out a winner below that bar.
How long should an A/B test run?
Until it reaches significance with an adequate sample, and through at least one full business cycle so weekday, weekend, and campaign traffic are all represented. For most sites that means two to four weeks. Stopping early because a variant looks good is the most common way teams ship false positives.
Can I A/B test without developers?
Yes. That's the point of prompt-built A/B testing. If you can describe the change in a sentence, RunPivot can build and launch the experiment. Your engineering team stays focused on the product.
What should I test first?
Your highest-traffic page with the clearest conversion problem. Headlines, calls to action, and checkout or form friction are the classic first tests because everyone sees them and small changes move real numbers.
Ready to ship your first winner?
Describe a test in plain English. RunPivot builds it, runs it, and ships the winner.
