A/B testing
Change one thing, let your visitors vote. Janus runs the test end to end: a four-step wizard to set it up, honest math while it runs, and verdicts written in plain English, no statistics degree required.
How Janus tests work
- One variable per test. A test changes exactly one thing: the headline, the trigger, the opener. That's what makes the result mean something.
- Up to four challenger versions. One variable can have several versions: headline B vs C vs D, all against your current popup (the control). Five arms total, max.
- Sticky assignment. The same visitor always sees the same version: no flicker, no double counting.
- The winner metric is locked up front. You pick what decides the winner before the test starts, and it can't be changed mid-test. No moving the goalposts.
Starting a test: the wizard
Open it with New test, from a campaign's ⋯ menu, its A/B Testing tab, or a list row. Four steps:
1 · What do you want to test?
Two tabs:
- One-click tests: ready-made recipes. Pick one and Janus builds the challenger for you:
Recipe The challenger it builds Popup vs no popup The holdout test: part of your traffic sees no popup at all. See below. Trigger delay Same popup, different opening delay. Trigger type Time delay vs exit intent vs scroll depth vs a trigger you build yourself. Header and button copy New headline or button text, edited live on the real popup. Opener A different first screen: mini quiz, spin to win, scratch card, yes/no and more. Countdown timer Adds an urgency countdown. X button position Moves the close button to the opposite corner. Popup design An exact copy of your popup you're free to change however you like. - Test ideas: a gallery of fourteen field-proven ideas (offer type, popup format, multi-step design, visual style, background contrast, second-chance re-shows and more). Ideas that map to a recipe set themselves up in one click.
2 · Set it up
Configure your challenger versions: + Add Variant for up to four. What you see depends on the recipe:
- Copy tests get a live mini-editor: the real popup renders beside the fields, updates as you type, and clicking any text in the preview jumps to its field. Multi-step popups get a step switcher.
- Opener tests pick from the full opener lineup (quiz, spin, scratch, plinko, slots, pick-a-card, yes/no…).
- Trigger tests can use a build-your-own trigger: combine Time on page, Scroll depth, Exit intent, Mouse leaves the page, Quick scroll-up, User goes idle, Item added to cart and Cart value ≥ $ with any/all logic.
Janus blocks identical versions and unconfigured picks with a plain-English hint, so a broken test can't start.
3 · Name & traffic
- Name (prefilled) and an optional hypothesis: "What do you expect to happen?" Writing it down keeps you honest later.
- What decides the winner?: Email signups, SMS signups, Clicks, Orders from the popup, or Store conversion. Locked once the test starts.
- Traffic allocation: Classic (fixed split) keeps the split you set; Smart (shift traffic to the winner) continuously moves traffic toward the better performer while still giving stragglers a chance to recover. Choose Smart when you'd rather bank conversions than wait for a formal verdict.
- Traffic split: a percentage per arm (must total 100; Split evenly does the math).
- Duration: 7 days, 14 days (default), 30 days, 8 weeks, or no end date.
- How should the test end?: You end it (default), or End it automatically: when Janus is at least 90% sure a version is best, it goes live for everyone.
4 · Review & start
Every arm renders side by side: your current popup, each challenger (with ✎ Edit design first if you want to polish it in the builder), and a "No popup" pane for holdouts. Start test goes live; Save as draft keeps the challengers without starting anything.
Reading the results
Janus reads out each test's state in a sentence, not a p-value:
- Collecting data, too early to call. Every arm needs a real audience (at least ~100 views each) before any verdict is offered.
- "B is ahead, 82% sure it's best. Keep it running." A lead, not yet a win, with an estimate of how many more visitors and days the test needs at its current pace.
- "B is the likely winner, 94% sure." The confidence bar is 90%. Time to act.
Under the hood the verdicts are Bayesian ("how likely is this version the best one?"), which stays honest even when you peek at the results every day. Classic significance (the 95% z-test) is shown alongside for the statistically inclined, but it never drives the verdict.
The Tests page
Tests in the nav lists every test: Running first (each card with arm thumbnails, the day count, its traffic mode and its current verdict), then Finished, a permanent history with winner, lift, and duration.
Click into a test for the full detail page:
- The verdict card, with End test…, ⏸ Pause / ▶ Resume, Edit split, and the test's facts: judged on, traffic mode, end condition, hypothesis.
- A confidence chart: each arm's plausible range for the primary metric. When ranges stop overlapping, you have your answer.
- The comparison matrix: every metric as a row, every arm as a column: views, email and SMS sign-up rates, click rate, orders, sales, store conversion, bounce rate, each with its change vs the control. A version that wins sign-ups but hurts the store shows its cost here.
- A cumulative sign-up-rate chart over the life of the test.
- Slices: cut a running test by UTM source, country or a saved segment. An even test overall can hide a mobile or paid-traffic winner.
Ending a test
End test asks which version won (each option shows its rate, views and lift) or No winner, keep my current popup. A winning challenger can also be crowned the campaign's new design (the old design stays safe in version history). Ending on a version that isn't a clear winner gets a warning first: stopping early can crown a false winner. Final numbers are saved to history forever; losing versions become drafts; nothing is deleted.
Tests with a duration flag themselves when time's up. Tests running past 8 weeks get a seasonality warning: at that point the traffic that started the test isn't the traffic finishing it.
Popup vs no popup
The most honest test there is, and a first-class option, no configuration tricks. Part of your traffic sees no popup at all; Janus counts those visitors at the exact moment the popup would have opened, then judges the test on sitewide conversion and revenue, not sign-ups. Expect the no-popup group to collect zero emails; that's the point: does the popup earn its interruption? Holdouts always run on a fixed split, so the comparison stays clean.
Autopilot: a schedule of tests that runs itself
One test answers one question. Autopilot asks the next ten for you. From any campaign's ⋯ menu, Autopilot reads your popup and builds an improvement plan: a queue of tests generated from your own design (a game opener, curiosity copy, a countdown, better timing), each one previewable side by side with your current popup before anything runs.
- Only what applies. Already have a countdown? That test isn't offered, and the page says why.
- Your voice, pre-filled. Headlines and questions are drafted from your own store's copy and products. Every field is editable before you start.
- One tap, then hands off. Tests run one at a time as normal A/B tests. When one concludes, the winner becomes your live popup and the next test starts against it; each test builds on the last winner.
- Plain-language notes. "Mystery copy won (+9%); it's your live popup now. Now testing: countdown timer." The campaign list shows Autopilot · test 2 of 6 the whole way.
- Nothing reckless. Winners only roll out at 90% confidence; a variant that's clearly losing ends early and the schedule moves on; anything that would change your offer's cost is never started without your OK. Pause or take over any time; every Autopilot test is a regular test on the Tests page with the full numbers.