A/B testing

Change one thing, let your visitors vote. Janus runs the test end to end: a four-step wizard to set it up, honest math while it runs, and verdicts written in plain English, no statistics degree required.

How Janus tests work

Starting a test: the wizard

Open it with New test, from a campaign's menu, its A/B Testing tab, or a list row. Four steps:

1 · What do you want to test?

Two tabs:

2 · Set it up

Configure your challenger versions: + Add Variant for up to four. What you see depends on the recipe:

Janus blocks identical versions and unconfigured picks with a plain-English hint, so a broken test can't start.

3 · Name & traffic

4 · Review & start

Every arm renders side by side: your current popup, each challenger (with ✎ Edit design first if you want to polish it in the builder), and a "No popup" pane for holdouts. Start test goes live; Save as draft keeps the challengers without starting anything.

Reading the results

Janus reads out each test's state in a sentence, not a p-value:

Under the hood the verdicts are Bayesian ("how likely is this version the best one?"), which stays honest even when you peek at the results every day. Classic significance (the 95% z-test) is shown alongside for the statistically inclined, but it never drives the verdict.

The Tests page

Tests in the nav lists every test: Running first (each card with arm thumbnails, the day count, its traffic mode and its current verdict), then Finished, a permanent history with winner, lift, and duration.

Click into a test for the full detail page:

Ending a test

End test asks which version won (each option shows its rate, views and lift) or No winner, keep my current popup. A winning challenger can also be crowned the campaign's new design (the old design stays safe in version history). Ending on a version that isn't a clear winner gets a warning first: stopping early can crown a false winner. Final numbers are saved to history forever; losing versions become drafts; nothing is deleted.

Tests with a duration flag themselves when time's up. Tests running past 8 weeks get a seasonality warning: at that point the traffic that started the test isn't the traffic finishing it.

Popup vs no popup

The most honest test there is, and a first-class option, no configuration tricks. Part of your traffic sees no popup at all; Janus counts those visitors at the exact moment the popup would have opened, then judges the test on sitewide conversion and revenue, not sign-ups. Expect the no-popup group to collect zero emails; that's the point: does the popup earn its interruption? Holdouts always run on a fixed split, so the comparison stays clean.

Good hygieneTest one thing at a time, let a test run through at least one full week (weekend traffic behaves differently), and treat "too early to call" as exactly that: most "losers" at day two are just noise.

Autopilot: a schedule of tests that runs itself

One test answers one question. Autopilot asks the next ten for you. From any campaign's menu, Autopilot reads your popup and builds an improvement plan: a queue of tests generated from your own design (a game opener, curiosity copy, a countdown, better timing), each one previewable side by side with your current popup before anything runs.