Recently, a client revamped their checkout — conversion dropped by 20%. We ran an A/B test that showed the old version performed 15% better. At full rollout, this would have cost the business $24,000 monthly. Split testing is the only way to make design decisions based on data, not intuition. Over our work, we've conducted more than 200 experiments for e-commerce stores, landing pages, and SaaS products. The average conversion lift is 15–30%. Some tests brought significant additional profit.
Without A/B tests, every change is a lottery. One of our clients spent $30,000 on a new homepage design that dropped conversion by 8%. A test would have shown this in 2 weeks, saving the entire budget.
Suppose you change a landing page without a test. If conversion drops by 10% with 1,000 daily visitors, that's 100 lost leads per day. In a month — 3,000 leads, each costing an average of $6 — losses are $18,000. The cost of the test is much lower: typically $600 to $1,800.
What problems does A/B testing solve?
Often, teams are confident that a new design or CTA will improve conversion, but statistics show the opposite. For example, we tested a button color change — expecting a 20% increase, we got a 5% drop. Technical errors also occur: incorrect user segmentation, data leakage between variants, improper tracking. Once we found that due to faulty implementation, 30% of users were in both variants — the test had to be restarted. And the classic trap: premature test stopping when the difference seems obvious but the sample size hasn't been reached. According to Nielsen Norman Group, 73% of tests are stopped early, leading to false conclusions.
How we do it
Each experiment follows the scheme: analytics → design → implementation → tracking → analysis. We use a modern stack: React 18, Next.js 14, TypeScript, Node.js, Docker. For data storage — PostgreSQL and Redis. Growthbook allows iterating hypotheses 3x faster compared to VWO: you write logic on the client or server, not through a visual editor. Bayesian statistics further improve decision-making under uncertainty.
Implementation via Vercel Edge Middleware
// middleware.ts import { NextResponse } from 'next/server'; import type { NextRequest } from 'next/server'; const EXPERIMENT_COOKIE = 'exp_checkout_v2'; const VARIANTS = ['control', 'variant-a', 'variant-b']; function assignVariant(): string { const rand = Math.random(); if (rand < 0.34) return 'control'; if (rand < 0.67) return 'variant-a'; return 'variant-b'; } export function middleware(request: NextRequest) { const response = NextResponse.next(); const existing = request.cookies.get(EXPERIMENT_COOKIE)?.value; if (existing && VARIANTS.includes(existing)) { return response; } const variant = assignVariant(); response.cookies.set(EXPERIMENT_COOKIE, variant, { maxAge: 60 * 60 * 24 * 30, httpOnly: true, sameSite: 'lax', }); response.headers.set('x-ab-checkout', variant); return response; } export const config = { matcher: ['/checkout/:path*'], }; // app/checkout/page.tsx import { cookies, headers } from 'next/headers'; export default function CheckoutPage() { const variant = headers().get('x-ab-checkout') ?? cookies().get('exp_checkout_v2')?.value ?? 'control'; return ( <> {variant === 'control' && <CheckoutV1 />} {variant === 'variant-a' && <CheckoutV2OneStep />} {variant === 'variant-b' && <CheckoutV2TwoStep />} <ABTracker experiment="checkout_v2" variant={variant} /> </> ); } Tracking results
// components/ABTracker.tsx (Client Component) 'use client'; import { useEffect } from 'react'; export function ABTracker({ experiment, variant }: { experiment: string; variant: string; }) { useEffect(() => { gtag('event', 'experiment_impression', { experiment_id: experiment, variant_id: variant, }); posthog.capture('$experiment_started', { '$experiment_id': experiment, '$variant_key': variant, }); }, [experiment, variant]); return null; } function trackConversion(variant: string) { gtag('event', 'purchase', { experiment_id: 'checkout_v2', variant_id: variant, value: orderTotal, }); } Statsig: fast integration
// Statsig SDK (server and client parts) import Statsig from 'statsig-node'; await Statsig.initialize(process.env.STATSIG_SERVER_KEY!); const experiment = Statsig.getExperiment( { userID: userId, email: userEmail }, 'checkout_redesign' ); const checkoutLayout = experiment.get('layout', 'single-page'); const ctaColor = experiment.get('cta_color', 'blue'); // Client side (React SDK) import { useExperiment } from 'statsig-react'; function PricingCTA() { const { config } = useExperiment('pricing_cta'); const buttonText = config.get('button_text', 'Get Started'); const buttonVariant = config.get('button_variant', 'primary'); return ( <Button variant={buttonVariant} onClick={() => { statsig.logEvent('cta_clicked', buttonText); }}> {buttonText} </Button> ); } Why statistical significance is critical
Without it, you risk mistaking random fluctuation for a win. Before launch, we calculate the required sample size using the frequentist approach:
# Python: sample size calculation from statsmodels.stats.power import zt_ind_solve_power baseline_rate = 0.03 expected_effect = 0.15 lift = baseline_rate * expected_effect n = zt_ind_solve_power( effect_size=lift / (baseline_rate * (1 - baseline_rate)) ** 0.5, alpha=0.05, power=0.8, ) print(f"Sample size per variant: {int(n)}") # ~12,000 Rule: do not stop the test before reaching the planned sample size, even if results look good. For a test with a 5% conversion and expected improvement of 10%, you need 6,500 users per variant — that's 2–3 weeks of traffic for an average site.
Tip: don't peek at results daily — it skews statistics. Automatically calculate p-value and stop the test only when the planned sample size is achieved. Use sequential testing if intermediate decisions are needed.
How to choose an A/B testing tool
| Tool | Type | Best for |
|---|---|---|
| Growthbook | Open source / SaaS | Technical teams, self-hosted |
| Statsig | SaaS | Quick start, analytics integration |
| Optimizely | Enterprise SaaS | Large companies, complex experiments |
| VWO | SaaS | Marketing teams without dev |
| Vercel Edge Experiments | PaaS | Next.js on Vercel |
| Custom implementation | - | Full control, minimal overhead |
Which metrics to track in an A/B test
| Metric | Type | Example |
|---|---|---|
| Primary | Target action | Conversion to purchase, sign-up |
| Secondary | Engagement | Time on site, page views |
| Business | Revenue, LTV | Average order value, retention |
| Guardrail | Risk | Bounce rate, errors |
All metrics must be defined before the experiment starts. Track them in GA4: use events experiment_impression and experiment_conversion.
What's included in the work
- Setting up an A/B testing tool for your stack (Growthbook, Statsig, VWO, Optimizely, or custom).
- Implementing variant distribution on backend/edge with consistency guarantees.
- Integrating event tracking into GA4, PostHog, Amplitude.
- Calculating required sample size and test duration.
- Documenting results and recommendations for further experiments.
- Training your team on how to run tests and interpret results.
Our process
- Analytics: study current metrics, identify bottlenecks, formulate a hypothesis.
- Design: choose the tool, define variants and success metrics.
- Implementation: integrate distribution and tracking, set up dashboards.
- Launch: start the test, monitor data correctness.
- Analysis: after reaching sample size — statistical checking, report generation.
Timeline: 2 to 4 business days for a simple test, 5–10 days for a complex one with custom logic. Cost is calculated individually.
Order A/B testing setup and get data-driven conversion growth. Contact us for a consultation on tool selection and experiment execution.







