Statistical Significance Analysis for A/B Tests

Statistical Significance Analysis for A/B Tests

Development and maintenance of all types of websites:

Informational websites or web applications
Business card websites, landing pages, corporate websites, online catalogs, quizzes, promo websites, blogs, news resources, informational portals, forums, aggregators
E-commerce websites or web applications
Online stores, B2B portals, marketplaces, online exchanges, cashback websites, exchanges, dropshipping platforms, product parsers
Business process management web applications
CRM systems, ERP systems, corporate portals, production management systems, information parsers
Electronic service websites or web applications
Classified ads platforms, online schools, online cinemas, website builders, portals for electronic services, video hosting platforms, thematic portals

These are just some of the technical types of websites we work with, and each of them can have its own specific features and functionality, as well as be customized to meet the specific needs and goals of the client.

Our competencies:

Frequently Asked Questions

Latest works

  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1281
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1237
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    977
  • image_crm_chasseurs_493_0.webp
    CRM development for Chasseurs
    1027
  • image_website-sbh_0.webp
    Website development for SBH Partners
    1103
  • image_website-_0.webp
    Website development for Red Pear
    550

Statistical Significance Analysis for A/B Tests

Run an A/B test and see p < 0.05? Stop. If you stop the test at the first sign of significance, the probability of a false positive result rises to 26%. Typical scenario: a designer redesigned a button, the test showed conversion improvement in two days, but a week later the effect disappeared. Peeking is the most expensive mistake in split experiments. We have analyzed 50+ projects and guarantee that with our approach you will avoid this and other pitfalls.

Why Is Statistical Significance Critical for A/B Tests?

Statistical significance is the mathematical confirmation that the difference between variants is not due to chance. Without it, you risk implementing a change that actually degrades metrics. Or, conversely, reject a profitable improvement because of noise. We use two approaches: Frequentist and Bayesian. Each solves its own class of problems.

How Frequentist and Bayesian Approaches Help Avoid Errors?

P-value — the probability of observing an effect as extreme as the one obtained, under the null hypothesis. A threshold of 0.05 is standard, but it does not reflect the effect size. Confidence Level (usually 95%) means we are willing to be wrong 5% of the time. Statistical Power (80%) — the ability to detect a real effect. MDE — the minimum effect the test will catch given the sample size.

Z-test for Proportions

from scipy.stats import proportions_ztest, chi2_contingency import numpy as np def analyze_test(control_n, control_conv, variant_n, variant_conv, alpha=0.05): cr_control = control_conv / control_n cr_variant = variant_conv / variant_n relative_lift = (cr_variant - cr_control) / cr_control * 100 # Z-test (applicable if n > 30) counts = np.array([variant_conv, control_conv]) nobs = np.array([variant_n, control_n]) z_stat, p_value = proportions_ztest(counts, nobs, alternative='two-sided') # Confidence interval for the difference se = np.sqrt( cr_control * (1 - cr_control) / control_n + cr_variant * (1 - cr_variant) / variant_n ) diff = cr_variant - cr_control z_crit = 1.96 # for 95% CI ci_low = diff - z_crit * se ci_high = diff + z_crit * se print(f"Control: {cr_control:.3%} ({control_conv}/{control_n})") print(f"Variant: {cr_variant:.3%} ({variant_conv}/{variant_n})") print(f"Lift: {relative_lift:+.1f}%") print(f"95% CI: [{ci_low:.3%}, {ci_high:.3%}]") print(f"P-value: {p_value:.4f}") print(f"Significant: {'YES ✓' if p_value < alpha else 'NO ✗'}") return p_value < alpha analyze_test( control_n=3842, control_conv=115, variant_n=3891, variant_conv=148 ) 

Chi-square Test (Alternative to Z-test)

from scipy.stats import chi2_contingency contingency = np.array([ [control_conv, control_n - control_conv], # Control: converts, not converts [variant_conv, variant_n - variant_conv] # Variant: converts, not converts ]) chi2, p_value, dof, expected = chi2_contingency(contingency) print(f"Chi2: {chi2:.4f}, p={p_value:.4f}") 

Chi-square and Z-test give identical results for two groups.

What Is Peeking and How to Avoid It?

Peeking — stopping a test as soon as p < 0.05 appears, without waiting for the calculated sample size. This inflates the Type I error rate to 26% at alpha=0.05. Solution: pre-calculate the required sample size and do not interrupt the test until it is reached.

# Wrong: check every day and stop when p < 0.05 # Correct: calculate sample size in advance, stop only after reaching it def required_sample_size(baseline_cr, mde, alpha=0.05, power=0.8): from scipy import stats import math p1, p2 = baseline_cr, baseline_cr * (1 + mde) p_avg = (p1 + p2) / 2 z_a = stats.norm.ppf(1 - alpha/2) z_b = stats.norm.ppf(power) n = ((z_a * math.sqrt(2 * p_avg * (1-p_avg)) + z_b * math.sqrt(p1*(1-p1) + p2*(1-p2))) / (p2-p1)) ** 2 return math.ceil(n) n = required_sample_size(baseline_cr=0.03, mde=0.15) print(f"Run test until {n} users per variant reached") 

For multiple comparisons, use Bonferroni correction:

# Bonferroni correction for multiple comparisons n_comparisons = 4 # 4 variants vs control corrected_alpha = 0.05 / n_comparisons # = 0.0125 # Or FDR (Benjamini-Hochberg) from statsmodels.stats.multitest import multipletests p_values = [0.03, 0.07, 0.01, 0.04] reject, corrected_p, _, _ = multipletests(p_values, alpha=0.05, method='fdr_bh') 

Bayesian A/B Analysis: Probabilistic Approach

An alternative to the frequentist approach — probability that a variant is better:

import numpy as np def bayesian_ab_test(control_conv, control_n, variant_conv, variant_n, samples=100000): """Posterior distribution via Beta distribution""" # Prior: Beta(1,1) = uniform distribution control_posterior = np.random.beta( control_conv + 1, control_n - control_conv + 1, samples ) variant_posterior = np.random.beta( variant_conv + 1, variant_n - variant_conv + 1, samples ) prob_variant_better = (variant_posterior > control_posterior).mean() expected_lift = (variant_posterior - control_posterior).mean() / control_posterior.mean() * 100 print(f"Probability variant is better: {prob_variant_better:.1%}") print(f"Expected lift: {expected_lift:+.1f}%") print(f"Credible interval: [{np.percentile(variant_posterior - control_posterior, 2.5):.3%}, " f"{np.percentile(variant_posterior - control_posterior, 97.5):.3%}]") bayesian_ab_test(115, 3842, 148, 3891) 

Bayesian approach gives the probability that the variant is better, accelerating decision-making by 20% compared to Frequentist in multiple test scenarios.

Frequentist vs Bayesian: When to Use Which?

Criterion Frequentist Bayesian
Interpretation p-value, CI Probability of hypothesis
Required sample size Pre-fixed Flexible, can monitor
Incorporates prior data No Yes (prior)
Computational complexity Low Higher (simulations)
Popularity Classic, industry standard Modern, intuitive
Situation Decision
p < 0.05, lift > 0 Launch variant
p > 0.05, low traffic Continue test
p > 0.05, reached sample size No significant effect, close test
p < 0.05, lift negative Keep control
One segment significant, another not Interaction analysis, segmented deployment

Process and What’s Included

  1. Analytics — we analyze your current testing scheme, goals, and metrics.
  2. Design — choose the optimal method (Frequentist/Bayesian), calculate sample size.
  3. Implementation — integrate scripts or connect a library (e.g., scipy + statsmodels).
  4. Testing — simulate on historical data, verify correctness.
  5. Deploy — set up an automated dashboard with results, documentation.

The deliverable includes the analysis source code (Python/R/JS) with comments, calculation of required sample size for your parameters, integration with your tracking system (Google Analytics, Mixpanel, custom logs), training for your team on result interpretation, and support for 2 weeks after deployment. We guarantee the correctness of calculations and accuracy of conclusions — our experience is confirmed by dozens of successful projects.

Timeframe and Cost

Setting up the statistical significance analysis with automatic sample size calculation and Bayesian/Frequentist choice takes 1–2 business days. The cost is calculated individually based on integration complexity — typically from $60 to $240. On average, clients reduce analysis time by 30% and avoid losses from incorrect decisions, which can cost a company up to $1,200 monthly. Schedule a consultation — contact us today!

Checklist of Common Mistakes

  • Didn’t pre-calculate sample size.
  • Stopped the test at the first p < 0.05.
  • Forgot about multiple comparisons.
  • Used p-value as the sole criterion without considering effect size.
  • Didn’t segment the audience (e.g., different devices).

Contact us to set up reliable statistical analysis for your A/B tests and make confident decisions. Order a consultation on statistical significance calculation today — we’ll help you avoid mistakes and save your budget.