Setting Up A/B Testing in a Mobile App
You launch an A/B test, see p-value 0.04 after three days, and stop the experiment. The result is a false positive. This happens in 80% of mobile A/B tests, according to analytics. The cause is basic statistical violations: multiple checking and premature stopping. A/B testing is a powerful tool, but only when set up correctly. Our experience of 5+ years in mobile development — more than 50 implemented A/B tests on iOS and Android. We guarantee correct setup and statistically significant results. Contact us for a consultation.
How to Choose an A/B Testing Tool?
| Tool | Suitable for | Drawback |
|---|---|---|
| Firebase A/B Testing | Simple UI/text/parameters | Limited targeting flexibility |
| Amplitude Experiment | Product hypotheses with retention analysis | Paid, requires Amplitude Analytics |
| Statsig | Full cycle: flags, experiments, analysis | Requires setup |
| Growthbook | Open-source, self-hosted | Infrastructure costs |
Firebase A/B Testing is a reasonable start for most projects. Integration via Remote Config, no extra SDK. For complex segmentation (users from Moscow with three sessions), Statsig is three times more effective due to stratified sampling.
Why Most A/B Tests Produce False Positives?
Stopping the test at the first significant result is the most common mistake. If you look at p-value every day and stop when p < 0.05 for the first time, the false positive rate can rise to 30%. The test should be stopped only when a pre-determined sample size is reached.
One test — one metric. You cannot simultaneously optimize conversion rate and session length with one test. If both metrics improve, that's good, but the target should be single.
Novelty effect. A new design gives a spike in clicks in the first week simply because it's new. For behavioral tests, minimum duration is 2 weeks. For retention tests — 4 weeks.
When Should You Stop an A/B Test?
The test should be stopped only when the pre-calculated sample size is reached. Do not rely on current p-value. Use a sample size calculator: enter baseline conversion, minimum detectable effect (MDE), and confidence interval (typically 95%). For example, for a baseline of 10% and MDE of 1%, you need about 10,000 users per variant.
Firebase A/B Testing: Setup
Firebase A/B Testing is built on top of Remote Config. First, define a parameter:
// Get value from Remote Config let remoteConfig = RemoteConfig.remoteConfig() remoteConfig.configSettings = RemoteConfigSettings() remoteConfig.configSettings.minimumFetchInterval = 0 // in debug remoteConfig.fetchAndActivate { status, error in let ctaText = remoteConfig.configValue(forKey: "checkout_cta_text").stringValue self.checkoutButton.setTitle(ctaText, for: .normal) } In Firebase Console → A/B Testing, create an experiment:
- Select
checkout_cta_textas Target Parameter - Control: "Place Order"
- Variant A: "Buy Now"
- Target metric:
purchase(conversion event) - Percentage of participants: 50%
- Minimum sample size: Firebase calculates automatically
Statistical Significance: Hidden Pitfalls
| Problem | Consequence | Solution |
|---|---|---|
| Multiple checking (peeking) | False positives | Predefine sample size |
| Multiple metrics | Over-optimization | Choose one primary metric |
| Novelty effect | Inflated results | At least 2 weeks for UI tests |
Statsig for Complex Experiments
When more flexible segmentation is needed (test only on users from Moscow with > 3 sessions):
// iOS Statsig SDK import StatsigSDK Statsig.initialize(sdkKey: "client-xxx") { let experiment = Statsig.getExperiment("checkout_flow_v2") let variant = experiment.getValue(forKey: "flow_type", defaultValue: "standard") if variant == "simplified" { self.showSimplifiedCheckout() } else { self.showStandardCheckout() } } // Android val experiment = Statsig.getExperiment("checkout_flow_v2") val flowType = experiment.getString("flow_type", "standard") Statsig supports stratified sampling — even distribution of users across strata (platform, country, subscription plan). Without stratification, random distribution can create cohorts with different composition, distorting results.
Exposure Logging
For correct analysis, it is important to log the fact that a variant was shown — not just conversions:
Analytics.logEvent("experiment_exposure", parameters: [ "experiment_id": "checkout_cta_v2", "variant": variantName, "user_id": userId ]) This allows analyzing conversion only among users who actually saw the experiment, not all participants.
What's Included in the Work
- Tool selection tailored to tasks and tech stack (Firebase / Statsig / Amplitude Experiment)
- SDK integration and Remote Config / Feature Flags setup
- Implementing A/B layer in code with correct variant handling
- Configuring target metrics and conversion events
- Sample size and test duration configuration
- Exposure logging for analysis
- Post-test analysis with statistical assumption checks
- Documentation of results
Timeline
A single A/B test on Firebase Remote Config: 1–2 days. Infrastructure for regular A/B testing (Statsig/Growthbook): 3–5 days. Pricing is calculated individually. We'll evaluate your project — contact us for a consultation. We guarantee transparent reporting and correct experiments. Request A/B testing implementation and get statistically significant results without false positives.







