Imagine you launched an A/B test via Optimizely but after a week noticed LCP increased by 300 ms due to a third-party script. Or you cannot get raw data — only aggregated graphs. And the license cost for 10,000 visitors is a significant monthly expense. A custom A/B testing platform solves all these problems: you control the code, data, and budget. Savings on licensing fees can reach up to 70%, with a payback period of 3–6 months. Moreover, a custom platform runs 2–3 times faster because there are no external scripts.
We have accumulated over 5 years of experience developing such systems for online stores with 1M+ monthly visitors — more than 50 projects. Our platform is built on a modular principle and easily adapts to any architecture.
Advantages of a Custom A/B Testing Platform
Ready-made tools accelerate the start but become expensive and inflexible at scale. A custom platform pays off after just a few tests — in one project, conversion increased by 15% after the first experiment. With annual use, the savings on licensing can cover the development cost within a few months.
How We Solve Deterministic Assignment
The key requirement of an A/B test: a user must always fall into the same experiment variant. To achieve this, we use hash-based assignment. A hash of user_id and experiment name modulo 100% determines the variant number. The result is saved in the database, ensuring consistency across multiple visits.
Here’s the table schema for storing experiments and assignments:
CREATE TABLE experiments ( id SERIAL PRIMARY KEY, slug VARCHAR(100) UNIQUE NOT NULL, name VARCHAR(255) NOT NULL, description TEXT, status VARCHAR(20) DEFAULT 'draft', -- draft, running, paused, completed traffic SMALLINT DEFAULT 100, -- % of traffic participating in the experiment start_at TIMESTAMPTZ, end_at TIMESTAMPTZ, created_at TIMESTAMPTZ DEFAULT NOW(), updated_at TIMESTAMPTZ DEFAULT NOW() ); CREATE TABLE experiment_variants ( id SERIAL PRIMARY KEY, experiment_id INTEGER REFERENCES experiments(id), slug VARCHAR(100) NOT NULL, name VARCHAR(255), weight SMALLINT DEFAULT 50, config JSONB DEFAULT '{}', UNIQUE(experiment_id, slug) ); CREATE TABLE user_assignments ( user_id BIGINT NOT NULL, experiment_id INTEGER REFERENCES experiments(id), variant_id INTEGER REFERENCES experiment_variants(id), assigned_at TIMESTAMPTZ DEFAULT NOW(), PRIMARY KEY (user_id, experiment_id) ); The PHP service implements deterministic distribution using crc32 and database persistence:
class ExperimentAssignmentService { public function getVariant(int $userId, string $experimentSlug): ?string { $experiment = $this->getActiveExperiment($experimentSlug); if (!$experiment) return null; $existing = $this->assignmentRepo->find($userId, $experiment['id']); if ($existing) return $existing['variant_slug']; $trafficBucket = $this->hashToBucket($userId, $experimentSlug . '_traffic'); if ($trafficBucket >= $experiment['traffic']) return null; $variantBucket = $this->hashToBucket($userId, $experimentSlug); $variant = $this->selectVariant($experiment['variants'], $variantBucket); $this->assignmentRepo->assign($userId, $experiment['id'], $variant['id']); $this->eventTracker->track($userId, 'experiment.assigned', [ 'experiment' => $experimentSlug, 'variant' => $variant['slug'], ]); return $variant['slug']; } private function hashToBucket(int $userId, string $salt): int { $hash = crc32($userId . '_' . $salt); return abs($hash) % 100; } private function selectVariant(array $variants, int $bucket): array { $cumulative = 0; foreach ($variants as $variant) { $cumulative += $variant['weight']; if ($bucket < $cumulative) return $variant; } return end($variants); } } Event Tracking in ClickHouse
We send all significant user actions with experiment context. To avoid slowing down the user experience, events are written asynchronously via a queue.
class ExperimentEventTracker { public function track(int $userId, string $event, array $properties = []): void { $activeVariants = $this->assignmentRepo->getUserVariants($userId); $payload = [ 'event' => $event, 'user_id' => $userId, 'session_id' => session_id(), 'occurred_at' => now()->toIso8601String(), 'experiments' => $activeVariants, 'properties' => $properties, ]; $this->queue->push(new TrackExperimentEvent($payload)); } } Data is stored in ClickHouse — a columnar DBMS optimized for analytical queries. This allows fast conversion calculations and report generation even with millions of events.
Results Computation: Z-test
After collecting data, we use a two-sided Z-test for proportions (Wikipedia). It indicates whether the difference between control and test group conversions is statistically significant. The minimum detectable effect (MDE) is configured in advance — for example, 5% at 80% power.
import numpy as np from scipy import stats def calculate_significance(control, treatment): p1 = control['conversions'] / control['users'] p2 = treatment['conversions'] / treatment['users'] n1, n2 = control['users'], treatment['users'] p_pool = (control['conversions'] + treatment['conversions']) / (n1 + n2) se = np.sqrt(p_pool * (1 - p_pool) * (1/n1 + 1/n2)) if se == 0: return {'error': 'Insufficient data'} z = (p2 - p1) / se p_value = 2 * (1 - stats.norm.cdf(abs(z))) diff = p2 - p1 se_diff = np.sqrt(p1*(1-p1)/n1 + p2*(1-p2)/n2) ci = [diff - 1.96*se_diff, diff + 1.96*se_diff] return {'significant': p_value < 0.05, 'p_value': round(p_value, 6), 'lift': round((p2-p1)/p1*100,2) if p1>0 else None} To ensure reliable results, we also check for Sample Ratio Mismatch (SRM) — whether the actual user distribution deviates from expected. If the chi-square test p-value is below 0.01, the data is flagged as unreliable.
| Parameter | Calculation |
|---|---|
| Conversion | conversions / users |
| Lift | (p2-p1)/p1 * 100% |
| Confidence Interval | p ± 1.96 * SE |
| Stage | Duration | Result |
|---|---|---|
| Analytics | 1–2 days | Goals, metrics, architecture |
| Design | 2–3 days | DB schema, API, contracts |
| Implementation | 7–10 days | Code, tests |
| Testing | 2 days | Unit, integration, load |
| Deployment | 1–2 days | Release, monitoring |
How Feature Flags Work in A/B Tests?
A/B testing and feature flags are related concepts. We integrate them as follows: an experiment variant contains a JSON configuration that influences feature behavior. For example, variant treatment_a enables {"checkout_steps": 1, "show_trust_badges": true}. The frontend or backend code simply reads this config and changes behavior.
$variant = $experimentService->getVariant($userId, 'checkout-redesign'); $config = $experimentService->getVariantConfig('checkout-redesign', $variant); $checkoutSteps = $config['checkout_steps'] ?? 3; What's Included
- Development of assignment service with hash-based distribution and unit tests
- Event tracking with asynchronous ClickHouse writes
- Statistical significance computation (Z-test, confidence intervals)
- Admin panel for launching and monitoring experiments
- Integration documentation and team training
- First month of pilot launch support
Our Work Process
- Analytics — we analyze your goals, metrics, current architecture (1–2 days)
- Design — prepare DB schema, API, contracts (2–3 days)
- Implementation — write code, write tests (7–10 days)
- Testing — unit tests, integration testing, load (2 days)
- Deployment — deploy to your environment, configure monitoring (1–2 days)
For any inquiries, contact us — we will help estimate the work volume and calculate savings.
Timeline and Budget
Estimated development time: 14 to 21 days. The cost is calculated individually depending on integration complexity and additional requirements. Order a custom platform today — we will provide a proposal within 2 business days.
We guarantee code quality: we use code reviews, test coverage of at least 80%, and provide a 30-day bug fix warranty after launch. Get a consultation and find out how a custom A/B testing platform can improve your conversions without performance trade-offs.







