We faced a challenge: a SaaS product with $1M MRR had a 5% monthly churn. Each percentage point reduction means $120K additional ARR annually. But without an accurate model, retention efforts are blind—blanket discounts burn margins. Churn prediction solves this: a model identifies high-risk customers before they churn. We build systems that reduce churn by 20% in practice. Our guaranteed methodology delivers a proven track record with over 5 years of experience and 30+ successful projects.
Problems we solve
Blurred targeting. In non-contractual scenarios (e-commerce, games), there is no explicit churn label—you must define an inactivity threshold. For example, if a customer hasn't made a purchase for 90 days, consider them churned. Choosing the threshold is critical: at 30 days, 20% of customers get a label; at 90 days, only 5%.
Imbalanced classes. 2-10% churners vs. 90% non-churners. Without correction, a model achieves 90% accuracy but zero recall on churners.
Feature engineering. RFM metrics are the foundation, but you also need trends (activity change over 30 days), feature adoption rate, and support tickets. We use rolling window aggregations and diff features.
How we do it: stack and a case
Stack: LightGBM (baseline)—LightGBM is 10x faster than LSTM on tabular data with comparable quality. CatBoost for categorical features, LSTM if event sequences are critical. Feature store—PostgreSQL with pgvector for embeddings. MLflow for experiments, SHAP for interpretation.
Detailed case from our practice: Client—B2B SaaS with 50K users. Baseline LightGBM gave PR-AUC 0.31. After adding trend features (login frequency change over 30 days) — 0.41, +32%. Adding a sequence model (LSTM on event sequences) pushed it to 0.49, but with 4x latency. Production solution: an ensemble of LightGBM + LSTM with cascading scoring. Implementation cost: $15,000. Savings: $30,000–$50,000 per 10,000 customers, ROI of 10x in 6 months.
How to define churn in non-contractual scenarios?
Define an inactivity period after which a customer is considered churned. We choose X based on analysis of inter-purchase interval distribution. Typical values: 60-90 days for B2B SaaS, 90-180 for e-commerce. A wrong choice introduces noise into the target variable.
Why LightGBM is a good baseline for churn prediction?
LightGBM handles missing values, works with categorical features, and captures non-linear dependencies. On standard churn tasks, it beats logistic regression by 0.15–0.25 AUC-ROC and is 2-3x faster than XGBoost.
Development and deployment
Feature Engineering
RFM metrics (most important predictors):
- Recency: days since last action/transaction
- Frequency: number of sessions/purchases in 30/90/180 days
- Monetary: total spend over period
Behavioral features:
- Trend features: activity increase/decrease over last 30 days vs. previous 30
- Feature adoption rate: % of key product features used by the customer
- Support tickets: number, type, NPS after resolution
Contractual/demographic:
- Time since onboarding
- Plan type
- Segment (SMB / Enterprise)
- Acquisition channel
Algorithm selection
| Algorithm | When to use | Accuracy | Interpretability |
|---|---|---|---|
| Logistic Regression | Baseline, interpretability needed | Medium | High |
| LightGBM / XGBoost | Tabular data, no time series | High | Medium (SHAP) |
| CatBoost | Many categorical features | High | Medium |
| LSTM / Transformer | Event sequences matter | Very High | Low |
Recommendation: start with LightGBM as baseline, add Sequence Model if behavioral patterns are important.
Handling imbalanced classes
Methods to counter imbalance include using class weights (class_weight='balanced' in sklearn)—simplest fix; SMOTE generates synthetic minority examples but can introduce noise; Focal Loss in neural networks downweights easy examples; threshold tuning on Precision-Recall curve (not 0.5)—a free way to improve Precision@K. For evaluation, use weighted F1-score as primary metric, AUC-ROC for ranking, Precision@K for marketing—precision among top-K at-risk customers is most important. See Wikipedia on imbalanced classes.
Deployment and usage
Batch scoring: weekly model run on the entire customer base. Output: a table with churn probability for each customer. Segmentation: high risk (>0.7), medium risk (0.4-0.7), low risk (<0.4).
Real-time scoring: API endpoint POST /score, <100ms response, score updated in CRM in real time. Retention by segment:
- High risk: personal call from Customer Success or a discount
- Medium risk: automated email campaign with value reminders
- Low risk: no action (don't waste resources)
Measuring business impact
Uplift modeling is the correct way to measure real system value. A standard A/B test: 50% of high-risk customers receive retention (treatment), 50% do not (control). Measure the difference in churn rate. Companies using churn prediction reduce churn by 15-20%. Average savings from implementation: $30–50K per 10K customers, with implementation costs starting at $15,000.
Process and timeline
Process
- Analytics: collect and clean data, define churn, analyze distribution.
- Feature engineering: RFM, trends, adoption, contractual data.
- Modeling: baseline (LightGBM), experiments (CatBoost, LSTM), threshold tuning.
- Testing: offline (AUC, F1, Precision@K), online A/B uplift test.
- Deploy: weekly batch scoring, real-time API (<100ms), CRM integration.
- Monitor: data drift, model drift, automatic retraining.
What's included
- Report on churn definition (target selection)
- Baseline model (LightGBM) + SHAP report
- Feature and pipeline documentation
- Batch scoring integration into your CRM
- Team training (2 hours)
- 3 months post-deployment support
Estimated timeline
| Phase | Duration |
|---|---|
| Baseline model (RFM features) | 2-3 weeks |
| Full system with monitoring | 8-10 weeks |
First model with basic RFM features: 2-3 weeks. Full system with feature store, drift monitoring, and CRM integration: 8-10 weeks. We are a team with 5+ years of ML production experience and 30+ successful churn prediction projects. Contact us to evaluate your project and get precise timelines. Get a consultation on implementing churn prediction.
Example code for computing RFM features
import pandas as pd
def rfm_features(transactions, as_of_date):
"""Compute Recency, Frequency, Monetary for each customer."""
rfm = transactions.groupby('customer_id').agg(
recency=('transaction_date', lambda x: (as_of_date - x.max()).days),
frequency=('transaction_id', 'nunique'),
monetary=('amount', 'sum')
).reset_index()
return rfm







