AI System for Customer Receipt Analysis
A customer leaves a store, scans the receipt through the Federal Tax Service app, and that's all you get—no data, no insights. The retail network loses granular details: purchase timing, item specifics, and more. We turn millions of unstructured transactions into actionable insights for increasing LTV and optimizing inventory. The system analyzes the basket, predicts churn, and suggests personalized offers—without manual processing. Order a pilot project for your network starting from $15,000. Implementation costs range from $15,000 to $50,000 depending on scale, with typical ROI achieved within 6 months.
Data Sources for Receipts
| Source | Format | Characteristics | Data Completeness Accuracy |
|---|---|---|---|
| Mobile app | JSON (FTS API) | Low transaction percentage, but high quality | 60–70% |
| Loyalty program | XML/CSV | Full coverage for identified customers | 95%+ |
| OFD (fiscal data operator) | JSON (API) | Streaming, 100% of receipts, but no customer identification | 99.9% |
| Taxpayer's personal account | JSON (FTS API) | Only individual receipts, legal entities not included | 50–80% |
Receipt Data Structure
A fiscal document in Russia follows a standardized format (FFD 1.05/1.1):
class Receipt(BaseModel): receipt_id: str fiscal_number: str date_time: datetime store_name: str store_inn: str cashier: str | None items: list[ReceiptItem] total_amount: Decimal payment_type: str # cash / card / QR class ReceiptItem(BaseModel): name: str # product name quantity: float price: Decimal amount: Decimal product_code: str | None # barcode ETL Pipeline Details
Data from all sources is aggregated in Apache Airflow. Deduplication is based on fiscal_number and receipt_id. Missing attributes (e.g., product_code) are recovered via fuzzy matching with the product catalog. The pipeline handles up to 10 million receipts per day.How do we normalize product names from chaos?
Product names in receipts are a mess: "MILK PAST.2.5% 1kg", "Milk past. 2.5% tetra pack 1L", "MILK HOUSE 2.5 1L"—all the same product. Standardization via LLM achieves >98% accuracy for frequent products, but latency is high. For bulk processing, we train a BERT classifier on a normalized corpus—inference is 100× faster than an LLM. The model is deployed in ONNX Runtime for CPU inference with p99 latency < 50 ms.
def normalize_product_name(raw_name: str) -> NormalizedProduct: return llm.parse(f"""Normalize the product name from a cash receipt. Extract: category, brand, volume/weight, fat content (for dairy). Raw name: {raw_name}""", response_format=NormalizedProduct ) How do we combine FP-Growth with gradient boosting?
Classic algorithms like Apriori and FP-Growth find static associations but ignore seasonality, trends, and price elasticity. We combine FP-Growth with gradient boosting (CatBoost) on features: day of week, promotions, weather. This increases recommendation lift by 15–20% compared to pure association rules. According to a Gartner report, implementing such hybrid analytics increases LTV by 20%.
Analytics based on transaction data includes:
- Market basket analysis: FP-Growth with filtering by support 0.01% and confidence 20%.
- Customer lifetime value: BG/NBD model with transaction prediction for 12 weeks.
- Price elasticity: bandit algorithms (Thompson Sampling) for A/B testing discounts.
- Category management: share of wallet dynamics by category with seasonal decomposition.
- Attrition prediction: gradient boosting on features: purchase frequency, receipt total, category diversity. ROC-AUC > 0.85.
Comparison of Basket Analysis Approaches
| Approach | Performance | Trend Consideration | Lift |
|---|---|---|---|
| Apriori | Slow on large data | No | 1.0x |
| FP-Growth | Fast, but no dynamics | No | 1.0x |
| FP-Growth + CatBoost | Moderate | Yes (temperature, promos) | 1.2x |
Visualization: Superset with dashboards for retail analysts without SQL—ready-to-use metrics by brands, categories, cohorts. Savings from assortment personalization can reach 15% of revenue.
What's Included in the Work
When ordering the system, you receive the following deliverables:
- Architectural documentation for pipelines and models.
- Configured dashboards in Superset with key metrics.
- Training for 2–3 employees on using the system.
- Support for 1 month after launch: model drift monitoring, alert configuration.
Implementation Stages
- ETL pipeline for receipts from OFD, mobile app, and loyalty program (Apache Airflow, Kafka – optional).
- ML models for standardization, categorization, forecasting (PyTorch, Hugging Face Transformers, CatBoost).
- Dashboards in Superset/Metabase: KPIs for sales, basket, assortment, churn.
- Team training (2–3 employees) and architecture documentation.
- Support for 1 month after launch: model drift monitoring, alert configuration.
Contact us and we will assess your project within 2 business days. Get a consultation on timelines and scope. We'll select a solution for your scale.
Our Experience and Guarantees
Years of experience in retail (from grocery chains to DIY). We guarantee NDA compliance and adherence to Federal Law 152-FZ. We hold certifications in PyTorch and MLflow. We can sign an SLA for model uptime of 99.9%.







