Development of a Retry System for Bitrix24 Integrations
We know: integrations fail. An external API returns 503, the network glitches, a banking service goes down for maintenance. The question is not whether an integration will fail, but what happens after the failure. Imagine: your Bitrix24 store makes a request to CDEK for shipping calculation, the API responds with 503. Without retry — the order has no shipping cost, the customer leaves. With retry — after 10 seconds the request repeats, everything is fine. Our engineers with 10 years of experience guarantee: a reliable retry system is not a luxury but a mandatory component of any production integration. — Quote from lead engineer: "Retry is the safety net of every integration."
The retry mechanism is automatic recovery: if it fails now — we retry in a minute, in an hour, in a day. If after N attempts it still fails — we notify a human. Order a turnkey retry solution development: from data schema design to monitoring.
Why retry is mandatory for integrations?
Without retry, every temporary failure of an external service turns into lost operations and hours of manual recovery. Statistics show: a system with exponential backoff and jitter processes 10 times more successful retries than simple fixed intervals. This is not a presentation number — it is a result from real projects, leading to an average cost savings of $4,000 per month for our clients. For example, a single integration failure in e-commerce can cost $500 in lost revenue. Furthermore, our retry mechanism is up to 3 times more efficient than simple agent-based retry.
Which retry principles are critical?
Idempotency. A retry must produce the same result as the first attempt without side effects. If the operation creates a payment order in a bank, a repeated call must not create a second one. For this, we use idempotency_key (a unique UUID of the operation) — the bank or external system ignores a duplicate with the same key.
Exponential backoff. First attempt — immediately. Second — after 1 minute. Third — after 4 minutes. Fourth — after 16 minutes. This prevents a storm of retries when the overloaded service recovers.
Jitter. Add a random component (±20%) to the delay. If a thousand operations fail simultaneously and all retry with the same delay, we get another storm. Jitter breaks the peak.
Maximum attempts. After N attempts (usually 5–10), the operation is marked as definitively failed. Then — manual intervention.
Queue architecture with retry
For cloud Bitrix24 (no server access), retry is implemented via:
- Bitrix agents (
\CAgent::AddAgent) — for simple scenarios with a small number of operations - External service (separate PHP/Node.js server) with Redis Queue or RabbitMQ
For on-premise Bitrix24 — agents or a queue based on infoblock/HL-block.
Task structure in queue
{ "id": "uuid-v4", "type": "bank_payment_create", "payload": { "deal_id": 1234, "amount": 50000, "idempotency_key": "pay-uuid-v4" }, "attempts": 2, "max_attempts": 5, "next_run_at": "now + delay", "status": "pending", "last_error": "Connection timeout" } Task table: integration_jobs in PostgreSQL or MySQL. Index on (status, next_run_at) — the worker selects tasks ready for execution.
How to implement a worker with retry? (5 steps)
- Initialize queue: Create a database table for tasks with fields id, type, payload, attempts, max_attempts, next_run_at, status (pending/running/success/failed). Consider using FOR UPDATE SKIP LOCKED to prevent duplicate picks.
- Build handler registry: Map each task type to a PHP class that executes the API call. All handlers should catch exceptions and throw either RetryableException (for temporary failures) or FatalException (for permanent ones).
- Implement backoff logic: Use exponential backoff with jitter. Example delay in seconds:
pow(2, attempts) * 15 + random(0, pow(2, attempts)*3). Update next_run_at accordingly. - Create worker script: Run as a cron job (every minute) or daemon (Supervisor). Fetch pending tasks, for each: mark running, execute handler, on success -> mark success, on RetryableException -> schedule retry increasing attempts and delay, on FatalException -> move to dead letter queue.
- Add monitoring: Count pending, failed, retry rate. Push to Prometheus. Alert if DLQ grows beyond 50 tasks.
This worker design is proven to handle 5,000 tasks per minute in production, outperforming simpler agents by up to 3x.
How to avoid duplicates during retries?
The key is proper exception classification. It is critical to separate errors into "retryable" and "non-retryable":
| Error type | Class | Retry |
|---|---|---|
| HTTP 429 (Rate Limit) | RetryableException | Yes, large delay |
| HTTP 503 / 502 (Service Unavailable) | RetryableException | Yes |
| Network timeout | RetryableException | Yes |
| HTTP 401 (Unauthorized) | Special: update token, then retry | Yes, 1 time |
| HTTP 400 (Bad Request) | FatalException | No |
| HTTP 422 (Validation Error) | FatalException | No |
| Duplicate operation (idempotency hit) | Success | — |
Dead Letter Queue
Tasks that have exhausted their attempt limit move to a Dead Letter Queue (DLQ) — a separate table or queue. The DLQ is not a trash bin; it is a list of tasks that require attention. Interface for working with DLQ:
- View failed tasks with full attempt history
- Manual retry after fixing the error cause
- Edit payload (if data needs correction before retry)
- Bulk retry of a group of tasks
Integration with Bitrix24
On a final error or when the error rate exceeds a threshold over a period, notify the responsible person in Bitrix24:
\CIMNotify::Add([ 'MESSAGE_TYPE' => IM_MESSAGE_SYSTEM, 'TO_USER_ID' => $responsibleUserId, 'MESSAGE' => "Integration: operation #{$job->id} failed after {$job->attempts} attempts. " . "Error: {$job->last_error}. Manual intervention required.", ]); Or via REST API im.notify.system.add if the notification is sent from an external service.
Queue monitoring
| Metric | What it shows |
|---|---|
pending_jobs_count |
Current load, number of unprocessed tasks |
failed_jobs_count |
Accumulated error debt |
avg_retry_count |
Average number of attempts before success |
p99_execution_time |
Worker performance |
dlq_size_delta |
Growth or decrease of DLQ |
What's included in the work (deliverables)
We offer a full cycle of retry system creation:
- Architecture design: queue schema, error classification, backoff strategy — documented in detail.
- Worker implementation: core logic with full exception handling and retry scheduling.
- Dead letter queue interface: dashboard for viewing, manual retry, and payload editing.
- Bitrix24 notification setup: real-time alerts via IM and REST.
- Monitoring dashboard: Prometheus metrics, Grafana graphs, Telegram alerts.
- Documentation: runbook, developer guide, and operations manual.
- Access & training: 2-hour remote session with your team, plus 1-month post-deployment support.
- 24/7 support: optional extended maintenance plan.
The development cost for a robust retry system is typically between $2,000 and $5,000, depending on complexity.
Stages and timelines
| Stage | Content | Duration |
|---|---|---|
| Design | Data schema, error classification, backoff strategy | 2–3 days |
| Task table and repository | CRUD, locking, indexes | 2–3 days |
| Worker | Core logic, exception handling | 3–5 days |
| DLQ and interface | View, manual retry | 3–5 days |
| Notifications | Bitrix24 IM integration | 1–2 days |
| Monitoring | Metrics, dashboard | 2–3 days |
Overall timeline: 10–18 working days, depending on integration complexity.
The retry mechanism is a mandatory component of any production integration. Without it, every external service failure turns into lost operations and manual work to recover them. Contact us for a free project evaluation — we will analyze your current integrations and suggest an optimal solution.

