Imagine this: your Bitrix online store runs fine, but around lunchtime the site starts slowing down — pages take 10–15 seconds to load, managers complain, some visitors leave. You launch htop: CPU is pegged at 100%, but you can't tell what's eating it. An hour later, load drops. Next day, it repeats. Without historical data, you can't find the cause; you need to catch it during the peak. We've seen this dozens of times. A properly configured monitoring setup captures anomalies and ties them to events — agents, cron jobs, deploys. This lets you fix the problem permanently, not just fight fires. One hour of downtime for an online store can cost from 30,000 to 100,000 rubles, and for large e-commerce up to 200,000 rubles, so monitoring setup quickly pays for itself. Additionally, our clients save an average of 50,000 rubles monthly on server costs.
In 1C-Bitrix projects, the main CPU consumers are php-fpm, mysqld, and agents (b_agent). PHP-FPM executes code that renders pages and handles user actions. MySQL handles database queries — unoptimized queries can hog a core for a long time, leading to high I/O wait. Agents run background tasks: updating the search index, catalog export, order processing. Each agent adds microseconds, but with many or bad intervals they create constant load. It's not enough to look at current usage; you need to analyze trends. For that we use Prometheus + Grafana, and for quick checks — htop and mpstat. Using Prometheus, you spot CPU spikes 10 times faster than relying on htop alone. Set up CPU alerts in Prometheus to notify your team immediately when load exceeds thresholds.
Why CPU Monitoring Is Critical for a Bitrix Server
High CPU usage is a symptom, not the cause. Without monitoring, you only see the symptom but don't know which process caused it. For example, if MySQL is using CPU, the problem is in the database: slow queries, missing indexes, suboptimal structures. If php-fpm, look at the code and application settings. Bitrix agents can spike CPU at each run. According to official Bitrix documentation, misconfigured agents cause up to 30% of extra server load. Monitoring with alerts lets you know about such events in real time. Load average is a system metric showing the average number of processes waiting to run. In 70% of cases, a problematic server has a load average above the number of cores. We also monitor CPU steal time in virtualized environments and check for context switching spikes with pidstat -w. Additional technical considerations include NUMA locality issues (monitor with numactl --hardware), inode usage on high-traffic sites, and cgroups to limit CPU per service.
How to Tell PHP Problems from MySQL Problems
For quick diagnostics, we use these commands:
htop -d 20 mpstat -P ALL 5 pidstat -p <PID> 5 vmstat 1 # for context switches and interrupts If mysqld consumes >50% CPU, check slow queries:
SHOW FULL PROCESSLIST; SELECT * FROM information_schema.PROCESSLIST WHERE TIME > 5; PHP problems are often hidden in code: php-fpm status shows the number of active processes. If they're constantly at max, you need more workers. In 40% of cases, the fix is increasing pm.max_children. High context switching indicates excessive process turnover; use pidstat -w to confirm. Also, adjust oom_score of PHP processes to prevent OOM kills.
Which CPU Monitoring Tools Are Most Effective for Bitrix?
For on-the-spot diagnostics, htop is best — it shows processes and a tree view. mpstat gives per-core breakdown and is useful for detecting I/O wait issues. For long-term analysis and graphing, we use Prometheus + Grafana: they store months of metrics and reveal trends. atop saves history for the last few days and requires no setup — great for retrospection. We've tried dozens of tools and settled on this combo; it covers 95% of cases. For Prometheus, configure node_exporter's CPU collector with collect[] parameter to gather per-CPU metrics.
Step-by-Step Guide to Setting Up CPU Monitoring
Installing Prometheus + Grafana takes 2 to 4 hours, including alert configuration. Here's a typical plan:
-
Audit current state: check
uptime,cat /proc/loadavg, find top processes withps aux --sort=-%cpu | head -10. - Install node_exporter on the server: download and run it with permissions to read CPU, memory, disk metrics.
-
Configure Prometheus: in
prometheus.yml, add a target for node_exporter (e.g.,localhost:9100). - Import a dashboard in Grafana: use the Node Exporter Full template (ID 1860) or create your own with CPU utilisation, load average, I/O wait.
-
Configure Alertmanager: add a rule for high CPU load, e.g.,
(node_load1 / count(count(node_cpu_seconds_total) by (cpu))) > 2for 10 minutes. -
Test alerts: simulate load with
stress --cpu 4and verify triggering.
Typical Thresholds
| Metric | Critical Level | Action |
|---|---|---|
| Load average / cores | > 2.0 | Check processes, alert |
| I/O wait | > 10% | Look at disk usage, slow queries |
| CPU utilization mysqld | > 80% | Optimize queries |
| CPU utilization php-fpm | > 70% | Increase max_children or optimize code |
Quick Diagnosis Example
In one project, load average stayed at 3.5 with 2 cores. Analysis showed the search index update agent ran every 5 minutes. Moving it to cron with nice reduced load to 1.2. In other cases, load average decreased from 3.5 to 0.8, and page load time improved from 15 to 2 seconds.Quick Diagnosis Checklist for High CPU
- Check load average via
uptimeorcat /proc/loadavg. - For quick slowdown diagnosis, find top CPU processes:
ps aux --sort=-%cpu | head -10. - If mysqld leads, enable slow query log and analyze queries.
- If php-fpm, check
pm.status_pathfor a queue. - Examine agents:
SELECT NAME, LAST_EXEC, NEXT_EXEC FROM b_agent WHERE ACTIVE='Y' ORDER BY LAST_EXEC LIMIT 10; - Move heavy agents to cron with
nice -n 19and batch them. - Monitor CPU steal time, context switching rates, and NUMA effects.
- Use cgroups to throttle high-CPU processes.
Average server resource savings after monitoring implementation is 20–40% — by identifying inefficient queries and agents. For example, moving one heavy agent to proper cron can halve peak load. One hour of online store downtime can cost tens of thousands of rubles, so monitoring setup quickly pays off. On average, monitoring saves 20,000 rubles monthly on server resources.
What's Included in Our Monitoring Setup Service
We offer a turnkey service: from audit to documentation handover. Our guarantee: a measurable improvement in server performance within one week. A typical project includes:
- Server performance audit (CPU, memory, disk, PHP and MySQL settings).
- Installation and configuration of Prometheus + node_exporter + Grafana.
- Dashboard with key metrics (CPU, load average, I/O wait, top processes).
- Alert configuration for Telegram/Slack.
- Optimization of php-fpm (max_children, pm.max_requests) and agents (move to cron, intervals).
- Handover of documentation with instructions and access.
- One month of support after deployment.
Timeline: 2 to 5 days depending on complexity. Pricing is customized — we'll assess your project in one business day. Contact us to discuss details.
How We Do It: A Real-World Example
Recently, an online store with a 50,000-product catalog came to us. Every night, an agent updated the search index, loading CPU to 100% for an hour. This caused background task failures. We moved the agent to cron with nice 19 and batched updates per 1,000 products. Peak load dropped from 100% to 30%, and execution time fell from 60 to 12 minutes. We also set an alert for load average > 1.5 — now the team learns about issues before they affect users.
Work Stages
| Stage | Duration | Result |
|---|---|---|
| Express audit | 1 day | Report on current CPU metrics, bottlenecks |
| Design | 1 day | Optimization plan, tool selection |
| Monitoring setup | 1-2 days | Prometheus + Grafana, alerts |
| Optimization | 1-2 days | php-fpm, agents, MySQL (if needed) |
| Handover & training | 1 day | Documentation, dashboards, access |
Average savings on server rental after monitoring implementation is 20–40%, which for an average project means tens of thousands of rubles monthly. Get a consultation on configuring monitoring for your server — we'll propose the optimal solution for your infrastructure. Order an audit, and we'll identify the bottlenecks in your Bitrix server. Our experience: over 10 years and 500+ projects in Bitrix optimization and support.

