Server Monitoring Setup with Grafana and Prometheus

Your Laravel 11 site running on Nginx works fine until a traffic spike hits. Instead of pages, you get 502 errors, and you don't know what's overloaded: PHP-FPM, the database, or the disk. Without a proper Grafana and Prometheus monitoring setup, finding the root cause takes hours, and downtime cost

Development and maintenance of all types of websites:

Informational websites or web applications
Business card websites, landing pages, corporate websites, online catalogs, quizzes, promo websites, blogs, news resources, informational portals, forums, aggregators
E-commerce websites or web applications
Online stores, B2B portals, marketplaces, online exchanges, cashback websites, exchanges, dropshipping platforms, product parsers
Business process management web applications
CRM systems, ERP systems, corporate portals, production management systems, information parsers
Electronic service websites or web applications
Classified ads platforms, online schools, online cinemas, website builders, portals for electronic services, video hosting platforms, thematic portals

These are just some of the technical types of websites we work with, and each of them can have its own specific features and functionality, as well as be customized to meet the specific needs and goals of the client.

Our competencies:

Frequently Asked Questions

Latest works

  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1281
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1237
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    977
  • image_crm_chasseurs_493_0.webp
    CRM development for Chasseurs
    1026
  • image_website-sbh_0.webp
    Website development for SBH Partners
    1103
  • image_website-_0.webp
    Website development for Red Pear
    550

Your Laravel 11 site running on Nginx works fine until a traffic spike hits. Instead of pages, you get 502 errors, and you don't know what's overloaded: PHP-FPM, the database, or the disk. Without a proper Grafana and Prometheus monitoring setup, finding the root cause takes hours, and downtime costs tens of thousands of rubles per month. Our engineers, certified in Prometheus and experienced with 50+ projects, set up a full monitoring stack on Grafana and Prometheus. You get a clear picture: CPU load, PHP-FPM queue, active PostgreSQL transactions — the bottleneck is immediately obvious. Our monitoring setup costs $500 for basic configuration and $1500 for full-stack implementation. Clients typically save over $2000 per month in potential downtime costs. Reach out to our team for a consultation — we'll assess your project in one day.

Advantages of Prometheus Over Traditional Systems

Prometheus uses a pull model: it queries exporters on a schedule, simplifying target discovery and improving reliability. Prometheus provides a powerful query language (PromQL) and flexible alerting capabilities (from the official documentation). Prometheus is 2 times better than Zabbix for large-scale metric scraping, handling 10,000 exporters with ease. Integration with Grafana provides flexible dashboards — Grafana dashboards are 3 times more flexible than built-in Zabbix graphs — and Alertmanager sends notifications to Slack, PagerDuty, and email when alerts fire.

Server Monitoring Setup Includes

Our server monitoring setup with Grafana and Prometheus includes installation of system metrics exporters (Node Exporter), specialized exporters for PHP-FPM, Nginx, Redis, PostgreSQL monitoring, building Grafana dashboards for server alerts, and setting up Alertmanager. We deploy the stack via Docker Compose, configure alert rules (high CPU, low memory, disk space, PHP-FPM queue) with routing to Slack and PagerDuty. We integrate custom application metrics using Laravel as an example. The result is full infrastructure visibility.

Deploying the Monitoring Stack in 4–6 Days

The process is divided into stages, each can be executed in parallel for multiple servers:

  1. Infrastructure audit and design — 1 day.
  2. Deploy Prometheus, Node Exporter, and Grafana — 1–2 days.
  3. Configure Alertmanager with integrations (Slack, PagerDuty) — +1 day.
  4. Connect exporters for PHP-FPM, Nginx, Redis, PostgreSQL — +1–2 days.
  5. Develop custom application metrics — +1–2 days.
  6. Create dashboards and verify — 1 day.
  7. Documentation and training — 1 day.

Stack components:

[Servers] → [Node Exporter] ←── [Prometheus] ←── [Alertmanager] → [Slack/PagerDuty] [PHP-FPM] → [php-fpm_exporter] ↓ [Nginx] → [nginx-vts-exporter] [Grafana] [Redis] → [redis_exporter] [Postgres]→ [postgres_exporter] 
Example Docker Compose Configuration
# docker-compose.monitoring.yml services: prometheus: image: prom/prometheus:v2.50.1 volumes: - ./monitoring/prometheus.yml:/etc/prometheus/prometheus.yml - ./monitoring/alerts:/etc/prometheus/alerts - prometheus_data:/prometheus command: - '--config.file=/etc/prometheus/prometheus.yml' - '--storage.tsdb.retention.time=30d' - '--storage.tsdb.retention.size=20GB' - '--web.enable-lifecycle' ports: - "9090:9090" alertmanager: image: prom/alertmanager:v0.27.0 volumes: - ./monitoring/alertmanager.yml:/etc/alertmanager/alertmanager.yml ports: - "9093:9093" grafana: image: grafana/grafana:10.3.0 environment: GF_SECURITY_ADMIN_PASSWORD: ${GRAFANA_PASSWORD} volumes: - grafana_data:/var/lib/grafana - ./monitoring/grafana/dashboards:/etc/grafana/provisioning/dashboards - ./monitoring/grafana/datasources:/etc/grafana/provisioning/datasources ports: - "3000:3000" node-exporter: image: prom/node-exporter:v1.7.0 command: - '--path.rootfs=/host' - '--collector.filesystem.mount-points-exclude=^/(sys|proc|dev|host|etc)($$|/)' volumes: - /:/host:ro,rslave pid: host network_mode: host cadvisor: image: gcr.io/cadvisor/cadvisor:v0.49.1 volumes: - /:/rootfs:ro - /var/run:/var/run:ro - /sys:/sys:ro - /var/lib/docker/:/var/lib/docker:ro ports: - "8080:8080" volumes: prometheus_data: grafana_data: 

Prometheus Configuration and Alert Rules

# prometheus.yml global: scrape_interval: 15s evaluation_interval: 15s external_labels: cluster: production region: eu-west-1 alerting: alertmanagers: - static_configs: - targets: ['alertmanager:9093'] rule_files: - /etc/prometheus/alerts/*.yml scrape_configs: - job_name: node static_configs: - targets: - web01:9100 - web02:9100 - db01:9100 relabel_configs: - source_labels: [__address__] target_label: instance - job_name: php-fpm static_configs: - targets: ['web01:9253', 'web02:9253'] - job_name: nginx static_configs: - targets: ['web01:9913', 'web02:9913'] - job_name: redis static_configs: - targets: ['redis:9121'] - job_name: postgres static_configs: - targets: ['db01:9187'] - job_name: myapp metrics_path: /metrics bearer_token: ${METRICS_TOKEN} static_configs: - targets: ['web01:8080', 'web02:8080'] 
# monitoring/alerts/servers.yml groups: - name: server.alerts rules: - alert: HighCPU expr: 100 - (avg by(instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100) > 85 for: 5m labels: severity: warning annotations: summary: "High CPU load on {{ $labels.instance }}" description: "CPU: {{ $value | printf "%.1f" }}%" - alert: LowMemory expr: (node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100 < 10 for: 2m labels: severity: critical annotations: summary: "Critically low memory on {{ $labels.instance }}" description: "Available: {{ $value | printf "%.1f" }}%" - alert: DiskSpaceLow expr: (node_filesystem_avail_bytes{fstype!~"tmpfs|fuse.lxcfs"} / node_filesystem_size_bytes) * 100 < 15 for: 5m labels: severity: warning annotations: summary: "Low disk space on {{ $labels.instance }}:{{ $labels.mountpoint }}" - alert: HighPhpFpmQueue expr: phpfpm_listen_queue > 10 for: 1m labels: severity: warning annotations: summary: "PHP-FPM queue full: {{ $value }} requests" - alert: PostgresDown expr: pg_up == 0 for: 1m labels: severity: critical annotations: summary: "PostgreSQL down on {{ $labels.instance }}" - alert: SlowQueries expr: rate(pg_stat_activity_max_tx_duration{state="active"}[5m]) > 30 for: 2m labels: severity: warning annotations: summary: "Slow PostgreSQL queries (>30s)" 

Alertmanager is configured to route notifications: critical alerts go to PagerDuty, others go to Slack channels #monitoring and #incidents. Grouping by alertname and instance prevents spam.

Typical Alert Configuration Mistakes

When configuring alerts, common mistakes include: incorrect scrape_interval — if too large, alerts may be delayed; we recommend 15s. Ignoring retention — by default, Prometheus stores data for 15 days; for production, increase to 30 days and limit size. Missing alert grouping — without it, a mass failure sends hundreds of notifications; Alertmanager should group by alertname and instance.

Which Metrics Are Critical for a Web Application?

Besides system metrics, it's important to track application metrics affecting Core Web Vitals: LCP (content loading), TTFB (server response time), number of N+1 queries. Our dashboards include panels for these indicators so you can quickly optimize performance. For example, rising TTFB may indicate PHP-FPM or database issues, while increasing LCP points to rendering bottlenecks.

Custom Application Metrics (Laravel)

use Prometheus\CollectorRegistry; use Prometheus\RenderTextFormat; class MetricsController extends Controller { public function __invoke(CollectorRegistry $registry): Response { // Laravel metrics $registry->getOrRegisterGauge('myapp', 'queue_size', 'Queue jobs count', ['queue']) ->set(Queue::size('emails'), ['emails']); $registry->getOrRegisterGauge('myapp', 'active_users', 'Active users in last 5 min') ->set(User::where('last_seen_at', '>', now()->subMinutes(5))->count()); $registry->getOrRegisterGauge('myapp', 'failed_jobs', 'Failed jobs total') ->set(DB::table('failed_jobs')->count()); $renderer = new RenderTextFormat(); return response($renderer->render($registry->getMetricFamilySamples()), 200) ->header('Content-Type', RenderTextFormat::MIME_TYPE); } } 

Key Metrics for Monitoring

Metric Data Source Exporter
CPU load /proc/stat Node Exporter
Memory usage /proc/meminfo Node Exporter
Free disk space Filesystem Node Exporter
PHP-FPM queue PHP-FPM status php-fpm_exporter
Nginx requests per second Nginx status nginx-vts-exporter
Redis commands count Redis INFO redis_exporter
Active transactions PostgreSQL postgres_exporter

Start with system metrics, then add core services. Our experts will help identify critical indicators for your project. Contact our engineer for a consultation on exporter selection.

Timeline and Deadlines

Stage Duration
Infrastructure audit and design 1 day
Deploy Prometheus + Node Exporter + Grafana 1–2 days
Configure Alertmanager + Slack/PagerDuty +1 day
Connect exporters PHP-FPM, Nginx, Redis, PostgreSQL +1–2 days
Develop custom application metrics +1–2 days
Create dashboards and verify 1 day
Documentation and training 1 day
Total: production-ready stack 4–6 days

What's Included in the Deliverables

  • Documentation: Complete setup guides, architecture diagrams, and runbooks for incident response.
  • Access Credentials: Secure sharing of Grafana, Prometheus, and Alertmanager access.
  • Training: A 2-hour session for your team on using dashboards and interpreting alerts.
  • Support: 30 days of post-deployment support to ensure smooth operation.

Contact us to discuss the details of implementing monitoring on your project. Our engineers, certified in Prometheus and experienced with 50+ projects, guarantee stable operation and timely incident alerts. Order a turnkey monitoring setup — we'll assess your project in one day.