Server Monitoring Setup with Grafana and Prometheus

When servers run without failures, you don't think about it. But a traffic spike can turn your site into a 502 error, and every hour of downtime hits your business. We set up monitoring with Grafana and Prometheus so you see the real picture of load and bottlenecks. Our team delivers the project turnkey—from installing exporters to alerts in messengers—with ongoing support.

Development and maintenance of all types of websites:

Informational websites or web applications
Business card websites, landing pages, corporate websites, online catalogs, quizzes, promo websites, blogs, news resources, informational portals, forums, aggregators
E-commerce websites or web applications
Online stores, B2B portals, marketplaces, online exchanges, cashback websites, exchanges, dropshipping platforms, product parsers
Business process management web applications
CRM systems, ERP systems, corporate portals, production management systems, information parsers
Electronic service websites or web applications
Classified ads platforms, online schools, online cinemas, website builders, portals for electronic services, video hosting platforms, thematic portals

These are just some of the technical types of websites we work with, and each of them can have its own specific features and functionality, as well as be customized to meet the specific needs and goals of the client.

Our competencies:

Frequently Asked Questions

Latest works

  • Development of a web application for FEEDME
    Development of a web application for FEEDME
    1342
  • Development of an online store for the company FURNORO
    Development of an online store for the company FURNORO
    1304
  • Development of a web application for Enviok
    Development of a web application for Enviok
    1047
  • CRM development for Chasseurs
    CRM development for Chasseurs
    1094
  • Website development for SBH Partners
    Website development for SBH Partners
    1169
  • Website development for Red Pear
    Website development for Red Pear
    593

Your Laravel 11 site running on Nginx works fine until a traffic spike hits. Instead of pages, you get 502 errors, and you don't know what's overloaded: PHP-FPM, the database, or the disk. Without a proper Grafana and Prometheus monitoring setup, finding the root cause takes hours, and downtime costs tens of about $9–13 in savings per month. Our engineers, certified in Prometheus and experienced with 50+ projects, set up a full monitoring stack on Grafana and Prometheus. You get a clear picture: CPU load, PHP-FPM queue, active PostgreSQL transactions — the bottleneck is immediately obvious. Our monitoring setup costs $500 for basic configuration and $1500 for full-stack implementation. Clients typically save over $2000 per month in potential downtime costs. Reach out to our team for a consultation — we'll assess your project in one day.

Advantages of Prometheus Over Traditional Systems

Prometheus uses a pull model: it queries exporters on a schedule, simplifying target discovery and improving reliability. Prometheus provides a powerful query language (PromQL) and flexible alerting capabilities (from the official documentation). Prometheus is 2 times better than Zabbix for large-scale metric scraping, handling 10,000 exporters with ease. Integration with Grafana provides flexible dashboards — Grafana dashboards are 3 times more flexible than built-in Zabbix graphs — and Alertmanager sends notifications to Slack, PagerDuty, and email when alerts fire.

Server Monitoring Setup Includes

Our server monitoring setup with Grafana and Prometheus includes installation of system metrics exporters (Node Exporter), specialized exporters for PHP-FPM, Nginx, Redis, PostgreSQL monitoring, building Grafana dashboards for server alerts, and setting up Alertmanager. We deploy the stack via Docker Compose, configure alert rules (high CPU, low memory, disk space, PHP-FPM queue) with routing to Slack and PagerDuty. We integrate custom application metrics using Laravel as an example. The result is full infrastructure visibility.

Deploying the Monitoring Stack in 4–6 Days

The process is divided into stages, each can be executed in parallel for multiple servers:

  1. Infrastructure audit and design — 1 day.
  2. Deploy Prometheus, Node Exporter, and Grafana — 1–2 days.
  3. Configure Alertmanager with integrations (Slack, PagerDuty) — +1 day.
  4. Connect exporters for PHP-FPM, Nginx, Redis, PostgreSQL — +1–2 days.
  5. Develop custom application metrics — +1–2 days.
  6. Create dashboards and verify — 1 day.
  7. Documentation and training — 1 day.

Stack components:

[Servers] → [Node Exporter] ←── [Prometheus] ←── [Alertmanager] → [Slack/PagerDuty]
[PHP-FPM] → [php-fpm_exporter]
↓
[Nginx] → [nginx-vts-exporter]
[Grafana]
[Redis] → [redis_exporter]
[Postgres]→ [postgres_exporter]
Example Docker Compose Configuration
# docker-compose.monitoring.yml
services:
  prometheus:
    image: prom/prometheus:v2.50.1
    volumes:
      - ./monitoring/prometheus.yml:/etc/prometheus/prometheus.yml
      - ./monitoring/alerts:/etc/prometheus/alerts
      - prometheus_data:/prometheus
    command:
      - '--config.file=/etc/prometheus/prometheus.yml'
      - '--storage.tsdb.retention.time=30d'
      - '--storage.tsdb.retention.size=20GB'
      - '--web.enable-lifecycle'
    ports:
      - "9090:9090"
  alertmanager:
    image: prom/alertmanager:v0.27.0
    volumes:
      - ./monitoring/alertmanager.yml:/etc/alertmanager/alertmanager.yml
    ports:
      - "9093:9093"
  grafana:
    image: grafana/grafana:10.3.0
    environment:
      GF_SECURITY_ADMIN_PASSWORD: ${GRAFANA_PASSWORD}
    volumes:
      - grafana_data:/var/lib/grafana
      - ./monitoring/grafana/dashboards:/etc/grafana/provisioning/dashboards
      - ./monitoring/grafana/datasources:/etc/grafana/provisioning/datasources
    ports:
      - "3000:3000"
  node-exporter:
    image: prom/node-exporter:v1.7.0
    command:
      - '--path.rootfs=/host'
      - '--collector.filesystem.mount-points-exclude=^/(sys|proc|dev|host|etc)($$|/)'
    volumes:
      - /:/host:ro,rslave
    pid: host
    network_mode: host
  cadvisor:
    image: gcr.io/cadvisor/cadvisor:v0.49.1
    volumes:
      - /:/rootfs:ro
      - /var/run:/var/run:ro
      - /sys:/sys:ro
      - /var/lib/docker/:/var/lib/docker:ro
    ports:
      - "8080:8080"
volumes:
  prometheus_data:
  grafana_data:

Prometheus Configuration and Alert Rules

# prometheus.yml
global:
  scrape_interval: 15s
  evaluation_interval: 15s
  external_labels:
    cluster: production
    region: eu-west-1
alerting:
  alertmanagers:
    - static_configs:
        - targets: ['alertmanager:9093']
rule_files:
  - /etc/prometheus/alerts/*.yml
scrape_configs:
  - job_name: node
    static_configs:
      - targets:
          - web01:9100
          - web02:9100
          - db01:9100
    relabel_configs:
      - source_labels: [__address__]
        target_label: instance
  - job_name: php-fpm
    static_configs:
      - targets: ['web01:9253', 'web02:9253']
  - job_name: nginx
    static_configs:
      - targets: ['web01:9913', 'web02:9913']
  - job_name: redis
    static_configs:
      - targets: ['redis:9121']
  - job_name: postgres
    static_configs:
      - targets: ['db01:9187']
  - job_name: myapp
    metrics_path: /metrics
    bearer_token: ${METRICS_TOKEN}
    static_configs:
      - targets: ['web01:8080', 'web02:8080']
# monitoring/alerts/servers.yml
groups:
  - name: server.alerts
    rules:
      - alert: HighCPU
        expr: 100 - (avg by(instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100) > 85
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "High CPU load on {{ $labels.instance }}"
          description: "CPU: {{ $value | printf "%.1f" }}%"
      - alert: LowMemory
        expr: (node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100 < 10
        for: 2m
        labels:
          severity: critical
        annotations:
          summary: "Critically low memory on {{ $labels.instance }}"
          description: "Available: {{ $value | printf "%.1f" }}%"
      - alert: DiskSpaceLow
        expr: (node_filesystem_avail_bytes{fstype!~"tmpfs|fuse.lxcfs"} / node_filesystem_size_bytes) * 100 < 15
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "Low disk space on {{ $labels.instance }}:{{ $labels.mountpoint }}"
      - alert: HighPhpFpmQueue
        expr: phpfpm_listen_queue > 10
        for: 1m
        labels:
          severity: warning
        annotations:
          summary: "PHP-FPM queue full: {{ $value }} requests"
      - alert: PostgresDown
        expr: pg_up == 0
        for: 1m
        labels:
          severity: critical
        annotations:
          summary: "PostgreSQL down on {{ $labels.instance }}"
      - alert: SlowQueries
        expr: rate(pg_stat_activity_max_tx_duration{state="active"}[5m]) > 30
        for: 2m
        labels:
          severity: warning
        annotations:
          summary: "Slow PostgreSQL queries (>30s)"

Alertmanager is configured to route notifications: critical alerts go to PagerDuty, others go to Slack channels #monitoring and #incidents. Grouping by alertname and instance prevents spam.

Typical Alert Configuration Mistakes

When configuring alerts, common mistakes include: incorrect scrape_interval — if too large, alerts may be delayed; we recommend 15s. Ignoring retention — by default, Prometheus stores data for 15 days; for production, increase to 30 days and limit size. Missing alert grouping — without it, a mass failure sends hundreds of notifications; Alertmanager should group by alertname and instance.

Which Metrics Are Critical for a Web Application?

Besides system metrics, it's important to track application metrics affecting Core Web Vitals: LCP (content loading), TTFB (server response time), number of N+1 queries. Our dashboards include panels for these indicators so you can quickly optimize performance. For example, rising TTFB may indicate PHP-FPM or database issues, while increasing LCP points to rendering bottlenecks.

Custom Application Metrics (Laravel)

use Prometheus\CollectorRegistry;
use Prometheus\RenderTextFormat;

class MetricsController extends Controller
{
    public function __invoke(CollectorRegistry $registry): Response
    {
        // Laravel metrics
        $registry->getOrRegisterGauge('myapp', 'queue_size', 'Queue jobs count', ['queue'])
            ->set(Queue::size('emails'), ['emails']);

        $registry->getOrRegisterGauge('myapp', 'active_users', 'Active users in last 5 min')
            ->set(User::where('last_seen_at', '>', now()->subMinutes(5))->count());

        $registry->getOrRegisterGauge('myapp', 'failed_jobs', 'Failed jobs total')
            ->set(DB::table('failed_jobs')->count());

        $renderer = new RenderTextFormat();
        return response($renderer->render($registry->getMetricFamilySamples()), 200)
            ->header('Content-Type', RenderTextFormat::MIME_TYPE);
    }
}

Key Metrics for Monitoring

Metric Data Source Exporter
CPU load /proc/stat Node Exporter
Memory usage /proc/meminfo Node Exporter
Free disk space Filesystem Node Exporter
PHP-FPM queue PHP-FPM status php-fpm_exporter
Nginx requests per second Nginx status nginx-vts-exporter
Redis commands count Redis INFO redis_exporter
Active transactions PostgreSQL postgres_exporter

Start with system metrics, then add core services. Our experts will help identify critical indicators for your project. Contact our engineer for a consultation on exporter selection.

Timeline and Deadlines

Stage Duration
Infrastructure audit and design 1 day
Deploy Prometheus + Node Exporter + Grafana 1–2 days
Configure Alertmanager + Slack/PagerDuty +1 day
Connect exporters PHP-FPM, Nginx, Redis, PostgreSQL +1–2 days
Develop custom application metrics +1–2 days
Create dashboards and verify 1 day
Documentation and training 1 day
Total: production-ready stack 4–6 days

What's Included in the Deliverables

  • Documentation: Complete setup guides, architecture diagrams, and runbooks for incident response.
  • Access Credentials: Secure sharing of Grafana, Prometheus, and Alertmanager access.
  • Training: A 2-hour session for your team on using dashboards and interpreting alerts.
  • Support: 30 days of post-deployment support to ensure smooth operation.

Contact us to discuss the details of implementing monitoring on your project. Our engineers, certified in Prometheus and experienced with 50+ projects, guarantee stable operation and timely incident alerts. Order a turnkey monitoring setup — we'll assess your project in one day.