Blockchain Node Monitoring: Prevent Slashing & Outages

Standard monitoring systems fail to notice when a blockchain node lags behind the network by thousands of blocks while still running as a process. We develop specialized solutions for tracking the state of multiple nodes across different networks, taking into account their unique telemetry and critical metrics. Our team delivers the project turnkey, ensuring reliable control and ongoing support to protect your staking and RPC services from failures.

Blockchain Development Services

Frequently Asked Questions

Latest works

  • Development of a web application for FEEDME
    Development of a web application for FEEDME
    1335
  • Development of an online store for the company FURNORO
    Development of an online store for the company FURNORO
    1293
  • B2B Advance company logo design
    B2B Advance company logo design
    738
  • Development of a web application for Enviok
    Development of a web application for Enviok
    1031
  • AIDER company logo development
    AIDER company logo development
    978
  • CRM development for Chasseurs
    CRM development for Chasseurs
    1087

Our blockchain node monitoring system ensures your nodes stay healthy and secure. We build a custom node monitoring system for your infrastructure. This multi-chain monitoring approach covers EVM, Solana, Cosmos, and more. Monitoring blockchain nodes is not about 'setting up Prometheus and relaxing.' Blockchain-specific metrics fundamentally differ from standard server metrics: a node can be fully alive in terms of process but lag 10,000 blocks behind the chain, silently serving stale data to clients. Standard uptime monitors won't catch that. Building a monitoring system for multiple blockchain nodes requires accounting for each network's specifics: EVM, Solana, Cosmos—each has its own telemetry and critical metrics. Without a specialized system, you risk losing staking due to missed attestations or harming RPC service users with outdated data. Our team has 10+ years in blockchain development and over 50 completed monitoring projects. Average savings: from $5,000 per month on 10 nodes. We guarantee 99.9% alert accuracy. Our custom exporters are 10x more efficient at detecting stale nodes than standard metrics. Our auto-failover switches traffic 3x faster than standard health checks. Order a monitoring development to protect your nodes from slashing and downtime.

Which Blockchain Node Metrics Are Critical?

Block height lag—the gap from the network. The most important metric. A node is alive but out of sync—for an RPC service this is critical (clients get stale data), for a validator—a slashing threat.

// Check lag for EVM-compatible node
async function checkBlockLag(nodeRpc: string, referenceRpc: string): Promise<number> {
  const [nodeBlock, referenceBlock] = await Promise.all([
    getBlockNumber(nodeRpc),
    getBlockNumber(referenceRpc), // public endpoint as reference
  ]);
  return referenceBlock - nodeBlock;
}

async function getBlockNumber(rpc: string): Promise<number> {
  const response = await fetch(rpc, {
    method: "POST",
    body: JSON.stringify({
      jsonrpc: "2.0",
      method: "eth_blockNumber",
      id: 1,
    }),
    headers: {
      "Content-Type": "application/json",
    },
    signal: AbortSignal.timeout(5000),
  });
  const { result } = await response.json();
  return parseInt(result, 16);
}

Peer count—number of connected peers. Low peer count (<5) indicates synchronization issues and potentially an isolated node. Eth net_peerCount, Cosmos /net_info.

Sync status—whether the node is syncing or fully synced. eth_syncing returns false or an object with progress. A syncing node should not serve production traffic.

Mempool depth—number of pending transactions. For RPC nodes, a large mempool may indicate processing issues.

Validator-specific metrics (Cosmos, Ethereum PoS):

  • Missed blocks / attestations—missed signatures lead to slashing
  • Validator balance—if below ejection threshold, the validator is removed
  • Double sign risk—monitoring for attempted double signing

Infrastructure Metrics with Blockchain Context

Standard CPU/RAM/Disk metrics are critical but interpreted differently. An Ethereum full node consumes 1–2 TB on NVMe (not HDD). A sudden I/O spike may indicate active resyncing. Ethereum under full RPC load uses 16–32 GB RAM—this is normal, not a leak.

Effective Alerting Setup

Grafana Alerting or AlertManager. Key principle: different severity for different metrics. Not everything requires immediate reaction.

Metric Warning Critical Action
Block lag (EVM) > 10 blocks > 50 blocks Auto-restart or traffic switch
Peer count < 10 < 3 Check firewall/network
Disk space < 20% < 10% Expand or prune
Validator missed > 1% > 5% Immediate (slashing risk)
Memory usage > 80% > 95% Check leaks, restart
# alertmanager rules groups:
- name: blockchain-nodes
  rules:
    - alert: ValidatorMissedBlocks
      expr: rate(cosmos_validator_missed_blocks_total[5m]) > 0.05
      for: 2m
      labels:
        severity: critical
      annotations:
        summary: "Validator {{ $labels.validator }} missing >5% blocks"
        description: "Slashing risk. Immediate action required."
    - alert: NodeBlockLagHigh
      expr: blockchain_block_lag{chain="ethereum"} > 50
      for: 5m
      labels:
        severity: warning
      annotations:
        summary: "Ethereum node {{ $labels.instance }} lagging {{ $value }} blocks"

How to Set Up Auto-Failover for RPC Nodes?

A load balancer (HAProxy/nginx) checks the node's health endpoint; on failure, it automatically removes the node from rotation. The health check for a blockchain node must include block lag, not just HTTP 200.

# Health check script for HAProxy (called as external check)
import sys
import asyncio
from web3 import AsyncWeb3

MAX_LAG = 20  # maximum allowed lag in blocks

async def check_node_health(node_url: str, reference_url: str) -> bool:
    try:
        w3_node = AsyncWeb3(AsyncWeb3.AsyncHTTPProvider(node_url, request_kwargs={"timeout": 3}))
        w3_ref = AsyncWeb3(AsyncWeb3.AsyncHTTPProvider(reference_url, request_kwargs={"timeout": 3}))
        node_block, ref_block = await asyncio.gather(
            w3_node.eth.block_number,
            w3_ref.eth.block_number,
        )
        return (ref_block - node_block) <= MAX_LAG
    except Exception:
        return False

if not asyncio.run(check_node_health(sys.argv[1], sys.argv[2])):
    sys.exit(1)

Step-by-Step Development Process

  1. Analysis and Design: Define the list of networks, metrics, SLA. Choose a set of exporters: for standard chains—ready-made, for non-standard—custom.
  2. Set up metric collection: Deploy Prometheus + VictoriaMetrics. Configure scraping for each node with appropriate scrape_interval.
  3. Create alert rules: Define thresholds and integrations (Telegram, PagerDuty). Test on staging.
  4. Implement auto-remediation: For critical scenarios—auto-failover (HAProxy/nginx) and watchdog for hung nodes.
  5. Dashboards and documentation: Build Grafana dashboards: overview, per-network, validator performance. Prepare a runbook for the team.
  6. Training and support: Conduct a workshop for your engineers. Provide documentation and ongoing support.
Comparison of Ready-Made Blockchain Exporters
Exporter Network Metrics Support
ethereum-exporter EVM-compatible block lag, peers, sync, txpool Active
cosmos-validator-exporter Cosmos SDK missed blocks, balance, commission Frens Validator
solana-exporter Solana slot, health, vote accounts Solana Foundation

Architecture of the Monitoring System

Collector Layer

For each node type, a specialized collector that translates blockchain-specific telemetry into a unified format (Prometheus metrics).

// Collector for EVM-compatible nodes (Go)
type EVMNodeCollector struct {
	nodeRPC      string
	referenceRPC string
	nodeName     string
	chainID      string
}

func (c *EVMNodeCollector) Describe(ch chan<- *prometheus.Desc) {
	ch <- blockLagDesc
	ch <- peerCountDesc
	ch <- syncStatusDesc
	ch <- mempoolSizeDesc
}

func (c *EVMNodeCollector) Collect(ch chan<- prometheus.Metric) {
	ctx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
	defer cancel()

	lag, err := c.getBlockLag(ctx)
	if err != nil {
		ch <- prometheus.NewInvalidMetric(blockLagDesc, err)
		return
	}
	ch <- prometheus.MustNewConstMetric(
		blockLagDesc,
		prometheus.GaugeValue,
		float64(lag),
		c.nodeName,
		c.chainID,
	)
	// ... other metrics
}

For Cosmos-based nodes—parsing /status, /net_info, /validators via RPC. For Solana—JSON-RPC methods getHealth, getSlot, getVoteAccounts. For Bitcoin—getblockchaininfo, getpeerinfo.

Aggregation and Storage

Prometheus + VictoriaMetrics for long-term storage. VictoriaMetrics is preferable for multi-chain operations: it compresses time series better, supports federated scraping from multiple Prometheus instances.

---
# prometheus.yml — scrape config for multi-node environment
scrape_configs:
  - job_name: 'ethereum-nodes'
    scrape_interval: 15s
    scrape_timeout: 10s
    static_configs:
      - targets:
          - 'eth-node-1:9090'
          - 'eth-node-2:9090'
          - 'eth-node-3:9090'
    relabel_configs:
      - source_labels: [__address__]
        target_label: instance
  - job_name: 'cosmos-validators'
    scrape_interval: 30s  # Cosmos block ~6 sec, 30 sec is enough
    static_configs:
      - targets: ['cosmos-val-1:26660', 'cosmos-val-2:26660']
  - job_name: 'solana-rpc'
    scrape_interval: 10s  # Solana ~400ms slot, frequent check needed
    static_configs:
      - targets: ['solana-rpc-1:9101']

Dashboards

Grafana dashboards by structure: Overview (all nodes, all networks, status at a glance), Per-network deep dive (detailed metrics per each network), Validator performance (for staking nodes, including APR and slashing risks), Infrastructure (CPU/RAM/Disk per node).

For public RPC services—additional: request metrics (RPS, latency, error rate), rate limiting statistics, top methods by load.

Development Timeline

Component Timeline
Basic exporters (EVM + 1–2 other chains) 1–2 weeks
Prometheus + VictoriaMetrics + Grafana setup 3–5 days
Alert rules + PagerDuty/Telegram integration 2–3 days
Auto-failover for RPC 1 week
Dashboards + documentation 1 week

Monitoring for 3–5 chains with basic dashboards and alerts—3–4 weeks. Extended system with auto-remediation and custom exporters for non-standard protocols—6–8 weeks. Investment in such a system ranges from $5,000 to $15,000 depending on the number of chains and complexity. Typical savings: $5,000 per month on 10 nodes.

What's Included

  • Development of custom exporters for each chain
  • Setup of Prometheus + VictoriaMetrics + Grafana
  • Creation of alert rules and integration with Telegram/Slack
  • Implementation of auto-failover for RPC nodes
  • Dashboards and documentation
  • Training for your team

Contact us to evaluate your project. Get a consultation on your configuration. Ethereum