Skip to content

Grafana & Alerting

Ready-to-import Grafana dashboards and Prometheus alerting rules built on the proxy and agent metrics. For the full metric reference and what each panel watches, see Monitoring.

Prerequisite

Metrics must be enabled (--metrics / METRICS_ENABLED=true) and a Prometheus datasource must be scraping the proxy's and the agents' internal /metrics endpoints, each agent in its own job and through the proxy. See Scraping Internal Metrics.

Dashboards

Two dashboards ship in the repository:

File Covers
grafana/prometheus-proxy.json Proxy health, throughput, latency, payload sizes, internal state, errors
grafana/prometheus-agents.json Agent health, connections, scrape activity, per-agent latency

Importing

In Grafana (10.0+): Dashboards → Import, then either upload the JSON file or paste the raw contents. Then open the dashboard and pick your Prometheus datasource in its Datasource drop-down; it starts on Grafana's default datasource:

# Proxy dashboard
curl -O https://raw.githubusercontent.com/pambrose/prometheus-proxy/master/grafana/prometheus-proxy.json

# Agents dashboard
curl -O https://raw.githubusercontent.com/pambrose/prometheus-proxy/master/grafana/prometheus-agents.json

What each dashboard section watches is described under Grafana Dashboards.

Alerting rules

The rules below ship as grafana/alerts.yml in the repository: add that file to Prometheus' rule_files. The thresholds are starting points — tune them to your environment. Each rule is grounded in a metric documented in Monitoring, except ProxyDown, which uses Prometheus' own up series and expects the proxy's scrape job to be named prometheus-proxy.

# Prometheus alerting rules for the proxy and agent. Load this file from Prometheus' rule_files and tune the
# thresholds to your environment. Each rule is built on a metric documented on the website's Monitoring page, except
# ProxyDown, which uses Prometheus' own up series;
# `promtool check rules grafana/alerts.yml` validates it, and CI runs that check.
groups:
  - name: prometheus-proxy
    rules:
      # The rules below need the proxy's metrics, so this one covers the proxy itself going away. The job name is the
      # one in the documented scrape config; change it if yours differs.
      - alert: ProxyDown
        expr: up{job="prometheus-proxy"} == 0
        for: 2m
        labels: { severity: critical }
        annotations:
          summary: "Prometheus cannot scrape the proxy at {{ $labels.instance }}"

      - alert: ProxyScrapeSuccessRateLow
        expr: |
          sum(rate(proxy_scrape_requests_total{type="success"}[5m]))
            / sum(rate(proxy_scrape_requests_total[5m])) < 0.99
        for: 10m
        labels: { severity: warning }
        annotations:
          summary: "Proxy scrape success rate below 99%"

      - alert: ProxyScrapeLatencyHigh
        expr: |
          histogram_quantile(0.99,
            sum by (le) (rate(proxy_scrape_request_latency_seconds_bucket[5m]))) > 2
        for: 10m
        labels: { severity: warning }
        annotations:
          summary: "Proxy P99 scrape latency above 2s"

      - alert: ProxyNoAgentsConnected
        expr: proxy_agent_map_size == 0
        for: 5m
        labels: { severity: critical }
        annotations:
          summary: "No agents connected to the proxy"

      - alert: ProxyBacklogGrowing
        expr: proxy_cumulative_agent_backlog_size > 100
        for: 10m
        labels: { severity: warning }
        annotations:
          summary: "Proxy agent scrape backlog is large ({{ $value }} queued)"

      # Covers both size limits: content_too_large is the agent's maxContentLengthMBytes rejecting
      # the target response, payload_too_large is the proxy's unzipped-size guard.
      - alert: ProxyPayloadTooLarge
        expr: rate(proxy_scrape_requests_total{type=~"payload_too_large|content_too_large"}[5m]) > 0
        for: 5m
        labels: { severity: warning }
        annotations:
          summary: "Scrapes are being rejected for exceeding the size limit ({{ $labels.type }})"

      - alert: ProxyFrequentEvictions
        expr: rate(proxy_eviction_count_total[5m]) > 0
        for: 15m
        labels: { severity: warning }
        annotations:
          summary: "Proxy is evicting stale agents repeatedly"

  - name: prometheus-agent
    rules:
      - alert: AgentConnectFailures
        expr: rate(agent_connect_count_total{type="failure"}[5m]) > 0
        for: 10m
        labels: { severity: warning }
        annotations:
          summary: "Agent {{ $labels.job }} is failing to connect to the proxy"

      - alert: AgentBacklogGrowing
        expr: agent_scrape_backlog_size > 50
        for: 10m
        labels: { severity: warning }
        annotations:
          summary: "Agent {{ $labels.job }} scrape backlog is large"

Detecting restarts

proxy_start_time_seconds and agent_start_time_seconds carry a per-process launch_id; a change in launch_id (or a sudden reset of *_start_time_seconds) flags a restart on a dashboard or alert.

See also