server-monitoring

v2026.09.24

Production server monitoring stack covering Prometheus, Node Exporter, Grafana, Alertmanager, Loki, and Promtail on bare-metal or VM Linux hosts. USE WHEN: - Setting up monitoring for a new production server or VPS - Configuring Prometheus scrape targets for application or system metrics - Creating Grafana dashboards and datasource provisioning - Writing Alertmanager routing rules with email/Slack notifications - Implementing the PLG stack (Promtail + Loki + Grafana) for log aggregation - Performing live system diagnostics with htop, iotop, nethogs, ss, vmstat, iostat - Setting up uptime monitoring with UptimeRobot or healthchecks.io DO NOT USE FOR: - Kubernetes-native observability (use the kubernetes skill instead) - Application-level APM (distributed tracing with Jaeger/Tempo — use observability skill) - Cloud-managed monitoring (CloudWatch, GCP Monitoring, Azure Monitor) - Windows Server monitoring

GitHub
安装命令
npx skhub add claude-dev-suite/server-monitoring
Markdown
SKILL.md

Server Monitoring — Production Stack

Stack Overview

LayerToolPurpose
Metrics collectionNode ExporterOS/hardware metrics from /proc and /sys
Metrics scrapingPrometheusPull-based time-series database
VisualizationGrafanaDashboards and alerting UI
AlertingAlertmanagerRoute, deduplicate, silence alerts
Log shippingPromtailTail logs → push to Loki
Log aggregationLokiLog storage with label-based indexing
Uptime (external)UptimeRobotExternal HTTP/TCP reachability checks
Cron monitoringhealthchecks.ioDetect silent cron job failures

Node Exporter Installation (systemd)

Download the latest release from https://github.com/prometheus/node_exporter/releases.

NODE_EXPORTER_VERSION=1.8.2
wget https://github.com/prometheus/node_exporter/releases/download/v${NODE_EXPORTER_VERSION}/node_exporter-${NODE_EXPORTER_VERSION}.linux-amd64.tar.gz
tar xvf node_exporter-${NODE_EXPORTER_VERSION}.linux-amd64.tar.gz
sudo cp node_exporter-${NODE_EXPORTER_VERSION}.linux-amd64/node_exporter /usr/local/bin/
sudo useradd --no-create-home --shell /bin/false node_exporter

/etc/systemd/system/node_exporter.service:

[Unit]
Description=Node Exporter
Wants=network-online.target
After=network-online.target

[Service]
User=node_exporter
Group=node_exporter
Type=simple
ExecStart=/usr/local/bin/node_exporter \
  --collector.systemd \
  --collector.processes \
  --collector.diskstats \
  --web.listen-address=127.0.0.1:9100
Restart=always
RestartSec=5

[Install]
WantedBy=multi-user.target
sudo systemctl daemon-reload
sudo systemctl enable --now node_exporter
sudo systemctl status node_exporter
curl -s http://127.0.0.1:9100/metrics | head -20

Bind Node Exporter to 127.0.0.1:9100 — never expose directly on 0.0.0.0. Prometheus scrapes it locally; use SSH tunnel or VPN for remote Prometheus.


Prometheus Installation and Configuration

PROMETHEUS_VERSION=2.53.0
wget https://github.com/prometheus/prometheus/releases/download/v${PROMETHEUS_VERSION}/prometheus-${PROMETHEUS_VERSION}.linux-amd64.tar.gz
tar xvf prometheus-${PROMETHEUS_VERSION}.linux-amd64.tar.gz
sudo cp prometheus-${PROMETHEUS_VERSION}.linux-amd64/{prometheus,promtool} /usr/local/bin/
sudo mkdir -p /etc/prometheus /var/lib/prometheus
sudo cp -r prometheus-${PROMETHEUS_VERSION}.linux-amd64/{consoles,console_libraries} /etc/prometheus/
sudo useradd --no-create-home --shell /bin/false prometheus
sudo chown -R prometheus:prometheus /etc/prometheus /var/lib/prometheus

/etc/prometheus/prometheus.yml:

global:
  scrape_interval: 15s           # Default scrape interval
  evaluation_interval: 15s       # Rule evaluation interval
  scrape_timeout: 10s

alerting:
  alertmanagers:
    - static_configs:
        - targets: ['localhost:9093']

rule_files:
  - /etc/prometheus/rules/*.yml

scrape_configs:
  # Node Exporter — OS metrics
  - job_name: node
    static_configs:
      - targets: ['localhost:9100']
        labels:
          server: 'prod-web-01'
          env: production

  # Application metrics — assumes /metrics on port 3000
  - job_name: app
    metrics_path: /metrics
    static_configs:
      - targets: ['localhost:3000']
        labels:
          app: myapp
          env: production

  # Blackbox Exporter — probe HTTP endpoints
  - job_name: blackbox_http
    metrics_path: /probe
    params:
      module: [http_2xx]
    static_configs:
      - targets:
          - https://example.com
          - https://example.com/api/health
    relabel_configs:
      - source_labels: [__address__]
        target_label: __param_target
      - source_labels: [__param_target]
        target_label: instance
      - target_label: __address__
        replacement: localhost:9115   # Blackbox Exporter address

  # Prometheus self-monitoring
  - job_name: prometheus
    static_configs:
      - targets: ['localhost:9090']

/etc/systemd/system/prometheus.service:

[Unit]
Description=Prometheus
Wants=network-online.target
After=network-online.target

[Service]
User=prometheus
Group=prometheus
Type=simple
ExecStart=/usr/local/bin/prometheus \
  --config.file=/etc/prometheus/prometheus.yml \
  --storage.tsdb.path=/var/lib/prometheus \
  --storage.tsdb.retention.time=30d \
  --storage.tsdb.retention.size=10GB \
  --web.listen-address=127.0.0.1:9090 \
  --web.enable-lifecycle
Restart=always
RestartSec=5

[Install]
WantedBy=multi-user.target
sudo systemctl daemon-reload
sudo systemctl enable --now prometheus
promtool check config /etc/prometheus/prometheus.yml

Alert Rules

/etc/prometheus/rules/server.yml:

groups:
  - name: server_alerts
    interval: 1m
    rules:

      - alert: InstanceDown
        expr: up == 0
        for: 2m
        labels:
          severity: critical
        annotations:
          summary: "Instance {{ $labels.instance }} is down"
          description: "{{ $labels.job }}/{{ $labels.instance }} has been unreachable for more than 2 minutes."

      - alert: HighCPU
        expr: 100 - (avg by(instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100) > 85
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "High CPU usage on {{ $labels.instance }}"
          description: "CPU usage is {{ printf \"%.1f\" $value }}% (threshold 85%)."

      - alert: HighMemory
        expr: |
          (node_memory_MemTotal_bytes - node_memory_MemAvailable_bytes)
          / node_memory_MemTotal_bytes * 100 > 90
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "High memory usage on {{ $labels.instance }}"
          description: "Memory usage is {{ printf \"%.1f\" $value }}% (threshold 90%)."

      - alert: DiskAlmostFull
        expr: |
          (node_filesystem_size_bytes{fstype!="tmpfs"} - node_filesystem_free_bytes{fstype!="tmpfs"})
          / node_filesystem_size_bytes{fstype!="tmpfs"} * 100 > 85
        for: 10m
        labels:
          severity: warning
        annotations:
          summary: "Disk almost full on {{ $labels.instance }}"
          description: "Filesystem {{ $labels.mountpoint }} is {{ printf \"%.1f\" $value }}% full."

      - alert: HighLoad
        expr: node_load15 / count without(cpu, mode)(node_cpu_seconds_total{mode="idle"}) > 2.0
        for: 10m
        labels:
          severity: warning
        annotations:
          summary: "High system load on {{ $labels.instance }}"
          description: "15-minute load average per CPU core is {{ printf \"%.2f\" $value }} (threshold 2.0)."

Validate rules: promtool check rules /etc/prometheus/rules/server.yml


Alertmanager

ALERTMANAGER_VERSION=0.27.0
wget https://github.com/prometheus/alertmanager/releases/download/v${ALERTMANAGER_VERSION}/alertmanager-${ALERTMANAGER_VERSION}.linux-amd64.tar.gz
tar xvf alertmanager-${ALERTMANAGER_VERSION}.linux-amd64.tar.gz
sudo cp alertmanager-${ALERTMANAGER_VERSION}.linux-amd64/{alertmanager,amtool} /usr/local/bin/
sudo mkdir -p /etc/alertmanager /var/lib/alertmanager

/etc/alertmanager/alertmanager.yml:

global:
  smtp_smarthost: 'smtp.gmail.com:587'
  smtp_from: 'alerts@example.com'
  smtp_auth_username: 'alerts@example.com'
  smtp_auth_password: 'app-specific-password'   # Use App Password, not account password
  smtp_require_tls: true
  resolve_timeout: 5m

route:
  group_by: ['alertname', 'instance']
  group_wait: 30s        # Wait before sending first notification for a new group
  group_interval: 5m     # How long to wait before sending alert for new alerts in the same group
  repeat_interval: 4h    # How often to re-send unresolved alerts
  receiver: 'team-ops'
  routes:
    - match:
        severity: critical
      receiver: 'pagerduty-critical'
      repeat_interval: 1h

receivers:
  - name: 'team-ops'
    email_configs:
      - to: 'ops-team@example.com'
        send_resolved: true
    slack_configs:
      - api_url: 'https://hooks.slack.com/services/<T_ID>/<B_ID>/<WEBHOOK_TOKEN>'
        channel: '#alerts'
        send_resolved: true
        title: '{{ if eq .Status "firing" }}:red_circle:{{ else }}:white_check_mark:{{ end }} {{ .CommonAnnotations.summary }}'
        text: '{{ range .Alerts }}{{ .Annotations.description }}{{ end }}'

  - name: 'pagerduty-critical'
    pagerduty_configs:
      - service_key: 'your-pagerduty-integration-key'
        send_resolved: true

inhibit_rules:
  - source_match:
      severity: critical
    target_match:
      severity: warning
    equal: ['alertname', 'instance']

/etc/systemd/system/alertmanager.service:

[Unit]
Description=Alertmanager
After=network-online.target

[Service]
User=prometheus
Group=prometheus
Type=simple
ExecStart=/usr/local/bin/alertmanager \
  --config.file=/etc/alertmanager/alertmanager.yml \
  --storage.path=/var/lib/alertmanager \
  --web.listen-address=127.0.0.1:9093
Restart=always
RestartSec=5

[Install]
WantedBy=multi-user.target

Grafana Installation and Datasource Provisioning

sudo apt-get install -y apt-transport-https software-properties-common
wget -q -O - https://apt.grafana.com/gpg.key | gpg --dearmor | sudo tee /etc/apt/keyrings/grafana.gpg > /dev/null
echo "deb [signed-by=/etc/apt/keyrings/grafana.gpg] https://apt.grafana.com stable main" | sudo tee /etc/apt/sources.list.d/grafana.list
sudo apt-get update && sudo apt-get install -y grafana
sudo systemctl enable --now grafana-server

Datasource provisioning (/etc/grafana/provisioning/datasources/prometheus.yml):

apiVersion: 1
datasources:
  - name: Prometheus
    type: prometheus
    access: proxy
    url: http://localhost:9090
    isDefault: true
    editable: false

  - name: Loki
    type: loki
    access: proxy
    url: http://localhost:3100
    editable: false

Import Node Exporter Full dashboard (ID 1860) via Grafana UI: Dashboards → Import → enter 1860 → select Prometheus datasource. Other useful dashboard IDs:

  • 3662 — Prometheus 2.0 Stats
  • 13659 — Node Exporter for Prometheus Dashboard
  • 10991 — Blackbox Exporter

PLG Stack: Promtail + Loki

Loki Installation

LOKI_VERSION=3.1.0
wget https://github.com/grafana/loki/releases/download/v${LOKI_VERSION}/loki-linux-amd64.zip
unzip loki-linux-amd64.zip
sudo mv loki-linux-amd64 /usr/local/bin/loki
sudo mkdir -p /etc/loki /var/lib/loki

/etc/loki/loki-config.yml:

auth_enabled: false

server:
  http_listen_port: 3100
  grpc_listen_port: 9096

common:
  instance_addr: 127.0.0.1
  path_prefix: /var/lib/loki
  storage:
    filesystem:
      chunks_directory: /var/lib/loki/chunks
      rules_directory: /var/lib/loki/rules
  replication_factor: 1
  ring:
    kvstore:
      store: inmemory

schema_config:
  configs:
    - from: 2024-01-01
      store: tsdb
      object_store: filesystem
      schema: v13
      index:
        prefix: loki_index_
        period: 24h

limits_config:
  retention_period: 30d           # Requires compactor below

compactor:
  working_directory: /var/lib/loki/compactor
  retention_enabled: true
  delete_request_cancel_period: 24h

Promtail Configuration

/etc/promtail/promtail-config.yml:

server:
  http_listen_port: 9080
  grpc_listen_port: 0

positions:
  filename: /var/lib/promtail/positions.yaml

clients:
  - url: http://localhost:3100/loki/api/v1/push

scrape_configs:
  # Systemd journal — captures all systemd unit logs
  - job_name: journal
    journal:
      max_age: 12h
      labels:
        job: systemd-journal
        host: prod-web-01
    relabel_configs:
      - source_labels: ['__journal__systemd_unit']
        target_label: unit
      - source_labels: ['__journal_priority_keyword']
        target_label: level

  # Application log files
  - job_name: app_logs
    static_configs:
      - targets:
          - localhost
        labels:
          job: myapp
          host: prod-web-01
          __path__: /var/log/myapp/*.log
    pipeline_stages:
      - json:
          expressions:
            level: level
            msg: message
      - labels:
          level:
      - timestamp:
          source: timestamp
          format: RFC3339Nano

  # Nginx access logs
  - job_name: nginx
    static_configs:
      - targets:
          - localhost
        labels:
          job: nginx
          host: prod-web-01
          __path__: /var/log/nginx/access.log

LogQL Basics

# Filter by label
{job="myapp", level="error"}

# Filter by content
{job="nginx"} |= "500"

# Pattern extraction
{job="nginx"} | pattern `<ip> - - [<_>] "<method> <path> <_>" <status> <_>`

# Rate of error log lines per minute
rate({job="myapp", level="error"}[1m])

# Count errors by unit over last hour
sum by(unit) (count_over_time({job="systemd-journal", level="err"}[1h]))

System-Level Diagnostic Tools

ToolKey UsageNotes
htopInteractive process viewerF5 tree view, F6 sort, F9 kill, u filter by user
iotopI/O per processsudo iotop -o (only active), -a accumulated
nethogsBandwidth per processsudo nethogs eth0
ssSocket statistics (replaces netstat)ss -tlnp TCP listening, ss -s summary, ss -o state time-wait
vmstatMemory/swap/CPU overviewvmstat 1 10 (10 samples, 1s interval)
iostatDisk I/O statsiostat -xz 1 extended, iostat -m in MB/s
dstatCombined resource statsdstat -cdngy CPU/disk/net/page/sys
sarHistorical performance (sysstat)sar -u 1 5 CPU, sar -r 1 5 memory, sar -b 1 5 I/O
free -hMemory and swap summaryCheck buff/cache vs actual available
df -hDisk space by filesystemAdd -i for inode usage
du -shDirectory size`du -sh /var/log/*

Uptime Monitoring

UptimeRobot (Free Tier)

  • Create HTTP(s) monitor: URL, check interval (5 min on free tier), keyword match
  • Alert contacts: email + Slack webhook
  • Status page: public URL for incident communication
  • TCP monitors for non-HTTP services (database ports, SMTP)

healthchecks.io (Cron Job Monitoring)

Add a curl ping at the end of every cron script:

#!/bin/bash
set -euo pipefail

# ... backup/maintenance logic here ...

# Signal success to healthchecks.io
curl -fsS --retry 3 https://hc-ping.com/YOUR-CHECK-UUID > /dev/null

For jobs that should ping start + finish:

curl -fsS --retry 3 https://hc-ping.com/YOUR-CHECK-UUID/start > /dev/null
# ... job logic ...
curl -fsS --retry 3 https://hc-ping.com/YOUR-CHECK-UUID > /dev/null

Anti-Patterns

Anti-PatternProblemFix
No alerting configuredIncidents discovered by users, not ops teamSet up Alertmanager with email + Slack from day one
Scrape interval < 10sHigh cardinality load on Prometheus, noisy dataUse 15s–60s depending on metric volatility
No retention limitsPrometheus disk fills up and crashesSet --storage.tsdb.retention.time=30d and --storage.tsdb.retention.size
Monitoring server on the same host being monitoredSingle point of failure — if server dies, so does the monitorRun Prometheus on a dedicated monitoring host or use external uptime service
Exposing Node Exporter on 0.0.0.0Metrics data leaked publiclyBind to 127.0.0.1:9100, use SSH tunnel or VPN for remote scraping
Catching all logs without labelsCannot filter Loki queries efficientlyAdd job, host, level labels in Promtail config
Alert rules without for durationFlapping alerts on transient spikesAlways add for: 2m or longer to avoid noise
No inhibit_rules in AlertmanagerWarning floods when critical alert also firingInhibit warnings when critical alert matches same instance
SMTP password in alertmanager.yml committed to gitCredential leakUse environment variable substitution or external secret management
No Grafana datasource provisioningDashboard import fails after Grafana reinstallProvision datasources as YAML in /etc/grafana/provisioning/

Troubleshooting

SymptomLikely CauseDiagnostic & Fix
Metrics not appearing in PrometheusNode Exporter not running or wrong portcurl http://127.0.0.1:9100/metrics — check systemd status, verify prometheus.yml target
up == 0 for a targetPrometheus cannot reach scrape endpointCheck firewall (ss -tlnp), test with curl from Prometheus host, verify target address
Alert not firing despite condition metfor duration not elapsed, or rule file not loadedpromtool check rules — verify rule file path in prometheus.yml, check Prometheus logs
Alert firing but no notificationAlertmanager not configured or unreachableTest Alertmanager with amtool alert add — check SMTP credentials, Slack webhook URL
Grafana blank dashboard after importWrong datasource selectedEdit dashboard → Panel → Query → verify datasource dropdown matches provisioned name
Loki not ingesting logsPromtail cannot connect or wrong URLcurl http://localhost:3100/ready — check Promtail logs (journalctl -u promtail)
Promtail not tailing journalMissing journal permissionsAdd promtail user to systemd-journal group: usermod -aG systemd-journal promtail
High cardinality error in LokiToo many unique label combinationsAvoid high-cardinality labels (request IDs, user IDs) — use log content for those
Prometheus OOM killedToo many metrics or insufficient retention pruningAdd --storage.tsdb.retention.size limit, reduce scrape targets or intervals
ALERTS metric not visiblePrometheus evaluating no rulesConfirm rule_files path glob matches actual files: ls /etc/prometheus/rules/*.yml
发现
标签

此技能尚未发布标签。

版本
最新版本元数据

版本

v2026.09.24

发布时间

2026年9月24日

分类

未分类

许可证

MIT

源路径

skills/infrastructure/server-monitoring

默认分支

main

最新提交

9496306

Tree SHA

fe4e2f1