GuidesDevelopers and hostingPrometheus and Grafana

Monitor your servers with Prometheus and Grafana

Prometheus and Grafana on Debian 13 in rootless Podman behind Caddy, with node_exporter on each server, a ready-made dashboard, alerts sent to your phone through ntfy, and Grafana's calls home switched off.

Tested on Prometheus 3.15.0, Grafana 13.2.3, Alertmanager 0.34.1 and node_exporter 1.9.0 on Debian 13 (trixie) on a Melonslab server Updated October 2, 2026

Recommended server for this guide

VC-S Micro · 2 vCPU · 8 GB Memory · 250 GB Storage

Month to month, no lock-in 7-day money-back guarantee

€7.99/mo

Deploy now
On this page

What you will set up

Four programs work together here:

  • node_exporter runs on each server you watch and reports its CPU, memory, disk and network figures.
  • Prometheus fetches those figures every 15 seconds, stores them, and checks your alert rules.
  • Alertmanager sends a message when a rule fires, here to your phone through ntfy.
  • Grafana draws the figures as dashboards in your browser.

Prometheus, Alertmanager and Grafana run in a pod under a user of their own called monitoring, behind Caddy from the Podman guide. Only Grafana can be reached, through Caddy. We also have a guide for Zabbix: pick Zabbix if you want one program that you set up by clicking in its web interface, and Prometheus and Grafana if you would rather keep your setup in text files and use Grafana's dashboards.

Prometheus, Alertmanager and node_exporter are open source under the Apache 2.0 licence, made by the Prometheus project, a graduated project of the Cloud Native Computing Foundation, which is part of the US non-profit Linux Foundation. They contact no outside service by default. Grafana is open source under the GNU AGPL 3.0, and is developed by Grafana Labs (Raintank Inc.), a company in New York, USA. By default, Grafana sends usage counters to stats.grafana.org every 24 hours, asks grafana.com for new versions of itself and its plugins, loads a news feed, and downloads plugins at startup: on our test server it updated its Prometheus plugin from grafana.com within 20 seconds of its first start. Step 5 switches all of that off.

Every step below was run on a Melonslab server with Debian 13:

  • node_exporter answered Prometheus from the pod, and port 9100 timed out from three test locations on the internet.
  • A second node_exporter answered only over HTTPS with a password: without the password it returned 401, and without the right certificate the connection was refused.
  • The default admin/admin login was refused from the internet, and so were sign-ups and visitors who are not logged in. Grafana logged each visitor's real address.
  • The Node Exporter Full dashboard showed CPU, memory, disk and network for both servers.
  • With the disk filled to 93 %, the alert arrived in the ntfy topic after five minutes, and a Resolved message followed when the space was freed. A second server that stopped answering was reported within three minutes, and resolved within five minutes of answering again.
  • After the settings in step 5, Grafana looked up grafana.com only when we imported a dashboard or opened the plugin list.
  • A backup of Grafana's data was restored with its dashboard intact, and everything came back by itself after a reboot.

The second server in these tests was a container on the same Melonslab server, with the same node_exporter package and settings as step 9. The firewall rules for a real second server in step 9 were not tested.

The pod used between 320 and 500 MB of memory while watching two servers. Most of that is Grafana, which used between 250 and 410 MB; Prometheus used 30 to 40 MB and Alertmanager about 15 MB. node_exporter used about 25 MB on each server.

Before you start

You need:

  • a server set up as in the Podman guide, with Caddy running, and ufw from the security guide;
  • an A record and an AAAA record for grafana.example.com pointing at your server;
  • somewhere to receive alerts: an ntfy topic, on your own ntfy server or on the public ntfy.sh, and the ntfy app on your phone subscribed to it. On ntfy.sh, anyone who knows the topic name can read its messages, so pick a long name that cannot be guessed, such as the output of openssl rand -hex 12.

The examples use grafana.example.com for Grafana, 203.0.113.10 for your server's IPv4 address, which ip -brief address show eth0 shows, 2001:db8::10 for its IPv6 address, web1.example.com for a second server you want to watch, and ntfy.example.com/servers for the ntfy topic. Replace them throughout.

1. Install node_exporter

As root, install node_exporter from Debian:

apt update
apt install -y --no-install-recommends prometheus-node-exporter curl

Debian's package runs as a user of its own, starts at boot, and gets Debian's security updates. --no-install-recommends leaves out a set of extra collectors, such as SMART and IPMI readers, that a virtual server has no use for. Check that it answers:

curl -s http://127.0.0.1:9100/metrics | grep -m1 node_uname_info

node_exporter listens on port 9100 on every address. It has to: Prometheus in its rootless pod reaches the server from the server's own IPv4 address, not from 127.0.0.1. ufw keeps the port closed to the internet, so do not allow 9100.

2. Create the user

useradd -m -s /bin/bash monitoring
loginctl enable-linger monitoring
machinectl shell monitoring@

3. Create Grafana's admin password

As monitoring:

openssl rand -base64 18 | tr -d '\n' | podman secret create grafana-admin-password -
podman secret inspect --showsecret --format '{{.SecretData}}' grafana-admin-password

The second line shows the password. Store it in your password manager. Grafana sets it for the user admin on its first start, so its well-known default password admin is never active, not even for a minute. That matters: in our test, scanners found the new name through its certificate and requested Grafana's start page within a minute of Caddy getting the certificate.

4. Write the configuration

mkdir -p ~/monitoring
cd ~/monitoring

Create prometheus.yml, which tells Prometheus what to fetch:

global:
  scrape_interval: 15s

rule_files:
  - /etc/prometheus/rules.yml

alerting:
  alertmanagers:
    - static_configs:
        - targets: ["localhost:9093"]

scrape_configs:
  - job_name: node
    static_configs:
      - targets: ["host.containers.internal:9100"]
        labels:
          instance: monitor

Inside the pod, host.containers.internal is the server outside it, where node_exporter runs. instance: monitor is the name the server gets in Grafana; pick your own.

Create rules.yml, with two alerts:

groups:
  - name: servers
    rules:
      - alert: ServerDown
        expr: up == 0
        for: 2m
        labels:
          severity: critical
        annotations:
          summary: "{{ $labels.instance }} is not answering"
      - alert: DiskAlmostFull
        expr: 1 - node_filesystem_avail_bytes{fstype!~"tmpfs|ramfs"} / node_filesystem_size_bytes > 0.85
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "{{ $labels.instance }}: {{ $labels.mountpoint }} is {{ $value | humanizePercentage }} full"

ServerDown fires when Prometheus has not reached a server for two minutes, and DiskAlmostFull when a disk has been more than 85 % full for five minutes.

Create alertmanager.yml, with your ntfy topic:

route:
  receiver: ntfy
  group_by: [alertname, instance]

receivers:
  - name: ntfy
    webhook_configs:
      - url: https://ntfy.example.com/servers?template=alertmanager

?template=alertmanager (ntfy 2.14 or newer) makes ntfy turn Alertmanager's message into a readable notification with a title. group_by sends one notification per alert and server. Without it, Alertmanager puts every alert into one notification.

Create datasources.yml, which connects Grafana to Prometheus:

apiVersion: 1
datasources:
  - name: Prometheus
    type: prometheus
    uid: prometheus
    url: http://localhost:9090
    isDefault: true

5. Define the pod

mkdir -p ~/.config/containers/systemd
cd ~/.config/containers/systemd

Create monitoring.pod:

[Pod]
PodName=monitoring
# Grafana, reached through Caddy only
PublishPort=127.0.0.1:8102:3000

[Install]
WantedBy=default.target

Prometheus and Alertmanager publish no port. Grafana reaches them inside the pod, on localhost.

Create prometheus.container:

[Container]
ContainerName=prometheus
Image=quay.io/prometheus/prometheus:latest
Pod=monitoring.pod
Volume=%h/monitoring/prometheus.yml:/etc/prometheus/prometheus.yml:ro
Volume=%h/monitoring/rules.yml:/etc/prometheus/rules.yml:ro
Volume=prometheus-data:/prometheus
Exec=--config.file=/etc/prometheus/prometheus.yml --storage.tsdb.path=/prometheus --storage.tsdb.retention.time=30d --storage.tsdb.retention.size=5GB
AutoUpdate=registry

[Service]
Restart=always

Create alertmanager.container:

[Container]
ContainerName=alertmanager
Image=quay.io/prometheus/alertmanager:latest
Pod=monitoring.pod
Volume=%h/monitoring/alertmanager.yml:/etc/alertmanager/alertmanager.yml:ro
Volume=alertmanager-data:/alertmanager
# One Alertmanager, no cluster
Exec=--config.file=/etc/alertmanager/alertmanager.yml --storage.path=/alertmanager --cluster.listen-address=
AutoUpdate=registry

[Service]
Restart=always

The empty --cluster.listen-address= switches off clustering, which you do not need for one Alertmanager. Without it, Alertmanager stops at once in a rootless pod.

And grafana.container:

[Container]
ContainerName=grafana
Image=docker.io/grafana/grafana:latest
Pod=monitoring.pod
Volume=grafana-data:/var/lib/grafana
Volume=%h/monitoring/datasources.yml:/etc/grafana/provisioning/datasources/datasources.yml:ro
Secret=grafana-admin-password
Environment=GF_SECURITY_ADMIN_PASSWORD__FILE=/run/secrets/grafana-admin-password
Environment=GF_SERVER_ROOT_URL=https://grafana.example.com/
# No usage reports, update checks, news feed, plugin downloads or Gravatar
Environment=GF_ANALYTICS_REPORTING_ENABLED=false GF_ANALYTICS_CHECK_FOR_UPDATES=false
Environment=GF_ANALYTICS_CHECK_FOR_PLUGIN_UPDATES=false GF_PLUGINS_PREINSTALL_DISABLED=true
Environment=GF_NEWS_NEWS_FEED_ENABLED=false GF_SECURITY_DISABLE_GRAVATAR=true
Environment=GF_PLUGINS_PUBLIC_KEY_RETRIEVAL_DISABLED=true
AutoUpdate=registry

[Service]
Restart=always

What the GF_ lines do:

  • GF_SERVER_ROOT_URL is Grafana's public address, which it uses in links.
  • GF_ANALYTICS_REPORTING_ENABLED=false stops the usage counters sent to stats.grafana.org every 24 hours. Grafana Labs' privacy policy says such reports can include the operating system, country, location, version and how the software is used.
  • The two CHECK_FOR lines stop the version checks against grafana.com.
  • GF_PLUGINS_PREINSTALL_DISABLED=true stops Grafana from downloading and updating its bundled app plugins from grafana.com at every start.
  • GF_NEWS_NEWS_FEED_ENABLED=false removes the news feed, and GF_SECURITY_DISABLE_GRAVATAR=true stops users' pictures from being loaded from Gravatar.
  • GF_PLUGINS_PUBLIC_KEY_RETRIEVAL_DISABLED=true makes Grafana use its built-in key to check plugin signatures, instead of fetching it from grafana.com every ten days.

Anonymous access and sign-ups are off in Grafana by default, so they need no line here. Step 7 checks both.

6. Start it

systemctl --user daemon-reload
systemctl --user start monitoring-pod
podman ps --format '{{.Names}} {{.Status}}'

All four, monitoring-infra, prometheus, alertmanager and grafana, show Up. Check that Prometheus reaches node_exporter:

podman exec prometheus wget -qO- localhost:9090/api/v1/targets | grep -o '"health":"[a-z]*"'

It prints "health":"up".

7. Put Caddy in front

Go back to root with exit, switch to machinectl shell caddy@, and add this block at the end of ~/Caddyfile:

grafana.example.com {
    reverse_proxy 127.0.0.1:8102
}

Restart Caddy with systemctl --user restart caddy. Then, from your own computer, check that the default login and sign-ups are refused:

curl -s -H 'Content-Type: application/json' -d '{"user":"admin","password":"admin"}' https://grafana.example.com/login
curl -s -H 'Content-Type: application/json' -d '{"email":"test@example.com"}' https://grafana.example.com/api/user/signup
curl -s https://grafana.example.com/api/search

They answer Invalid username or password, User signup is disabled and Unauthorized. Grafana logs refused requests, such as these three, with the visitor's real address, which it reads from Caddy's X-Forwarded-For header. As monitoring, podman logs grafana | grep remote_addr shows them with your own address. Grafana's port is only open on 127.0.0.1, so nobody on the internet can reach Grafana past Caddy to fake that header.

8. Log in and add the dashboard

Open https://grafana.example.com and log in as admin with the password from step 3.

The Node Exporter Full dashboard is a widely used ready-made dashboard for node_exporter, published on grafana.com under the ID 1860. To import it:

  • Open Dashboards, choose New, then Import dashboard.
  • Enter 1860 in the field Grafana.com dashboard URL or ID and choose Load.
  • Choose Import.

Importing by ID makes your server fetch the dashboard from grafana.com. If you would rather not, download the JSON file from the dashboard's page on grafana.com and use Upload dashboard JSON file instead.

The dashboard opens on your server. Job and Instance at the top choose the server; the first rows show CPU, memory, network traffic and disk space, and the collapsed rows below hold the details.

9. Watch another server

node_exporter has no password and no encryption by default, so on a server that Prometheus reaches over the internet, give it both. On the other server, as root:

apt update
apt install -y --no-install-recommends prometheus-node-exporter apache2-utils
mkdir -p /etc/node-exporter
cd /etc/node-exporter
openssl req -x509 -newkey ec -pkeyopt ec_paramgen_curve:prime256v1 -days 3650 -nodes \
  -keyout node-exporter.key -out node-exporter.crt \
  -subj "/CN=web1.example.com" -addext "subjectAltName=DNS:web1.example.com"
chgrp prometheus node-exporter.key
chmod 640 node-exporter.key

This makes a certificate of its own for the server, valid for ten years. Now choose a long password for Prometheus to log in with, and keep it for the next step. htpasswd asks for it twice, and only its bcrypt hash goes into node_exporter's configuration:

HASH=$(htpasswd -nBC 12 "" | tr -d ':\n')
cat > web.yml <<EOF
tls_server_config:
  cert_file: /etc/node-exporter/node-exporter.crt
  key_file: /etc/node-exporter/node-exporter.key
basic_auth_users:
  prometheus: $HASH
EOF
chgrp prometheus web.yml
chmod 640 web.yml

Point node_exporter at the file, and allow your monitoring server, by both its addresses, to reach the port:

sed -i 's|^ARGS=.*|ARGS="--web.config.file=/etc/node-exporter/web.yml"|' /etc/default/prometheus-node-exporter
systemctl restart prometheus-node-exporter
ufw allow from 203.0.113.10 to any port 9100 proto tcp
ufw allow from 2001:db8::10 to any port 9100 proto tcp
cat node-exporter.crt

Prometheus connects over IPv6 when web1.example.com has an AAAA record, which is why both addresses are needed. Copy the certificate that cat printed. If both servers are Melonslab servers, they cannot reach each other directly by default: connect them first. If you would rather not open a port at all, a WireGuard VPN between the servers works too, and then Prometheus fetches over the VPN address instead.

On the monitoring server, as monitoring, save the certificate and the password next to the configuration:

cd ~/monitoring
nano web1.crt        # paste the certificate
nano web1.password   # the password, on one line
chmod 644 web1.crt web1.password

Prometheus runs as a different user inside its container, so it needs the files to be readable. Your home directory stays closed to other users on the server, which ls -ld ~ shows as drwx------. Add a job at the end of prometheus.yml:

  - job_name: web1
    scheme: https
    tls_config:
      ca_file: /etc/prometheus/web1.crt
    basic_auth:
      username: prometheus
      password_file: /etc/prometheus/web1.password
    static_configs:
      - targets: ["web1.example.com:9100"]
        labels:
          instance: web1

And in ~/.config/containers/systemd/prometheus.container, add two lines below the other Volume= lines:

Volume=%h/monitoring/web1.crt:/etc/prometheus/web1.crt:ro
Volume=%h/monitoring/web1.password:/etc/prometheus/web1.password:ro

Restart Prometheus and check the configuration:

systemctl --user daemon-reload
systemctl --user restart prometheus
podman exec prometheus promtool check config /etc/prometheus/prometheus.yml

promtool prints SUCCESS for the configuration and the rules. The target check from step 6 now prints "health":"up" twice, and in Grafana, Job web1 shows the second server.

Prometheus checks the server's certificate against the one you copied, so nobody can stand in for the server, and the password keeps everyone else out. From any other address, the port does not answer at all.

10. Test the alerts

Stop node_exporter on the other server for a moment:

systemctl stop prometheus-node-exporter

Within about three minutes, your phone shows Alert: ServerDown with Instance: web1: the rule waits two minutes, Prometheus checks it once a minute, and Alertmanager waits 30 seconds for other alerts to send with it. Start it again with systemctl start prometheus-node-exporter. A Resolved: ServerDown message follows within five minutes, because Alertmanager sends updates to an open alert at most every five minutes.

While an alert stays active, Alertmanager repeats it every four hours. Alerting, then Alert rules in Grafana shows both rules and their state. The Source link in the message points inside the pod and does not open from your phone.

The ServerDown rule does not cover the monitoring server itself: if it is down, nothing is left to send the alert. To be told about that, watch Grafana's address from somewhere else, for example with Uptime Kuma on another server.

11. Check what Grafana contacts

To see which names the server looks up, as root:

apt install -y tcpdump
timeout 600 tcpdump -i eth0 -n -l 'udp port 53' | grep -o 'A? [^ ]*'

Leave it running while you use Grafana for a few minutes. With the settings from step 5, Grafana looked up grafana.com only when we imported a dashboard by its ID, and when we opened Administration, Plugins and data, Plugins: that page lists the plugin catalogue from grafana.com. Your alerts show up as lookups of your ntfy server. Prometheus and Alertmanager looked up nothing else, before or after a reboot.

12. Disk space

Prometheus keeps its figures for 30 days, and never lets them grow past 5 GB, whichever limit comes first. These are the --storage.tsdb.retention options in step 5. Our two servers gave Prometheus about 1,800 series, and after 40 minutes its data took 2.2 MB. Prometheus' documentation puts a stored sample at 1 to 2 bytes, which for two servers every 15 seconds comes to roughly 10 to 20 MB a day, so 30 days stay well under the 5 GB limit; we did not run it for 30 days. To see how much it uses, as monitoring:

podman unshare du -sh ~/.local/share/containers/storage/volumes/prometheus-data/_data

13. Keep it up to date

As monitoring, turn on Podman's daily updates:

systemctl --user enable --now podman-auto-update.timer

The latest tags follow every new release, including new major versions. Prometheus 3 and Grafana both publish upgrade notes, which are worth reading when the major version changes. To stay on one Grafana release line instead, use a tag such as 13.2 in grafana.container, and change it by hand when you want to move on. node_exporter comes from Debian, so the unattended upgrades from the security guide install its security updates.

14. Back up

Grafana keeps its users, dashboards and settings in a database in the volume grafana-data. Stop it for a moment to copy the volume, as monitoring:

mkdir -p ~/backup
systemctl --user stop grafana
podman volume export grafana-data > ~/backup/grafana-data.tar
systemctl --user start grafana
tar -czf ~/backup/monitoring-config.tar.gz monitoring .config/containers/systemd

The second file holds your configuration from steps 4, 5 and 9, including the other server's password. Copy ~/backup to another machine, and keep it safe.

Prometheus' figures are left out. They are usually expendable: they rebuild themselves from the moment Prometheus runs again, and you only lose the history. If you want to keep them, copy the volume prometheus-data the same way, with Prometheus stopped.

To restore Grafana's data on a new server, after steps 1 to 6:

systemctl --user stop grafana
podman volume rm grafana-data
podman volume create grafana-data
podman volume import grafana-data ~/backup/grafana-data.tar
podman unshare chown 472:0 "$(podman volume inspect --format '{{.Mountpoint}}' grafana-data)"
systemctl --user start grafana

The chown gives the volume back to Grafana's user inside the container, 472.

Troubleshooting

Alertmanager does not start, and podman logs alertmanager says no private IP address found, and explicit IP not provided. The --cluster.listen-address= option from step 5 is missing. In a rootless pod, Alertmanager only sees the server's public address, and refuses to start a cluster on it.

A target shows "health":"down", and podman exec prometheus wget -qO- localhost:9090/api/v1/targets shows server returned HTTP status 401 Unauthorized. The password in web1.password does not match the one you gave htpasswd. Check that the file holds only the password.

The target error says x509: certificate signed by unknown authority. web1.crt is not the certificate from the other server, for example after you made a new one there. Copy it again from /etc/node-exporter/node-exporter.crt, and restart Prometheus.

Grafana does not start after a restore, and its log says GF_PATHS_DATA='/var/lib/grafana' is not writable. The podman unshare chown line from step 14 is missing.

Grafana's login page loads, but admin and your password are refused. The password is only set when Grafana creates its database. If Grafana started once before the secret was in place, the password is still admin. Set it to the secret's password, as monitoring:

podman exec grafana grafana cli admin reset-admin-password "$(podman secret inspect --showsecret --format '{{.SecretData}}' grafana-admin-password)"

It ends with Admin password changed successfully.

Run it on your own server

VC-S Micro

€7.99/mo

vCPU
2
Memory
8 GB
Storage
250 GB
Transfer
10 TB
Standard
HDD · RAID 10
  • Full root access
  • Native /64 IPv6
  • RAID-protected storage
  • Malmö, Sweden
  • Month to month, no lock-in
  • 7-day money-back guarantee
All guides