Skip to content

How-to Guides

Task-oriented recipes for running Grafana: installing, configuring for production, scaling, securing, managing dashboards and alerts as code, collecting telemetry with Alloy, upgrading, and troubleshooting. Commands target Grafana 13.x unless marked otherwise. Look-up tables (config defaults, ports, roles, limits) are in Reference; background is in Explanation.

Grafana 13 command and image changes

  • Use grafana server and grafana cli; the grafana-server and grafana-cli binaries were removed in 13.0.
  • Use the grafana/grafana image; grafana/grafana-oss is no longer updated from 12.4.0.
  • Use GF_PLUGINS_PREINSTALL instead of the older GF_INSTALL_PLUGINS.
  • The Grafana Helm chart now lives in grafana-community/helm-charts.

Install Grafana

Choose an Installation Method

Method Recommended for Command or notes
Docker Dev, CI, small production docker run -d -p 3000:3000 grafana/grafana
Helm (Kubernetes) Production grafana-community/grafana chart
APT / YUM / Zypper VMs and bare metal Packages grafana or grafana-enterprise from apt.grafana.com / rpm.grafana.com
Homebrew Local dev on macOS brew install grafana
Standalone binary Air-gapped hosts Tarball from grafana.com/grafana/download
grafana/otel-lgtm Local OTel + LGTM sandbox All-in-one dev image, not for production
Grafana Cloud / Amazon Managed Grafana / Azure Managed Grafana Managed No servers to run; see Reference

Run with Docker

# Grafana OSS with a persistent volume
docker run -d --name grafana \
  -p 3000:3000 \
  -v grafana-data:/var/lib/grafana \
  grafana/grafana:13.2.2

# With environment overrides, preinstalled plugins, and PostgreSQL
docker run -d --name grafana \
  -p 3000:3000 \
  -e GF_SECURITY_ADMIN_PASSWORD__FILE=/run/secrets/admin_password \
  -e GF_PLUGINS_PREINSTALL=grafana-clock-panel,yesoreyeram-infinity-datasource \
  -e GF_DATABASE_TYPE=postgres \
  -e GF_DATABASE_HOST=postgres:5432 \
  -e GF_DATABASE_NAME=grafana \
  -e GF_DATABASE_USER=grafana \
  -e GF_DATABASE_PASSWORD__FILE=/run/secrets/db_password \
  grafana/grafana:13.2.2

The __FILE suffix reads a value from a file (Docker/Kubernetes secrets) instead of putting it in the environment. Pin a version tag in production rather than latest. Pin a plugin version with plugin-id@1.2.3.

Try the Whole Stack Locally

For a local OpenTelemetry sandbox with Grafana, Prometheus, Loki, Tempo, and Pyroscope in one container:

docker run --rm -ti -p 3000:3000 -p 4317:4317 -p 4318:4318 grafana/otel-lgtm

Or run the components separately with Docker Compose (development only; each backend needs its own config file):

# docker-compose.yml: minimal LGTM stack for development
services:
  grafana:
    image: grafana/grafana:13.2.2
    ports: ["3000:3000"]
    environment:
      - GF_SECURITY_ADMIN_PASSWORD=admin
    volumes:
      - grafana-data:/var/lib/grafana
      - ./provisioning:/etc/grafana/provisioning

  mimir:
    image: grafana/mimir:latest
    command: ["-config.file=/etc/mimir/config.yaml", "-target=all"]
    ports: ["9009:9009"]
    volumes:
      - ./mimir-config.yaml:/etc/mimir/config.yaml

  loki:
    image: grafana/loki:latest
    command: ["-config.file=/etc/loki/local-config.yaml"]
    ports: ["3100:3100"]

  tempo:
    image: grafana/tempo:latest
    command: ["-config.file=/etc/tempo/config.yaml"]
    ports:
      - "3200:3200"   # Tempo API
      - "4317:4317"   # OTLP gRPC
      - "4318:4318"   # OTLP HTTP
    volumes:
      - ./tempo-config.yaml:/etc/tempo/config.yaml

  alloy:
    image: grafana/alloy:latest
    command: ["run", "--server.http.listen-addr=0.0.0.0:12345", "/etc/alloy/config.alloy"]
    ports: ["12345:12345"]   # Alloy UI
    volumes:
      - ./config.alloy:/etc/alloy/config.alloy

volumes:
  grafana-data:

Install on Debian or Ubuntu

sudo apt-get install -y apt-transport-https wget gnupg
sudo mkdir -p /etc/apt/keyrings
sudo wget -O /etc/apt/keyrings/grafana.asc https://apt.grafana.com/gpg-full.key
echo "deb [signed-by=/etc/apt/keyrings/grafana.asc] https://apt.grafana.com stable main" \
  | sudo tee /etc/apt/sources.list.d/grafana.list
sudo apt-get update
sudo apt-get install grafana          # or grafana-enterprise
sudo systemctl enable --now grafana-server

The systemd unit is still named grafana-server; only the old wrapper binary was removed.

Install on Kubernetes with Helm

# Grafana chart (community-maintained since 2026-01-30)
helm repo add grafana-community https://grafana-community.github.io/helm-charts
helm repo update

helm install grafana grafana-community/grafana \
  --namespace monitoring --create-namespace \
  -f grafana-values.yaml
# or: helm install grafana oci://ghcr.io/grafana-community/helm-charts/grafana -n monitoring -f grafana-values.yaml

# Alloy and Mimir charts remain in the Grafana Labs repo
helm repo add grafana https://grafana.github.io/helm-charts
helm install alloy grafana/alloy -n monitoring -f alloy-values.yaml
helm install mimir grafana/mimir-distributed -n monitoring -f mimir-values.yaml

# Loki (OSS) and Tempo charts moved to grafana-community as well
helm install loki grafana-community/loki -n monitoring -f loki-values.yaml
helm install tempo grafana-community/tempo-distributed -n monitoring -f tempo-values.yaml

A minimal grafana-values.yaml for an HA setup with an external database:

replicas: 3
persistence:
  enabled: false            # state lives in PostgreSQL
envFromSecret: grafana-env  # GF_DATABASE_PASSWORD, GF_SECURITY_SECRET_KEY, ...
grafana.ini:
  server:
    root_url: https://grafana.example.com
  database:
    type: postgres
    host: postgres.monitoring.svc:5432
    name: grafana
    user: grafana
    ssl_mode: require
  unified_alerting:
    ha_peers: grafana-headless.monitoring.svc.cluster.local:9094

Backend (Mimir, Loki, Tempo) sizing and values are covered in the LGTM Stack how-to guides.

Configure Grafana for Production

Set Essential grafana.ini Options

# === Database (required for HA) ===
[database]
type = postgres
host = postgres.internal:5432
name = grafana
user = grafana
password = $__file{/run/secrets/db_password}
ssl_mode = require

# === Optional shared cache (sessions are stored in the database) ===
[remote_cache]
type = redis
connstr = addr=redis.internal:6379,pool_size=100,db=0

# === Server ===
[server]
http_port = 3000
domain = grafana.example.com
root_url = https://grafana.example.com

# === Security ===
[security]
admin_password = $__file{/run/secrets/admin_password}
secret_key = $__file{/run/secrets/secret_key}
cookie_secure = true
cookie_samesite = lax
content_security_policy = true
strict_transport_security = true

[auth]
login_maximum_inactive_lifetime_duration = 7d
login_maximum_lifetime_duration = 30d

# === Alerting HA ===
[unified_alerting]
enabled = true
ha_peers = grafana-0.grafana-headless:9094,grafana-1.grafana-headless:9094

# === Data proxy ===
[dataproxy]
timeout = 300
dialTimeout = 30
keep_alive_seconds = 30

# === Rendering (13.0+: token must not be empty or "-") ===
[rendering]
server_url = http://renderer:8081/render
callback_url = http://grafana:3000/
renderer_token = $__file{/run/secrets/renderer_token}
concurrent_render_request_limit = 30

Corrected: no Redis session store

Older versions of this note configured [sessions] provider = redis. That section no longer exists: Grafana stores session tokens in its database, so replicas only need a shared MySQL/PostgreSQL database. Redis is optional, for [remote_cache], alerting HA (ha_redis_address) or Grafana Live HA.

$__file{...} and $__env{...} expand secrets from files and environment variables inside grafana.ini.

Override Settings with Environment Variables

Every key can be set with GF_<SECTION>_<KEY> (uppercase; dots and dashes become underscores):

GF_DATABASE_TYPE=postgres
GF_SECURITY_ADMIN_PASSWORD__FILE=/run/secrets/admin_password
GF_AUTH_GENERIC_OAUTH_ENABLED=true
GF_FEATURE_TOGGLES_ENABLE=...   # deprecated in 13.0; set individual toggles instead

Scale and Make Grafana Highly Available

Horizontal Scaling Checklist

  • Move from SQLite to PostgreSQL or MySQL 8.0+ (SQLite cannot be shared).
  • Run 2+ replicas behind a load balancer (no session affinity needed).
  • Configure alerting HA (ha_peers over port 9094, or ha_redis_address) so notifications are sent once; consider ha_single_node_evaluation (13.0+).
  • Run the Image Renderer as its own Deployment if you render images or reports.
  • Terminate TLS at the ingress.
  • Provision data sources and dashboards (files, Git Sync, or Terraform) so replicas are identical.
  • Set resource requests/limits and an HPA on CPU.

The diagram below shows a typical HA topology.

flowchart TB
    LB["Ingress / load balancer<br/>(TLS termination)"]
    subgraph Grafana["Grafana replicas"]
        G1["grafana-0"]
        G2["grafana-1"]
        G3["grafana-2"]
    end
    PG[("PostgreSQL<br/>(HA: RDS / Cloud SQL)")]
    Rend["Image Renderer<br/>Deployment"]
    Redis[("Redis (optional)<br/>remote cache, alerting HA")]
    LB --> G1
    LB --> G2
    LB --> G3
    G1 --> PG
    G2 --> PG
    G3 --> PG
    G1 <-->|"alerting gossip :9094"| G2
    G2 <-->|"alerting gossip :9094"| G3
    G1 -.-> Redis
    G1 --> Rend

Configure Authentication

Set Up OAuth2 / OIDC

Generic OAuth example with Okta groups mapped to roles:

[auth.generic_oauth]
enabled = true
name = Okta
client_id = $__env{OKTA_CLIENT_ID}
client_secret = $__file{/run/secrets/okta_client_secret}
scopes = openid profile email groups
auth_url = https://your-org.okta.com/oauth2/v1/authorize
token_url = https://your-org.okta.com/oauth2/v1/token
api_url = https://your-org.okta.com/oauth2/v1/userinfo
role_attribute_path = contains(groups[*], 'grafana-admins') && 'Admin' || contains(groups[*], 'grafana-editors') && 'Editor' || 'Viewer'
role_attribute_strict = true
allow_sign_up = true
use_pkce = true

Google as a pre-configured provider:

[auth.google]
enabled = true
client_id = <client-id>
client_secret = <client-secret>
allowed_domains = example.com
allow_sign_up = true
auto_login = false
scopes = openid email profile
  • allow_sign_up creates users on first login; auto_login skips the login page.
  • Team sync from IdP groups is an Enterprise/Cloud feature.
  • SSO providers can also be configured in the UI (Administration > Authentication) or through the SSO settings API.

Set Up LDAP

[auth.ldap]
enabled = true
config_file = /etc/grafana/ldap.toml
allow_sign_up = true

LDAP maps directory groups to org roles, supports multiple servers with fallback, and (Enterprise) scheduled background sync. Use TLS, a read-only bind account, ssl_skip_verify = false, and TLS 1.2+.

Set Up SAML (Enterprise / Cloud)

[auth.saml]
enabled = true
name = Corporate IdP
private_key_path = /etc/grafana/saml/private_key.pem
certificate_path = /etc/grafana/saml/certificate.cert
idp_metadata_url = https://idp.example.com/saml/metadata
assertion_attribute_name = DisplayName
assertion_attribute_login = Login
assertion_attribute_email = Email
assertion_attribute_groups = Group

SAML and cookies

The IdP posts back cross-site, so SAML is sensitive to [security] cookie_samesite and cookie_secure. Serve Grafana over HTTPS with cookie_secure = true; check the SAML docs for the SameSite value your login flow (SP- or IdP-initiated) needs.

Put Grafana Behind an Authenticating Proxy

[auth.proxy]
enabled = true
header_name = X-WEBAUTH-USER
header_property = username
auto_sign_up = true
enable_login_token = true
whitelist = 10.0.0.10   # only accept the header from the proxy's address
# Apache example
<Proxy *>
    AuthType Basic
    AuthName GrafanaAuthProxy
    AuthBasicProvider file
    AuthUserFile /etc/apache2/grafana_htpasswd
    Require valid-user
    RewriteEngine On
    RewriteRule .* - [E=PROXY_USER:%{LA-U:REMOTE_USER},NS]
    RequestHeader set X-WEBAUTH-USER "%{PROXY_USER}e"
</Proxy>

Auth proxy trust

Grafana trusts the header completely. Make sure only the proxy can reach Grafana (network policy plus the whitelist setting), otherwise anyone can log in as anyone by sending the header.

Secure Grafana

Harden Authentication

  1. Prefer SSO (OIDC/SAML) and enforce MFA at the IdP; set [auth] disable_login_form = true once SSO works, or enable [auth.basic] password_policy = true.
  2. Keep [auth.anonymous] enabled = false and [users] allow_sign_up = false.
  3. Shorten sessions: login_maximum_lifetime_duration = 12h under [auth].
  4. Serve HTTPS only; set cookie_secure, HSTS, and CSP.
  5. Change the default admin password and secret_key before the first production start.

The full checklist is in Reference.

Enable Content-Security-Policy and Embedding Safely

[security]
content_security_policy = true
# default template (13.x) — customize connect-src/img-src as needed
content_security_policy_template = """script-src 'self' 'unsafe-eval' 'unsafe-inline' 'strict-dynamic' $NONCE;object-src 'none';font-src 'self';style-src 'self' 'unsafe-inline' blob:;img-src * data: blob:;base-uri 'self';connect-src 'self' grafana.com ws://$ROOT_PATH wss://$ROOT_PATH;manifest-src 'self';media-src 'none';form-action 'self';"""

# Only if dashboards must be iframed from another site:
allow_embedding = true
cookie_samesite = none
cookie_secure = true

Prefer shared (public) dashboards over cross-site iframes; if you must iframe, restrict embedding origins with a CSP frame-ancestors directive at the proxy.

Enable Anonymous Viewing for a Kiosk Org

[auth.anonymous]
enabled = true
org_name = Public
org_role = Viewer

Create a separate org for anonymous users and give it only the data sources it needs.

Manage Permissions

  1. Organize dashboards into folders per team; grant folder permissions to teams synced from the IdP rather than to individuals.
  2. Restrict sensitive data sources with data source permissions (Enterprise/Cloud).
  3. Use custom RBAC roles for patterns like "can edit dashboards in folder X only".
  4. Keep Grafana Admin to the platform team.
# Set folder permissions (legacy API)
curl -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  "$GRAFANA_URL/api/folders/$FOLDER_UID/permissions" \
  -d '{"items":[{"role":"Viewer","permission":1},{"role":"Editor","permission":2},{"teamId":5,"permission":4}]}'

Keep Secrets out of Config

  • Use $__file{} / __FILE env vars, Kubernetes Secrets, or External Secrets Operator.
  • Use read-only database users for SQL data sources.
  • Enterprise: store data source credentials in a secrets keeper (Vault, AWS).
  • Enable auditing (Enterprise):
[auditing]
enabled = true

Manage Grafana as Code

Sync Dashboards with Git Sync (13.0+)

  1. Create a repository (GitHub, GitHub Enterprise, GitLab, Bitbucket, or any Git server) and a token or GitHub App with contents and pull-request write access.
  2. In Grafana go to Administration > Provisioning (Git Sync), add the repository, choose the branch and path, and pick folder or folderless (whole-instance) sync.
  3. Choose whether UI saves commit directly or open a pull request; enable branch protection for review.
  4. Configure webhooks so pushes sync immediately (otherwise Grafana polls every 60 s); set root_url so the provider can reach the webhook.
  5. Use Migrate to GitOps to export existing dashboards into the repository.

Self-managed limits are controlled in grafana.ini:

[provisioning]
enabled = true
max_repositories = 10
max_resources_per_repository = 0   # 0 = unlimited; keep each connection under ~1,000

Pull and Push Resources with gcx

brew install gcx
gcx login local --server http://localhost:3000 --token "$SA_TOKEN"

gcx resources pull dashboards -p ./resources -o yaml
gcx resources pull folders -p ./resources -o yaml
gcx resources validate -p ./resources
gcx resources push -p ./resources --dry-run
gcx resources push -p ./resources

gcx works with Grafana 12+; writing Grafana-managed alert rules through it needs Grafana 13 (or the kubernetesAlertingRules toggle on 12.x). It replaces grafanactl, archived on 2026-06-01.

Provision Data Sources from Files

# /etc/grafana/provisioning/datasources/datasources.yaml
apiVersion: 1
datasources:
  - name: Mimir
    type: prometheus
    uid: mimir
    access: proxy
    url: http://mimir-query-frontend:8080/prometheus
    isDefault: true
    jsonData:
      httpMethod: POST
      exemplarTraceIdDestinations:
        - name: traceID
          datasourceUid: tempo

  - name: Loki
    type: loki
    uid: loki
    access: proxy
    url: http://loki-gateway:3100
    jsonData:
      derivedFields:
        - datasourceUid: tempo
          matcherRegex: '"traceID":"(\w+)"'
          name: TraceID
          url: '$${__value.raw}'

  - name: Tempo
    type: tempo
    uid: tempo
    access: proxy
    url: http://tempo-query-frontend:3200
    jsonData:
      tracesToMetrics:
        datasourceUid: mimir
      tracesToLogsV2:
        datasourceUid: loki
        tags: [{ key: 'service.name', value: 'service_name' }]

Set explicit uids so cross-links (datasourceUid) and dashboards stay stable; Grafana 12 enforces a stricter UID format.

Provision Dashboards from Files

# /etc/grafana/provisioning/dashboards/provider.yaml
apiVersion: 1
providers:
  - name: Infrastructure
    orgId: 1
    folder: Infrastructure
    type: file
    editable: false
    options:
      path: /etc/grafana/provisioning/dashboards/infra
      foldersFromFilesStructure: true

File provisioning accepts v1 JSON and v2 resources (v2 support added in 12.4).

Provision Alert Rules and Contact Points

# /etc/grafana/provisioning/alerting/rules.yaml
apiVersion: 1
groups:
  - orgId: 1
    name: infrastructure-alerts
    folder: Infrastructure Alerts
    interval: 1m
    rules:
      - uid: high-cpu-alert
        title: High CPU Usage
        condition: A
        data:
          - refId: A
            relativeTimeRange: { from: 600, to: 0 }
            datasourceUid: mimir
            model:
              expr: '100 - (avg by(instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100) > 80'
              refId: A
        for: 5m
        keepFiringFor: 5m
        labels:
          severity: warning
          team: infra
        annotations:
          summary: "CPU usage above 80% on {{ $labels.instance }}"
# /etc/grafana/provisioning/alerting/contactpoints.yaml
apiVersion: 1
contactPoints:
  - orgId: 1
    name: slack-oncall
    receivers:
      - uid: slack-receiver
        type: slack
        settings:
          recipient: '#alerts-oncall'
          token: '$SLACK_BOT_TOKEN'
          title: '{{ template "slack.default.title" . }}'
          text: '{{ template "slack.default.text" . }}'

Manage Grafana with Terraform

terraform {
  required_providers {
    grafana = {
      source  = "grafana/grafana"
      version = ">= 3.0.0"
    }
  }
}

provider "grafana" {
  url  = "https://grafana.example.com"
  auth = var.grafana_service_account_token
}

resource "grafana_folder" "infrastructure" {
  title = "Infrastructure"
}

resource "grafana_dashboard" "node_overview" {
  config_json = file("${path.module}/dashboards/node-overview.json")
  folder      = grafana_folder.infrastructure.uid
  overwrite   = true
}

resource "grafana_data_source" "mimir" {
  type       = "prometheus"
  name       = "Mimir"
  url        = "http://mimir-query-frontend:8080/prometheus"
  is_default = true
  json_data_encoded = jsonencode({ httpMethod = "POST" })
}

Collect Telemetry with Alloy

Forward OTLP to Grafana Cloud (Alloy Syntax)

// config.alloy: OTLP in, batch, OTLP/HTTP out to Grafana Cloud
otelcol.receiver.otlp "default" {
  grpc { endpoint = "0.0.0.0:4317" }
  http { endpoint = "0.0.0.0:4318" }
  output {
    metrics = [otelcol.processor.batch.default.input]
    logs    = [otelcol.processor.batch.default.input]
    traces  = [otelcol.processor.batch.default.input]
  }
}

otelcol.processor.batch "default" {
  output {
    metrics = [otelcol.exporter.otlphttp.grafana_cloud.input]
    logs    = [otelcol.exporter.otlphttp.grafana_cloud.input]
    traces  = [otelcol.exporter.otlphttp.grafana_cloud.input]
  }
}

otelcol.auth.basic "grafana_cloud" {
  username = sys.env("GRAFANA_CLOUD_INSTANCE_ID")
  password = sys.env("GRAFANA_CLOUD_API_KEY")
}

otelcol.exporter.otlphttp "grafana_cloud" {
  client {
    endpoint = sys.env("GRAFANA_CLOUD_OTLP_ENDPOINT")
    auth     = otelcol.auth.basic.grafana_cloud.handler
  }
}

sys.env() is the current standard-library name for the older top-level env() function.

Scrape Kubernetes Pods into Mimir

discovery.kubernetes "pods" {
  role = "pod"
}

prometheus.scrape "kubernetes_pods" {
  targets    = discovery.kubernetes.pods.targets
  forward_to = [prometheus.remote_write.mimir.receiver]
}

prometheus.remote_write "mimir" {
  endpoint {
    url = "http://mimir-distributor:8080/api/v1/push"
  }
}

Run Upstream OTel Collector YAML (Experimental)

alloy otel validate --config=otel-config.yaml
alloy otel --config=otel-config.yaml

Migrate from Grafana Agent

Grafana Agent reached end of life on 2025-11-01. Convert existing configs with the built-in converter, then review the output:

alloy convert --source-format=static --output=config.alloy agent-static.yaml
alloy convert --source-format=prometheus --output=config.alloy prometheus.yaml
alloy convert --source-format=otelcol --output=config.alloy otelcol.yaml

Agent Flow configs are largely compatible with Alloy syntax.

Use the API

# Service account token (recommended; API keys were removed in 12.1)
curl -H "Authorization: Bearer $TOKEN" "$GRAFANA_URL/api/dashboards/home"

# Create a service account and a token
curl -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"name": "automation-sa", "role": "Editor"}' "$GRAFANA_URL/api/serviceaccounts"
curl -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"name": "ci-token"}' "$GRAFANA_URL/api/serviceaccounts/$SA_ID/tokens"

# List dashboards
curl -s -H "Authorization: Bearer $TOKEN" \
  "$GRAFANA_URL/api/search?type=dash-db" | jq '.[] | {title, uid, folderTitle}'

# Get a dashboard as an app-platform resource (13.x)
curl -s -H "Authorization: Bearer $TOKEN" \
  "$GRAFANA_URL/apis/dashboard.grafana.app/v1/namespaces/default/dashboards/$UID" | jq .

# Export a dashboard (legacy API) and re-import into a folder
curl -s -H "Authorization: Bearer $TOKEN" "$GRAFANA_URL/api/dashboards/uid/$UID" \
  | jq '.dashboard' > dashboard-export.json
curl -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d "{\"dashboard\": $(cat dashboard-export.json), \"overwrite\": true, \"folderUid\": \"$FOLDER_UID\"}" \
  "$GRAFANA_URL/api/dashboards/db"

# Data sources (reference by UID; numeric-ID APIs are disabled by default in 13.0)
curl -s -H "Authorization: Bearer $TOKEN" "$GRAFANA_URL/api/datasources" | jq '.[] | {name, uid, type, url}'
curl -s -H "Authorization: Bearer $TOKEN" "$GRAFANA_URL/api/datasources/uid/mimir/health"
curl -s -H "Authorization: Bearer $TOKEN" \
  "$GRAFANA_URL/api/datasources/proxy/uid/mimir/api/v1/query?query=up"

# Alerting (legacy provisioning API; alert-rule endpoints deprecated in 13.0)
curl -s -H "Authorization: Bearer $TOKEN" "$GRAFANA_URL/api/v1/provisioning/alert-rules" | jq .
curl -s -H "Authorization: Bearer $TOKEN" "$GRAFANA_URL/api/v1/provisioning/contact-points" | jq .

# Contact points via the app-platform API (13.x)
curl -s -H "Authorization: Bearer $TOKEN" \
  "$GRAFANA_URL/apis/notifications.alerting.grafana.app/v1beta1/namespaces/default/receivers" | jq .

Useful One-Liners

# Back up every dashboard (legacy API)
for uid in $(curl -s -H "Authorization: Bearer $TOKEN" "$GRAFANA_URL/api/search?type=dash-db" | jq -r '.[].uid'); do
  curl -s -H "Authorization: Bearer $TOKEN" "$GRAFANA_URL/api/dashboards/uid/$uid" | jq '.dashboard' > "backup-$uid.json"
done

# Health and version
curl -s "$GRAFANA_URL/api/health" | jq .
curl -s -H "Authorization: Bearer $TOKEN" "$GRAFANA_URL/api/frontend/settings" | jq '.buildInfo'

# Count dashboards per folder
curl -s -H "Authorization: Bearer $TOKEN" "$GRAFANA_URL/api/search?type=dash-db" \
  | jq 'group_by(.folderTitle) | map({folder: .[0].folderTitle, count: length})'

Manage Plugins and the Admin Account

grafana cli plugins install grafana-clock-panel
grafana cli plugins install grafana-clock-panel <version>  # pin a specific version
grafana cli plugins ls
grafana cli plugins list-versions grafana-clock-panel
grafana cli plugins update-all
grafana cli plugins uninstall grafana-clock-panel
grafana cli --pluginUrl https://example.com/custom-plugin.zip plugins install custom-plugin

# Reset the admin password (read it from stdin to keep it out of shell history)
grafana cli --homepath /usr/share/grafana --config /etc/grafana/grafana.ini \
  admin reset-admin-password --password-from-stdin

# Move plaintext data source passwords into secure_json_data
grafana cli admin data-migration encrypt-datasource-passwords

# Re-encrypt secrets after rotating the encryption key
grafana cli admin secrets-migration re-encrypt

Bake plugins into a custom image (or use [plugins] preinstall) rather than installing at runtime on every pod start. The old Pie Chart and Worldmap Angular plugins no longer load since 12.0; use the built-in Pie chart and Geomap panels.

Upgrade Grafana

  1. Back up the database (and grafana.ini, plugins directory). On 13.0+ this is the only rollback path because of the unified storage migration.
  2. Upgrade to the latest patch of your current minor first, then update all plugins (React 19 compatibility for 13.x).
  3. Read the upgrade guide and breaking changes for each major you cross (v13.0 guide).
  4. For 13.x specifically:
    • skip 13.0.0 (withdrawn) and go to the latest 13.x patch;
    • replace grafana-cli / grafana-server in scripts and systemd overrides;
    • run the Image Renderer as a service and set a non-default renderer_token;
    • switch API calls from numeric data source IDs to UIDs (or temporarily enable datasourceLegacyIdApi);
    • on SQLite, if the migration hits database is locked, set [unified_storage] migration_parquet_buffer = true.
  5. Upgrade one replica, watch logs and grafana_unified_storage_migration_status, then roll the rest.

Operate Grafana Well

Dashboard Governance

  1. Use folders per team or domain, and team folders (13.0) to record ownership.
  2. Manage critical dashboards as code (Git Sync, provisioning, or Terraform) to prevent drift.
  3. Every dashboard has an owner; review quarterly and archive unused ones (deleted dashboards can be restored).
  4. Name dashboards with a team or domain prefix, for example [infra] Node Overview.
  5. Use variables (or quick filters in 13.1+) for environment, region and service.
  6. Aim for 8–12 panels on overview dashboards and 15–20 on detailed ones; use tabs and rows in dynamic dashboards to split larger ones.

Optimize Queries

  1. Filter early with precise label selectors in PromQL/LogQL.
  2. Avoid high-cardinality labels (user IDs, IPs, request IDs, raw paths). For Loki, keep roughly 10 or fewer indexed labels with low-cardinality values and put pod names or trace IDs in structured metadata; aim for fewer than ~10k active streams per tenant.
  3. Pre-compute expensive queries with recording rules in Mimir/Prometheus (or Grafana-managed recording rules).
  4. Set Max data points and use $__interval / $__rate_interval.
  5. Avoid refresh intervals under 10 s unless needed.
  6. Use saved queries (GA in 13.2 for Enterprise/Cloud) to share vetted queries.

Monitor Grafana Itself

Scrape /metrics from every replica (for example with Alloy or kube-prometheus-stack) and alert on request latency, data proxy latency, alert evaluation failures, and grafana_alerting_scheduler_behind_seconds. Metric names are listed in Reference.

Troubleshooting Tools

  1. Query inspector: raw query, response, timing per panel.
  2. Server logs: stdout in containers or /var/log/grafana/grafana.log; raise [log] level = debug temporarily.
  3. Grafana Advisor (GA in 13.0): automated checks for plugins, data sources, and configuration.
  4. Tracing: Grafana can export its own traces over OTLP, or to a file as OTLP/JSON (13.2).
  5. Alloy UI: http://localhost:12345 for the live pipeline graph and component health.

Troubleshooting Playbook

Symptom Likely cause Fix
Dashboard loads slowly Expensive queries, too many panels, short refresh Query inspector, recording rules, fewer panels, longer refresh
"Data source is not available" Wrong URL, network, or proxy mode Check the URL from the Grafana host, use the data source health check
Alerts not firing Rule paused, evaluation errors, or contact point broken Check rule health, grafana_alerting_rule_evaluation_failures_total, test the contact point
Duplicate notifications in HA Alerting HA not configured Set ha_peers (port 9094 reachable) or ha_redis_address
Login loop root_url mismatch or cookie_secure over plain HTTP Fix root_url, serve over HTTPS
"database is locked" SQLite with several replicas, or SQLite during the 13.0 migration Move to PostgreSQL/MySQL; for the migration enable migration_parquet_buffer
Plugin not loading Unsigned plugin or Angular plugin on 12.0+ Allow-list the ID with allow_loading_unsigned_plugins, or replace Angular plugins
Scripts fail with grafana-cli: command not found 13.0 removed the binary Use grafana cli
Reports and PNG exports fail after 13.0 Renderer plugin removed / token not set Run the renderer service and set renderer_token
API returns 404 for /api/datasources/<id> Numeric-ID APIs disabled in 13.0 Use /api/datasources/uid/<uid>
High memory on Grafana pods Many concurrent viewers or heavy SQL data sources Scale out, enable query caching (Enterprise), reduce refresh

Migrate from ELK to Loki and Grafana

  1. Deploy Alloy next to existing Logstash/Filebeat and dual-ship logs to Loki (Grafana recommends Alloy over the Logstash Loki output plugin).
  2. Choose a small set of low-cardinality labels; move other fields into structured metadata.
  3. Configure object storage (S3/GCS/Azure) and retention for Loki.
  4. Rebuild Kibana dashboards in Grafana (no automatic converter) and rewrite queries in LogQL.
  5. Keep Elasticsearch read-only for historical data until its retention expires, then decommission.

Sources