How-to Guides¶
Task-oriented recipes for running Grafana: installing, configuring for production, scaling, securing, managing dashboards and alerts as code, collecting telemetry with Alloy, upgrading, and troubleshooting. Commands target Grafana 13.x unless marked otherwise. Look-up tables (config defaults, ports, roles, limits) are in Reference; background is in Explanation.
Grafana 13 command and image changes
- Use
grafana serverandgrafana cli; thegrafana-serverandgrafana-clibinaries were removed in 13.0. - Use the
grafana/grafanaimage;grafana/grafana-ossis no longer updated from 12.4.0. - Use
GF_PLUGINS_PREINSTALLinstead of the olderGF_INSTALL_PLUGINS. - The Grafana Helm chart now lives in
grafana-community/helm-charts.
Install Grafana¶
Choose an Installation Method¶
| Method | Recommended for | Command or notes |
|---|---|---|
| Docker | Dev, CI, small production | docker run -d -p 3000:3000 grafana/grafana |
| Helm (Kubernetes) | Production | grafana-community/grafana chart |
| APT / YUM / Zypper | VMs and bare metal | Packages grafana or grafana-enterprise from apt.grafana.com / rpm.grafana.com |
| Homebrew | Local dev on macOS | brew install grafana |
| Standalone binary | Air-gapped hosts | Tarball from grafana.com/grafana/download |
grafana/otel-lgtm |
Local OTel + LGTM sandbox | All-in-one dev image, not for production |
| Grafana Cloud / Amazon Managed Grafana / Azure Managed Grafana | Managed | No servers to run; see Reference |
Run with Docker¶
# Grafana OSS with a persistent volume
docker run -d --name grafana \
-p 3000:3000 \
-v grafana-data:/var/lib/grafana \
grafana/grafana:13.2.2
# With environment overrides, preinstalled plugins, and PostgreSQL
docker run -d --name grafana \
-p 3000:3000 \
-e GF_SECURITY_ADMIN_PASSWORD__FILE=/run/secrets/admin_password \
-e GF_PLUGINS_PREINSTALL=grafana-clock-panel,yesoreyeram-infinity-datasource \
-e GF_DATABASE_TYPE=postgres \
-e GF_DATABASE_HOST=postgres:5432 \
-e GF_DATABASE_NAME=grafana \
-e GF_DATABASE_USER=grafana \
-e GF_DATABASE_PASSWORD__FILE=/run/secrets/db_password \
grafana/grafana:13.2.2
The __FILE suffix reads a value from a file (Docker/Kubernetes secrets) instead of putting it in the environment. Pin a version tag in production rather than latest. Pin a plugin version with plugin-id@1.2.3.
Try the Whole Stack Locally¶
For a local OpenTelemetry sandbox with Grafana, Prometheus, Loki, Tempo, and Pyroscope in one container:
Or run the components separately with Docker Compose (development only; each backend needs its own config file):
# docker-compose.yml: minimal LGTM stack for development
services:
grafana:
image: grafana/grafana:13.2.2
ports: ["3000:3000"]
environment:
- GF_SECURITY_ADMIN_PASSWORD=admin
volumes:
- grafana-data:/var/lib/grafana
- ./provisioning:/etc/grafana/provisioning
mimir:
image: grafana/mimir:latest
command: ["-config.file=/etc/mimir/config.yaml", "-target=all"]
ports: ["9009:9009"]
volumes:
- ./mimir-config.yaml:/etc/mimir/config.yaml
loki:
image: grafana/loki:latest
command: ["-config.file=/etc/loki/local-config.yaml"]
ports: ["3100:3100"]
tempo:
image: grafana/tempo:latest
command: ["-config.file=/etc/tempo/config.yaml"]
ports:
- "3200:3200" # Tempo API
- "4317:4317" # OTLP gRPC
- "4318:4318" # OTLP HTTP
volumes:
- ./tempo-config.yaml:/etc/tempo/config.yaml
alloy:
image: grafana/alloy:latest
command: ["run", "--server.http.listen-addr=0.0.0.0:12345", "/etc/alloy/config.alloy"]
ports: ["12345:12345"] # Alloy UI
volumes:
- ./config.alloy:/etc/alloy/config.alloy
volumes:
grafana-data:
Install on Debian or Ubuntu¶
sudo apt-get install -y apt-transport-https wget gnupg
sudo mkdir -p /etc/apt/keyrings
sudo wget -O /etc/apt/keyrings/grafana.asc https://apt.grafana.com/gpg-full.key
echo "deb [signed-by=/etc/apt/keyrings/grafana.asc] https://apt.grafana.com stable main" \
| sudo tee /etc/apt/sources.list.d/grafana.list
sudo apt-get update
sudo apt-get install grafana # or grafana-enterprise
sudo systemctl enable --now grafana-server
The systemd unit is still named grafana-server; only the old wrapper binary was removed.
Install on Kubernetes with Helm¶
# Grafana chart (community-maintained since 2026-01-30)
helm repo add grafana-community https://grafana-community.github.io/helm-charts
helm repo update
helm install grafana grafana-community/grafana \
--namespace monitoring --create-namespace \
-f grafana-values.yaml
# or: helm install grafana oci://ghcr.io/grafana-community/helm-charts/grafana -n monitoring -f grafana-values.yaml
# Alloy and Mimir charts remain in the Grafana Labs repo
helm repo add grafana https://grafana.github.io/helm-charts
helm install alloy grafana/alloy -n monitoring -f alloy-values.yaml
helm install mimir grafana/mimir-distributed -n monitoring -f mimir-values.yaml
# Loki (OSS) and Tempo charts moved to grafana-community as well
helm install loki grafana-community/loki -n monitoring -f loki-values.yaml
helm install tempo grafana-community/tempo-distributed -n monitoring -f tempo-values.yaml
A minimal grafana-values.yaml for an HA setup with an external database:
replicas: 3
persistence:
enabled: false # state lives in PostgreSQL
envFromSecret: grafana-env # GF_DATABASE_PASSWORD, GF_SECURITY_SECRET_KEY, ...
grafana.ini:
server:
root_url: https://grafana.example.com
database:
type: postgres
host: postgres.monitoring.svc:5432
name: grafana
user: grafana
ssl_mode: require
unified_alerting:
ha_peers: grafana-headless.monitoring.svc.cluster.local:9094
Backend (Mimir, Loki, Tempo) sizing and values are covered in the LGTM Stack how-to guides.
Configure Grafana for Production¶
Set Essential grafana.ini Options¶
# === Database (required for HA) ===
[database]
type = postgres
host = postgres.internal:5432
name = grafana
user = grafana
password = $__file{/run/secrets/db_password}
ssl_mode = require
# === Optional shared cache (sessions are stored in the database) ===
[remote_cache]
type = redis
connstr = addr=redis.internal:6379,pool_size=100,db=0
# === Server ===
[server]
http_port = 3000
domain = grafana.example.com
root_url = https://grafana.example.com
# === Security ===
[security]
admin_password = $__file{/run/secrets/admin_password}
secret_key = $__file{/run/secrets/secret_key}
cookie_secure = true
cookie_samesite = lax
content_security_policy = true
strict_transport_security = true
[auth]
login_maximum_inactive_lifetime_duration = 7d
login_maximum_lifetime_duration = 30d
# === Alerting HA ===
[unified_alerting]
enabled = true
ha_peers = grafana-0.grafana-headless:9094,grafana-1.grafana-headless:9094
# === Data proxy ===
[dataproxy]
timeout = 300
dialTimeout = 30
keep_alive_seconds = 30
# === Rendering (13.0+: token must not be empty or "-") ===
[rendering]
server_url = http://renderer:8081/render
callback_url = http://grafana:3000/
renderer_token = $__file{/run/secrets/renderer_token}
concurrent_render_request_limit = 30
Corrected: no Redis session store
Older versions of this note configured [sessions] provider = redis. That section no longer exists: Grafana stores session tokens in its database, so replicas only need a shared MySQL/PostgreSQL database. Redis is optional, for [remote_cache], alerting HA (ha_redis_address) or Grafana Live HA.
$__file{...} and $__env{...} expand secrets from files and environment variables inside grafana.ini.
Override Settings with Environment Variables¶
Every key can be set with GF_<SECTION>_<KEY> (uppercase; dots and dashes become underscores):
GF_DATABASE_TYPE=postgres
GF_SECURITY_ADMIN_PASSWORD__FILE=/run/secrets/admin_password
GF_AUTH_GENERIC_OAUTH_ENABLED=true
GF_FEATURE_TOGGLES_ENABLE=... # deprecated in 13.0; set individual toggles instead
Scale and Make Grafana Highly Available¶
Horizontal Scaling Checklist¶
- Move from SQLite to PostgreSQL or MySQL 8.0+ (SQLite cannot be shared).
- Run 2+ replicas behind a load balancer (no session affinity needed).
- Configure alerting HA (
ha_peersover port 9094, orha_redis_address) so notifications are sent once; considerha_single_node_evaluation(13.0+). - Run the Image Renderer as its own Deployment if you render images or reports.
- Terminate TLS at the ingress.
- Provision data sources and dashboards (files, Git Sync, or Terraform) so replicas are identical.
- Set resource requests/limits and an HPA on CPU.
The diagram below shows a typical HA topology.
flowchart TB
LB["Ingress / load balancer<br/>(TLS termination)"]
subgraph Grafana["Grafana replicas"]
G1["grafana-0"]
G2["grafana-1"]
G3["grafana-2"]
end
PG[("PostgreSQL<br/>(HA: RDS / Cloud SQL)")]
Rend["Image Renderer<br/>Deployment"]
Redis[("Redis (optional)<br/>remote cache, alerting HA")]
LB --> G1
LB --> G2
LB --> G3
G1 --> PG
G2 --> PG
G3 --> PG
G1 <-->|"alerting gossip :9094"| G2
G2 <-->|"alerting gossip :9094"| G3
G1 -.-> Redis
G1 --> Rend
Configure Authentication¶
Set Up OAuth2 / OIDC¶
Generic OAuth example with Okta groups mapped to roles:
[auth.generic_oauth]
enabled = true
name = Okta
client_id = $__env{OKTA_CLIENT_ID}
client_secret = $__file{/run/secrets/okta_client_secret}
scopes = openid profile email groups
auth_url = https://your-org.okta.com/oauth2/v1/authorize
token_url = https://your-org.okta.com/oauth2/v1/token
api_url = https://your-org.okta.com/oauth2/v1/userinfo
role_attribute_path = contains(groups[*], 'grafana-admins') && 'Admin' || contains(groups[*], 'grafana-editors') && 'Editor' || 'Viewer'
role_attribute_strict = true
allow_sign_up = true
use_pkce = true
Google as a pre-configured provider:
[auth.google]
enabled = true
client_id = <client-id>
client_secret = <client-secret>
allowed_domains = example.com
allow_sign_up = true
auto_login = false
scopes = openid email profile
allow_sign_upcreates users on first login;auto_loginskips the login page.- Team sync from IdP groups is an Enterprise/Cloud feature.
- SSO providers can also be configured in the UI (Administration > Authentication) or through the SSO settings API.
Set Up LDAP¶
LDAP maps directory groups to org roles, supports multiple servers with fallback, and (Enterprise) scheduled background sync. Use TLS, a read-only bind account, ssl_skip_verify = false, and TLS 1.2+.
Set Up SAML (Enterprise / Cloud)¶
[auth.saml]
enabled = true
name = Corporate IdP
private_key_path = /etc/grafana/saml/private_key.pem
certificate_path = /etc/grafana/saml/certificate.cert
idp_metadata_url = https://idp.example.com/saml/metadata
assertion_attribute_name = DisplayName
assertion_attribute_login = Login
assertion_attribute_email = Email
assertion_attribute_groups = Group
SAML and cookies
The IdP posts back cross-site, so SAML is sensitive to [security] cookie_samesite and cookie_secure. Serve Grafana over HTTPS with cookie_secure = true; check the SAML docs for the SameSite value your login flow (SP- or IdP-initiated) needs.
Put Grafana Behind an Authenticating Proxy¶
[auth.proxy]
enabled = true
header_name = X-WEBAUTH-USER
header_property = username
auto_sign_up = true
enable_login_token = true
whitelist = 10.0.0.10 # only accept the header from the proxy's address
# Apache example
<Proxy *>
AuthType Basic
AuthName GrafanaAuthProxy
AuthBasicProvider file
AuthUserFile /etc/apache2/grafana_htpasswd
Require valid-user
RewriteEngine On
RewriteRule .* - [E=PROXY_USER:%{LA-U:REMOTE_USER},NS]
RequestHeader set X-WEBAUTH-USER "%{PROXY_USER}e"
</Proxy>
Auth proxy trust
Grafana trusts the header completely. Make sure only the proxy can reach Grafana (network policy plus the whitelist setting), otherwise anyone can log in as anyone by sending the header.
Secure Grafana¶
Harden Authentication¶
- Prefer SSO (OIDC/SAML) and enforce MFA at the IdP; set
[auth] disable_login_form = trueonce SSO works, or enable[auth.basic] password_policy = true. - Keep
[auth.anonymous] enabled = falseand[users] allow_sign_up = false. - Shorten sessions:
login_maximum_lifetime_duration = 12hunder[auth]. - Serve HTTPS only; set
cookie_secure, HSTS, and CSP. - Change the default admin password and
secret_keybefore the first production start.
The full checklist is in Reference.
Enable Content-Security-Policy and Embedding Safely¶
[security]
content_security_policy = true
# default template (13.x) — customize connect-src/img-src as needed
content_security_policy_template = """script-src 'self' 'unsafe-eval' 'unsafe-inline' 'strict-dynamic' $NONCE;object-src 'none';font-src 'self';style-src 'self' 'unsafe-inline' blob:;img-src * data: blob:;base-uri 'self';connect-src 'self' grafana.com ws://$ROOT_PATH wss://$ROOT_PATH;manifest-src 'self';media-src 'none';form-action 'self';"""
# Only if dashboards must be iframed from another site:
allow_embedding = true
cookie_samesite = none
cookie_secure = true
Prefer shared (public) dashboards over cross-site iframes; if you must iframe, restrict embedding origins with a CSP frame-ancestors directive at the proxy.
Enable Anonymous Viewing for a Kiosk Org¶
Create a separate org for anonymous users and give it only the data sources it needs.
Manage Permissions¶
- Organize dashboards into folders per team; grant folder permissions to teams synced from the IdP rather than to individuals.
- Restrict sensitive data sources with data source permissions (Enterprise/Cloud).
- Use custom RBAC roles for patterns like "can edit dashboards in folder X only".
- Keep Grafana Admin to the platform team.
# Set folder permissions (legacy API)
curl -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
"$GRAFANA_URL/api/folders/$FOLDER_UID/permissions" \
-d '{"items":[{"role":"Viewer","permission":1},{"role":"Editor","permission":2},{"teamId":5,"permission":4}]}'
Keep Secrets out of Config¶
- Use
$__file{}/__FILEenv vars, Kubernetes Secrets, or External Secrets Operator. - Use read-only database users for SQL data sources.
- Enterprise: store data source credentials in a secrets keeper (Vault, AWS).
- Enable auditing (Enterprise):
Manage Grafana as Code¶
Sync Dashboards with Git Sync (13.0+)¶
- Create a repository (GitHub, GitHub Enterprise, GitLab, Bitbucket, or any Git server) and a token or GitHub App with contents and pull-request write access.
- In Grafana go to Administration > Provisioning (Git Sync), add the repository, choose the branch and path, and pick folder or folderless (whole-instance) sync.
- Choose whether UI saves commit directly or open a pull request; enable branch protection for review.
- Configure webhooks so pushes sync immediately (otherwise Grafana polls every 60 s); set
root_urlso the provider can reach the webhook. - Use Migrate to GitOps to export existing dashboards into the repository.
Self-managed limits are controlled in grafana.ini:
[provisioning]
enabled = true
max_repositories = 10
max_resources_per_repository = 0 # 0 = unlimited; keep each connection under ~1,000
Pull and Push Resources with gcx¶
brew install gcx
gcx login local --server http://localhost:3000 --token "$SA_TOKEN"
gcx resources pull dashboards -p ./resources -o yaml
gcx resources pull folders -p ./resources -o yaml
gcx resources validate -p ./resources
gcx resources push -p ./resources --dry-run
gcx resources push -p ./resources
gcx works with Grafana 12+; writing Grafana-managed alert rules through it needs Grafana 13 (or the kubernetesAlertingRules toggle on 12.x). It replaces grafanactl, archived on 2026-06-01.
Provision Data Sources from Files¶
# /etc/grafana/provisioning/datasources/datasources.yaml
apiVersion: 1
datasources:
- name: Mimir
type: prometheus
uid: mimir
access: proxy
url: http://mimir-query-frontend:8080/prometheus
isDefault: true
jsonData:
httpMethod: POST
exemplarTraceIdDestinations:
- name: traceID
datasourceUid: tempo
- name: Loki
type: loki
uid: loki
access: proxy
url: http://loki-gateway:3100
jsonData:
derivedFields:
- datasourceUid: tempo
matcherRegex: '"traceID":"(\w+)"'
name: TraceID
url: '$${__value.raw}'
- name: Tempo
type: tempo
uid: tempo
access: proxy
url: http://tempo-query-frontend:3200
jsonData:
tracesToMetrics:
datasourceUid: mimir
tracesToLogsV2:
datasourceUid: loki
tags: [{ key: 'service.name', value: 'service_name' }]
Set explicit uids so cross-links (datasourceUid) and dashboards stay stable; Grafana 12 enforces a stricter UID format.
Provision Dashboards from Files¶
# /etc/grafana/provisioning/dashboards/provider.yaml
apiVersion: 1
providers:
- name: Infrastructure
orgId: 1
folder: Infrastructure
type: file
editable: false
options:
path: /etc/grafana/provisioning/dashboards/infra
foldersFromFilesStructure: true
File provisioning accepts v1 JSON and v2 resources (v2 support added in 12.4).
Provision Alert Rules and Contact Points¶
# /etc/grafana/provisioning/alerting/rules.yaml
apiVersion: 1
groups:
- orgId: 1
name: infrastructure-alerts
folder: Infrastructure Alerts
interval: 1m
rules:
- uid: high-cpu-alert
title: High CPU Usage
condition: A
data:
- refId: A
relativeTimeRange: { from: 600, to: 0 }
datasourceUid: mimir
model:
expr: '100 - (avg by(instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100) > 80'
refId: A
for: 5m
keepFiringFor: 5m
labels:
severity: warning
team: infra
annotations:
summary: "CPU usage above 80% on {{ $labels.instance }}"
# /etc/grafana/provisioning/alerting/contactpoints.yaml
apiVersion: 1
contactPoints:
- orgId: 1
name: slack-oncall
receivers:
- uid: slack-receiver
type: slack
settings:
recipient: '#alerts-oncall'
token: '$SLACK_BOT_TOKEN'
title: '{{ template "slack.default.title" . }}'
text: '{{ template "slack.default.text" . }}'
Manage Grafana with Terraform¶
terraform {
required_providers {
grafana = {
source = "grafana/grafana"
version = ">= 3.0.0"
}
}
}
provider "grafana" {
url = "https://grafana.example.com"
auth = var.grafana_service_account_token
}
resource "grafana_folder" "infrastructure" {
title = "Infrastructure"
}
resource "grafana_dashboard" "node_overview" {
config_json = file("${path.module}/dashboards/node-overview.json")
folder = grafana_folder.infrastructure.uid
overwrite = true
}
resource "grafana_data_source" "mimir" {
type = "prometheus"
name = "Mimir"
url = "http://mimir-query-frontend:8080/prometheus"
is_default = true
json_data_encoded = jsonencode({ httpMethod = "POST" })
}
Collect Telemetry with Alloy¶
Forward OTLP to Grafana Cloud (Alloy Syntax)¶
// config.alloy: OTLP in, batch, OTLP/HTTP out to Grafana Cloud
otelcol.receiver.otlp "default" {
grpc { endpoint = "0.0.0.0:4317" }
http { endpoint = "0.0.0.0:4318" }
output {
metrics = [otelcol.processor.batch.default.input]
logs = [otelcol.processor.batch.default.input]
traces = [otelcol.processor.batch.default.input]
}
}
otelcol.processor.batch "default" {
output {
metrics = [otelcol.exporter.otlphttp.grafana_cloud.input]
logs = [otelcol.exporter.otlphttp.grafana_cloud.input]
traces = [otelcol.exporter.otlphttp.grafana_cloud.input]
}
}
otelcol.auth.basic "grafana_cloud" {
username = sys.env("GRAFANA_CLOUD_INSTANCE_ID")
password = sys.env("GRAFANA_CLOUD_API_KEY")
}
otelcol.exporter.otlphttp "grafana_cloud" {
client {
endpoint = sys.env("GRAFANA_CLOUD_OTLP_ENDPOINT")
auth = otelcol.auth.basic.grafana_cloud.handler
}
}
sys.env() is the current standard-library name for the older top-level env() function.
Scrape Kubernetes Pods into Mimir¶
discovery.kubernetes "pods" {
role = "pod"
}
prometheus.scrape "kubernetes_pods" {
targets = discovery.kubernetes.pods.targets
forward_to = [prometheus.remote_write.mimir.receiver]
}
prometheus.remote_write "mimir" {
endpoint {
url = "http://mimir-distributor:8080/api/v1/push"
}
}
Run Upstream OTel Collector YAML (Experimental)¶
Migrate from Grafana Agent¶
Grafana Agent reached end of life on 2025-11-01. Convert existing configs with the built-in converter, then review the output:
alloy convert --source-format=static --output=config.alloy agent-static.yaml
alloy convert --source-format=prometheus --output=config.alloy prometheus.yaml
alloy convert --source-format=otelcol --output=config.alloy otelcol.yaml
Agent Flow configs are largely compatible with Alloy syntax.
Use the API¶
# Service account token (recommended; API keys were removed in 12.1)
curl -H "Authorization: Bearer $TOKEN" "$GRAFANA_URL/api/dashboards/home"
# Create a service account and a token
curl -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
-d '{"name": "automation-sa", "role": "Editor"}' "$GRAFANA_URL/api/serviceaccounts"
curl -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
-d '{"name": "ci-token"}' "$GRAFANA_URL/api/serviceaccounts/$SA_ID/tokens"
# List dashboards
curl -s -H "Authorization: Bearer $TOKEN" \
"$GRAFANA_URL/api/search?type=dash-db" | jq '.[] | {title, uid, folderTitle}'
# Get a dashboard as an app-platform resource (13.x)
curl -s -H "Authorization: Bearer $TOKEN" \
"$GRAFANA_URL/apis/dashboard.grafana.app/v1/namespaces/default/dashboards/$UID" | jq .
# Export a dashboard (legacy API) and re-import into a folder
curl -s -H "Authorization: Bearer $TOKEN" "$GRAFANA_URL/api/dashboards/uid/$UID" \
| jq '.dashboard' > dashboard-export.json
curl -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
-d "{\"dashboard\": $(cat dashboard-export.json), \"overwrite\": true, \"folderUid\": \"$FOLDER_UID\"}" \
"$GRAFANA_URL/api/dashboards/db"
# Data sources (reference by UID; numeric-ID APIs are disabled by default in 13.0)
curl -s -H "Authorization: Bearer $TOKEN" "$GRAFANA_URL/api/datasources" | jq '.[] | {name, uid, type, url}'
curl -s -H "Authorization: Bearer $TOKEN" "$GRAFANA_URL/api/datasources/uid/mimir/health"
curl -s -H "Authorization: Bearer $TOKEN" \
"$GRAFANA_URL/api/datasources/proxy/uid/mimir/api/v1/query?query=up"
# Alerting (legacy provisioning API; alert-rule endpoints deprecated in 13.0)
curl -s -H "Authorization: Bearer $TOKEN" "$GRAFANA_URL/api/v1/provisioning/alert-rules" | jq .
curl -s -H "Authorization: Bearer $TOKEN" "$GRAFANA_URL/api/v1/provisioning/contact-points" | jq .
# Contact points via the app-platform API (13.x)
curl -s -H "Authorization: Bearer $TOKEN" \
"$GRAFANA_URL/apis/notifications.alerting.grafana.app/v1beta1/namespaces/default/receivers" | jq .
Useful One-Liners¶
# Back up every dashboard (legacy API)
for uid in $(curl -s -H "Authorization: Bearer $TOKEN" "$GRAFANA_URL/api/search?type=dash-db" | jq -r '.[].uid'); do
curl -s -H "Authorization: Bearer $TOKEN" "$GRAFANA_URL/api/dashboards/uid/$uid" | jq '.dashboard' > "backup-$uid.json"
done
# Health and version
curl -s "$GRAFANA_URL/api/health" | jq .
curl -s -H "Authorization: Bearer $TOKEN" "$GRAFANA_URL/api/frontend/settings" | jq '.buildInfo'
# Count dashboards per folder
curl -s -H "Authorization: Bearer $TOKEN" "$GRAFANA_URL/api/search?type=dash-db" \
| jq 'group_by(.folderTitle) | map({folder: .[0].folderTitle, count: length})'
Manage Plugins and the Admin Account¶
grafana cli plugins install grafana-clock-panel
grafana cli plugins install grafana-clock-panel <version> # pin a specific version
grafana cli plugins ls
grafana cli plugins list-versions grafana-clock-panel
grafana cli plugins update-all
grafana cli plugins uninstall grafana-clock-panel
grafana cli --pluginUrl https://example.com/custom-plugin.zip plugins install custom-plugin
# Reset the admin password (read it from stdin to keep it out of shell history)
grafana cli --homepath /usr/share/grafana --config /etc/grafana/grafana.ini \
admin reset-admin-password --password-from-stdin
# Move plaintext data source passwords into secure_json_data
grafana cli admin data-migration encrypt-datasource-passwords
# Re-encrypt secrets after rotating the encryption key
grafana cli admin secrets-migration re-encrypt
Bake plugins into a custom image (or use [plugins] preinstall) rather than installing at runtime on every pod start. The old Pie Chart and Worldmap Angular plugins no longer load since 12.0; use the built-in Pie chart and Geomap panels.
Upgrade Grafana¶
- Back up the database (and
grafana.ini, plugins directory). On 13.0+ this is the only rollback path because of the unified storage migration. - Upgrade to the latest patch of your current minor first, then update all plugins (React 19 compatibility for 13.x).
- Read the upgrade guide and breaking changes for each major you cross (v13.0 guide).
- For 13.x specifically:
- skip 13.0.0 (withdrawn) and go to the latest 13.x patch;
- replace
grafana-cli/grafana-serverin scripts and systemd overrides; - run the Image Renderer as a service and set a non-default
renderer_token; - switch API calls from numeric data source IDs to UIDs (or temporarily enable
datasourceLegacyIdApi); - on SQLite, if the migration hits
database is locked, set[unified_storage] migration_parquet_buffer = true.
- Upgrade one replica, watch logs and
grafana_unified_storage_migration_status, then roll the rest.
Operate Grafana Well¶
Dashboard Governance¶
- Use folders per team or domain, and team folders (13.0) to record ownership.
- Manage critical dashboards as code (Git Sync, provisioning, or Terraform) to prevent drift.
- Every dashboard has an owner; review quarterly and archive unused ones (deleted dashboards can be restored).
- Name dashboards with a team or domain prefix, for example
[infra] Node Overview. - Use variables (or quick filters in 13.1+) for environment, region and service.
- Aim for 8–12 panels on overview dashboards and 15–20 on detailed ones; use tabs and rows in dynamic dashboards to split larger ones.
Optimize Queries¶
- Filter early with precise label selectors in PromQL/LogQL.
- Avoid high-cardinality labels (user IDs, IPs, request IDs, raw paths). For Loki, keep roughly 10 or fewer indexed labels with low-cardinality values and put pod names or trace IDs in structured metadata; aim for fewer than ~10k active streams per tenant.
- Pre-compute expensive queries with recording rules in Mimir/Prometheus (or Grafana-managed recording rules).
- Set Max data points and use
$__interval/$__rate_interval. - Avoid refresh intervals under 10 s unless needed.
- Use saved queries (GA in 13.2 for Enterprise/Cloud) to share vetted queries.
Monitor Grafana Itself¶
Scrape /metrics from every replica (for example with Alloy or kube-prometheus-stack) and alert on request latency, data proxy latency, alert evaluation failures, and grafana_alerting_scheduler_behind_seconds. Metric names are listed in Reference.
Troubleshooting Tools¶
- Query inspector: raw query, response, timing per panel.
- Server logs: stdout in containers or
/var/log/grafana/grafana.log; raise[log] level = debugtemporarily. - Grafana Advisor (GA in 13.0): automated checks for plugins, data sources, and configuration.
- Tracing: Grafana can export its own traces over OTLP, or to a file as OTLP/JSON (13.2).
- Alloy UI:
http://localhost:12345for the live pipeline graph and component health.
Troubleshooting Playbook¶
| Symptom | Likely cause | Fix |
|---|---|---|
| Dashboard loads slowly | Expensive queries, too many panels, short refresh | Query inspector, recording rules, fewer panels, longer refresh |
| "Data source is not available" | Wrong URL, network, or proxy mode | Check the URL from the Grafana host, use the data source health check |
| Alerts not firing | Rule paused, evaluation errors, or contact point broken | Check rule health, grafana_alerting_rule_evaluation_failures_total, test the contact point |
| Duplicate notifications in HA | Alerting HA not configured | Set ha_peers (port 9094 reachable) or ha_redis_address |
| Login loop | root_url mismatch or cookie_secure over plain HTTP |
Fix root_url, serve over HTTPS |
| "database is locked" | SQLite with several replicas, or SQLite during the 13.0 migration | Move to PostgreSQL/MySQL; for the migration enable migration_parquet_buffer |
| Plugin not loading | Unsigned plugin or Angular plugin on 12.0+ | Allow-list the ID with allow_loading_unsigned_plugins, or replace Angular plugins |
Scripts fail with grafana-cli: command not found |
13.0 removed the binary | Use grafana cli |
| Reports and PNG exports fail after 13.0 | Renderer plugin removed / token not set | Run the renderer service and set renderer_token |
API returns 404 for /api/datasources/<id> |
Numeric-ID APIs disabled in 13.0 | Use /api/datasources/uid/<uid> |
| High memory on Grafana pods | Many concurrent viewers or heavy SQL data sources | Scale out, enable query caching (Enterprise), reduce refresh |
Migrate from ELK to Loki and Grafana¶
- Deploy Alloy next to existing Logstash/Filebeat and dual-ship logs to Loki (Grafana recommends Alloy over the Logstash Loki output plugin).
- Choose a small set of low-cardinality labels; move other fields into structured metadata.
- Configure object storage (S3/GCS/Azure) and retention for Loki.
- Rebuild Kibana dashboards in Grafana (no automatic converter) and rewrite queries in LogQL.
- Keep Elasticsearch read-only for historical data until its retention expires, then decommission.