Explanation¶
How Grafana works and why it is built the way it is: the server architecture, the query path, plugins, the new app platform and dashboard schema, Git Sync, alerting, Grafana's place in the LGTM stack, the collector (Alloy), AI features, and the security model. Look-up tables (versions, config keys, roles, limits, pricing) are in Reference; tasks are in How-to Guides.
Architecture Overview¶
Grafana is a query, transform, and visualize server. It stores configuration (users, dashboards, folders, alert rules, data source definitions) in its own database but does not store telemetry: every panel query is forwarded to an external data source. The backend is Go; the frontend is TypeScript/React.
The diagram below shows the main components of a Grafana 13 server and what they talk to.
flowchart TB
subgraph Browser["Browser (TypeScript / React)"]
direction LR
DashUI["Dashboards<br/>(Scenes, schema v2)"]
ExploreUI["Explore and<br/>Drilldown apps"]
AlertUI["Alerting UI"]
Assist["Grafana Assistant"]
end
subgraph Server["Grafana server (Go)"]
direction TB
HTTP["HTTP server<br/>/api (legacy) and /apis (app platform)"]
AuthN["AuthN / AuthZ<br/>(sessions, SSO, RBAC)"]
QS["Query service<br/>(data frames, expressions)"]
Ngalert["Alerting engine<br/>(scheduler, state manager,<br/>embedded Alertmanager)"]
Prov["Provisioning and Git Sync<br/>(repository jobs)"]
US["Unified storage<br/>(resource API)"]
PM["Plugin manager<br/>(backend plugins over gRPC)"]
Live["Grafana Live<br/>(WebSocket streaming)"]
end
subgraph State["State"]
DB[("SQL database<br/>SQLite / MySQL / PostgreSQL")]
RC[("Remote cache<br/>database / Redis / Memcached")]
end
subgraph Ext["External systems"]
DS["Data sources<br/>Prometheus, Mimir, Loki, Tempo,<br/>Pyroscope, SQL, CloudWatch"]
Git["Git provider<br/>GitHub, GitLab, Bitbucket"]
Rend["Image Renderer service"]
Notif["Contact points<br/>Slack, PagerDuty, email"]
end
Browser --> HTTP
HTTP --> AuthN
HTTP --> QS
HTTP --> US
QS --> PM
PM -->|gRPC| DS
QS --> DS
Ngalert --> QS
Ngalert --> Notif
Prov <--> Git
Prov --> US
US --> DB
AuthN --> DB
AuthN --> RC
Ngalert --> DB
HTTP --> Rend
Live --> Browser
Key Architectural Properties¶
| Property | Detail |
|---|---|
| Stateless servers | All persistent state is in the SQL database; any replica can serve any request (sessions are DB-backed tokens, no sticky sessions needed) |
| Proxy, not store | Telemetry stays in the data sources; Grafana only caches (optional query caching in Enterprise) |
| Plugin isolation | Backend plugins run as separate processes over gRPC (HashiCorp go-plugin) |
| Two API families | Legacy /api (deprecated in 13.0) and Kubernetes-style /apis/<group>/<version>/namespaces/<ns>/... |
| Multi-org | One instance hosts isolated organizations (namespaces default, org-<id>) |
| As-code friendly | Provisioning files, Git Sync, Terraform, gcx, and the Foundation SDK all target the same resource APIs |
Request Lifecycle¶
When a user opens a dashboard, the browser loads the dashboard resource, then each panel issues queries through the backend, which calls the data source and returns data frames.
- The browser loads the dashboard (schema v2 resource, or v1 JSON model for older dashboards).
- Each panel's query runner sends a
/api/ds/queryrequest with the panel's queries, time range and variables. - The query service authorizes the request (data source permissions), resolves variables, and dispatches to the data source plugin (in-process for core sources, gRPC for external backend plugins).
- Optional server-side expressions (math, reduce, resample, SQL expressions) run on the returned frames.
- Frames return to the browser, pass through panel transformations, and are rendered by the panel plugin.
sequenceDiagram
participant B as Browser (panel query runner)
participant G as Grafana server
participant P as Data source plugin (gRPC)
participant D as Backend (e.g. Mimir)
B->>G: GET dashboard resource
G-->>B: Dashboard spec (panels, variables)
B->>G: POST /api/ds/query (PromQL, time range)
G->>G: AuthZ check and variable interpolation
G->>P: QueryData request
P->>D: HTTP query_range
D-->>P: Matrix result
P-->>G: Data frames
G->>G: Server-side expressions (optional)
G-->>B: Data frames (JSON / Arrow)
B->>B: Transformations and render
Data Frames¶
Grafana uses a unified, typed, columnar structure called a data frame (similar to a Pandas DataFrame). Every data source plugin must return frames, so any panel can render data from any source and transformations work uniformly.
Plugin Architecture¶
Grafana's extensibility comes from four plugin types (data source, panel, app, renderer; see Reference).
Plugin Lifecycle¶
- Discovery: Grafana scans the plugin directories on startup (and installs anything listed in
[plugins] preinstall). - Bootstrap: reads
plugin.jsonmetadata (ID, type, dependencies, Grafana version range). - Validation: checks the plugin signature (Grafana-signed, community, commercial, private, or unsigned).
- Initialization: loads the frontend module and, for backend plugins, starts the Go binary and connects over gRPC.
Backend Isolation¶
Backend plugins run as separate processes and talk to the server over gRPC:
- a crashing plugin does not crash Grafana,
- plugins can implement their own auth, caching, streaming, and alerting support,
- secrets (
secureJsonData) stay server-side and are never sent to the browser.
Key SDK packages: @grafana/data (data structures), @grafana/ui (React component library), @grafana/runtime (runtime services), @grafana/scenes (dashboard-like app pages), and the Grafana Plugin SDK for Go. AngularJS plugins stopped working in Grafana 12.0, and Grafana 13 moved the frontend to React 19, so older plugins may need updates before an upgrade.
App Platform and Unified Storage¶
Since Grafana 12, Grafana has been re-platforming its resources onto a Kubernetes-style API server (the "app platform"):
- Each resource type (dashboards, folders, alert rules, contact points, playlists, repositories) is served under an API group such as
dashboard.grafana.appwith explicit versions (v1alpha1→v1beta1→v1). - Resources have Kubernetes-like metadata (
metadata.name,namespace, labels, annotations) and support watch, dry-run and generated clients. - Unified storage is the backend for these resources. In Grafana 13.0, folders and dashboards are automatically migrated from the legacy SQL tables (
dashboard,dashboard_version,folder, and so on) into unified storage on first start; the legacy tables are deprecated.
Why downgrades are risky after 13.0
After the unified storage migration, an older Grafana reads the stale legacy tables and does not see changes made in unified storage. Rolling back requires restoring the pre-upgrade database backup.
This platform is what makes Git Sync, gcx, and the v2 dashboard API possible: they all read and write the same versioned resources.
Dashboards: Schema v2 and Dynamic Dashboards¶
Dashboards were long a single JSON model (now called v1). Grafana 12 introduced a v2 schema and "dynamic dashboards"; both reached GA with Grafana 13 (April 2026), and the new layout engine is on by default for every dashboard.
| Aspect | v1 JSON model | v2 schema / dynamic dashboards |
|---|---|---|
| Layout | Absolute gridPos grid, rows |
Layout kinds: grid, auto grid, rows, tabs; nesting (depth 4 and nested tabs in 13.2) |
| Structure | Panels embed queries and options | Separate elements (panels, library panels) and layout referencing them |
| Variables | Dashboard-wide | Dashboard-wide plus section-level variables on rows and tabs (13.0/13.1) |
| Filtering | Ad hoc filters | "Filters" (renamed ad hoc filters), quick filters and grouping without template variables (13.1) |
| API | /api/dashboards/db |
/apis/dashboard.grafana.app/... resource API; v1 dashboards are converted on the fly |
| Frontend | Scenes-based renderer (the pre-Scenes architecture toggle was removed in 13.0) | Scenes-based renderer with side-pane editing |
The frontend runtime for both is Scenes (@grafana/scenes), a scene-graph framework of objects (SceneQueryRunner, SceneDataTransformer, layouts, variable sets, time ranges) that sync state to the URL. App plugins use the same library to build dashboard-like pages, so plugin authors get variables, time-range inheritance and URL sync without wiring React state by hand.
Observability as Code and Git Sync¶
Grafana now has a single as-code story built on the resource APIs:
| Tool | Role |
|---|---|
| Git Sync | Bidirectional sync between Grafana and a Git repository (GA in 13.0) |
gcx |
CLI that pulls/pushes resources as files (gcx resources pull / push), plus queries and Assistant access; replaces grafanactl |
| Foundation SDK | Typed builders (Go, TypeScript, Python, Java, PHP) that generate dashboard resources |
| File provisioning | Classic YAML/JSON provisioning from disk (still supported, including v2 dashboards) |
| Terraform / Ansible / Grafana Operator / Crossplane | Infrastructure-as-code integrations |
Git Sync connects a folder (or, since 13.1, the whole instance with "folderless" sync) to a repository path. Saving a dashboard in the UI can commit directly or open a pull request, and Grafana posts dashboard previews and diffs on the PR. Changes pushed to Git are pulled back by polling (default 60 s) or immediately via webhooks.
The sequence below shows the UI-to-Git round trip with pull-request review enabled.
sequenceDiagram
participant U as Editor (Grafana UI)
participant G as Grafana (provisioning service)
participant R as Git repository
participant CI as Reviewers / CI
U->>G: Save dashboard (commit to new branch)
G->>R: Push commit and open pull request
R->>G: Webhook (PR opened)
G->>R: Comment with preview link and diff
CI->>R: Review and merge to main
R->>G: Webhook (push to main)
G->>G: Sync job updates unified storage
G-->>U: Dashboard shows as provisioned and in sync
Design choices worth knowing:
- Provisioned resources are owned by the repository. Grafana blocks changing manager properties and marks such dashboards with a "provisioned" badge; unmanaged resources outside the synced folder are untouched.
- Limits are deliberately low for now (10 connections, ~1,000 resources per connection recommended) because each sync job loads the database; shard large repositories across several connections. See Reference.
- Auditability: 13.1 added verified (signed) commits and branch-protection checks; 13.2 added commit authoring as the signing user and user attribution for jobs.
13.0.0 upgrade bug
Self-managed 12.x instances that had the Git Sync feature flags enabled could lose or revert dashboards when upgrading to 13.0.0. That build was withdrawn; upgrade to 13.0.1 or later, and restore from backup first if you were affected.
Alerting¶
Since Grafana 9, alerting is unified: Grafana-managed rules can query any backend data source, and the same UI can manage data-source-managed rules in Mimir, Loki or Prometheus rulers. Legacy dashboard alerting was removed in Grafana 11.
The pipeline below shows how a Grafana-managed rule becomes a notification.
flowchart LR
subgraph Rules["Alert rules (in folders)"]
R1["PromQL: cpu > 80%"]
R2["LogQL: error rate"]
end
Sched["Scheduler<br/>(evaluation interval per group)"]
QS["Query service and<br/>expressions"]
SM["State manager<br/>(one instance per series)"]
AM["Embedded Alertmanager"]
subgraph Routing["Notification policies"]
Tree["Routing trees<br/>(label matchers, multiple trees since 13.x)"]
end
subgraph CP["Contact points"]
PD["PagerDuty"]
Slack["Slack"]
Email["Email"]
WH["Webhook"]
end
Rules --> Sched --> QS --> SM --> AM --> Tree
Tree -->|"severity=critical"| PD
Tree -->|"team=backend"| Slack
Tree -->|default| Email
Tree -->|custom| WH
Each series returned by a rule becomes an alert instance with its own state. The state machine below includes the Recovering state added in Grafana 12.0 for keep_firing_for.
stateDiagram-v2
[*] --> Normal
Normal --> Pending: condition true
Pending --> Normal: condition false
Pending --> Alerting: pending period elapsed
Alerting --> Recovering: condition false and keep_firing_for set
Alerting --> Normal: condition false
Recovering --> Alerting: condition true again
Recovering --> Normal: keep_firing_for elapsed
Normal --> NoData: query returns no data
Normal --> Error: evaluation fails
NoData --> Normal: data returns
Error --> Normal: evaluation succeeds
Key Concepts¶
- Alert rules define queries, expressions, a condition, the pending period, and labels/annotations. Grafana-managed recording rules also exist and write to a Prometheus-compatible target.
- Labels on instances drive routing (
severity=critical,team=infra). - Notification policies form routing trees; Grafana 13 added multiple named routing trees ("managed routes") with their own access control, and rules can select a policy through notification settings.
- Contact points hold integrations (Slack, PagerDuty, email, webhook, Opsgenie, Teams, and more).
- Mute timings, silences, and inhibition rules suppress notifications.
- State history can be written to Loki for long-term analysis.
- Import to Grafana-managed alerting (GMA): since 12.0 a wizard converts Prometheus/Mimir rule groups and Alertmanager configuration into Grafana-managed resources; 13.2 added staged imports, auto-sync and template import.
High Availability¶
Every replica evaluates every rule by default; replicas share notification state over memberlist gossip (port 9094) or Redis so each notification is sent once. Grafana 13.0 added single-node evaluation mode (ha_single_node_evaluation) so only one node evaluates while the others stand by. Load distribution of evaluation across nodes is not supported.
API Direction¶
Alerting is moving to app-platform APIs: notification resources are served under notifications.alerting.grafana.app/v1beta1 (receivers, routing trees, templates, time intervals, inhibition rules), several legacy Alertmanager-config endpoints were removed or restricted in 13.0, and the provisioning endpoints for alert rules are deprecated in favour of the rules.alerting.grafana.app API.
Grafana in the LGTM Stack¶
Grafana is the "G" in Grafana Labs' LGTM stack. Each signal has a purpose-built backend; the deep internals (write/read paths, hash rings, compaction, per-tenant limits, benchmarks) are documented in the LGTM Stack topic and are only summarized here.
| Signal | Backend | Query language | Storage |
|---|---|---|---|
| Metrics | Grafana Mimir (or Prometheus) | PromQL | TSDB blocks in object storage |
| Logs | Grafana Loki | LogQL | Compressed chunks + label index in object storage |
| Traces | Grafana Tempo | TraceQL | Apache Parquet blocks in object storage |
| Profiles | Grafana Pyroscope | Prometheus-style label selectors on profile types (historically called FlameQL) | Object storage |
| Collection | Grafana Alloy | Alloy configuration syntax (formerly "River") or OTel Collector YAML | Pipeline agent |
The diagram below shows how telemetry reaches the backends and how Grafana queries them.
flowchart TB
subgraph Collection["Grafana Alloy"]
direction LR
R["Receivers<br/>OTLP, Prometheus scrape, Loki sources"]
P["Processors<br/>batch, filter, transform"]
E["Exporters<br/>remote_write, loki.write, OTLP"]
R --> P --> E
end
subgraph Backends["LGTM backends"]
Mimir["Mimir<br/>(metrics)"]
Loki["Loki<br/>(logs)"]
Tempo["Tempo<br/>(traces)"]
Pyroscope["Pyroscope<br/>(profiles)"]
end
S3[("Object storage<br/>S3 / GCS / Azure Blob")]
subgraph Grafana["Grafana"]
Dash["Dashboards"]
Drill["Explore and Drilldown apps"]
Alert["Alerting"]
end
Collection -->|remote_write| Mimir
Collection -->|push| Loki
Collection -->|OTLP| Tempo
Collection -->|push| Pyroscope
Mimir --> S3
Loki --> S3
Tempo --> S3
Pyroscope --> S3
Tempo -->|"metrics-generator (RED)"| Mimir
Grafana -.->|PromQL| Mimir
Grafana -.->|LogQL| Loki
Grafana -.->|TraceQL| Tempo
Grafana -.->|profile queries| Pyroscope
Backend design choices in one line each (details in the LGTM architecture overview):
- Mimir 3.x: distributors shard series to ingesters; since Mimir 3.0 the default "ingest storage" architecture puts a Kafka-compatible log in front of the ingesters (the classic RF=3 replicated-ingester mode still exists). Blocks land in object storage and are served by store-gateways.
- Loki 3.x: indexes only labels, not log content, so it is much cheaper than full-text indexing but needs label-first queries and low-cardinality labels (high-cardinality fields belong in structured metadata). Promtail was removed in favour of Alloy.
- Tempo 3.x: no traditional index; spans are stored as Parquet columns and TraceQL reads only the needed columns. Tempo 3.0 (2026-05) replaced ingesters with a Kafka-based write path (block-builders and live-stores). The metrics-generator derives RED metrics and service graphs.
- Pyroscope 2.x: continuous profiling with a new v2 storage architecture; linked to traces through span profiles.
Cross-Signal Correlation¶
Grafana's value in the stack is linking signals:
- Exemplars: metric samples carry trace IDs, so a latency spike links to a trace in Tempo.
- Trace to logs / metrics / profiles: span attributes map to Loki stream labels, Mimir queries, and Pyroscope profiles.
- Derived fields and correlations: trace IDs parsed from log lines link back to Tempo.
- Drilldown apps (Metrics, Logs, Traces, Profiles Drilldown; GA in Grafana 12): query-less exploration built on these links.
Collection: Grafana Alloy¶
Grafana Alloy (announced April 2024) is Grafana Labs' OpenTelemetry Collector distribution with built-in Prometheus pipelines and native Loki and Pyroscope components. It is the successor to Grafana Agent, which reached end of life on 2025-11-01 (Agent Flow was the basis for Alloy).
- Component model: pipelines are built from components (
prometheus.scrape,loki.source.file,otelcol.receiver.otlp, ...) wired together in the Alloy configuration syntax, which was previously called River. - OpenTelemetry Engine (experimental in 2026):
alloy otelruns standard OTel Collector YAML inside the Alloy binary, for teams that want upstream-compatible config. - Operations: clustering for scrape-target distribution, a debug UI and live pipeline graph on port 12345, remote configuration via Fleet Management.
- Releases: roughly every 5–7 weeks in 2026; latest 1.20.0 (2026-09-25). License Apache-2.0.
AI Features¶
| Feature | What it is | Availability |
|---|---|---|
| Grafana Assistant | Context-aware LLM agent in the Grafana UI: writes and explains queries, builds dashboards, analyzes alerts and flame graphs, uses SQL expressions | GA in Grafana Cloud 2025-10-08; available on-prem with Grafana 13 (April 2026); pre-installed in Grafana Enterprise 13.1 |
| Assistant Investigations | Background agents that fan out across metrics, logs, traces and profiles and return a root-cause report | Preview October 2025; GA in July 2026 ("AI Week") |
Assistant Automations, Workspace, Cloud MCP server, gcx, Agent Observability |
Agentic operations tooling announced GA together in July 2026 | Grafana Cloud |
| Sift | Earlier automated diagnostic checks launched from Explore, dashboards or Incident; its role is now largely covered by Assistant Investigations | Still in Grafana Cloud through the Machine Learning app, but being phased out: IRM no longer starts Sift from an incident, and mcp-grafana removed its deprecated Sift tools (Sift in IRM docs, mcp-grafana PR #1191; checked 2026-09-27) |
| Grafana Advisor | Health checks for a Grafana instance (plugins, data sources, config) | GA in 13.0 |
Assistant is billed per active AI user plus token overage (see Reference: Grafana Cloud plans); Grafana Labs has not published which model vendors back it (third-party coverage noted this in April 2026).
Security Model¶
Identity and Authentication¶
Grafana authenticates users through built-in basic auth (bcrypt-hashed passwords in the Grafana DB), OAuth2/OIDC (generic plus Google, GitHub, GitLab, Azure AD/Entra ID, Okta, Keycloak), LDAP, SAML (Enterprise/Cloud), JWT, or an authenticating reverse proxy (auth proxy). SSO settings can also be managed through the UI and API (SSO settings API). SCIM user and team provisioning arrived in Grafana 12 for Enterprise/Cloud.
Once authenticated, the user gets a server-side session token stored in the database and rotated every 10 minutes by default, which is why replicas need a shared database but no sticky sessions. Configuration recipes are in How-to Guides.
The diagram shows the identity and authorization path through a Grafana deployment.
flowchart TD
subgraph Outside["External"]
Browser["Browser / API client"]
IdP["Identity provider<br/>OIDC / SAML / LDAP / SCIM"]
end
subgraph Edge["Edge"]
Proxy["Reverse proxy / ingress<br/>(TLS, optional auth proxy)"]
end
subgraph GF["Grafana server"]
AuthN["AuthN clients<br/>(session, SA token, JWT, OAuth)"]
Sess["Session tokens<br/>(cookie_secure, SameSite)"]
RBAC["RBAC evaluator<br/>(basic roles, fixed and custom roles)"]
DSP["Data source permissions"]
Audit["Audit log (Enterprise)"]
AI["Grafana Assistant<br/>(runs as the user)"]
end
subgraph Stores["Backends"]
DB[("Grafana DB")]
Keeper[("Secrets keeper<br/>Vault / AWS (Enterprise)")]
end
Browser --> Proxy --> AuthN
AuthN <--> IdP
AuthN --> Sess --> RBAC
RBAC --> DSP
RBAC --> AI
DSP --> Audit
Sess --> DB
AuthN -.-> Keeper
Authorization¶
- Basic roles (Viewer, Editor, Admin, plus "No basic role") are assigned per organization; Grafana Admin is server-wide.
- Folders are the main permission boundary; dashboards inherit folder permissions and can add their own. Org Admins cannot be restricted.
- RBAC (Enterprise/Cloud) adds fixed roles, custom roles built from actions and scopes, and data source permissions, enabling least privilege (query-only access to one data source, folder-scoped editing, alerting managed by one team).
- Team folders (13.0) tie folders to owning teams.
Role and permission tables are in Reference.
Service Accounts¶
Service accounts are the identity for automation (Terraform, CI/CD, gcx, Git Sync jobs). They are org-scoped, can hold multiple tokens, survive the deletion of the user who created them, and can be given basic and RBAC roles. API keys were migrated to service accounts automatically in 11.6 and their endpoints were removed in 12.1.
Embedding, Anonymous Access, and Browser Protections¶
- Embedding is off by default (
allow_embedding = false). Turning it on for cross-site iframes also needscookie_samesite = nonewithcookie_secure = true, which weakens CSRF protection; prefer public/shared dashboards or an allow-listed CSPframe-ancestors. - Content-Security-Policy can be enabled with a per-request
$NONCE, so only Grafana-generated inline scripts run. - Anonymous access grants a fixed role to everyone who can reach the server; isolate it in a dedicated org and restrict data sources.
- Brute-force protection locks accounts after 5 failed logins by default.
Enterprise Security Features¶
- Audit logging of logins, permission changes, resource changes, and (optionally) data source queries; since 13.0 request and response bodies are no longer logged by default.
- Request security restricts outbound requests from Grafana (SSRF protection).
- Secrets management: external secrets keepers (HashiCorp Vault, AWS) and a secrets management UI (public preview in 13.0).
Grafana Assistant Security¶
The Assistant runs with the invoking user's permissions: it can only query dashboards and data sources that user can access, and in Grafana Cloud LLM calls are proxied through Grafana Cloud infrastructure. On self-managed Grafana (13.0+) it uses a hybrid model: the Assistant app runs in your Grafana, is paired with a Grafana Cloud stack through an authorization flow, and forwards LLM requests to that stack, where the backend, usage limits and billing live. Pairing sends your Grafana URL, including internal hostnames, to Grafana Cloud. Features that need the full Cloud backend (infrastructure memory, messaging integrations, Cloud MCP connections) are not available on-prem (self-managed Assistant docs, checked 2026-09-27).
Critical advisory
CVE-2025-41115 (CVSS 10.0) let a SCIM client impersonate users on Grafana Enterprise 12.0.0–12.2.1 with SCIM enabled. Patched versions and other advisories are listed in Reference.
Data Model¶
Grafana's own configuration data is organized by organization. The diagram shows the main entities (not the storage schema).
erDiagram
ORG ||--o{ USER_MEMBERSHIP : "has members with a basic role"
ORG ||--o{ TEAM : contains
ORG ||--o{ SERVICE_ACCOUNT : contains
ORG ||--o{ DATASOURCE : defines
ORG ||--o{ FOLDER : contains
FOLDER ||--o{ FOLDER : "nests (subfolders)"
FOLDER ||--o{ DASHBOARD : contains
FOLDER ||--o{ ALERT_RULE : contains
FOLDER ||--o{ LIBRARY_PANEL : contains
DASHBOARD ||--o{ DASHBOARD_VERSION : "keeps history"
ALERT_RULE }o--|| DATASOURCE : queries
TEAM ||--o{ FOLDER : "owns (team folders)"
REPOSITORY ||--o{ FOLDER : "manages via Git Sync"
Lifecycles¶
Server Lifecycle¶
- Startup: loads
defaults.inithengrafana.ini/env overrides, runs SQL migrations (and on 13.0+ the one-time unified storage migration), installs pre-install plugins, loads plugins. - Runtime: serves HTTP, evaluates alert rules, runs provisioning/Git Sync jobs, rotates session tokens.
- Shutdown: drains connections and persists alert state.
Dashboard Lifecycle¶
- Created in the UI, through the API, from a template or suggested dashboard, or by provisioning/Git Sync.
- Versioned: each save creates a version; deleted dashboards can be restored (GA in 13.0).
- Provisioned (optional): owned by files on disk or a Git repository; UI edits either commit back (Git Sync) or are blocked (file provisioning with
editable: false). - Exported as JSON (v1) or v2 resources for sharing or as-code workflows.
Performance and Cost Considerations¶
Grafana itself is rarely the bottleneck: query latency is dominated by the data source, and panel count multiplied by refresh rate determines query load. The practical levers are recording rules in the metrics backend, maxDataPoints and $__interval, longer refresh intervals, query caching (Enterprise/Cloud), and scaling replicas. Image rendering and large numbers of short-interval alert rules are the most common reasons a deployment outgrows its initial size. Guideline tables are in Reference.
For cost, the trade-off is operational effort versus usage billing:
| Model (mid-size: ~500k–1M active series, ~100 GB/day logs, ~50M spans/day) | Rough monthly cost | Trade-off |
|---|---|---|
| Self-hosted LGTM on Kubernetes | ~$700–2,000 infrastructure | Full control; needs roughly 1–3 engineers of operational attention at scale |
| Grafana Cloud Pro | ~$4,400–7,800+ at list price (metrics ~$3,200–6,400 at $6.50 per 1,000 billable series, logs ~$1,200–1,300; traces extra), before discounts or Adaptive Metrics | Managed; usage-billed, so cardinality and log volume drive cost |
| Datadog (equivalent) | ~$4,500–15,000+ | Fully managed, highest cost |
Estimates
The self-hosted and Datadog ranges are illustrative figures carried over from earlier research (2026-04); no vendor or independent source publishes them. The Grafana Cloud row was recomputed from the list prices in Reference on 2026-09-27 (the earlier ~$1,000–3,300 range was below list price). All of them vary widely with retention, cardinality, and discounts. Use the Grafana Cloud pricing table and the LGTM cost comparison for inputs.
History and Licensing¶
- 2013–2014: Torkel Ödegaard forked the Kibana 3 frontend to build a Graphite dashboard; Grafana 1.0 shipped in January 2014.
- 2014–2017: the company Raintank formed around Grafana and later renamed itself Grafana Labs.
- April 2021: Grafana, Loki and Tempo were relicensed from Apache-2.0 to AGPL-3.0 (Grafana 8.0), to stop cloud providers offering modified hosted versions without contributing back. SDKs, plugins, agents and some packages stay Apache-2.0.
- 2022–2024: Mimir (2022) and Pyroscope (acquired 2023) completed the LGTM(P) stack; Alloy replaced Grafana Agent (2024).
- 2025–2026: Grafana 12 (May 2025) introduced observability as code, Drilldown apps and dynamic dashboards; Grafana 13 (April 2026) made Git Sync and dynamic dashboards GA and brought the Assistant on-prem. Grafana Labs reported $400M+ ARR and 7,000+ customers in September 2025 and was reported (SiliconANGLE, February 2026) to be raising at about a $9B valuation, up from $6B in August 2024.
AGPL practicalities: using unmodified Grafana (even commercially, even as an internal SaaS) triggers no obligations; if you modify Grafana and let users interact with it over a network, you must offer the modified source under AGPL.
Sources¶
- Grafana CHANGELOG.md
- What's new in Grafana v13.0 and v12.0
- Upgrade to Grafana v13.0
- New API structure in Grafana (
/apis) - Introduction to Git Sync
- Observability as code
- Dynamic dashboards GA (what's new)
- Git Sync GA (what's new)
- Alert rule state and health
- Set up Grafana for high availability
- Grafana Assistant GA press release (2025-10-08)
- Grafana Labs AI Week press release (2026-07-27)
- Grafana Alloy docs and OpenTelemetry Engine
- Grafana Agent README (EOL notice)
- Grafana authentication docs
- Grafana Enterprise features
- Service accounts
- CVE-2025-41115 advisory
- Grafana 2021 relicensing announcement
- Grafana Labs $400M ARR press release (2025-09-30)
- SiliconANGLE: Grafana Labs reportedly raising at $9B valuation