OpenObserve How-to Guides¶
What this page covers
Task recipes for OpenObserve (O2) v1.0.x: install a single node or an HA cluster, connect object storage and PostgreSQL, send logs, metrics, and traces, query data, tune search, secure the deployment, manage configuration as code, and upgrade. Defaults, ports, and endpoint lists are in Reference. The reasons behind each setting are in Explanation.
Replace example credentials
The commands below use the upstream example root user root@example.com / Complexpass#123. Use your own values, and keep real credentials in a secret store, not in shell history or Git.
Run a Single Node with Docker¶
Goal: a local instance for evaluation or light production.
docker run -d --name openobserve \
-v "$PWD/data:/data" \
-e ZO_DATA_DIR="/data" \
-e ZO_ROOT_USER_EMAIL="root@example.com" \
-e ZO_ROOT_USER_PASSWORD="Complexpass#123" \
-p 5080:5080 -p 5081:5081 \
public.ecr.aws/zinclabs/openobserve:latest
# UI: http://localhost:5080
ZO_LOCAL_MODEalready defaults totrue(SQLite plus local disk), so you do not need to set it.- The root user variables are needed on first start only.
- For the Enterprise build, which is free up to 50 GB/day, use
o2cr.ai/openobserve/openobserve-enterprise:latest. - On AVX-512 (Intel) or NEON (ARM) hosts, the
latest-simdtag is faster. Pin a version tag such asv1.0.4in production.
ECR pull errors
If pulling from public.ecr.aws fails with an auth error, log in first:
aws ecr-public get-login-password --region us-east-1 | docker login --username AWS --password-stdin public.ecr.aws
Run a Single Node with Docker Compose¶
services:
openobserve:
image: public.ecr.aws/zinclabs/openobserve:latest
restart: unless-stopped
ports:
- "5080:5080"
- "5081:5081"
environment:
ZO_DATA_DIR: /data
ZO_ROOT_USER_EMAIL: root@example.com
ZO_ROOT_USER_PASSWORD: Complexpass#123
volumes:
- o2data:/data
volumes:
o2data:
The top-level version: key is obsolete in Compose v2, so it is omitted.
Run the Native Binary¶
Download the binary for your platform from the downloads page, then run:
chmod +x openobserve
ZO_ROOT_USER_EMAIL="root@example.com" ZO_ROOT_USER_PASSWORD="Complexpass#123" ./openobserve
If you see version GLIBC_2.27 not found, use the -linux-musl build. It is slightly slower but has no glibc dependency.
Keep a Single Node's Data in Object Storage¶
Goal: one node, but data that survives loss of the local disk.
docker run -d --name openobserve \
-v "$PWD/data:/data" -e ZO_DATA_DIR="/data" \
-e ZO_ROOT_USER_EMAIL="root@example.com" \
-e ZO_ROOT_USER_PASSWORD="Complexpass#123" \
-e ZO_LOCAL_MODE_STORAGE="s3" \
-e ZO_S3_BUCKET_NAME="my-o2-bucket" \
-e ZO_S3_REGION_NAME="us-east-1" \
-e ZO_S3_ACCESS_KEY="$AWS_ACCESS_KEY_ID" \
-e ZO_S3_SECRET_KEY="$AWS_SECRET_ACCESS_KEY" \
-p 5080:5080 public.ecr.aws/zinclabs/openobserve:latest
Metadata stays in SQLite under /data, so keep that volume persistent too.
Deploy an HA Cluster with Helm¶
Goal: a production cluster on Kubernetes. HA mode needs object storage, PostgreSQL, and NATS. Local disk storage is not supported.
- Create the bucket in advance. The chart does not create it.
-
Install the CloudNativePG operator. The chart uses it to create a PostgreSQL cluster with one primary and one replica.
-
Download the chart values and edit them. See the next two sections.
-
Install:
-
Verify. Expect about 12 pods (NATS x3, postgres x2, OpenFGA, router, ingester, querier, compactor, scheduler, and an OpenFGA init job), all
Running.
Chart defaults: Enterprise image, older app version
The openobserve chart defaults to enterprise.enabled: true, so it runs o2cr.ai/openobserve/openobserve-enterprise, which is free up to 50 GB/day. Set enterprise.enabled: false for the AGPL OSS image. Chart 1.0.2 also defaults to app v1.0.1. To run the latest patch, set image.enterprise.tag (or image.oss.tag), for example to v1.0.4.
Scale components¶
Replica counts sit under replicaCount in values.yaml:
Turn on ingester.persistence (WAL) and querier.persistence (disk cache) for production. Ingesters need fast disks; the vendor suggests about 3,000 IOPS.
Configure Object Storage for HA¶
Pick the block that matches your store and merge it into values.yaml.
Amazon EKS + S3 (IRSA)¶
config:
ZO_S3_BUCKET_NAME: "my-o2-bucket"
serviceAccount:
annotations:
eks.amazonaws.com/role-arn: arn:aws:iam::111122223333:role/zo-s3-eks
Give the role s3:PutObject, s3:GetObject, s3:ListBucket, and s3:DeleteObject on the bucket.
S3 with static keys¶
auth:
ZO_S3_ACCESS_KEY: "<access-key>"
ZO_S3_SECRET_KEY: "<secret-key>"
config:
ZO_S3_BUCKET_NAME: "my-o2-bucket"
ZO_S3_REGION_NAME: "us-west-1"
MinIO¶
config:
ZO_S3_SERVER_URL: "http://minio.minio.svc:9000"
ZO_S3_BUCKET_NAME: "my-o2-bucket"
ZO_S3_REGION_NAME: "us-west-1"
ZO_S3_PROVIDER: "minio"
Set auth.ZO_S3_ACCESS_KEY and auth.ZO_S3_SECRET_KEY as in the static-keys example. The HA docs also cover RustFS (ZO_S3_PROVIDER: "s3", path-style), OpenStack Swift, and Civo.
Google Cloud Storage (HMAC keys)¶
config:
ZO_S3_SERVER_URL: "https://storage.googleapis.com"
ZO_S3_BUCKET_NAME: "my-o2-bucket"
ZO_S3_REGION_NAME: "auto"
ZO_S3_PROVIDER: "s3"
ZO_S3_FEATURE_HTTP1_ONLY: "true"
Use an External PostgreSQL¶
config:
ZO_META_STORE: "postgres"
auth:
ZO_META_POSTGRES_DSN: "postgres://openobserve:<password>@pg.example.internal:5432/openobserve?sslmode=require"
postgres:
enabled: false # disable the bundled CloudNativePG cluster
Create the database first. Cluster mode accepts only PostgreSQL. MySQL is rejected at startup in v1.0.4 even though the HA docs page still describes it.
Send Logs¶
JSON over HTTP¶
curl -u 'root@example.com:Complexpass#123' \
-H 'Content-Type: application/json' \
http://localhost:5080/api/default/app_logs/_json \
-d '[{"level":"info","message":"hello from curl","service":"checkout"}]'
The URL pattern is /api/{org}/{stream}/_json. The stream is created on first write.
Elasticsearch bulk API¶
cat > logs.ndjson <<'EOF'
{"index":{"_index":"app_logs"}}
{"level":"error","message":"payment failed","service":"checkout"}
EOF
curl -u 'root@example.com:Complexpass#123' \
-H 'Content-Type: application/x-ndjson' \
http://localhost:5080/api/default/_bulk --data-binary @logs.ndjson
Beats and Vector use /api/{org}/ as their Elasticsearch base path. Filebeat example:
output.elasticsearch:
hosts: ["http://localhost:5080"]
path: "/api/default/"
index: "default"
username: "root@example.com"
password: "Complexpass#123"
Send OpenTelemetry Data¶
Build the Basic auth header first:
OTLP/HTTP exporter (the endpoint must not end with /, because the collector appends /v1/logs, /v1/metrics, and /v1/traces):
exporters:
otlphttp/openobserve:
endpoint: http://openobserve:5080/api/default
headers:
Authorization: "Basic <base64>"
stream-name: default
service:
pipelines:
logs:
receivers: [otlp]
exporters: [otlphttp/openobserve]
traces:
receivers: [otlp]
exporters: [otlphttp/openobserve]
OTLP/gRPC exporter on port 5081. The vendor measures 60-100% more ingest throughput with gRPC than with HTTP JSON:
exporters:
otlp/openobserve:
endpoint: openobserve:5081
headers:
Authorization: "Basic <base64>"
organization: default
stream-name: default
tls:
insecure: true # only for in-cluster plaintext; enable TLS otherwise
To collect Kubernetes logs, metrics, and traces with a prebuilt collector, use the openobserve-collector chart. Its README shows how to compute the Authorization header in CI.
Send Prometheus Metrics¶
Add a remote_write block to prometheus.yml:
remote_write:
- url: http://openobserve:5080/api/default/prometheus/api/v1/write
queue_config:
max_samples_per_send: 10000
basic_auth:
username: root@example.com
password: Complexpass#123
Query the metrics with PromQL or SQL in the Metrics explorer or in dashboards.
Query Data¶
SQL in the Logs UI¶
The stream name is the table name. Pick the time range in the UI or API rather than in WHERE.
-- Full-text search on indexed text fields, narrowed by a field filter
SELECT * FROM "app_logs"
WHERE service = 'checkout' AND match_all('timeout')
ORDER BY _timestamp DESC
LIMIT 100;
-- Error count per minute
SELECT histogram(_timestamp, '1 minute') AS ts, COUNT(*) AS errors
FROM "app_logs"
WHERE level = 'error'
GROUP BY ts
ORDER BY ts;
-- Noisiest services
SELECT service, COUNT(*) AS cnt
FROM "app_logs"
GROUP BY service
ORDER BY cnt DESC
LIMIT 20;
Search API¶
start_time and end_time are required and are in microseconds.
END=$(( $(date +%s) * 1000000 )); START=$(( END - 3600 * 1000000 ))
curl -u 'root@example.com:Complexpass#123' \
-H 'Content-Type: application/json' \
http://localhost:5080/api/default/_search \
-d "{\"query\":{\"sql\":\"SELECT level, COUNT(*) AS c FROM \\\"app_logs\\\" GROUP BY level\",\"start_time\":$START,\"end_time\":$END,\"from\":0,\"size\":100}}"
Speed Up Search¶
Work through these in order and measure after each step:
- Narrow the time range. Every query benefits, because data is partitioned by hour.
- Add a bloom filter on high-cardinality equality fields (
trace_id,request_id,user_id) in Stream settings. - Mark full-text fields (
body,message,log) and keep the inverted index on (ZO_ENABLE_INVERTED_INDEX=true, the default). Then usematch_all()together with a field filter. Expect about 25% extra storage. - Add a partition key only for a low-cardinality field you filter on often (namespace, host). Use hash partitioning (8-128 buckets) if the values are skewed.
- Use the SIMD image (
latest-simd) on AVX-512 or NEON hardware. - Separate background work. Give some queriers
ZO_NODE_ROLE_GROUP=backgroundso alerts and reports do not compete with interactive searches.
Partition keys are permanent
You cannot change or remove a partition key once it is set, and too many keys produce small files with poor compression. If you get one wrong, you have to create a new stream. Keep each Parquet file above about 5 MB.
Transform Data with a Pipeline¶
- Open Pipelines and create a real-time pipeline with the source stream as the source node.
- Add a condition node, for example
level != debug, so debug lines are not passed on. -
Add a function node with VRL, for example to mask email addresses with the VRL
redactfunction: -
Add a destination node: another stream (for example
app_logs_clean), or a metrics stream for logs-to-metrics. - Save, then compare ingest rates. VRL at ingest uses ingester CPU.
Secure the Deployment¶
Enable TLS¶
Either terminate TLS at your ingress or load balancer, or let OpenObserve serve TLS itself:
config:
ZO_HTTP_TLS_ENABLED: "true"
ZO_HTTP_TLS_CERT_PATH: "/certs/tls.crt"
ZO_HTTP_TLS_KEY_PATH: "/certs/tls.key"
ZO_HTTP_TLS_MIN_VERSION: "1.3"
ZO_GRPC_TLS_ENABLED: "true"
ZO_GRPC_TLS_CERT_DOMAIN: "openobserve.example.com"
ZO_GRPC_TLS_CERT_PATH: "/certs/tls.crt"
ZO_GRPC_TLS_KEY_PATH: "/certs/tls.key"
A minimal NGINX proxy in front of the HTTP port:
server {
listen 443 ssl;
ssl_certificate /etc/nginx/certs/server.crt;
ssl_certificate_key /etc/nginx/certs/server.key;
client_max_body_size 200m; # match ZO_PAYLOAD_LIMIT
location / {
proxy_pass http://openobserve:5080;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
}
}
Keep secrets out of values files¶
kubectl -n openobserve create secret generic o2-root \
--from-literal=ZO_ROOT_USER_EMAIL=admin@example.com \
--from-literal=ZO_ROOT_USER_PASSWORD="$(openssl rand -base64 24)" \
--from-literal=ZO_ROOT_USER_TOKEN="$(openssl rand -hex 32)"
Then point the chart at it with externalSecret or your secrets operator. Put the PostgreSQL DSN under auth.ZO_META_POSTGRES_DSN; the chart renders auth values into a Secret. Also set ZO_INTERNAL_GRPC_TOKEN so that nodes must authenticate to each other.
Lock down the bucket¶
Use IRSA (see Configure Object Storage for HA), turn on SSE-KMS default encryption, and add a deny-all-except-role bucket policy:
{
"Effect": "Deny",
"Principal": "*",
"Action": "s3:*",
"Resource": ["arn:aws:s3:::my-o2-bucket", "arn:aws:s3:::my-o2-bucket/*"],
"Condition": {"StringNotEquals": {"aws:PrincipalArn": "arn:aws:iam::111122223333:role/zo-s3-eks"}}
}
Give collectors their own service accounts¶
- Go to IAM > Service Accounts > Add Service Account. Copy the token; it is shown only once.
- On Enterprise, assign a role (for example a custom
ingest_onlyrole) under IAM > Roles. Service accounts start with no permissions. - Authenticate with Basic auth using the service account's email as the username and the token as the password.
Enable RBAC and SSO (Enterprise)¶
RBAC needs the Enterprise image, HA mode, and OpenFGA. The current chart already defaults to enterprise.enabled: true (Enterprise image) and enterprise.openfga.enabled: true. SSO also needs Dex, which is off by default:
enterprise:
enabled: true # chart default; set false to run the OSS image
openfga:
enabled: true # chart default; RBAC decisions
dex:
enabled: true # chart default is false; OIDC broker for SSO
For SSO, set the O2_DEX_* values (client ID and secret, base URL, redirect and callback URLs, scopes, and the group and role attributes). Then configure Dex connectors for LDAP, OIDC, or SAML, and map IdP groups to organizations and roles.
Manage Configuration as Code (Enterprise)¶
O2 CLI¶
brew tap openobserve/tap && brew install o2
o2 configure --profile prod # endpoint, org, credentials
o2 list organizations --profile prod
o2 list alert --folder default --enabled-only --profile prod
o2 get template MyAlert --profile dev -o json > tested.json
o2 create template -f tested.json --profile prod
o2 alert import-prometheus -f prometheus-rules.yaml --destination my-slack --dry-run
Kubernetes operator¶
o2-k8s-operator manages alerts, templates, destinations, pipelines, functions, and dashboards as CRDs. It does not install or scale OpenObserve itself.
apiVersion: openobserve.ai/v1alpha1
kind: Alert
metadata:
name: high-error-rate
spec:
configRef:
name: production # connection config CR pointing at the O2 endpoint
name: high-error-rate-alert
streamName: default
streamType: logs
enabled: true
queryCondition:
type: custom
sql: "SELECT COUNT(*) as count FROM default WHERE level='error'"
aggregation:
function: count
having: {column: count, operator: GreaterThan, value: 100}
duration: 5
frequency: 1
destinations: [slack-alerts]
Since v1.0, alerts and SLOs can also be exported to Terraform or OpenTofu from the UI.
Check Health and Troubleshoot¶
# Liveness (the same path the Helm probes use)
curl -s http://localhost:5080/healthz
# List streams in the org, with their stats
curl -s -u 'root@example.com:Complexpass#123' http://localhost:5080/api/default/streams
# Pods and logs by role (chart label: role=<component>)
kubectl -n openobserve get pods
kubectl -n openobserve logs -l role=ingester --tail=100
kubectl -n openobserve logs -l role=querier --tail=100
# Restart ingesters one at a time (the WAL is replayed on start)
kubectl -n openobserve rollout restart statefulset o2-openobserve-ingester
| Symptom | Likely cause | Fix |
|---|---|---|
| Startup error: "Meta store only supports postgres in cluster mode" | ZO_LOCAL_MODE=false with SQLite |
Set ZO_META_STORE=postgres and a DSN |
| Startup error: "We don't support MySQL anymore" | ZO_META_STORE=mysql |
Migrate metadata to PostgreSQL |
| OTLP/HTTP returns 404 | Trailing / on the exporter endpoint |
Use http://host:5080/api/<org> with no trailing slash |
| Data not visible | Wrong time range, org, or stream; timestamps older than ZO_INGEST_ALLOWED_UPTO |
Widen the range, check the org selector, fix the timestamps |
| Records rejected or fields missing | More than ZO_COLS_PER_RECORD_LIMIT fields, or deep JSON |
Flatten upstream, raise the limit, or tune ZO_INGEST_FLATTEN_LEVEL |
| Slow full-text search | No inverted index or full-text fields set | See Speed Up Search |
| Many tiny Parquet files | Too many partition keys | Create a new stream without them |
ImagePullBackOff |
No outbound access to o2cr.ai |
Mirror the image or allow egress |
Upgrade¶
- Read the release notes for each version you skip. They include metadata migrations and renamed settings, such as
alertmanagerbecomingscheduler. - Back up PostgreSQL, for example with a CloudNativePG backup. The bucket needs no backup for the upgrade itself.
- Upgrade.
# Docker
docker pull public.ecr.aws/zinclabs/openobserve:v1.0.4
docker stop openobserve && docker rm openobserve
# re-run the docker run command with the same volume and the new tag
# Helm
helm repo update openobserve
helm diff upgrade o2 openobserve/openobserve -n openobserve -f values.yaml # needs the helm-diff plugin
helm upgrade o2 openobserve/openobserve -n openobserve -f values.yaml \
--set image.enterprise.tag=v1.0.4 # or image.oss.tag when enterprise.enabled=false
Check the chart's UPGRADING.md for chart-level breaking changes. Recent examples: the PostgreSQL DSN moved into the Secret, the cert-manager issuer annotation is no longer hardcoded, and openobserve-standalone now honours persistence.enabled: false.