Explanation¶
How Apache SkyWalking works: the OAP analysis pipeline, the OAL/MAL/LAL/MQE languages and the 10.4 V2 engine, BanyanDB's storage design, the Horizon UI split in 11.0, eBPF profiling with Rover, Kubernetes operation with SWCK, OpenTelemetry integration, and the security model. Version facts and benchmark tables live in Reference; tasks live in How-to Guides.
System Architecture¶
SkyWalking is a backend-centric APM. Probes (language agents, eBPF, service-mesh telemetry, OpenTelemetry) push data to the OAP (Observability Analysis Platform) server. OAP turns raw telemetry into topology, metrics, records and alarms. It writes them to pluggable storage and serves them through GraphQL, PromQL, LogQL, TraceQL and Zipkin query APIs. Since 11.0.0 the UI is a separate project, Horizon UI, which reads OAP's query port (12800) and the new admin host (17128).
The diagram below shows the 11.0 component layout, from probes through OAP modules to storage and consumers.
flowchart TB
subgraph Probes["Probes and data sources"]
JA["Java agent<br/>(bytecode injection)"]
LA["Python / Go / NodeJS / PHP<br/>Rust / Ruby / LUA agents"]
ROVER["Rover<br/>(eBPF profiler)"]
SAT["Satellite<br/>(edge proxy)"]
OTEL["OTel SDK / Collector<br/>(OTLP gRPC + HTTP)"]
MESH["Envoy ALS / metrics<br/>(Istio)"]
ZIP["Zipkin reporters"]
end
subgraph OAP["OAP server cluster (Java)"]
direction TB
RECV["Receivers<br/>gRPC 11800 / HTTP 12800"]
subgraph DSL["Analysis languages"]
OAL["OAL<br/>(native traces/mesh to metrics)"]
MAL["MAL<br/>(meter / OTel / Prometheus metrics)"]
LAL["LAL<br/>(log parsing, log to metrics)"]
end
L1["L1 aggregation<br/>(BatchQueue)"]
L2["L2 aggregation + persistence"]
ALARM["Alarm core<br/>(rules, baselines, hooks)"]
QUERY["Query layer<br/>GraphQL, MQE, PromQL,<br/>LogQL, TraceQL, Zipkin"]
ADMIN["admin-server 17128<br/>(runtime rules, debugger,<br/>inspect, UI templates)"]
end
subgraph Storage["Storage (one selected)"]
BDB["BanyanDB<br/>(default)"]
ES["Elasticsearch /<br/>OpenSearch"]
SQL["MySQL /<br/>PostgreSQL"]
end
subgraph Consumers["Consumers"]
HUI["Horizon UI"]
GRAF["Grafana<br/>(PromQL / LogQL / Tempo)"]
CLI["swctl / MCP clients"]
end
Probes --> RECV
RECV --> DSL
DSL --> L1 --> L2 --> Storage
L2 --> ALARM
QUERY --> Storage
HUI --> QUERY
HUI --> ADMIN
GRAF --> QUERY
CLI --> QUERY
OAP Roles and Clustering¶
OAP nodes are stateless with respect to data; state lives in storage. A node runs in one of three roles (SW_CORE_ROLE): Mixed (default), Receiver (receives and does L1 aggregation), or Aggregator (L2 aggregation and persistence). Nodes discover each other through a cluster coordinator (standalone, kubernetes, zookeeper, consul, etcd, nacos) and forward L1-aggregated metrics to the owning aggregator over the core gRPC port. This two-level aggregation keeps per-minute metric writes to storage roughly proportional to the number of entities, not to the request rate.
Write Path for a Trace Segment¶
The sequence below follows a Java agent trace segment from the application to BanyanDB.
sequenceDiagram
participant App as Java app + agent
participant R as OAP receiver node
participant A as OAP aggregator node
participant L as BanyanDB liaison
participant D as BanyanDB data node
App->>R: TraceSegmentReportService (gRPC 11800)
R->>R: Segment analysis, OAL sources
R->>R: L1 aggregation (BatchQueue, per minute)
R->>A: Forward metrics by entity hash (core gRPC)
R->>L: Write segment to trace group (bulk, flushInterval 15s)
A->>A: L2 merge and downsample (minute, hour, day)
A->>L: Bulk write measures (persistentPeriod 25s)
L->>D: Route by shard key (group, entity)
D->>D: Memory part, flush, merge into parts
Note over A,D: Alarm core evaluates metrics after L2
Analysis Languages¶
| Language | Input | Output | Where defined |
|---|---|---|---|
| OAL (Observability Analysis Language) | Native sources from traces, service mesh and eBPF (Service, Endpoint, ServiceRelation, ...) |
Metrics such as service_cpm, endpoint_resp_time |
config/oal/*.oal |
| MAL (Meter Analysis Language) | Sample families from native meters, OTLP metrics, Prometheus (via OTel Collector), Telegraf, Zabbix | Metrics per layer and scope | config/otel-rules/, config/meter-analyzer-config/ |
| LAL (Log Analysis Language) | Logs from agents, OTLP, Envoy ALS, Kafka | Parsed/filtered logs, derived metrics, sampled records (slow SQL, sampled traces) | config/lal/*.yaml |
| MQE (Metrics Query Expression) | Stored metrics | Computed query results (arithmetic, top_n, aggregation, baselines) for UI and alarms |
Query-time, and in alarm rules |
| Hierarchy rules | Service metadata | Cross-layer service hierarchy (for example K8s service to mesh service to agent service) | config/hierarchy-definition.yml |
Layers group services by technology (GENERAL, MESH, K8S_SERVICE, MYSQL, GENAI, VIRTUAL_GENAI, AIRFLOW, IOS, ...). Since 11.0.0 operators can declare custom layers in layer-extensions.yml or inline in MAL/LAL rule files, with no OAP code change.
The V2 Engine (10.4.0)¶
Before 10.4.0, MAL, LAL and hierarchy rules ran as Groovy scripts, and OAL was compiled from ANTLR4 grammar into Java classes at startup. 10.4.0 replaced both paths:
- OAL V2: immutable AST, type-safe operator enums, and error reporting with file, line and column.
- MAL / LAL / Hierarchy V2: ANTLR4 parsers generate Javassist bytecode at startup. The Groovy runtime dependency was removed from the OAP backend. Syntax and type errors fail fast at boot instead of at first execution. Generated classes hold no
ThreadLocalor shared mutable state, and class names follow{yamlFileName}_L{lineNo}_{ruleName}so stack traces point back at the rule. - A cross-version checker validated V1/V2 equivalence across 1,290+ expressions. JMH figures are in Reference.
LAL breaking change in 10.4.0
The slowSql {} and sampledTrace {} sub-DSLs were removed. A rule now declares outputType: SlowSQL or outputType: SampledTrace, assigns output fields directly in the extractor, and must contain an explicit sink {} block or nothing is persisted. Custom LALOutputBuilder.init() implementations gained an Optional<Object> extraLog parameter.
Runtime Rules and the Live Debugger (11.0.0)¶
11.0.0 added an admin-server module (HTTP 17128, internal gRPC 17129) and moved all write/debug APIs there:
- Runtime rule hot-update for MAL and LAL (and the native
meter-analyzer-configcatalog). A structural change goes through a cluster-wide phase machine (pending, DDL, fencing, rolling-out, applied). BanyanDB schema changes are fenced so every OAP node applies the new shape before writes resume. - Live debugger for MAL, LAL and OAL (SWIP-13). It captures per-statement samples on each node, and for LAL it reports the drop reason when a parse step fails.
- Inspect API, UI-management API (dashboard templates as REST) and the relocated
/status/*and/debugging/*routes.
The admin host has no built-in authentication and can change running analysis code. That is why it runs on its own port, separate from the agent-facing 11800.
BanyanDB¶
BanyanDB is SkyWalking's own observability database, written in Go and created in 2022. It stores metrics (measures), logs and records (streams), traces (trace model, added for 10.3), and metadata (properties). It became the default OAP storage when H2 was removed in 10.2.0. Its versions are still 0.x: 0.11.1 is current, and OAP pins an exact API version.
Cluster Architecture¶
The diagram below shows the 0.11 cluster shape. The etcd meta nodes of earlier releases were removed in 0.11.0 in favour of a property-based schema registry.
flowchart LR
subgraph OAPC["OAP cluster"]
O1["OAP"]
O2["OAP"]
end
subgraph BYDB["BanyanDB cluster (0.11)"]
direction LR
subgraph Liaison["Liaison tier (gRPC 17912, HTTP 17913)"]
LW["Liaison (write)"]
LRD["Liaison (read)"]
end
subgraph Hot["Hot data nodes"]
H1["Data node<br/>shards 0..n"]
H2["Data node"]
end
subgraph Warm["Warm nodes (type=warm)"]
W1["Data node"]
end
subgraph Cold["Cold nodes (type=cold)"]
C1["Data node"]
end
REG["Property-based schema registry<br/>(port 17916, no etcd)"]
FODC["FODC agent / proxy<br/>(diagnostics, metrics)"]
end
O1 -->|"bulk write"| LW
O2 -->|"BydbQL query"| LRD
LW -->|"shard routing"| H1
LW --> H2
LRD -->|"map phase"| H1
LRD --> H2
H1 -->|"lifecycle migration"| W1
W1 -->|"lifecycle migration"| C1
REG -.-> LW
REG -.-> H1
FODC -.-> H1
Key design points:
- Liaison nodes are stateless gateways. They authenticate, enforce TTL, route writes by shard key (resource name plus entity), and run the reduce phase of distributed queries. Upstream recommends separate liaison sets for write-heavy collectors and read-heavy UIs.
- Data nodes own shards and store data as
group / segment / shard / part. Segments are time buckets (segmentInterval). Parts are immutable columnar files that background merges combine, in the same spirit as an LSM tree. - Node discovery is
none(standalone),dns(SRV records) orfile. Since 0.11.0 the only schema registry mode isproperty, and all--etcd-*flags are gone. - Replication: groups have a
replicassetting. The clustering doc still says durability is mainly delegated to the underlying block or object storage rather than heavy application-level replication. There is no cross-datacenter replication. - Standalone mode (
banyand standalone) runs liaison, data and registry in one process. It is used by docker-compose quickstarts and small installs.
Data Model¶
| Concept | Description |
|---|---|
| Group | Namespace with its own shardNum, segmentInterval, TTL, replicas and optional warm/cold stages. OAP creates sw_records, sw_trace, sw_metricsMinute, and similar groups |
| Measure | Metric series: entity tags plus numeric fields, columnar and compressed. Supports TopN pre-aggregation and (0.12, unreleased) tag aggregation and time buckets |
| Stream | Append-only time-ordered elements (logs, sampled records) |
| Trace | Span storage keyed by trace ID with secondary indexes (sidx); introduced for OAP 10.3's trace model |
| Property | Mutable key-value documents (UI templates, runtime rules, schema registry) with a repair mechanism |
| IndexRule / IndexRuleBinding | Inverted-index definitions bound to tags for query acceleration |
Data Lifecycle and Tiering¶
Data cannot be deleted row by row. It expires when a whole segment passes the group's TTL. With warm/cold stages enabled, a lifecycle process migrates expiring segments to nodes labelled type=warm or type=cold, each with its own TTL and segment interval. Default TTLs are in Reference.
The state diagram below shows a segment's life when every stage is enabled.
stateDiagram-v2
[*] --> Hot: write into current segment
Hot --> Hot: memory part flush, part merge
Hot --> Warm: hot TTL reached, lifecycle migration
Hot --> Deleted: hot TTL reached, no warm stage
Warm --> Cold: warm TTL reached
Warm --> Deleted: no cold stage
Cold --> Deleted: cold TTL reached
Deleted --> [*]
Recent BanyanDB Changes¶
- 0.10.x (OAP 10.4): BanyanDB MCP server; map-reduce aggregation for distributed measure queries; property repair on by default;
nonenode discovery as default; Windows binaries dropped. - 0.11.0 (2026-08-28, OAP 11): etcd removed; vectorized columnar query paths enabled by default for measure, stream and trace (upgrade liaisons before data nodes); schema-consistency barriers with revision gates; trace tail sampling inside merges (
sw-trace-sampler,zipkin-trace-samplerplugins); FODC crash diagnostics; BydbQL positional parameters to prevent query injection; a data migration tool. - 0.11.1 (2026-09-20): a security fix enforcing the Canopy read-only role on
/monitoring/*. - 0.12.0 (in development): removes row-based query execution, adds a native inverted index encoder, and aggregation over tags.
Why a Custom Database¶
Upstream argues that general-purpose engines fit APM data poorly: traces and logs are write-heavy, append-only and queried by time plus a few entity keys, while metrics are pre-aggregated per minute. BanyanDB trades general search flexibility (no free-text relevance scoring, no ad hoc joins) for columnar compression, time-segmented TTL deletion and SkyWalking-aware sharding. The 2024 benchmark against Elasticsearch showed about 5x less memory and about 30% less disk at slightly higher data-node CPU. See Reference for the full table and its caveats.
Horizon UI and the 11.0 Split¶
Up to 10.4, OAP distributions bundled the Vue-based booster UI behind an Armeria proxy (apm-webapp). OAP stored dashboard templates and menus through GraphQL mutations. 11.0.0 deleted all of that:
- Horizon UI (
apache/skywalking-horizon-ui, first release 1.0.0 on 2026-08-28) is a separate project with its own BFF (backend-for-frontend), built-in authentication (local users with argon2id hashes, LDAP, API tokens, RBAC), dashboard library and release cadence. It ships inapache/skywalking-uiwithhorizon-<version>tags. - Dashboard templates move to OAP's
/ui-management/templatesREST API on the admin host (live mode), or Horizon renders its bundled templates (readonly mode, needed against OAP 10.x). - Horizon adds an optional AI assistant (bring-your-own OpenAI-compatible or Bedrock model, off by default), an MCP endpoint at
/api/mcpexposing investigation tools to agents, and admin screens for runtime rules, the live debugger and metric inspection.
Design consequence
Decoupling the UI means OAP and UI versions no longer move in lockstep. Only the OAP/BanyanDB pairing stays hard-locked.
Profiling and eBPF¶
SkyWalking offers several profiling modes, all triggered as tasks from the UI and reported back through OAP:
| Mode | Target | Mechanism |
|---|---|---|
| Trace profiling | Java, Python, Go agent services | Thread dumps sampled for slow endpoints |
| Async-profiler (10.2+) | Java | CPU, ALLOC, LOCK, WALL events via async-profiler |
| pprof (10.3+) | Go agent services | CPU, heap, block, goroutine and mutex profiles |
| eBPF on/off-CPU | Any process (C, C++, Go, Rust, JVM with symbols) | Rover attaches eBPF programs on the node |
| eBPF network profiling | Pods, processes | Rover captures TCP/TLS/HTTP traffic for process topology and metrics |
| Continuous profiling | Kubernetes processes | Rover triggers profiling automatically when thresholds (CPU, HTTP error rate) are crossed |
Rover also runs as a Kubernetes network monitor, producing access logs for mesh-less topology (including Istio ambient ztunnel detection in the 0.8.0 development line). Its latest release is 0.7.0 (2024-10). Development continues on the main branch, but releases have been infrequent.
Kubernetes Operation (SWCK)¶
SkyWalking Cloud on Kubernetes (SWCK, 0.11.0 on 2026-09-01) is the operator path, as an alternative to the Helm chart:
- CRDs:
OAPServer,UI,Storage,Satellite,Fetcher,JavaAgent, plus config CRDs (OAPServerConfig,OAPServerDynamicConfig). - Java agent injector: a mutating webhook that adds the agent to pods in namespaces labelled
swck-injection=enabled, for workloads labelledswck-java-agent-injected: "true". - Custom metrics adapter: exposes OAP metrics to the Kubernetes HPA.
- 0.11.0 wired BanyanDB storage (auth, TLS), made
UI kind: horizonthe only supported UI, references credentials from Secrets instead of copying them, and ships an operator Helm chart (oci://registry-1.docker.io/apache/skywalking-swck,-helmtag suffix). It requires cert-manager.
OpenTelemetry Integration¶
SkyWalking consumes OpenTelemetry. It is not an OpenTelemetry-native backend:
- Metrics: OTLP metrics flow through MAL rules (
otel-rules/) into layer-specific metrics. Dots in attribute names become underscores. This is how SkyWalking monitors MySQL, PostgreSQL, Redis, Kafka, ClickHouse, Kubernetes, VMs (node-exporter, and in 11.1 development also the Collectorhostmetricsreceiver), Airflow, Envoy AI Gateway and more. - Traces: in 11.0 OTLP traces are converted to Zipkin format and queried through the Zipkin API, Lens UI, Horizon's Zipkin tab, or TraceQL. They are not merged into native SkyWalking segments. The 11.1 development line adds native OTLP span storage (
otlpTraceStorage: otlp). - Logs: OTLP logs go through LAL.
- Transport: OTLP/gRPC on 11800, and since 11.0.0 OTLP/HTTP on 12800 (
/v1/traces,/v1/logs,/v1/metrics, protobuf or JSON). - The OTel Collector's SkyWalking exporter is deprecated upstream; use OTLP export to OAP instead.
The practical effect: native SkyWalking agents still give the richest data (service topology, endpoint metrics, profiling hooks), and OTel data is treated as an additional source.
GraalVM Distro¶
The experimental skywalking-graalvm-distro (0.4.0, 2026-09-10) compiles the unchanged OAP into a native image on JDK 25. All OAL/MAL/LAL/Hierarchy code generation and classpath scanning move to build time. Trade-offs: BanyanDB is the only storage, the module set is fixed at build time (no SPI discovery), and runtime rule hot-update does not work, although rule catalogs and the live debugger do. Upstream benchmarks report 5 ms startup and ~41 MiB idle memory against 635 ms / ~1.2 GiB for the JVM (Reference).
Performance Internals (10.4.0)¶
- BatchQueue replaced DataCarrier for L1 aggregation, L2 persistence, TopN persistence, exporters and the remote client. OAL and MAL metrics now share partitioned, self-draining queues with adaptive partitioning and throughput-weighted rebalancing.
- Virtual threads (JDK 25+) serve gRPC and Armeria HTTP handlers. Eleven pools share about nine carrier threads instead of up to 1,400+ platform threads. A 2-node cluster dropped from 150+ to about 72 OAP threads.
- The default Docker base image moved to JDK 25. OAP itself is still compiled for Java 11.
Security Model¶
SkyWalking has four trust boundaries. Each has different defaults.
flowchart LR
subgraph Untrusted["Application / edge network"]
AG["Agents, OTel SDKs,<br/>browser and mobile SDKs"]
end
subgraph OAPB["OAP"]
G["gRPC 11800<br/>(TLS + token optional)"]
H["HTTP 12800<br/>(TLS optional, no user auth)"]
AD["admin 17128<br/>(no auth, can change rules)"]
end
subgraph Data["Storage network"]
S["BanyanDB 17912<br/>(basic auth + TLS optional)"]
end
subgraph Users["Operators"]
UI["Horizon UI<br/>(login required, RBAC)"]
end
AG -->|"sw8 / OTLP"| G
AG -->|"browser, OTLP/HTTP"| H
UI -->|"GraphQL"| H
UI -->|"admin REST"| AD
G --> S
- Agents to OAP: the optional shared token (
SW_AUTHENTICATION/agent.authentication) is one secret for all agents, sent as gRPC metadata. It prevents casual injection but is not per-tenant identity. gRPC TLS and mTLS (gRPCSslTrustedCAPath) are available on the core and sharing servers. Client-side SDKs (browser, iOS, mini-programs) must reach OAP from the public internet, so upstream's security guide recommends a dedicated ingress path for them. - OAP HTTP: 12800 carries both queries and some receivers. 11.0.0 added hot-reloaded TLS for every HTTP server (server-side only, no mTLS), but OAP has no user authentication. Put it behind a gateway.
- Admin host: the most sensitive surface, because it can hot-update analysis rules and capture live data. Keep it on a private interface.
- UI: the booster UI had no login. Horizon requires configured users, but it does not fail closed: a pod with no users reports Ready and nobody can log in.
- Storage: BanyanDB supports username/password auth (
--auth-config-file) and TLS on liaison gRPC/HTTP; 0.11.1 also enforces a read-only role in its Canopy web console. Elasticsearch relies on its own security features. Data at rest encryption is left to the disk or cloud volume. - LAL test tool (
SW_QUERY_GRAPHQL_ENABLE_LOG_TEST_TOOL) executes untrusted scripts and is off by default.
Hardening steps are listed in Reference. Configuration recipes are in How-to Guides.
Sub-Project Ecosystem¶
The diagram below groups the ASF sub-projects around the core OAP repository.
flowchart LR
subgraph Core["Platform"]
OAPC["OAP server<br/>(apache/skywalking)"]
HUIC["Horizon UI"]
BDBC["BanyanDB<br/>(Go)"]
GRAAL["GraalVM Distro<br/>(experimental)"]
AIS["AI Sessionizer"]
end
subgraph Collect["Agents and collection"]
AGENTS["Java, Python, Go, NodeJS,<br/>PHP, Rust, Ruby, LUA,<br/>Kong, Client JS"]
ROVC["Rover (eBPF)"]
SATC["Satellite"]
KEE["K8s Event Exporter"]
end
subgraph Ops["Deployment and tools"]
HELM["skywalking-helm<br/>+ banyandb-helm"]
SWCKC["SWCK operator"]
SWCTL["swctl CLI"]
MCP["SkyWalking MCP"]
end
Collect --> OAPC
OAPC --> BDBC
HUIC --> OAPC
Ops --> Core