Apache Kafka¶
Summary
Apache Kafka is the de facto standard open-source event streaming platform: durable, partitioned, replicated logs with exactly-once transactions and a large ecosystem (Connect, Streams, Schema Registry, MirrorMaker 2). Since 4.0 (March 2025) it runs KRaft-only, with no ZooKeeper. The latest release is 4.3.1 (2026-06-25), and 4.4.0 is in its release process. Queue semantics (share groups, KIP-932) have been production-ready since 4.2. The project is Apache-2.0 licensed and governed by the ASF. Confluent, the main commercial vendor, has been an IBM subsidiary since March 2026.
Overview¶
Apache Kafka is an open-source distributed event streaming platform. It was built at LinkedIn, open-sourced in 2011, and has been an Apache top-level project since 2012. It stores events in durable, partitioned, replicated commit logs organized into topics. Producers append records, and consumers read at their own pace, tracking durable per-group offsets. Since Kafka 4.0 the platform runs KRaft-only: ZooKeeper support was removed and metadata lives in an internal Raft quorum of controllers.
Kafka's value rests on four properties:
- High throughput per broker, from sequential disk I/O, reliance on the page cache, and zero-copy
sendfile(). - Durable replication through the in-sync replica (ISR) protocol, with Eligible Leader Replicas (ELR, KIP-966) on by default for new clusters since 4.1.
- Exactly-once semantics (EOS) for read-process-write workloads, using the idempotent producer (on by default since 3.0) and transactions.
- A large ecosystem: Kafka Connect, Kafka Streams, Schema Registry implementations, MirrorMaker 2, tiered storage for cheap long retention, and share groups for queue-style consumption.
Key Facts¶
| Attribute | Detail |
|---|---|
| Website | kafka.apache.org |
| GitHub Stars | ~30k+ (apache/kafka) |
| Latest Version | 4.3.1 (2026-06-25) |
| Supported Lines | 4.3.1, 4.2.1 (2026-05-30), 4.1.2 (2026-03-17), per the downloads page |
| Next Release | 4.4.0, in the release process (not published as of 2026-09-25) |
| Release Cadence | Target of 3 minor releases a year (every 4 months). Bug-fix releases for supported lines only |
| Language | Java (plus Scala 2.13 in parts of the broker). Clients in Java and many community languages |
| Java Requirement | Brokers, Connect, tools: Java 17+. Clients and Streams: Java 11+. Tested with Java 17 and 25 |
| License | Apache License 2.0 |
| Governance | Apache Software Foundation (top-level project since 2012) |
| Primary Vendor | Confluent (founded 2014 by Kafka's original authors). IBM completed its acquisition of Confluent on 2026-03-17 |
| Coordination | KRaft (Raft metadata quorum). ZooKeeper removed in 4.0 |
| Wire Protocol | Binary TCP, versioned per API key |
| Storage Format | Append-only segmented log, magic v2 record batches |
| Container Images | apache/kafka:4.3.1 (JVM), apache/kafka-native:4.3.1 (GraalVM) |
Evaluation¶
Pros¶
| Pro | Detail |
|---|---|
| Large ecosystem | Connect, Streams, Schema Registry implementations, hundreds of connectors, Flink and Spark integration |
| Durability and replication | ISR protocol, configurable min.insync.replicas, ELR (KIP-966) |
| Exactly-once semantics | Idempotent producer plus transactions across topics and partitions, hardened by KIP-890 in 4.0 |
| High throughput | LinkedIn measured about 2M writes/s on 3 commodity brokers in 2014 (details) |
| Replayable history | Consumers can re-read from any retained offset, useful for compliance and reprocessing |
| Tiered storage (KIP-405) | Production-ready since 3.9 (early access in 3.6). Offloads closed segments to object storage through a plugin |
| Queues on the same platform | Share groups (KIP-932, production-ready in 4.2) add per-record acknowledgement and more consumers than partitions |
| Truly open | Apache 2.0, with no BSL or SSPL relicensing like Redis or Elastic went through |
| Multi-language clients | Official Java client plus mature community clients (librdkafka and its wrappers, franz-go, sarama) |
Cons¶
| Con | Detail |
|---|---|
| Operational complexity | Topic design, partition planning, ISR shrinkage, and rolling restarts all need expertise |
| JVM-bound brokers | Heap, GC, and page-cache sizing matter. Some competitors (Redpanda) avoid the JVM entirely |
| Partition count limits | Per-broker and per-cluster partition counts are a real ceiling that depends on hardware and version. Test your limits |
| Consumer rebalance pain | The classic protocol still rebalances group-wide. KIP-848 fixes this but is opt-in (group.protocol=consumer) and needs 4.0+ brokers |
| Cross-region replication | MirrorMaker 2 works but is not transparent: offset translation, lag, and cutover need care |
| No built-in schema management | Confluent Schema Registry, Apicurio, or Karapace must be deployed separately |
| Tiered storage needs a plugin | Apache ships no production RemoteStorageManager, and compacted topics cannot be tiered |
Architecture (Summary)¶
A compact view of a KRaft cluster: controllers own metadata, brokers store and replicate partitions, and closed segments can be offloaded to object storage. The full internals are in Explanation.
flowchart LR
subgraph Producers["Producers"]
P1["KafkaProducer<br/>(idempotent)"]
P2["KafkaProducer<br/>(transactional)"]
end
subgraph KafkaCluster["Kafka cluster (KRaft)"]
direction TB
subgraph ControllerQuorum["Controller quorum (Raft)"]
KC1["Controller 1"]
KC2["Controller 2 (active)"]
KC3["Controller 3"]
end
subgraph Brokers["Broker pool"]
KS1["Broker 1<br/>LogManager / ReplicaManager"]
KS2["Broker 2<br/>LogManager / ReplicaManager"]
KS3["Broker 3<br/>LogManager / ReplicaManager"]
end
Brokers -- "Fetch __cluster_metadata" --> ControllerQuorum
end
subgraph RemoteStorage["Tiered storage (KIP-405)"]
S3["RemoteStorageManager plugin<br/>(S3 / GCS / Azure)"]
end
subgraph Consumers["Consumers"]
CG1["Consumer group A"]
CG2["Consumer group B<br/>(read_committed)"]
SG["Share group<br/>(queue semantics)"]
end
P1 -- "Produce" --> KS1
P2 -- "Produce (transactional)" --> KS2
KS2 -- "Replica fetch" --> KS1
KS3 -- "Replica fetch" --> KS2
KS1 -- "Upload closed segments" --> S3
CG1 -- "Fetch" --> KS2
CG2 -- "Fetch" --> KS3
SG -- "ShareFetch / ShareAcknowledge" --> KS1
style ControllerQuorum fill:#1f3a5f,color:#fff
style KS2 fill:#0d6e0d,color:#fff
Detailed architecture, KRaft consensus, replication, the group protocols, tiered storage, and the security model are in Explanation.
Use Cases¶
| Use case | Why Kafka fits |
|---|---|
| Event sourcing | Durable, replayable, ordered per partition |
| Microservices messaging backbone | Decouples producers and consumers, supports fan-out and back-pressure |
| Real-time stream processing | Kafka Streams, Flink, Spark Structured Streaming |
| CDC (change data capture) | Debezium connectors stream Postgres, MySQL, and MongoDB change events into topics |
| Work queues | Share groups (4.2+) for independent jobs with per-record acknowledgement and retries |
| Log aggregation | Replaced Scribe and Flume with replicated storage and replay |
| Metrics and telemetry transport | OpenTelemetry Collector Kafka exporter and receiver, Grafana Tempo's Kafka-based ingest |
| Data lake ingestion | Connect S3 sink, Iceberg sink |
| Audit trail / immutable event log | Compacted topics plus retention policies |
| Activity stream / clickstream | The original LinkedIn use case: high-cardinality data partitioned by user or session |
Licensing & Pricing¶
- Apache Kafka: Apache License 2.0. Free, no usage restrictions, no telemetry call-home, vendor-neutral.
- Confluent Platform: Some components (Schema Registry, REST Proxy, ksqlDB) use the source-available Confluent Community License. Confluent Enterprise (commercial) adds RBAC, audit logs, Cluster Linking, and Control Center. Confluent has been a wholly owned IBM subsidiary since 2026-03-17.
- Confluent Cloud: Fully managed SaaS priced by cluster type (for example Basic, Standard, Enterprise, Dedicated, Freight), throughput, storage, and partitions. It runs on Kora, Confluent's cloud-native Kafka engine.
- Managed alternatives: Amazon MSK (provisioned and Serverless), Aiven for Apache Kafka, NetApp Instaclustr, and Azure Event Hubs (Kafka-protocol endpoint). Upstash Kafka was deprecated in September 2024 and shut down on 2025-03-11.
- Self-hosted on Kubernetes: Strimzi (CNCF Incubating, free). Bitnami's free public images and charts were largely retired in 2025, so check image availability before depending on them.
License clarity
Unlike HashiCorp Vault (BSL 1.1), Redis (RSALv2/SSPLv1, now also AGPLv3), Elastic (ELv2/SSPL, now also AGPLv3), or MongoDB (SSPL), Apache Kafka has never been relicensed and remains Apache 2.0. Confluent's add-ons are licensed separately. The IBM acquisition does not change the ASF project's license or governance.
Ecosystem¶
| Component | Purpose |
|---|---|
| Kafka Connect | Source and sink integration framework: JDBC, S3, Elasticsearch, MongoDB, Snowflake, BigQuery, Iceberg |
| Kafka Streams | Embedded JVM library for stateful stream processing (KStream, KTable, joins, windowing) |
| ksqlDB | SQL-on-streams engine (Confluent Community License) |
| Schema Registry | Avro, JSON Schema, and Protobuf schema storage with compatibility checks (Confluent, Apicurio, Karapace) |
| MirrorMaker 2 | Connect-based cross-cluster replication (KIP-382). MirrorMaker 1 was removed in 4.0 |
| Cruise Control | LinkedIn's automated rebalancing and self-healing |
| Strimzi | CNCF Incubating Kubernetes operator (KRaft-only, kafka.strimzi.io/v1 CRDs) |
| Debezium | CDC connectors built on Kafka Connect |
| librdkafka | C/C++ client library behind the Python, Go, and .NET Confluent clients |
| kcat (kafkacat) | CLI for produce, consume, and metadata |
| Kroxylicious | Kafka-protocol proxy with filters such as record encryption |
| Conduktor / Kafbat UI / AKHQ / Redpanda Console | Web UIs for cluster admin and topic browsing |
Compatibility & Requirements¶
- JDK: Java 17+ for brokers, controllers, Connect, and tools (Java 11 dropped for the server side in 4.0). Java 11+ for clients and Kafka Streams (Java 8 dropped in 4.0). Java 25 support was added in 4.2.
- Operating system: Linux strongly preferred for production (
sendfile, page cache tuning,epoll). macOS for development only. Windows brokers are not recommended for production. - Hardware: SSD or NVMe for log directories, 10 GbE or faster networking, 32 GiB or more RAM per broker for sizeable workloads (page cache).
- Filesystem: XFS recommended over ext4. Do not put active log directories on network-attached storage. Use tiered storage for cold data.
- Containers: Official
apache/kafkaandapache/kafka-nativeimages, KRaft-native, for every supported release. Strimzi and Confluent images are also common. - Wire compatibility: 4.x brokers need clients at 2.1 or later, and 4.x clients need brokers at 2.1 or later (KIP-896). Newer brokers accept older API versions within that range.
- Upgrades: Rolling upgrades are finalized with
kafka-features.sh upgrade --release-version <X.Y>(metadata version). The ZooKeeper-erainter.broker.protocol.versionno longer applies. See How-to Guides.
Latest Versions¶
| Version | Release | Highlights |
|---|---|---|
| 4.4.0 | Not yet released | Upgrade notes list a share-group DLQ (KIP-1191), broker.id deprecation (KIP-1232), controller unregistration (KIP-1312), and Streams static membership on the streams protocol |
| 4.3.1 | 2026-06-25 | Bug fix: Kafka Streams RocksDB native memory leak (KAFKA-20616) |
| 4.3.0 | 2026-05-22 | 25 KIPs. Log-directory cordoning (KIP-1066), follower fetch from tiered offset (KIP-1023), streams-scala deprecated, classic consumer protocol deprecation phase 1 (KIP-1274) |
| 4.2.1 | 2026-05-30 | Share-group deadlock fix, Streams protocol migration fix |
| 4.2.0 | 2026-02-17 | Share groups (KIP-932) production-ready. Streams Rebalance Protocol GA (core features). DLQ in Streams exception handlers. Java 25 |
| 4.1.x | 2025-09-02 to 2026-03-17 | Share groups preview. Streams Rebalance Protocol (KIP-1071) early access. ELR on by default for new clusters. Static-to-dynamic quorum upgrade |
| 4.0.0 | 2025-03-18 | KRaft-only (ZooKeeper removed). KIP-848 GA. ELR preview. Queues early access. Brokers need Java 17. Log4j2 |
| 3.9.x | 2024-11-06 to 2026-02-21 | Last line with ZooKeeper (the migration bridge). Tiered storage production-ready. Dynamic quorums (KIP-853) |
The full release and support matrix, KIP status, and CVE list are in Reference.
Alternatives¶
| Alternative | Style | When to prefer |
|---|---|---|
| Redpanda | Kafka API, C++ broker, no JVM | Lower tail latency, simpler ops, single-binary deploys |
| Apache Pulsar | Separate compute and storage (BookKeeper) | Built-in geo-replication, multi-tenancy, native tiered storage |
| NATS / JetStream | Lightweight pub/sub plus streams | Edge and IoT, very low overhead, simpler operations |
| RabbitMQ | AMQP broker (plus streams) | Per-message TTL, complex routing, priorities, classic work queues |
| AWS Kinesis Data Streams | Managed shard-based stream | All-in on AWS, smaller scale, no Kafka API |
| Google Pub/Sub | Managed at-least-once pub/sub | All-in on GCP, serverless scaling, weaker ordering |
| Azure Event Hubs | Managed, Kafka-protocol-compatible | All-in on Azure. The gateway speaks the Kafka wire protocol |
| WarpStream | Object-storage-native, Kafka-compatible (owned by Confluent since 2024) | Diskless brokers with data only in S3, at higher latency |
| AutoMQ | Object-storage-based Kafka fork | Cloud cost optimization with the Kafka API |
See Streaming Brokers Comparison for a head-to-head of Kafka, Redpanda, and Pulsar, and Messaging Patterns Comparison for queue versus log versus pub/sub.
Migration & Lock-in¶
- API surface: The Kafka wire protocol is open. Redpanda, WarpStream, AutoMQ, Azure Event Hubs, and others implement it. Moving clients to a wire-compatible alternative usually means changing
bootstrap.serversand security settings. - Cross-cluster migration: MirrorMaker 2 (Connect-based) is the standard tool. It replicates topic data, consumer offsets (through the checkpoint connector), heartbeats, and ACLs. Confluent Cluster Linking is a commercial alternative that preserves offsets exactly.
- ZooKeeper to KRaft: Only possible on 3.x (3.9.x recommended). 4.x cannot read ZooKeeper metadata.
- Connector lock-in: Some Confluent-licensed connectors do not run outside Confluent Platform. Open-source connectors (Debezium, JDBC, S3) port freely.
- Operational lock-in: Tiered-storage data is written in the format of the
RemoteStorageManagerplugin you chose. Switching plugins means re-tiering or a clean cutover. - Schema lock-in: Confluent Schema Registry's wire format adds a magic byte and a 4-byte schema ID prefix. Apicurio and Karapace implement the same format for compatibility.
Community Health¶
- Apache top-level project since 2012 and one of the most active ASF projects by commit volume. 4.3.0 had 147 contributors.
- Design changes go through the KIP (Kafka Improvement Proposal) process: public design documents voted on by the community. Recent flagship KIPs are KIP-848 (consumer rebalance), KIP-932 (queues), KIP-405 (tiered storage), KIP-966 (ELR), KIP-853 (dynamic quorums), and KIP-1071 (Streams rebalance).
- The project targets three minor releases a year and ships bug-fix releases for the supported lines, currently 4.3.x, 4.2.x, and 4.1.x.
- Confluent (now part of IBM) employs many committers, but the project is governed independently by the ASF PMC.
- Broad third-party ecosystem: Strimzi (CNCF Incubating), Debezium, Cruise Control, Aiven, Instaclustr, Amazon MSK, Azure Event Hubs.
Topic Map¶
- How-to Guides: dev broker, production planning, producer/broker/consumer tuning, upgrades, listener security, troubleshooting, tiered storage, Strimzi, monitoring.
- Reference: release and support matrix, Java requirements, ports, internal topics, KIP status, config defaults, CLI tools, advisories, hardening checklist, benchmarks.
- Explanation: KRaft quorum, the replicated log, ISR and Eligible Leader Replicas, group protocols, exactly-once transactions, tiered storage, security model.
Related Topics¶
- Messaging domain: Messaging overview, Redpanda, Pulsar, NATS, RabbitMQ
- Comparisons: Streaming Brokers Comparison, Messaging Patterns Comparison
- Other domains: Kubernetes (Strimzi runs Kafka on it), OpenTelemetry Collector (Kafka exporter and receiver), LGTM stack (Tempo's Kafka-based ingest), PostgreSQL (CDC source through Debezium)
Sources¶
- Apache Kafka: official documentation
- Apache Kafka downloads (release dates, supported releases)
- Apache Kafka 4.3.1 release announcement (2026-06-25)
- Apache Kafka 4.3.0 release announcement (2026-05-22)
- Apache Kafka 4.2.0 release announcement (2026-02-17)
- Apache Kafka 4.0.0 release announcement (2025-03-18)
- Apache Kafka upgrade notes (4.x)
- Apache Kafka CVE list
- Apache Kafka on GitHub
- KIP index (Apache wiki)
- IBM: IBM Completes Acquisition of Confluent (2026-03-17)
- Strimzi is now a CNCF Incubating project (2024-02-08)
- Upstash: Announcing Upstash Workflow and deprecating Upstash Kafka
- Confluent Platform documentation
- Confluent: Apache Kafka performance and test results
- LinkedIn: Benchmarking Apache Kafka, 2 million writes per second
- Strimzi (Kubernetes operator)
Questions¶
- Q: Are share groups ready to replace RabbitMQ for queue workloads? Production-ready since 4.2, but 4.2.1 fixed a critical share-group deadlock (KAFKA-20505), and a native DLQ only arrives in 4.4 (KIP-1191). Evaluate on staging, and prefer 4.2.1+ or 4.3.x. RabbitMQ still wins on routing, priorities, and per-message TTL.
- Q: When does tiered storage pay off compared with more local disk? Usually once retention runs to weeks at high ingest and most reads are tail reads, with occasional historical replays. The break-even depends on the plugin, object-store request pricing, and cross-AZ costs. No neutral public study found (searched 2026-09-28).
- Q: How do KRaft controller failovers compare with ZooKeeper failovers in large production clusters? Anecdotally faster and simpler, but few public head-to-head studies exist for very large clusters (more than 200 brokers).
- Q: What is the practical partitions-per-broker ceiling on KRaft 4.x? Public numbers vary by hardware and replication factor, and no current official per-broker limit is published. Load-test your own ceiling.
- Q: Will the IBM ownership of Confluent change Confluent's investment in upstream Kafka? Too early to tell as of 2026-09. Watch committer affiliation and KIP authorship.
Answered Questions¶
- Q: When will
group.protocol=consumerbecome the client default? In Apache Kafka 5.0. KIP-1274 runs in three phases: 4.3 logs a recommendation while the default staysclassic; 5.0 switches the default toconsumerand deprecates the classic protocol inKafkaConsumer; 6.0 removes it (checked 2026-09-28).