Skip to content

Apache Kafka

Summary

Apache Kafka is the de facto standard open-source event streaming platform: durable, partitioned, replicated logs with exactly-once transactions and a large ecosystem (Connect, Streams, Schema Registry, MirrorMaker 2). Since 4.0 (March 2025) it runs KRaft-only, with no ZooKeeper. The latest release is 4.3.1 (2026-06-25), and 4.4.0 is in its release process. Queue semantics (share groups, KIP-932) have been production-ready since 4.2. The project is Apache-2.0 licensed and governed by the ASF. Confluent, the main commercial vendor, has been an IBM subsidiary since March 2026.

Overview

Apache Kafka is an open-source distributed event streaming platform. It was built at LinkedIn, open-sourced in 2011, and has been an Apache top-level project since 2012. It stores events in durable, partitioned, replicated commit logs organized into topics. Producers append records, and consumers read at their own pace, tracking durable per-group offsets. Since Kafka 4.0 the platform runs KRaft-only: ZooKeeper support was removed and metadata lives in an internal Raft quorum of controllers.

Kafka's value rests on four properties:

  1. High throughput per broker, from sequential disk I/O, reliance on the page cache, and zero-copy sendfile().
  2. Durable replication through the in-sync replica (ISR) protocol, with Eligible Leader Replicas (ELR, KIP-966) on by default for new clusters since 4.1.
  3. Exactly-once semantics (EOS) for read-process-write workloads, using the idempotent producer (on by default since 3.0) and transactions.
  4. A large ecosystem: Kafka Connect, Kafka Streams, Schema Registry implementations, MirrorMaker 2, tiered storage for cheap long retention, and share groups for queue-style consumption.

Key Facts

Attribute Detail
Website kafka.apache.org
GitHub Stars ~30k+ (apache/kafka)
Latest Version 4.3.1 (2026-06-25)
Supported Lines 4.3.1, 4.2.1 (2026-05-30), 4.1.2 (2026-03-17), per the downloads page
Next Release 4.4.0, in the release process (not published as of 2026-09-25)
Release Cadence Target of 3 minor releases a year (every 4 months). Bug-fix releases for supported lines only
Language Java (plus Scala 2.13 in parts of the broker). Clients in Java and many community languages
Java Requirement Brokers, Connect, tools: Java 17+. Clients and Streams: Java 11+. Tested with Java 17 and 25
License Apache License 2.0
Governance Apache Software Foundation (top-level project since 2012)
Primary Vendor Confluent (founded 2014 by Kafka's original authors). IBM completed its acquisition of Confluent on 2026-03-17
Coordination KRaft (Raft metadata quorum). ZooKeeper removed in 4.0
Wire Protocol Binary TCP, versioned per API key
Storage Format Append-only segmented log, magic v2 record batches
Container Images apache/kafka:4.3.1 (JVM), apache/kafka-native:4.3.1 (GraalVM)

Evaluation

Pros

Pro Detail
Large ecosystem Connect, Streams, Schema Registry implementations, hundreds of connectors, Flink and Spark integration
Durability and replication ISR protocol, configurable min.insync.replicas, ELR (KIP-966)
Exactly-once semantics Idempotent producer plus transactions across topics and partitions, hardened by KIP-890 in 4.0
High throughput LinkedIn measured about 2M writes/s on 3 commodity brokers in 2014 (details)
Replayable history Consumers can re-read from any retained offset, useful for compliance and reprocessing
Tiered storage (KIP-405) Production-ready since 3.9 (early access in 3.6). Offloads closed segments to object storage through a plugin
Queues on the same platform Share groups (KIP-932, production-ready in 4.2) add per-record acknowledgement and more consumers than partitions
Truly open Apache 2.0, with no BSL or SSPL relicensing like Redis or Elastic went through
Multi-language clients Official Java client plus mature community clients (librdkafka and its wrappers, franz-go, sarama)

Cons

Con Detail
Operational complexity Topic design, partition planning, ISR shrinkage, and rolling restarts all need expertise
JVM-bound brokers Heap, GC, and page-cache sizing matter. Some competitors (Redpanda) avoid the JVM entirely
Partition count limits Per-broker and per-cluster partition counts are a real ceiling that depends on hardware and version. Test your limits
Consumer rebalance pain The classic protocol still rebalances group-wide. KIP-848 fixes this but is opt-in (group.protocol=consumer) and needs 4.0+ brokers
Cross-region replication MirrorMaker 2 works but is not transparent: offset translation, lag, and cutover need care
No built-in schema management Confluent Schema Registry, Apicurio, or Karapace must be deployed separately
Tiered storage needs a plugin Apache ships no production RemoteStorageManager, and compacted topics cannot be tiered

Architecture (Summary)

A compact view of a KRaft cluster: controllers own metadata, brokers store and replicate partitions, and closed segments can be offloaded to object storage. The full internals are in Explanation.

flowchart LR
    subgraph Producers["Producers"]
        P1["KafkaProducer<br/>(idempotent)"]
        P2["KafkaProducer<br/>(transactional)"]
    end

    subgraph KafkaCluster["Kafka cluster (KRaft)"]
        direction TB
        subgraph ControllerQuorum["Controller quorum (Raft)"]
            KC1["Controller 1"]
            KC2["Controller 2 (active)"]
            KC3["Controller 3"]
        end
        subgraph Brokers["Broker pool"]
            KS1["Broker 1<br/>LogManager / ReplicaManager"]
            KS2["Broker 2<br/>LogManager / ReplicaManager"]
            KS3["Broker 3<br/>LogManager / ReplicaManager"]
        end
        Brokers -- "Fetch __cluster_metadata" --> ControllerQuorum
    end

    subgraph RemoteStorage["Tiered storage (KIP-405)"]
        S3["RemoteStorageManager plugin<br/>(S3 / GCS / Azure)"]
    end

    subgraph Consumers["Consumers"]
        CG1["Consumer group A"]
        CG2["Consumer group B<br/>(read_committed)"]
        SG["Share group<br/>(queue semantics)"]
    end

    P1 -- "Produce" --> KS1
    P2 -- "Produce (transactional)" --> KS2
    KS2 -- "Replica fetch" --> KS1
    KS3 -- "Replica fetch" --> KS2
    KS1 -- "Upload closed segments" --> S3
    CG1 -- "Fetch" --> KS2
    CG2 -- "Fetch" --> KS3
    SG -- "ShareFetch / ShareAcknowledge" --> KS1

    style ControllerQuorum fill:#1f3a5f,color:#fff
    style KS2 fill:#0d6e0d,color:#fff

Detailed architecture, KRaft consensus, replication, the group protocols, tiered storage, and the security model are in Explanation.

Use Cases

Use case Why Kafka fits
Event sourcing Durable, replayable, ordered per partition
Microservices messaging backbone Decouples producers and consumers, supports fan-out and back-pressure
Real-time stream processing Kafka Streams, Flink, Spark Structured Streaming
CDC (change data capture) Debezium connectors stream Postgres, MySQL, and MongoDB change events into topics
Work queues Share groups (4.2+) for independent jobs with per-record acknowledgement and retries
Log aggregation Replaced Scribe and Flume with replicated storage and replay
Metrics and telemetry transport OpenTelemetry Collector Kafka exporter and receiver, Grafana Tempo's Kafka-based ingest
Data lake ingestion Connect S3 sink, Iceberg sink
Audit trail / immutable event log Compacted topics plus retention policies
Activity stream / clickstream The original LinkedIn use case: high-cardinality data partitioned by user or session

Licensing & Pricing

  • Apache Kafka: Apache License 2.0. Free, no usage restrictions, no telemetry call-home, vendor-neutral.
  • Confluent Platform: Some components (Schema Registry, REST Proxy, ksqlDB) use the source-available Confluent Community License. Confluent Enterprise (commercial) adds RBAC, audit logs, Cluster Linking, and Control Center. Confluent has been a wholly owned IBM subsidiary since 2026-03-17.
  • Confluent Cloud: Fully managed SaaS priced by cluster type (for example Basic, Standard, Enterprise, Dedicated, Freight), throughput, storage, and partitions. It runs on Kora, Confluent's cloud-native Kafka engine.
  • Managed alternatives: Amazon MSK (provisioned and Serverless), Aiven for Apache Kafka, NetApp Instaclustr, and Azure Event Hubs (Kafka-protocol endpoint). Upstash Kafka was deprecated in September 2024 and shut down on 2025-03-11.
  • Self-hosted on Kubernetes: Strimzi (CNCF Incubating, free). Bitnami's free public images and charts were largely retired in 2025, so check image availability before depending on them.

License clarity

Unlike HashiCorp Vault (BSL 1.1), Redis (RSALv2/SSPLv1, now also AGPLv3), Elastic (ELv2/SSPL, now also AGPLv3), or MongoDB (SSPL), Apache Kafka has never been relicensed and remains Apache 2.0. Confluent's add-ons are licensed separately. The IBM acquisition does not change the ASF project's license or governance.

Ecosystem

Component Purpose
Kafka Connect Source and sink integration framework: JDBC, S3, Elasticsearch, MongoDB, Snowflake, BigQuery, Iceberg
Kafka Streams Embedded JVM library for stateful stream processing (KStream, KTable, joins, windowing)
ksqlDB SQL-on-streams engine (Confluent Community License)
Schema Registry Avro, JSON Schema, and Protobuf schema storage with compatibility checks (Confluent, Apicurio, Karapace)
MirrorMaker 2 Connect-based cross-cluster replication (KIP-382). MirrorMaker 1 was removed in 4.0
Cruise Control LinkedIn's automated rebalancing and self-healing
Strimzi CNCF Incubating Kubernetes operator (KRaft-only, kafka.strimzi.io/v1 CRDs)
Debezium CDC connectors built on Kafka Connect
librdkafka C/C++ client library behind the Python, Go, and .NET Confluent clients
kcat (kafkacat) CLI for produce, consume, and metadata
Kroxylicious Kafka-protocol proxy with filters such as record encryption
Conduktor / Kafbat UI / AKHQ / Redpanda Console Web UIs for cluster admin and topic browsing

Compatibility & Requirements

  • JDK: Java 17+ for brokers, controllers, Connect, and tools (Java 11 dropped for the server side in 4.0). Java 11+ for clients and Kafka Streams (Java 8 dropped in 4.0). Java 25 support was added in 4.2.
  • Operating system: Linux strongly preferred for production (sendfile, page cache tuning, epoll). macOS for development only. Windows brokers are not recommended for production.
  • Hardware: SSD or NVMe for log directories, 10 GbE or faster networking, 32 GiB or more RAM per broker for sizeable workloads (page cache).
  • Filesystem: XFS recommended over ext4. Do not put active log directories on network-attached storage. Use tiered storage for cold data.
  • Containers: Official apache/kafka and apache/kafka-native images, KRaft-native, for every supported release. Strimzi and Confluent images are also common.
  • Wire compatibility: 4.x brokers need clients at 2.1 or later, and 4.x clients need brokers at 2.1 or later (KIP-896). Newer brokers accept older API versions within that range.
  • Upgrades: Rolling upgrades are finalized with kafka-features.sh upgrade --release-version <X.Y> (metadata version). The ZooKeeper-era inter.broker.protocol.version no longer applies. See How-to Guides.

Latest Versions

Version Release Highlights
4.4.0 Not yet released Upgrade notes list a share-group DLQ (KIP-1191), broker.id deprecation (KIP-1232), controller unregistration (KIP-1312), and Streams static membership on the streams protocol
4.3.1 2026-06-25 Bug fix: Kafka Streams RocksDB native memory leak (KAFKA-20616)
4.3.0 2026-05-22 25 KIPs. Log-directory cordoning (KIP-1066), follower fetch from tiered offset (KIP-1023), streams-scala deprecated, classic consumer protocol deprecation phase 1 (KIP-1274)
4.2.1 2026-05-30 Share-group deadlock fix, Streams protocol migration fix
4.2.0 2026-02-17 Share groups (KIP-932) production-ready. Streams Rebalance Protocol GA (core features). DLQ in Streams exception handlers. Java 25
4.1.x 2025-09-02 to 2026-03-17 Share groups preview. Streams Rebalance Protocol (KIP-1071) early access. ELR on by default for new clusters. Static-to-dynamic quorum upgrade
4.0.0 2025-03-18 KRaft-only (ZooKeeper removed). KIP-848 GA. ELR preview. Queues early access. Brokers need Java 17. Log4j2
3.9.x 2024-11-06 to 2026-02-21 Last line with ZooKeeper (the migration bridge). Tiered storage production-ready. Dynamic quorums (KIP-853)

The full release and support matrix, KIP status, and CVE list are in Reference.

Alternatives

Alternative Style When to prefer
Redpanda Kafka API, C++ broker, no JVM Lower tail latency, simpler ops, single-binary deploys
Apache Pulsar Separate compute and storage (BookKeeper) Built-in geo-replication, multi-tenancy, native tiered storage
NATS / JetStream Lightweight pub/sub plus streams Edge and IoT, very low overhead, simpler operations
RabbitMQ AMQP broker (plus streams) Per-message TTL, complex routing, priorities, classic work queues
AWS Kinesis Data Streams Managed shard-based stream All-in on AWS, smaller scale, no Kafka API
Google Pub/Sub Managed at-least-once pub/sub All-in on GCP, serverless scaling, weaker ordering
Azure Event Hubs Managed, Kafka-protocol-compatible All-in on Azure. The gateway speaks the Kafka wire protocol
WarpStream Object-storage-native, Kafka-compatible (owned by Confluent since 2024) Diskless brokers with data only in S3, at higher latency
AutoMQ Object-storage-based Kafka fork Cloud cost optimization with the Kafka API

See Streaming Brokers Comparison for a head-to-head of Kafka, Redpanda, and Pulsar, and Messaging Patterns Comparison for queue versus log versus pub/sub.

Migration & Lock-in

  • API surface: The Kafka wire protocol is open. Redpanda, WarpStream, AutoMQ, Azure Event Hubs, and others implement it. Moving clients to a wire-compatible alternative usually means changing bootstrap.servers and security settings.
  • Cross-cluster migration: MirrorMaker 2 (Connect-based) is the standard tool. It replicates topic data, consumer offsets (through the checkpoint connector), heartbeats, and ACLs. Confluent Cluster Linking is a commercial alternative that preserves offsets exactly.
  • ZooKeeper to KRaft: Only possible on 3.x (3.9.x recommended). 4.x cannot read ZooKeeper metadata.
  • Connector lock-in: Some Confluent-licensed connectors do not run outside Confluent Platform. Open-source connectors (Debezium, JDBC, S3) port freely.
  • Operational lock-in: Tiered-storage data is written in the format of the RemoteStorageManager plugin you chose. Switching plugins means re-tiering or a clean cutover.
  • Schema lock-in: Confluent Schema Registry's wire format adds a magic byte and a 4-byte schema ID prefix. Apicurio and Karapace implement the same format for compatibility.

Community Health

  • Apache top-level project since 2012 and one of the most active ASF projects by commit volume. 4.3.0 had 147 contributors.
  • Design changes go through the KIP (Kafka Improvement Proposal) process: public design documents voted on by the community. Recent flagship KIPs are KIP-848 (consumer rebalance), KIP-932 (queues), KIP-405 (tiered storage), KIP-966 (ELR), KIP-853 (dynamic quorums), and KIP-1071 (Streams rebalance).
  • The project targets three minor releases a year and ships bug-fix releases for the supported lines, currently 4.3.x, 4.2.x, and 4.1.x.
  • Confluent (now part of IBM) employs many committers, but the project is governed independently by the ASF PMC.
  • Broad third-party ecosystem: Strimzi (CNCF Incubating), Debezium, Cruise Control, Aiven, Instaclustr, Amazon MSK, Azure Event Hubs.

Topic Map

  • How-to Guides: dev broker, production planning, producer/broker/consumer tuning, upgrades, listener security, troubleshooting, tiered storage, Strimzi, monitoring.
  • Reference: release and support matrix, Java requirements, ports, internal topics, KIP status, config defaults, CLI tools, advisories, hardening checklist, benchmarks.
  • Explanation: KRaft quorum, the replicated log, ISR and Eligible Leader Replicas, group protocols, exactly-once transactions, tiered storage, security model.

Sources

Questions

  • Q: Are share groups ready to replace RabbitMQ for queue workloads? Production-ready since 4.2, but 4.2.1 fixed a critical share-group deadlock (KAFKA-20505), and a native DLQ only arrives in 4.4 (KIP-1191). Evaluate on staging, and prefer 4.2.1+ or 4.3.x. RabbitMQ still wins on routing, priorities, and per-message TTL.
  • Q: When does tiered storage pay off compared with more local disk? Usually once retention runs to weeks at high ingest and most reads are tail reads, with occasional historical replays. The break-even depends on the plugin, object-store request pricing, and cross-AZ costs. No neutral public study found (searched 2026-09-28).
  • Q: How do KRaft controller failovers compare with ZooKeeper failovers in large production clusters? Anecdotally faster and simpler, but few public head-to-head studies exist for very large clusters (more than 200 brokers).
  • Q: What is the practical partitions-per-broker ceiling on KRaft 4.x? Public numbers vary by hardware and replication factor, and no current official per-broker limit is published. Load-test your own ceiling.
  • Q: Will the IBM ownership of Confluent change Confluent's investment in upstream Kafka? Too early to tell as of 2026-09. Watch committer affiliation and KIP authorship.

Answered Questions

  • Q: When will group.protocol=consumer become the client default? In Apache Kafka 5.0. KIP-1274 runs in three phases: 4.3 logs a recommendation while the default stays classic; 5.0 switches the default to consumer and deprecates the classic protocol in KafkaConsumer; 6.0 removes it (checked 2026-09-28).