Explanation¶
Scope
How MinIO works and why: server pools, erasure sets, quorum, healing and bitrot protection, the read and write paths, IAM and STS, encryption through KES, site replication, the threat model, and how the open-source project turned into the commercial AIStor product between 2021 and 2026. The same core design (Go, erasure coding, S3 API) runs in the archived community edition, in AIStor, and in community forks such as PGSTY Silo. Exact limits and tables are in Reference; tasks are in How-to Guides.
MinIO is an S3-compatible object store written in Go. It runs as one minio server binary per node and uses erasure coding, not replication, as its basic resiliency mechanism. It targets local NVMe/SSD drives on commodity servers and scales out by adding server pools.
Architecture Overview¶
The diagram shows a two-pool deployment behind a load balancer, with external identity and key services. Every node can take any request and routes it internally to the right pool and erasure set.
flowchart TB
subgraph Clients["S3 clients"]
SDK["AWS SDKs / minio-go / minio-py"]
MC["mc CLI"]
APP["S3 tools: Spark, Trino, Velero, Loki"]
end
LB["Load balancer<br/>NGINX / HAProxy<br/>(passes headers unchanged for SigV4)"]
subgraph P1["Server pool 1: 4 nodes x 4 drives"]
N1["minio server node1"]
N2["minio server node2"]
N3["minio server node3"]
N4["minio server node4"]
ES1[("Erasure set 1<br/>16 drives, EC:4<br/>1 drive per node, striped")]
end
subgraph P2["Server pool 2: expansion"]
N5["minio server node5..8"]
ES2[("Erasure sets of pool 2")]
end
subgraph Ext["External services"]
IDP["OIDC / LDAP IdP"]
KES["KES (deprecated OSS) or<br/>AIStor MinKMS"]
KMS["Vault / AWS KMS / GCP KMS / Azure KV"]
end
SDK --> LB
MC --> LB
APP --> LB
LB --> N1
LB --> N2
LB --> N3
LB --> N4
LB --> N5
N1 --- ES1
N2 --- ES1
N3 --- ES1
N4 --- ES1
N5 --- ES2
N1 -.->|"internode RPC"| N5
N1 -->|"STS: AssumeRoleWith*"| IDP
N1 -->|"mTLS: generate/decrypt DEK"| KES
KES --> KMS
Server Pools¶
A production deployment starts with at least 4 homogeneous nodes (matching CPU, RAM, storage, and network). MinIO combines all nodes of the initial deployment into one server pool. Single-node single-drive and single-node multi-drive modes also exist, but they give little or no availability.
- Locally attached storage: MinIO expects direct-attached NVMe or SSD drives, formatted XFS, presented as JBOD with no RAID, pooling, or controller caching. Those layers hide failures from MinIO and add unpredictable latency.
- Any node can serve any request: every server knows the full topology. The node that receives a request routes it to the correct erasure set.
- Pool expansion: capacity grows by adding pools. Each pool has its own erasure sets. MinIO must ask each pool where an object lives, so every added pool adds some internode traffic per request. Existing data stays where it is; new objects are placed in pools in proportion to each pool's free space.
- Pool failure is global: each pool needs at least one erasure set with quorum. If a whole pool fails, MinIO cannot tell which pool an object belongs to, so it stops all I/O until the pool returns, even though data on healthy pools is safe.
Erasure Coding¶
How Erasure Coding Works¶
MinIO groups all drives in a pool into erasure sets, the basic unit of availability. It picks the set size (2 to 16 drives; up to 32 in recent AIStor releases) from the node and drive count, then stripes each set symmetrically across nodes: one drive per node, wrapping around when the set is larger than the node count. A 16-node, 8-drive pool therefore gets 8 erasure sets of 16 drives, each with exactly one drive on every node.
For every write, MinIO splits the object with Reed-Solomon coding into K data shards and M parity shards, where N = K + M is the erasure set size and M is the parity (EC:M). The example shows a 12-drive erasure set with EC:4: 8 data shards and 4 parity shards, one per drive.
graph LR
OBJ["Object<br/>(binary data)"]
subgraph EC["12-drive erasure set, EC:4 (K=8, M=4)"]
SH1["Data shard 1"]
SH2["Data shard 2"]
SH3["Data shard 3"]
SH4["Data shard 4"]
SH5["Data shard 5"]
SH6["Data shard 6"]
SH7["Data shard 7"]
SH8["Data shard 8"]
SH9["Parity shard 1"]
SH10["Parity shard 2"]
SH11["Parity shard 3"]
SH12["Parity shard 4"]
end
OBJ --> SH1
OBJ --> SH2
OBJ --> SH3
OBJ --> SH4
OBJ --> SH5
OBJ --> SH6
OBJ --> SH7
OBJ --> SH8
OBJ --> SH9
OBJ --> SH10
OBJ --> SH11
OBJ --> SH12
SH1 --> DRV1[(Drive 1)]
SH2 --> DRV2[(Drive 2)]
SH3 --> DRV3[(Drive 3)]
SH4 --> DRV4[(Drive 4)]
SH5 --> DRV5[(Drive 5)]
SH6 --> DRV6[(Drive 6)]
SH7 --> DRV7[(Drive 7)]
SH8 --> DRV8[(Drive 8)]
SH9 --> DRV9[(Drive 9)]
SH10 --> DRV10[(Drive 10)]
SH11 --> DRV11[(Drive 11)]
SH12 --> DRV12[(Drive 12)]
Parity Levels and Fault Tolerance¶
Parity is a trade-off between usable capacity and how many drives can fail. The default depends on the erasure set size (EC:4 for sets of 8 to 16 drives, lower for smaller sets), and the maximum is half the set. Parity is set deployment-wide with MINIO_STORAGE_CLASS_STANDARD (and MINIO_STORAGE_CLASS_RRS for the reduced-redundancy class); a client selects a class per object with the x-amz-storage-class header. There is no --parity server flag and no per-bucket parity setting. A change affects only objects written afterwards.
For a 16-drive set:
| Parity | Data / parity shards | Usable share of raw capacity | Drive losses tolerated for reads |
|---|---|---|---|
| EC:2 | 14 / 2 | 87.5% | 2 |
| EC:4 (default) | 12 / 4 | 75% | 4 |
| EC:8 (maximum) | 8 / 8 | 50% | 8 |
EC:0 (no parity) is allowed but leaves resiliency entirely to the underlying storage. AIStor documentation requires EC:3 or higher in production. The full table, including write quorum, is in Reference.
Read and Write Quorum¶
- Read quorum is
K, the number of data shards. WithEC:4on 16 drives, any 12 drives (data or parity) are enough to rebuild the object. - Write quorum is also
Kwhen parity is below half the set. It becomesK + 1only when parity is exactly half (EC:8on 16 drives needs 9 drives). The extra drive prevents two halves of a split set from both accepting writes. - Degraded writes get more parity: if drives are offline when an object is written, MinIO raises that object's parity (for example to
EC:6) so it still has the same failure tolerance as objects written to a healthy set. - An erasure set that loses more drives than its parity loses both read and write quorum. Other erasure sets keep working.
Erasure Set Selection¶
MinIO hashes the full object name (BUCKET/PREFIX/.../OBJECT) deterministically to choose the erasure set. The same name always maps to the same set, so no lookup table is needed. Within the set, the order of data and parity shards is shuffled per object, so no drive holds only parity and load spreads evenly.
Erasure Coding Math¶
For a 16-drive set with EC:8 (8 data + 8 parity):
- A 16 MiB object is split into 8 data shards of 2 MiB.
- 8 parity shards of 2 MiB are computed from the data shards.
- All 16 shards (32 MiB in total) are written, one per drive.
- Storage efficiency is 50%: 32 MiB of raw capacity holds 16 MiB of user data.
- Durability: any 8 drives can fail and the object can still be read. Writes need 9 drives.
With the default EC:4 on the same set, the object becomes 12 data shards plus 4 parity shards of about 1.33 MiB, about 21.3 MiB raw for 16 MiB of data (75% efficiency).
Object Healing¶
When drives fail or shards are corrupted, MinIO heals objects automatically:
- It finds damaged or missing shards during reads, during the background scanner cycle, or when a replaced drive is detected.
- It rebuilds lost shards from the remaining data and parity shards (possible while the object keeps read quorum).
- It writes the rebuilt shards to healthy or replacement drives.
- For a client read, the node rebuilds the object in memory first, so the client does not see the damage.
Healing and scanner speed can be tuned (heal max_sleep, max_io; scanner delay, max_wait, cycle). See How-to Guides.
Bitrot Protection¶
Bit rot is silent corruption on the media: decayed charge, firmware bugs, or flipped bits. MinIO computes a HighwayHash checksum for every shard on write and checks it on every read. A shard that fails the check is treated as missing: the object is rebuilt from the other shards and the bad shard is healed. MinIO's HighwayHash implementation is SIMD-accelerated; the docs cite more than 10 GB/s per core on Intel CPUs, much faster than SHA-256 or MD5.
Write Path¶
The sequence shows a PUT through a load balancer to a two-pool deployment.
sequenceDiagram
participant C as S3 client
participant LB as Load balancer
participant N as Receiving minio node
participant P as Other pools
participant ES as Erasure set drives
C->>LB: PUT /bucket/key (SigV4 signed)
LB->>N: Forward request, headers unchanged
N->>N: Verify SigV4 and IAM policy
N->>P: Ask which pool already holds bucket/key
P-->>N: Existing object location or none
N->>N: Hash bucket/key to erasure set, Reed-Solomon encode
N->>ES: Write K data + M parity shards with HighwayHash and xl.meta
ES-->>N: Acks from at least write quorum drives
N-->>C: 200 OK with ETag
Write Internals¶
- Authenticate: the node checks the AWS Signature V4 and evaluates the IAM and bucket policies.
- Pick the pool and set: for an existing object name, the pool that holds it; for a new name, a pool chosen in proportion to free space. Inside the pool, the name hash selects the erasure set.
- Encode and hash: the stream is split into blocks, Reed-Solomon encoded, and each shard gets a HighwayHash checksum.
- Parallel shard write: all shards are written to their drives at the same time.
- Quorum check: the write succeeds once write quorum is met (
Kdrives, orK + 1when parity is half the set). Missing shards are healed later. - Metadata: size, ETag, content type, user metadata, and version information go into an
xl.metafile next to the shards on each drive. Small objects can be stored inline inxl.meta.
Read Path¶
- The client sends
GET /bucket/key. - The node finds the pool and hashes the name to the erasure set.
- It reads
xl.metafrom the set's drives and uses the latest agreed version. - It reads enough shards to rebuild the object (the data shards when they are healthy) and checks each shard's HighwayHash.
- If a shard is missing or fails the check, the node reads parity shards, rebuilds the data, and queues the object for healing.
- It streams the object to the client.
MinIO does not rank drives by speed for reads. The community source only prefers drives local to the serving node (prefer[index] = disk.Hostname() == "" in cmd/erasure-object.go); the parallel reader tries those first and falls back to other shards when a read fails (erasure-object.go, checked 2026-09-27).
Identity and Access Management (IAM)¶
MinIO's IAM model mirrors AWS IAM:
- Root credentials come from
MINIO_ROOT_USERandMINIO_ROOT_PASSWORDat startup. They are like an AWS root account and should be kept for break-glass use. - Users (access key plus secret key) are created with
mc admin user add. Service accounts (access keys) belong to a parent user and can carry a narrower inline policy. - Groups collect users so a policy can be attached once.
- Policies are AWS-IAM-format JSON documents attached to users or groups. MinIO ships five built-in policies (
consoleAdmin,readonly,readwrite,diagnostics,writeonly; see Reference). - Policy variables are filled in when the policy is evaluated:
${aws:username}for the user name,${jwt:claim}for OIDC claims,${ldap:username}for LDAP users. They let one policy give each user a private prefix. - Bucket policies are attached to a bucket instead of a principal. They can allow anonymous access (
Principal: "*") and support condition keys such asaws:SourceIpandaws:SecureTransport. - External identity: OIDC and AD/LDAP providers authenticate users; MinIO maps claims or group DNs to policies. In the community edition, OIDC/LDAP login to the web console was removed in May 2025; programmatic STS still works.
Security Token Service (STS)¶
STS turns an identity (MinIO user, OIDC JWT, LDAP bind, client certificate, or custom token) into temporary credentials: an access key, secret key, and session token that expire after a set duration. Applications never hold long-lived keys, and expired credentials cannot be reused. STS is required whenever an external IdP is used, because it converts IdP identities into SigV4 credentials. The list of STS APIs is in Reference.
Bucket Notifications¶
MinIO publishes S3 event notifications for bucket operations:
- Event types:
s3:ObjectCreated:*,s3:ObjectRemoved:*,s3:ObjectAccessed:*, and replication and ILM events. - Targets: AMQP, Elasticsearch, Kafka, MQTT, MySQL, NATS, NSQ, PostgreSQL, Redis, and webhooks.
- Configure them with
mc event addor the S3 notification API. - Notification configuration is not copied between sites in site replication, by design.
Information Lifecycle Management (ILM)¶
ILM rules automate expiry and tiering:
- Transition rules move objects after N days to a remote tier (another MinIO, AWS S3, Azure Blob, or GCS) created with
mc ilm tier add. The rule's storage class is the tier name, not an AWS class such asGLACIER. MinIO storage classes (STANDARD,REDUCED_REDUNDANCY) only select parity. - Expiration rules delete objects, delete markers, or incomplete multipart uploads after a period.
- Noncurrent version expiration limits how long old versions stay in versioned buckets.
- ILM configuration is not copied between sites in site replication.
Site Replication¶
Site replication links independent MinIO deployments as active-active peers for disaster recovery and geo-local access.
graph TB
subgraph SiteA["Site A (peer)"]
MA[MinIO deployment A]
end
subgraph SiteB["Site B (peer)"]
MB[MinIO deployment B]
end
subgraph SiteC["Site C (peer)"]
MC[MinIO deployment C]
end
GLB["Global load balancer<br/>geo-local routing and failover"]
IDP["Shared OIDC / LDAP IdP"]
KMS["Shared KMS via KES"]
GLB --> MA
GLB --> MB
GLB --> MC
MA <-->|"objects, buckets, IAM"| MB
MB <-->|"objects, buckets, IAM"| MC
MA <-->|"objects, buckets, IAM"| MC
MA -.-> IDP
MA -.-> KMS
- All sites are peers: a write on any site replicates to all others, asynchronously.
- What replicates: buckets and objects, IAM users, groups, policies, STS credentials, service accounts, bucket policies, tags, object lock, and encryption configuration. Versioning is turned on for all buckets.
- What does not: bucket notifications and ILM rules, which are meant to differ per site.
- Prerequisites: only one site may hold data at setup, and all sites must share the IdP and (for SSE) the KMS.
- Latency: replication lag follows inter-site round-trip time. With a 100 ms round trip an object needs at least that long before peers have it.
- Failure handling: failed replications are queued and retried.
Security Model¶
MinIO layers IAM, server-side encryption (SSE), TLS, and audit logging on top of the S3 security model.
Server-Side Encryption (SSE)¶
MinIO uses envelope encryption: every object gets a unique data encryption key (DEK), and the DEK is sealed by a key from a KMS.
| Mode | Who holds the key | Notes |
|---|---|---|
| SSE-S3 | KMS master key, chosen by MinIO | Still needs a KMS (KES, or AIStor MinKMS). Can be the bucket default ("auto-encryption") |
| SSE-KMS | Named key in the external KMS, chosen per bucket or request | Adds per-key access control and KMS audit |
| SSE-C | Client sends a 256-bit key with every request (X-Amz-Server-Side-Encryption-Customer-Key) |
MinIO never stores the key. A lost key means lost data. Requires TLS |
Commands to enable each mode are in How-to Guides.
KES (Key Encryption Service)¶
KES is a stateless proxy between MinIO and the KMS. MinIO nodes authenticate to KES with mutual TLS; KES enforces which keys each identity may use and forwards key operations to the KMS (HashiCorp Vault, AWS KMS, GCP KMS, Azure Key Vault, and others). Because it is stateless, several KES replicas can sit behind a load balancer.
flowchart LR
MINIO["minio server nodes"]
KES["KES replicas"]
KMS["HashiCorp Vault<br/>AWS KMS / GCP KMS / Azure Key Vault"]
MINIO -->|"mTLS: generate / decrypt DEK"| KES
KES -->|"key operations"| KMS
Open-source KES is deprecated
The minio/kes repository says it is no longer maintained. Paying customers get an enterprise KES or MinKMS (AIStor Key Manager). Community deployments keep working with the last KES release but receive no fixes.
TLS¶
MinIO serves TLS when it finds public.crt and private.key in its certificate directory (~/.minio/certs by default, or --certs-dir). Extra CAs go in CAs/, and subdirectories allow several certificates chosen by SNI. An encrypted private key needs MINIO_CERT_PASSWD. TLS is required for SSE-C and strongly advised for any use of STS or SSE, because credentials and keys cross the network. Client certificates can also be used as an identity through AssumeRoleWithCertificate.
Audit Logging¶
Audit targets receive one JSON record per API call: user, action, bucket and object, source IP, request ID, status, and errors. MinIO supports webhook (audit_webhook) and Kafka (audit_kafka) targets, and several targets at once. In the community edition the console no longer shows audit or admin views after the May 2025 removal; AIStor has them.
Identity Provider Integration¶
- OIDC (Keycloak, Okta, Dex, Entra ID): users get a JWT from the provider and exchange it with
AssumeRoleWithWebIdentity. MinIO maps a claim (defaultpolicy, configurable withMINIO_IDENTITY_OPENID_CLAIM_NAME) to policies, or uses role policies. - AD/LDAP: users exchange LDAP credentials with
AssumeRoleWithLDAPIdentity. MinIO maps user and group DNs to policies.
Threat Model Summary¶
| Threat | Mitigation |
|---|---|
| Unauthorized S3 access | IAM policies, bucket policies, access key scoping |
| Data exposure at rest | SSE-S3, SSE-KMS, or SSE-C encryption |
| Key compromise | KMS-managed key lifecycle and rotation through KES or MinKMS |
| Network interception | TLS on all endpoints. mTLS for KES communication |
| Credential leakage | STS temporary credentials with short TTLs |
| Privilege escalation | Policy variable scoping, OIDC claim-based policies. Note CVE-2025-62506 (session-policy bypass, fixed 2025-10-15) |
| Denial of service | Unfixed in community: CVE-2026-39414 (S3 Select CSV memory exhaustion). Restrict s3:PutObject and block S3 Select at the proxy |
| Missing audit trail | Webhook or Kafka audit logging for all API operations |
| Replay attacks | AWS Signature V4 with timestamp validation |
| Unmaintained software | Archived community edition gets no patches; move to AIStor, a maintained fork, or another store |
Key Architectural Properties¶
- Strict S3 API compatibility: requests must be signed with AWS Signature V4 (V2 for old clients). Proxies must not change signed headers.
- Erasure coding by default: no separate replication layer inside a site. Erasure coding gives both resiliency and storage efficiency.
- Deterministic placement: the object-to-erasure-set mapping is a hash, not a metadata lookup. There is no separate metadata server.
- No RAID, no caching: MinIO wants raw XFS drives. RAID or drive-level caching adds unpredictable latency.
- Any-to-any routing: any node can take any request and route it internally.
- Pool-based horizontal scaling: capacity grows by adding pools. Existing data stays in place (pools can be decommissioned to move data off).
Performance Considerations¶
MinIO's throughput depends mostly on drive type, network bandwidth between nodes, and parity. Every write sends N shards over the network, so internode bandwidth usually limits write throughput before CPU does. Higher parity costs capacity and some write bandwidth. Each extra pool adds a lookup on each request. Indicative, unsourced performance figures are kept in Reference; measure real hardware with MinIO warp.
From Open Source to AIStor¶
MinIO began as an Apache-2.0 project and moved to GNU AGPLv3 between October 2019 and May 2021. AGPL made MinIO free for self-hosting but pushed companies that embed or resell it towards a commercial license. From 2025, MinIO, Inc. moved its engineering to the commercial AIStor product and took features and distribution channels away from the community edition, one step at a time. The dated list is in Reference.
The state diagram summarises the life cycle of the community edition.
stateDiagram-v2
[*] --> Apache2: launched under Apache 2.0
Apache2 --> AGPLv3: relicense completed 2021-05
AGPLv3 --> ConsoleStripped: RELEASE.2025-05-24 removes admin UI
ConsoleStripped --> SourceOnly: 2025-10, no binaries or images
SourceOnly --> MaintenanceMode: 2025-12-03
MaintenanceMode --> Archived: 2026-02, no longer maintained
Archived --> ImagesDeleted: 2026-09-11, Docker Hub repos removed
ImagesDeleted --> [*]
Archived --> Forks: community forks continue
Forks --> [*]
Key consequences:
- No community security fixes: the last fixed CVE (CVE-2025-62506) shipped on 2025-10-15. CVE-2026-39414 (published 2026-04-09) affects every community release and is fixed only in AIStor.
- No official community binaries or images: the community edition is built from source (
go install, Go 1.24+). The Docker Hub repositories were deleted on 2026-09-11, and anonymous pulls from Quay return 401 Unauthorized since about 2026-09-24 (Reference). - Tooling around it is frozen too: the Operator (v7.1.1),
mc, KES, the community docs, and the Helm chart are archived, deprecated, or offline. - Free path from the vendor: AIStor Free is single-node only. Multi-node use needs Enterprise Lite (below 400 TiB) or Enterprise, both quote-priced.
- Community response: forks such as PGSTY Silo (AGPLv3, restored console, packages, backported fixes) and the short-lived OpenMaxIO console fork, and more interest in independent stores such as RustFS (Apache 2.0), SeaweedFS, Garage, and Ceph RGW. Comparison facts are in Reference.
Why the architecture still matters
AIStor and the Silo fork keep the same on-disk format (xl.meta), erasure-set model, and S3 behaviour. Knowledge of the internals above carries over to whichever build a team runs.
Sources¶
- MinIO architecture (community docs source)
- Availability and resiliency (community docs source)
- Erasure coding (community docs source)
- AIStor erasure coding
- AIStor healing
- AIStor site replication
- MinIO STS quickstart and certificate STS
- MinIO KMS guide and KES deprecation notice
- AIStor IAM access control
- AIStor server-side encryption
- minio/minio README (source-only distribution and no-longer-maintained notice)
- Blocks and Files: admin UI removed from Community Edition