Argo CD Explanation¶
Scope
How Argo CD works and why it is built this way: its components, the reconciliation and diff model, sync phases, ApplicationSets, HA and sharding, the classic hub-and-spoke multi-cluster model compared with the newer Argo CD Agent, the Source Hydrator (rendered-manifest pattern), and the security model. Look-up tables live in Reference. Step-by-step tasks live in How-to Guides.
Component Overview¶
Argo CD is a set of Kubernetes controllers and services around three CRDs: Application, ApplicationSet and AppProject. Each component has one job:
| Component | Role |
|---|---|
argocd-server |
gRPC/REST API server. Serves the web UI and CLI, handles SSO callbacks and Git webhooks, and enforces RBAC |
argocd-repo-server |
Clones Git repos and pulls Helm/OCI artifacts, then renders manifests (Helm, Kustomize, Jsonnet, plain YAML, Config Management Plugins) |
argocd-application-controller |
Watches target clusters, computes drift between desired and live state, runs sync operations, assesses health (a StatefulSet so it can be sharded) |
argocd-redis |
Cache for rendered manifests, app resource trees and live-state summaries. Losing it costs performance, not data |
argocd-dex-server |
Optional OIDC broker. Needed for SAML, LDAP, GitHub OAuth and other non-OIDC connectors |
argocd-applicationset-controller |
Renders ApplicationSet generators and templates into Application resources |
argocd-notifications-controller |
Sends triggers/templates to Slack, Microsoft Teams Workflows, email, webhooks and similar targets |
argocd-commit-server |
Optional, used only by the Source Hydrator. Pushes rendered manifests to Git with dedicated write credentials |
System Architecture¶
The diagram shows how the control-plane components talk to each other, to sources and to target clusters. The ApplicationSet and notifications controllers work through the Kubernetes API (CRDs), not through argocd-server.
graph TB
subgraph Clients["Clients"]
UI["Web UI"]
CLI["argocd CLI"]
WH["Git webhook<br/>(GitHub/GitLab/Bitbucket)"]
end
subgraph CP["argocd namespace (control plane)"]
Server["argocd-server<br/>(API, UI, RBAC)"]
Dex["argocd-dex-server<br/>(optional OIDC broker)"]
RepoSvr["argocd-repo-server<br/>(clone + render)"]
AppCtrl["argocd-application-controller<br/>(StatefulSet, shardable)"]
AppSet["argocd-applicationset-controller"]
Notif["argocd-notifications-controller"]
Commit["argocd-commit-server<br/>(Source Hydrator only)"]
Redis[("argocd-redis<br/>(cache)")]
KAPI["Kubernetes API<br/>(Application / ApplicationSet / AppProject CRs,<br/>cluster + repo Secrets)"]
end
subgraph Sources["Sources"]
Git["Git repositories"]
Helm["Helm repos / OCI registries"]
end
subgraph Targets["Target clusters"]
InCluster["in-cluster<br/>(kubernetes.default.svc)"]
Remote["Remote clusters<br/>(credentials in cluster Secrets)"]
end
UI --> Server
CLI --> Server
WH --> Server
Server --> Dex
Server --> RepoSvr
Server --> Redis
Server --> KAPI
AppCtrl --> RepoSvr
AppCtrl --> Redis
AppCtrl --> KAPI
AppCtrl -->|"watch + apply"| InCluster
AppCtrl -->|"watch + apply"| Remote
AppCtrl -->|"hydrate request"| Commit
Commit -->|"push rendered manifests"| Git
RepoSvr --> Git
RepoSvr --> Helm
AppSet -->|"create/update Applications"| KAPI
Notif -->|"watch Applications"| KAPI
Reconciliation Model¶
The application controller keeps a watch-based cache of every managed cluster, so a reconciliation does not re-list the cluster. The loop fires on three events: the periodic refresh (timeout.reconciliation, 3 minutes by default, with jitter), a Git webhook, or a change to a watched live resource.
This sequence shows one refresh of one Application with auto-sync enabled:
sequenceDiagram
participant Trig as Refresh trigger (timer or webhook)
participant AppCtrl as application-controller
participant RepoSvr as repo-server
participant Redis as Redis
participant Git as Git or OCI source
participant Cache as Cluster cache (in controller)
participant K8s as Target kube-apiserver
Trig->>AppCtrl: Refresh Application
AppCtrl->>RepoSvr: GenerateManifests(repo, revision, path)
RepoSvr->>Git: ls-remote to resolve revision to commit SHA
RepoSvr->>Redis: Lookup cached manifests for SHA
alt Cache miss
RepoSvr->>Git: fetch (shallow if enabled)
RepoSvr->>RepoSvr: helm template / kustomize build / CMP
RepoSvr->>Redis: Store manifests (24h TTL default)
end
RepoSvr-->>AppCtrl: Desired manifests
AppCtrl->>Cache: Read live objects (kept current by watches)
AppCtrl->>AppCtrl: Diff (legacy 3-way or Server-Side Diff) and health
alt OutOfSync and auto-sync enabled
AppCtrl->>K8s: Apply in phase and wave order (client-side apply, or SSA if ServerSideApply=true)
K8s-->>Cache: Watch events update live state
else Synced
AppCtrl->>AppCtrl: Status Synced, health aggregated
end
Key behaviours:
- Sync status is one of
Synced,OutOfSyncorUnknown(comparison error). Health isHealthy,Progressing,Degraded,Suspended,MissingorUnknown. Since 3.4 an app isMissingonly when all its resources are missing. - Self-heal re-syncs when live state drifts. Pruning removes resources that are no longer in Git only when
prune: trueis set. - Rendered manifests are cached per commit SHA, so a timer refresh with no new commit is cheap. Helm charts changed without a version bump, or Kustomize remote bases, can stay stale until the cache expires (
--repo-cache-expiration).
Drift Detection and Diff Strategies¶
Argo CD compares the desired state (rendered manifests) with the live state (objects from the controller's watch cache) after normalisation:
- Legacy (client-side) diff: a three-way comparison using the
kubectl.kubernetes.io/last-applied-configurationannotation, similar tokubectl apply. This is the default. - Server-Side Diff: sends a Server-Side Apply dry-run to the kube-apiserver and diffs the result, so defaulting and mutating/validating webhooks count. Turn it on per app with
argocd.argoproj.io/compare-options: ServerSideDiff=trueor globally withcontroller.diff.server.side. From 3.6 it is also implied by theServerSideApply=truesync option, and the older Structured-Merge diff is removed.
Since 3.0, status fields are ignored for all resource kinds, and system-level ignoreDifferences rules also suppress resource-update churn. Fields owned by other controllers are ignored only when you configure them, for example ignoreDifferences.managedFieldsManagers or JSON pointers/JQ expressions.
Server-Side Diff and Secrets
Server-Side Diff runs dry-runs with the controller's privileges. In 3.2.0 to 3.2.10 and 3.3.0 to 3.3.8, that path could leak plaintext Secret data to users with read-only Argo CD access (CVE-2026-42880, CVSS 9.6). Stay on a patched release. See Reference.
Resource Tracking¶
Argo CD has to know which live objects belong to which Application. Since 3.0 the default is annotation-based tracking (argocd.argoproj.io/tracking-id), replacing the app.kubernetes.io/instance label. Tools such as Helm and Kustomize often copy labels onto generated objects, which made label tracking claim resources it did not own and confuse prune decisions. The annotation carries the app name, namespace, group/kind and name, so copied metadata no longer causes false ownership.
Sync Phases, Waves and Hooks¶
A sync runs in phases. Within each phase, resources run in ascending wave order (argocd.argoproj.io/sync-wave). Argo CD waits for each wave to become healthy before starting the next.
stateDiagram-v2
[*] --> PreSync
PreSync --> Sync: all PreSync hooks succeeded
PreSync --> SyncFail: hook failed
Sync --> PostSync: resources applied and Healthy
Sync --> SyncFail: apply or hook failed
PostSync --> Succeeded: PostSync hooks succeeded
PostSync --> SyncFail: hook failed
SyncFail --> Failed: SyncFail hooks run (cleanup, alert)
Succeeded --> [*]
Failed --> [*]
Within the Sync phase, waves apply in order. A typical layout is wave -1 for Namespaces and CRDs, 0 for ConfigMaps and Secrets, 1 for Deployments and Services, and 2 for Ingress. Deletion has its own hooks. PreDelete (3.3+) runs Jobs such as data export or traffic drain before an Application's resources are removed, and a failing hook blocks the deletion. PostDelete (2.10+) runs after removal. Hooks do not run during a selective sync.
ApplicationSet Architecture¶
An ApplicationSet pairs one or more generators that produce parameter sets with a template that turns each parameter set into an Application. The controller owns the Applications it generates. By default it creates, updates and deletes them to match the generator output, and spec.syncPolicy.applicationsSync (create-only, create-update, create-delete, sync) can restrict this, provided the controller allows per-ApplicationSet overrides (applicationsetcontroller.enable.policy.override). Generators compose: matrix crosses two generators, for example clusters × Git directories, and merge overlays per-cluster overrides. The full generator list is in Reference.
- Go templates (
goTemplate: true) are the recommended templating mode. They offer Sprig functions andtemplatePatch. The legacy{{name}}fasttemplate syntax still works. - Progressive Syncs (beta since 3.3) add a
RollingSyncstrategy. Generated apps are updated in label-selected steps, for example dev, then staging, then prod, withmaxUpdatelimits. - Security: templated
projectfields and SCM/PR generators can let repository contributors influence where apps deploy. Treat ApplicationSets as admin-level objects, or restrict them to specific namespaces ("ApplicationSets in any namespace" plus SCM provider allow-lists). - 3.5 added an ApplicationSet view to the UI with change previews (per the 3.5 release announcement).
High Availability and Sharding¶
Every component scales on its own axis:
graph TB
LB["Ingress / LoadBalancer"] --> Svr1["argocd-server replica 1"]
LB --> Svr2["argocd-server replica 2"]
subgraph Repo["argocd-repo-server (stateless, CPU/memory bound)"]
RepoSvr1["replica 1"]
RepoSvr2["replica 2"]
end
subgraph Controller["argocd-application-controller StatefulSet (ARGOCD_CONTROLLER_REPLICAS=2)"]
Shard0["shard 0<br/>(clusters hashed to 0)"]
Shard1["shard 1<br/>(clusters hashed to 1)"]
end
subgraph RedisHA["redis-ha"]
HAP["HAProxy"] --> R1[("Redis + Sentinel x3")]
end
Svr1 --> Repo
Svr2 --> Repo
Shard0 --> Repo
Shard1 --> Repo
Svr1 --> HAP
Svr2 --> HAP
Shard0 --> HAP
Shard1 --> HAP
argocd-server: stateless, runs 2+ replicas behind an Ingress.argocd-repo-server: stateless, butfork/execof Helm and Kustomize is CPU and memory hungry. Scale replicas and bound concurrency withreposerver.parallelism.limit. Clones live in/tmp, so large monorepos may need a bigger volume.argocd-application-controller: work is sharded by cluster. The shard count must matchARGOCD_CONTROLLER_REPLICAS. The default algorithm islegacy(UID hash, uneven);round-robinandconsistent-hashingare alpha. A single very large cluster cannot be split across shards, so memory tracks that cluster's object count.- Redis: the HA manifests ship
redis-ha(Sentinel) behind HAProxy. Redis holds only a cache. After a restart it is rebuilt, at the cost of a slow warm-up.
Multi-Cluster Models¶
Classic hub-and-spoke (push)¶
A single Argo CD instance stores a credential for each remote cluster in a Secret labelled argocd.argoproj.io/secret-type: cluster: a bearer token, exec-provider or cloud IAM (EKS IRSA/Pod Identity, GKE Workload Identity, AKS workload identity). The application controller opens watches against every cluster's kube-apiserver and pushes applies to it.
graph LR
subgraph Mgmt["Management cluster"]
Hub["Argo CD<br/>(controller holds watches)"]
CS[("cluster Secrets<br/>secret-type: cluster")]
end
subgraph Fleet["Workload clusters"]
ProdUS["prod-us kube-apiserver"]
ProdEU["prod-eu kube-apiserver"]
Dev1["dev-1 kube-apiserver"]
end
CS -.-> Hub
Hub -->|"bearer token / IAM"| ProdUS
Hub -->|"bearer token / IAM"| ProdEU
Hub -->|"exec provider"| Dev1
This gives one pane of glass and one place to manage RBAC. The trade-offs are that the hub needs network reach into every API server, holds credentials for all of them, and its memory grows with the total number of watched objects in the fleet.
Argo CD Agent (pull, hub-and-spoke)¶
argocd-agent (an argoproj-labs project, pre-GA 0.x as of 2026-09) reverses the connection. A principal runs on the control-plane ("hub") cluster next to the Argo CD API server and UI. An agent plus a slim Argo CD (at least an application controller) runs on each workload ("spoke") cluster. Agents dial out to the principal over a gRPC-based, bi-directional, mTLS-authenticated channel. The hub never holds spoke credentials and never reconciles clusters itself.
graph TB
subgraph Hub["Control-plane cluster (hub)"]
API["argocd-server + UI"]
Principal["argocd-agent principal<br/>(resource proxy, event queues)"]
HubCR[("Applications / AppProjects")]
end
subgraph SpokeA["Workload cluster A: managed mode"]
AgentA["agent"]
CtrlA["application-controller + repo-server"]
end
subgraph SpokeB["Workload cluster B: autonomous mode"]
AgentB["agent"]
CtrlB["application-controller + repo-server"]
GitB["Git (apps defined locally)"]
end
AgentA -->|"outbound gRPC + mTLS"| Principal
AgentB -->|"outbound gRPC + mTLS"| Principal
Principal --- HubCR
API --- Principal
AgentA --> CtrlA
CtrlB --> GitB
AgentB --> CtrlB
- Managed mode: the hub is the source of truth. Applications are created on the hub, and the agent copies them to the spoke and reports status back.
- Autonomous mode: the spoke is the source of truth. Applications are defined on the spoke, often by a local app-of-apps, and the hub gets observability only. This mode suits air-gapped or intermittently connected sites.
- Modes can be mixed across a fleet. The principal also proxies live-resource views and logs back to the central UI.
- Red Hat ships it in OpenShift GitOps (documented for 1.19 and 1.20).
When to consider the agent
Choose the agent when you have hundreds of clusters, edge or intermittent networks, or security rules that forbid inbound access to spoke API servers. For a handful of well-connected clusters, the classic model is simpler and fully GA.
Source Hydrator (Rendered-Manifest Pattern)¶
Helm and Kustomize keep configuration DRY but hide what actually gets applied. The Source Hydrator was introduced as alpha in 2.14 and became beta in 3.5. It is disabled by default and enabled with hydrator.enabled: "true" plus the commit-server. It renders a dry source and commits the plain YAML to a sync source branch. The cluster then syncs from that branch, so every deployed change is a readable Git diff.
sequenceDiagram
participant Dev as Developer PR
participant Dry as Dry branch (main)
participant Ctrl as application-controller
participant Repo as repo-server
participant Commit as commit-server
participant Env as environments/dev branch
participant K8s as Cluster
Dev->>Dry: Merge change to Helm values
Ctrl->>Repo: Render drySource at new SHA
Repo-->>Ctrl: Hydrated manifests
Ctrl->>Commit: Hydrate request (paths, dry SHA)
Commit->>Env: Push commit (repository-write credentials, git note records dry SHA)
Ctrl->>Env: syncSource refresh
Ctrl->>K8s: Sync hydrated manifests
spec.sourceHydrator.drySourcegives the repo, path and revision.syncSourcegivestargetBranchandpath. The optionalhydrateTonames a staging branch so a hydrated commit can be reviewed or promoted before deployment.- Push credentials come from Secrets labelled
argocd.argoproj.io/secret-type: repository-write, kept separate from pull credentials for isolation. - Since 3.3, hydration state is tracked with git notes (3.2 made root hydration paths invalid).
- GitOps Promoter (argoproj-labs) builds on hydrated branches. It promotes changes between environment branches through pull requests gated by commit statuses, and Argo CD ships health checks for its
PromotionStrategyandChangeTransferPolicyCRs.
Repo Server: Manifest Rendering¶
- Clones repositories into
/tmp(orTMPDIR), resolves symbolic revisions withgit ls-remote, and keeps the working tree clean between renders. Shallow clones (3.3+) cut fetch time on large histories. - Runs
helm template,kustomize build, Jsonnet, or a Config Management Plugin sidecar as child processes, with anARGOCD_EXEC_TIMEOUT(90s default) and SIGTERM then SIGKILL. - Caches output in Redis keyed by commit SHA and app parameters (24h default).
- Because it runs templating code from repositories, it is the component most exposed to malicious input. It has no Kubernetes API token by default and should have tightly restricted egress.
- 3.5 made Helm 4 the renderer and added opt-in mTLS between the repo-server and its clients.
Storage Model¶
Argo CD has no database. All durable state lives in Kubernetes objects, so Git plus a backup of the argocd namespace is enough for disaster recovery.
| Data | Storage |
|---|---|
| Application, ApplicationSet, project specs | Application / ApplicationSet / AppProject CRs (in argocd, or allowed namespaces with apps-in-any-namespace) |
| Cluster credentials | Secrets labelled argocd.argoproj.io/secret-type: cluster |
| Repository credentials | Secrets labelled repository or repo-creds (templates), plus repository-write for the hydrator |
| RBAC policies | argocd-rbac-cm ConfigMap and AppProject.spec.roles |
| Settings | argocd-cm, argocd-cmd-params-cm, argocd-secret |
| Runtime cache | Redis (ephemeral, rebuilt on restart) |
| Resource health (since 3.0) | Kept outside the Application CR status by default |
| Audit trail | Kubernetes events on Applications plus the kube-apiserver audit log |
Security Model¶
Threat Model¶
Argo CD holds write access to every cluster it manages and reads every repository it deploys from. That makes it a high-value target.
| Asset / path | Threat | Primary control |
|---|---|---|
API and UI (argocd-server) |
Unauthorised sync, secret exposure through diffs or APIs | SSO with group-mapped RBAC, deny-by-default policy.default, prompt patching (see CVEs) |
| Git / Helm / OCI sources | Malicious manifest or template injection | Branch protection, AppProject.sourceRepos, Source Integrity commit signatures (3.5+) |
| repo-server | Code execution through templating tools and plugins | No K8s token, egress NetworkPolicy, separate CMP sidecars, mTLS (3.5+) |
| application-controller credentials | Cluster-admin on every cluster | Sync impersonation (beta), AppProject destination/resource allow-lists, agent model for spokes |
Cluster/repo Secrets in argocd namespace |
Credential theft | Restrict namespace RBAC, prefer cloud IAM and short-lived tokens |
Authentication¶
- Built-in
admin: its bcrypt hash lives inargocd-secret, and the initial password is inargocd-initial-admin-secret. Disable it withadmin.enabled: "false"once SSO works. - Local accounts (
accounts.<name>inargocd-cm) withapiKeyand/orlogincapabilities, for automation. - API tokens are JWTs signed by Argo CD, issued per account (
argocd account generate-token) or per project role (argocd proj role create-token). - SSO: native OIDC (Okta, Entra ID, Keycloak and others), with the PKCE code flow handled by the server since 3.1 and background token refresh since 3.3. Bundled Dex covers SAML, LDAP and GitHub. Since 3.0, Dex users are identified by
federated_claims.user_idinstead of the unstablesub.
Configuration recipes are in How-to Guides.
Authorization¶
RBAC is Casbin-based (argocd-rbac-cm). AppProjects are the tenancy boundary. They restrict source repos, destination clusters and namespaces, allowed cluster-scoped and namespaced kinds, sync windows, and project roles with their own tokens. policy.default applies to every authenticated user without a mapping. The empty default denies everything, which is the recommended setting. Anonymous access exists only if users.anonymous.enabled is set.
In 3.0, update/delete on an Application stopped implying rights on its managed resources (update/*, delete/*), and logs became a first-class, always-enforced permission. Both changes close escalation paths where "can edit an app" silently meant "can edit or delete everything it deploys" or "can read pod logs". The full resource/action matrix is in Reference.
Encryption and Secrets¶
| Asset | Protection |
|---|---|
| Admin password | bcrypt hash in argocd-secret |
| SSO client secrets | Stored in argocd-secret (base64, not encrypted); referenced as $key from argocd-cm |
| Cluster and repository credentials | Kubernetes Secrets (base64), so rely on etcd encryption at rest and namespace RBAC |
| Redis | Password auth (Secret argocd-redis); cache contents are rebuilt on restart |
| Dex | Runs with in-memory storage in the default manifests; no persistent database |
In transit, argocd-server, argocd-repo-server and argocd-dex-server serve TLS, with self-signed certificates unless you supply your own. Since 3.5, repo-server connections can be upgraded to mutual TLS. Redis traffic is not TLS-encrypted in the default manifests, so rely on NetworkPolicies or configure Redis TLS yourself.
Argo CD deliberately does not manage secret values. The common patterns are:
| Approach | Mechanism | Trade-off |
|---|---|---|
| External Secrets Operator (ESO) | Argo CD deploys ExternalSecret/SecretStore, and ESO syncs values from Vault, AWS SM, GCP SM or Azure KV |
Best separation; secret values never pass through Argo CD |
| Argo CD Vault Plugin (AVP) | CMP replaces <path:...#key> placeholders at render time |
Secrets end up in the repo-server cache, and the upstream project discourages render-time injection |
| SOPS + KSOPS | Encrypted files in Git, decrypted by a Kustomize plugin during render | Everything in Git; key management is critical |
| Sealed Secrets | kubeseal-encrypted SealedSecret in Git, decrypted by an in-cluster controller |
Simple; the encrypted blobs are cluster-key specific |
Recommended pattern
Use External Secrets Operator for production. Argo CD manages only the ExternalSecret and SecretStore manifests, and the values stay in the external store.
Network Segmentation¶
- Restrict
argocd-repo-serveregress to the required Git, Helm and OCI hosts. A compromised template should not be able to reach internal services. - Allow Redis ingress only from Argo CD pods.
- Keep the repo-server without a Kubernetes API token, which is the default.
- Consider a separate Argo CD instance, or dedicated CMP sidecars, for untrusted repositories.
The NetworkPolicy recipe is in How-to Guides.
Known Pitfalls¶
| Pitfall | Impact | Mitigation |
|---|---|---|
policy.default: role:readonly with many SSO users |
Every authenticated user can read all app manifests and diffs, which may contain sensitive config | Set policy.default: "" and map groups explicitly |
| Static cluster bearer tokens that never expire | A compromised token gives persistent cluster access | Use short-lived tokens or cloud IAM, or the agent model |
Missing ignoreDifferences for operator-mutated fields |
Endless sync loops against cert-manager, Gatekeeper and similar | Add resource.customizations.ignoreDifferences |
| Unrestricted repo-server egress | A malicious repo can exfiltrate data or pivot | Strict egress NetworkPolicies |
| Running an unpatched minor | Known critical CVEs (e.g. CVE-2025-55190, CVE-2026-42880) | Stay within the three supported minors |
Why Argo CD 3.0 Changed Defaults¶
The 3.0 release (May 2025) contained few new features. Instead it made safer and faster behaviour the default: annotation tracking, fine-grained RBAC without inheritance, always-on logs RBAC, default exclusion of high-churn kinds such as Endpoints, EndpointSlice and Lease, and health kept out of the Application status. Each change cuts either API-server load and controller churn or an implicit privilege. The 3.x line since then has shipped on the quarterly cadence. PreDelete hooks, OIDC refresh and shallow clones arrived in 3.3. Cluster reconciliation pause and Teams Workflows notifications came in 3.4. 3.5 brought Helm 4, repo-server mTLS, Source Integrity, and beta status for impersonation and the Source Hydrator. The per-minor list is in Reference.