Skip to content

Argo CD Explanation

Scope

How Argo CD works and why it is built this way: its components, the reconciliation and diff model, sync phases, ApplicationSets, HA and sharding, the classic hub-and-spoke multi-cluster model compared with the newer Argo CD Agent, the Source Hydrator (rendered-manifest pattern), and the security model. Look-up tables live in Reference. Step-by-step tasks live in How-to Guides.

Component Overview

Argo CD is a set of Kubernetes controllers and services around three CRDs: Application, ApplicationSet and AppProject. Each component has one job:

Component Role
argocd-server gRPC/REST API server. Serves the web UI and CLI, handles SSO callbacks and Git webhooks, and enforces RBAC
argocd-repo-server Clones Git repos and pulls Helm/OCI artifacts, then renders manifests (Helm, Kustomize, Jsonnet, plain YAML, Config Management Plugins)
argocd-application-controller Watches target clusters, computes drift between desired and live state, runs sync operations, assesses health (a StatefulSet so it can be sharded)
argocd-redis Cache for rendered manifests, app resource trees and live-state summaries. Losing it costs performance, not data
argocd-dex-server Optional OIDC broker. Needed for SAML, LDAP, GitHub OAuth and other non-OIDC connectors
argocd-applicationset-controller Renders ApplicationSet generators and templates into Application resources
argocd-notifications-controller Sends triggers/templates to Slack, Microsoft Teams Workflows, email, webhooks and similar targets
argocd-commit-server Optional, used only by the Source Hydrator. Pushes rendered manifests to Git with dedicated write credentials

System Architecture

The diagram shows how the control-plane components talk to each other, to sources and to target clusters. The ApplicationSet and notifications controllers work through the Kubernetes API (CRDs), not through argocd-server.

graph TB
    subgraph Clients["Clients"]
        UI["Web UI"]
        CLI["argocd CLI"]
        WH["Git webhook<br/>(GitHub/GitLab/Bitbucket)"]
    end

    subgraph CP["argocd namespace (control plane)"]
        Server["argocd-server<br/>(API, UI, RBAC)"]
        Dex["argocd-dex-server<br/>(optional OIDC broker)"]
        RepoSvr["argocd-repo-server<br/>(clone + render)"]
        AppCtrl["argocd-application-controller<br/>(StatefulSet, shardable)"]
        AppSet["argocd-applicationset-controller"]
        Notif["argocd-notifications-controller"]
        Commit["argocd-commit-server<br/>(Source Hydrator only)"]
        Redis[("argocd-redis<br/>(cache)")]
        KAPI["Kubernetes API<br/>(Application / ApplicationSet / AppProject CRs,<br/>cluster + repo Secrets)"]
    end

    subgraph Sources["Sources"]
        Git["Git repositories"]
        Helm["Helm repos / OCI registries"]
    end

    subgraph Targets["Target clusters"]
        InCluster["in-cluster<br/>(kubernetes.default.svc)"]
        Remote["Remote clusters<br/>(credentials in cluster Secrets)"]
    end

    UI --> Server
    CLI --> Server
    WH --> Server
    Server --> Dex
    Server --> RepoSvr
    Server --> Redis
    Server --> KAPI
    AppCtrl --> RepoSvr
    AppCtrl --> Redis
    AppCtrl --> KAPI
    AppCtrl -->|"watch + apply"| InCluster
    AppCtrl -->|"watch + apply"| Remote
    AppCtrl -->|"hydrate request"| Commit
    Commit -->|"push rendered manifests"| Git
    RepoSvr --> Git
    RepoSvr --> Helm
    AppSet -->|"create/update Applications"| KAPI
    Notif -->|"watch Applications"| KAPI

Reconciliation Model

The application controller keeps a watch-based cache of every managed cluster, so a reconciliation does not re-list the cluster. The loop fires on three events: the periodic refresh (timeout.reconciliation, 3 minutes by default, with jitter), a Git webhook, or a change to a watched live resource.

This sequence shows one refresh of one Application with auto-sync enabled:

sequenceDiagram
    participant Trig as Refresh trigger (timer or webhook)
    participant AppCtrl as application-controller
    participant RepoSvr as repo-server
    participant Redis as Redis
    participant Git as Git or OCI source
    participant Cache as Cluster cache (in controller)
    participant K8s as Target kube-apiserver

    Trig->>AppCtrl: Refresh Application
    AppCtrl->>RepoSvr: GenerateManifests(repo, revision, path)
    RepoSvr->>Git: ls-remote to resolve revision to commit SHA
    RepoSvr->>Redis: Lookup cached manifests for SHA
    alt Cache miss
        RepoSvr->>Git: fetch (shallow if enabled)
        RepoSvr->>RepoSvr: helm template / kustomize build / CMP
        RepoSvr->>Redis: Store manifests (24h TTL default)
    end
    RepoSvr-->>AppCtrl: Desired manifests
    AppCtrl->>Cache: Read live objects (kept current by watches)
    AppCtrl->>AppCtrl: Diff (legacy 3-way or Server-Side Diff) and health
    alt OutOfSync and auto-sync enabled
        AppCtrl->>K8s: Apply in phase and wave order (client-side apply, or SSA if ServerSideApply=true)
        K8s-->>Cache: Watch events update live state
    else Synced
        AppCtrl->>AppCtrl: Status Synced, health aggregated
    end

Key behaviours:

  • Sync status is one of Synced, OutOfSync or Unknown (comparison error). Health is Healthy, Progressing, Degraded, Suspended, Missing or Unknown. Since 3.4 an app is Missing only when all its resources are missing.
  • Self-heal re-syncs when live state drifts. Pruning removes resources that are no longer in Git only when prune: true is set.
  • Rendered manifests are cached per commit SHA, so a timer refresh with no new commit is cheap. Helm charts changed without a version bump, or Kustomize remote bases, can stay stale until the cache expires (--repo-cache-expiration).

Drift Detection and Diff Strategies

Argo CD compares the desired state (rendered manifests) with the live state (objects from the controller's watch cache) after normalisation:

  1. Legacy (client-side) diff: a three-way comparison using the kubectl.kubernetes.io/last-applied-configuration annotation, similar to kubectl apply. This is the default.
  2. Server-Side Diff: sends a Server-Side Apply dry-run to the kube-apiserver and diffs the result, so defaulting and mutating/validating webhooks count. Turn it on per app with argocd.argoproj.io/compare-options: ServerSideDiff=true or globally with controller.diff.server.side. From 3.6 it is also implied by the ServerSideApply=true sync option, and the older Structured-Merge diff is removed.

Since 3.0, status fields are ignored for all resource kinds, and system-level ignoreDifferences rules also suppress resource-update churn. Fields owned by other controllers are ignored only when you configure them, for example ignoreDifferences.managedFieldsManagers or JSON pointers/JQ expressions.

Server-Side Diff and Secrets

Server-Side Diff runs dry-runs with the controller's privileges. In 3.2.0 to 3.2.10 and 3.3.0 to 3.3.8, that path could leak plaintext Secret data to users with read-only Argo CD access (CVE-2026-42880, CVSS 9.6). Stay on a patched release. See Reference.

Resource Tracking

Argo CD has to know which live objects belong to which Application. Since 3.0 the default is annotation-based tracking (argocd.argoproj.io/tracking-id), replacing the app.kubernetes.io/instance label. Tools such as Helm and Kustomize often copy labels onto generated objects, which made label tracking claim resources it did not own and confuse prune decisions. The annotation carries the app name, namespace, group/kind and name, so copied metadata no longer causes false ownership.

Sync Phases, Waves and Hooks

A sync runs in phases. Within each phase, resources run in ascending wave order (argocd.argoproj.io/sync-wave). Argo CD waits for each wave to become healthy before starting the next.

stateDiagram-v2
    [*] --> PreSync
    PreSync --> Sync: all PreSync hooks succeeded
    PreSync --> SyncFail: hook failed
    Sync --> PostSync: resources applied and Healthy
    Sync --> SyncFail: apply or hook failed
    PostSync --> Succeeded: PostSync hooks succeeded
    PostSync --> SyncFail: hook failed
    SyncFail --> Failed: SyncFail hooks run (cleanup, alert)
    Succeeded --> [*]
    Failed --> [*]

Within the Sync phase, waves apply in order. A typical layout is wave -1 for Namespaces and CRDs, 0 for ConfigMaps and Secrets, 1 for Deployments and Services, and 2 for Ingress. Deletion has its own hooks. PreDelete (3.3+) runs Jobs such as data export or traffic drain before an Application's resources are removed, and a failing hook blocks the deletion. PostDelete (2.10+) runs after removal. Hooks do not run during a selective sync.

ApplicationSet Architecture

An ApplicationSet pairs one or more generators that produce parameter sets with a template that turns each parameter set into an Application. The controller owns the Applications it generates. By default it creates, updates and deletes them to match the generator output, and spec.syncPolicy.applicationsSync (create-only, create-update, create-delete, sync) can restrict this, provided the controller allows per-ApplicationSet overrides (applicationsetcontroller.enable.policy.override). Generators compose: matrix crosses two generators, for example clusters × Git directories, and merge overlays per-cluster overrides. The full generator list is in Reference.

  • Go templates (goTemplate: true) are the recommended templating mode. They offer Sprig functions and templatePatch. The legacy {{name}} fasttemplate syntax still works.
  • Progressive Syncs (beta since 3.3) add a RollingSync strategy. Generated apps are updated in label-selected steps, for example dev, then staging, then prod, with maxUpdate limits.
  • Security: templated project fields and SCM/PR generators can let repository contributors influence where apps deploy. Treat ApplicationSets as admin-level objects, or restrict them to specific namespaces ("ApplicationSets in any namespace" plus SCM provider allow-lists).
  • 3.5 added an ApplicationSet view to the UI with change previews (per the 3.5 release announcement).

High Availability and Sharding

Every component scales on its own axis:

graph TB
    LB["Ingress / LoadBalancer"] --> Svr1["argocd-server replica 1"]
    LB --> Svr2["argocd-server replica 2"]

    subgraph Repo["argocd-repo-server (stateless, CPU/memory bound)"]
        RepoSvr1["replica 1"]
        RepoSvr2["replica 2"]
    end

    subgraph Controller["argocd-application-controller StatefulSet (ARGOCD_CONTROLLER_REPLICAS=2)"]
        Shard0["shard 0<br/>(clusters hashed to 0)"]
        Shard1["shard 1<br/>(clusters hashed to 1)"]
    end

    subgraph RedisHA["redis-ha"]
        HAP["HAProxy"] --> R1[("Redis + Sentinel x3")]
    end

    Svr1 --> Repo
    Svr2 --> Repo
    Shard0 --> Repo
    Shard1 --> Repo
    Svr1 --> HAP
    Svr2 --> HAP
    Shard0 --> HAP
    Shard1 --> HAP
  • argocd-server: stateless, runs 2+ replicas behind an Ingress.
  • argocd-repo-server: stateless, but fork/exec of Helm and Kustomize is CPU and memory hungry. Scale replicas and bound concurrency with reposerver.parallelism.limit. Clones live in /tmp, so large monorepos may need a bigger volume.
  • argocd-application-controller: work is sharded by cluster. The shard count must match ARGOCD_CONTROLLER_REPLICAS. The default algorithm is legacy (UID hash, uneven); round-robin and consistent-hashing are alpha. A single very large cluster cannot be split across shards, so memory tracks that cluster's object count.
  • Redis: the HA manifests ship redis-ha (Sentinel) behind HAProxy. Redis holds only a cache. After a restart it is rebuilt, at the cost of a slow warm-up.

Multi-Cluster Models

Classic hub-and-spoke (push)

A single Argo CD instance stores a credential for each remote cluster in a Secret labelled argocd.argoproj.io/secret-type: cluster: a bearer token, exec-provider or cloud IAM (EKS IRSA/Pod Identity, GKE Workload Identity, AKS workload identity). The application controller opens watches against every cluster's kube-apiserver and pushes applies to it.

graph LR
    subgraph Mgmt["Management cluster"]
        Hub["Argo CD<br/>(controller holds watches)"]
        CS[("cluster Secrets<br/>secret-type: cluster")]
    end
    subgraph Fleet["Workload clusters"]
        ProdUS["prod-us kube-apiserver"]
        ProdEU["prod-eu kube-apiserver"]
        Dev1["dev-1 kube-apiserver"]
    end
    CS -.-> Hub
    Hub -->|"bearer token / IAM"| ProdUS
    Hub -->|"bearer token / IAM"| ProdEU
    Hub -->|"exec provider"| Dev1

This gives one pane of glass and one place to manage RBAC. The trade-offs are that the hub needs network reach into every API server, holds credentials for all of them, and its memory grows with the total number of watched objects in the fleet.

Argo CD Agent (pull, hub-and-spoke)

argocd-agent (an argoproj-labs project, pre-GA 0.x as of 2026-09) reverses the connection. A principal runs on the control-plane ("hub") cluster next to the Argo CD API server and UI. An agent plus a slim Argo CD (at least an application controller) runs on each workload ("spoke") cluster. Agents dial out to the principal over a gRPC-based, bi-directional, mTLS-authenticated channel. The hub never holds spoke credentials and never reconciles clusters itself.

graph TB
    subgraph Hub["Control-plane cluster (hub)"]
        API["argocd-server + UI"]
        Principal["argocd-agent principal<br/>(resource proxy, event queues)"]
        HubCR[("Applications / AppProjects")]
    end
    subgraph SpokeA["Workload cluster A: managed mode"]
        AgentA["agent"]
        CtrlA["application-controller + repo-server"]
    end
    subgraph SpokeB["Workload cluster B: autonomous mode"]
        AgentB["agent"]
        CtrlB["application-controller + repo-server"]
        GitB["Git (apps defined locally)"]
    end
    AgentA -->|"outbound gRPC + mTLS"| Principal
    AgentB -->|"outbound gRPC + mTLS"| Principal
    Principal --- HubCR
    API --- Principal
    AgentA --> CtrlA
    CtrlB --> GitB
    AgentB --> CtrlB
  • Managed mode: the hub is the source of truth. Applications are created on the hub, and the agent copies them to the spoke and reports status back.
  • Autonomous mode: the spoke is the source of truth. Applications are defined on the spoke, often by a local app-of-apps, and the hub gets observability only. This mode suits air-gapped or intermittently connected sites.
  • Modes can be mixed across a fleet. The principal also proxies live-resource views and logs back to the central UI.
  • Red Hat ships it in OpenShift GitOps (documented for 1.19 and 1.20).

When to consider the agent

Choose the agent when you have hundreds of clusters, edge or intermittent networks, or security rules that forbid inbound access to spoke API servers. For a handful of well-connected clusters, the classic model is simpler and fully GA.

Source Hydrator (Rendered-Manifest Pattern)

Helm and Kustomize keep configuration DRY but hide what actually gets applied. The Source Hydrator was introduced as alpha in 2.14 and became beta in 3.5. It is disabled by default and enabled with hydrator.enabled: "true" plus the commit-server. It renders a dry source and commits the plain YAML to a sync source branch. The cluster then syncs from that branch, so every deployed change is a readable Git diff.

sequenceDiagram
    participant Dev as Developer PR
    participant Dry as Dry branch (main)
    participant Ctrl as application-controller
    participant Repo as repo-server
    participant Commit as commit-server
    participant Env as environments/dev branch
    participant K8s as Cluster

    Dev->>Dry: Merge change to Helm values
    Ctrl->>Repo: Render drySource at new SHA
    Repo-->>Ctrl: Hydrated manifests
    Ctrl->>Commit: Hydrate request (paths, dry SHA)
    Commit->>Env: Push commit (repository-write credentials, git note records dry SHA)
    Ctrl->>Env: syncSource refresh
    Ctrl->>K8s: Sync hydrated manifests
  • spec.sourceHydrator.drySource gives the repo, path and revision. syncSource gives targetBranch and path. The optional hydrateTo names a staging branch so a hydrated commit can be reviewed or promoted before deployment.
  • Push credentials come from Secrets labelled argocd.argoproj.io/secret-type: repository-write, kept separate from pull credentials for isolation.
  • Since 3.3, hydration state is tracked with git notes (3.2 made root hydration paths invalid).
  • GitOps Promoter (argoproj-labs) builds on hydrated branches. It promotes changes between environment branches through pull requests gated by commit statuses, and Argo CD ships health checks for its PromotionStrategy and ChangeTransferPolicy CRs.

Repo Server: Manifest Rendering

  • Clones repositories into /tmp (or TMPDIR), resolves symbolic revisions with git ls-remote, and keeps the working tree clean between renders. Shallow clones (3.3+) cut fetch time on large histories.
  • Runs helm template, kustomize build, Jsonnet, or a Config Management Plugin sidecar as child processes, with an ARGOCD_EXEC_TIMEOUT (90s default) and SIGTERM then SIGKILL.
  • Caches output in Redis keyed by commit SHA and app parameters (24h default).
  • Because it runs templating code from repositories, it is the component most exposed to malicious input. It has no Kubernetes API token by default and should have tightly restricted egress.
  • 3.5 made Helm 4 the renderer and added opt-in mTLS between the repo-server and its clients.

Storage Model

Argo CD has no database. All durable state lives in Kubernetes objects, so Git plus a backup of the argocd namespace is enough for disaster recovery.

Data Storage
Application, ApplicationSet, project specs Application / ApplicationSet / AppProject CRs (in argocd, or allowed namespaces with apps-in-any-namespace)
Cluster credentials Secrets labelled argocd.argoproj.io/secret-type: cluster
Repository credentials Secrets labelled repository or repo-creds (templates), plus repository-write for the hydrator
RBAC policies argocd-rbac-cm ConfigMap and AppProject.spec.roles
Settings argocd-cm, argocd-cmd-params-cm, argocd-secret
Runtime cache Redis (ephemeral, rebuilt on restart)
Resource health (since 3.0) Kept outside the Application CR status by default
Audit trail Kubernetes events on Applications plus the kube-apiserver audit log

Security Model

Threat Model

Argo CD holds write access to every cluster it manages and reads every repository it deploys from. That makes it a high-value target.

Asset / path Threat Primary control
API and UI (argocd-server) Unauthorised sync, secret exposure through diffs or APIs SSO with group-mapped RBAC, deny-by-default policy.default, prompt patching (see CVEs)
Git / Helm / OCI sources Malicious manifest or template injection Branch protection, AppProject.sourceRepos, Source Integrity commit signatures (3.5+)
repo-server Code execution through templating tools and plugins No K8s token, egress NetworkPolicy, separate CMP sidecars, mTLS (3.5+)
application-controller credentials Cluster-admin on every cluster Sync impersonation (beta), AppProject destination/resource allow-lists, agent model for spokes
Cluster/repo Secrets in argocd namespace Credential theft Restrict namespace RBAC, prefer cloud IAM and short-lived tokens

Authentication

  • Built-in admin: its bcrypt hash lives in argocd-secret, and the initial password is in argocd-initial-admin-secret. Disable it with admin.enabled: "false" once SSO works.
  • Local accounts (accounts.<name> in argocd-cm) with apiKey and/or login capabilities, for automation.
  • API tokens are JWTs signed by Argo CD, issued per account (argocd account generate-token) or per project role (argocd proj role create-token).
  • SSO: native OIDC (Okta, Entra ID, Keycloak and others), with the PKCE code flow handled by the server since 3.1 and background token refresh since 3.3. Bundled Dex covers SAML, LDAP and GitHub. Since 3.0, Dex users are identified by federated_claims.user_id instead of the unstable sub.

Configuration recipes are in How-to Guides.

Authorization

RBAC is Casbin-based (argocd-rbac-cm). AppProjects are the tenancy boundary. They restrict source repos, destination clusters and namespaces, allowed cluster-scoped and namespaced kinds, sync windows, and project roles with their own tokens. policy.default applies to every authenticated user without a mapping. The empty default denies everything, which is the recommended setting. Anonymous access exists only if users.anonymous.enabled is set.

In 3.0, update/delete on an Application stopped implying rights on its managed resources (update/*, delete/*), and logs became a first-class, always-enforced permission. Both changes close escalation paths where "can edit an app" silently meant "can edit or delete everything it deploys" or "can read pod logs". The full resource/action matrix is in Reference.

Encryption and Secrets

Asset Protection
Admin password bcrypt hash in argocd-secret
SSO client secrets Stored in argocd-secret (base64, not encrypted); referenced as $key from argocd-cm
Cluster and repository credentials Kubernetes Secrets (base64), so rely on etcd encryption at rest and namespace RBAC
Redis Password auth (Secret argocd-redis); cache contents are rebuilt on restart
Dex Runs with in-memory storage in the default manifests; no persistent database

In transit, argocd-server, argocd-repo-server and argocd-dex-server serve TLS, with self-signed certificates unless you supply your own. Since 3.5, repo-server connections can be upgraded to mutual TLS. Redis traffic is not TLS-encrypted in the default manifests, so rely on NetworkPolicies or configure Redis TLS yourself.

Argo CD deliberately does not manage secret values. The common patterns are:

Approach Mechanism Trade-off
External Secrets Operator (ESO) Argo CD deploys ExternalSecret/SecretStore, and ESO syncs values from Vault, AWS SM, GCP SM or Azure KV Best separation; secret values never pass through Argo CD
Argo CD Vault Plugin (AVP) CMP replaces <path:...#key> placeholders at render time Secrets end up in the repo-server cache, and the upstream project discourages render-time injection
SOPS + KSOPS Encrypted files in Git, decrypted by a Kustomize plugin during render Everything in Git; key management is critical
Sealed Secrets kubeseal-encrypted SealedSecret in Git, decrypted by an in-cluster controller Simple; the encrypted blobs are cluster-key specific

Recommended pattern

Use External Secrets Operator for production. Argo CD manages only the ExternalSecret and SecretStore manifests, and the values stay in the external store.

Network Segmentation

  • Restrict argocd-repo-server egress to the required Git, Helm and OCI hosts. A compromised template should not be able to reach internal services.
  • Allow Redis ingress only from Argo CD pods.
  • Keep the repo-server without a Kubernetes API token, which is the default.
  • Consider a separate Argo CD instance, or dedicated CMP sidecars, for untrusted repositories.

The NetworkPolicy recipe is in How-to Guides.

Known Pitfalls

Pitfall Impact Mitigation
policy.default: role:readonly with many SSO users Every authenticated user can read all app manifests and diffs, which may contain sensitive config Set policy.default: "" and map groups explicitly
Static cluster bearer tokens that never expire A compromised token gives persistent cluster access Use short-lived tokens or cloud IAM, or the agent model
Missing ignoreDifferences for operator-mutated fields Endless sync loops against cert-manager, Gatekeeper and similar Add resource.customizations.ignoreDifferences
Unrestricted repo-server egress A malicious repo can exfiltrate data or pivot Strict egress NetworkPolicies
Running an unpatched minor Known critical CVEs (e.g. CVE-2025-55190, CVE-2026-42880) Stay within the three supported minors

Why Argo CD 3.0 Changed Defaults

The 3.0 release (May 2025) contained few new features. Instead it made safer and faster behaviour the default: annotation tracking, fine-grained RBAC without inheritance, always-on logs RBAC, default exclusion of high-churn kinds such as Endpoints, EndpointSlice and Lease, and health kept out of the Application status. Each change cuts either API-server load and controller churn or an implicit privilege. The 3.x line since then has shipped on the quarterly cadence. PreDelete hooks, OIDC refresh and shallow clones arrived in 3.3. Cluster reconciliation pause and Teams Workflows notifications came in 3.4. 3.5 brought Helm 4, repo-server mTLS, Source Integrity, and beta status for impersonation and the Source Hydrator. The per-minor list is in Reference.

Sources