GitOps interview questions
Practice GitOps interview answers that go from desired-state basics to Argo CD/Flux architecture, progressive delivery, and multi-cluster operations.
Levels run Beginner → Intermediate → Advanced → Expert. Answers are phrased the way you would say them in an interview; Advanced and Expert answers add the deeper reasoning, a diagram where it helps, and the follow-up an interviewer often asks next.
Fundamentals
The short version is: Git holds the desired state for apps and infra, and an agent inside the cluster keeps reality matching that repo. A deploy isn't kubectl from CI — it's a reviewed commit the controller applies.
# commit a manifest change → controller detects → applies → cluster matches Git git commit -am "bump web to v1.2.3" && git pushLink to this question
I'd boil it down to four things: desired state is declarative, it's versioned in Git, an agent pulls it (CI doesn't push with cluster creds), and reconciliation runs continuously so drift gets caught and fixed.
declarative # YAML/Helm/Kustomize in Git versioned # every change is a commit pulled # agent inside the cluster reconciled # level-based loop, not one-shot applyLink to this question
Desired state is whatever Git — or a chart/OCI artifact — says should exist. Observed state is what's actually on the API server right now. Reconciliation is just closing that gap until they match.
argocd app get web # Sync / Health argocd app diff web # desired (Git) vs liveLink to this question
It's the object that tells the controller 'watch this Git/OCI path and keep that destination cluster/namespace matching it.' Argo calls it an Application; Flux usually pairs a source with a Kustomization or HelmRelease.
In Argo CD the source can be a repo path, a Helm chart or an OCI artifact at a pinned targetRevision, and the Application also carries the sync policy: manual, or automated with prune and selfHeal. Argo CD renders the manifests and reports two separate things, Sync status (does the cluster match Git) and Health (are the resources actually working).
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata: { name: web }
spec:
source:
repoURL: https://github.com/org/config
path: apps/web/overlays/prod
targetRevision: main
destination: { namespace: web, server: https://kubernetes.default.svc }
syncPolicy:
automated: { prune: true, selfHeal: true }Link to this questionSync is the reconcile step — apply what's in Git so the live cluster matches. OutOfSync means Git and cluster disagree. Healthy is separate: Synced just means the YAML matches, not that the app is happy.
argocd app get web # Sync: Synced | OutOfSync # Health: Healthy | Degraded | ProgressingLink to this question
An OCI artifact is an immutable, versioned snapshot of the manifests that lives in the same registry as the images, so you can pin it by digest, sign it, and mirror it into air-gapped or edge clusters. Git is still where changes are reviewed; CI publishes the artifact from the merged commit and the controller pulls from the registry.
flux push artifact oci://ghcr.io/org/config/app:$(git rev-parse --short HEAD) --path=./deploy --source=https://github.com/org/config --revision="main@sha1:$(git rev-parse HEAD)" # Flux: OCIRepository (source.toolkit.fluxcd.io/v1, optional spec.verify with cosign) + Kustomization sourceRef # Argo CD: source.repoURL: oci://ghcr.io/org/config/app, targetRevision: tag or digest
Interviewer often follows with: How would you make the controller verify the artifact signature before it applies anything?
Link to this questionCI still does the heavy lifting I'd expect: build, test, push images, and open or update the config commit. What it shouldn't need is prod kubeconfig. The in-cluster agent stays the continuous applier.
docker push registry/web:$SHA # update overlays/dev image tag in config repo # Argo/Flux applies — CI never kubectl apply -f to prodLink to this question
Push means CI holds kubeconfig and runs kubectl or Helm against the cluster. Pull means a controller inside the cluster watches Git and applies changes. I'd rather keep cluster credentials in the cluster, and I get continuous drift correction for free.
With push, every pipeline and every environment ends up holding cluster creds — a leaked CI token is basically cluster access. Pull keeps write access inside each cluster, makes every deploy a Git change, and turns reconcile into ongoing self-heal instead of a one-shot apply. That's why I push images from CI but let Argo or Flux own the apply.
kustomize edit set image app=app:v2 git commit -am "promote app:v2" && git push # agent reconciles — CI never ran kubectl against prodLink to this question
I'd describe it as the controller periodically comparing desired vs observed and converging them. It's level-based, not event-based — so a missed webhook or a hand-deleted resource still gets fixed on the next sync.
argocd app sync web # Flux: flux reconcile kustomization apps --with-sourceLink to this question
Drift is any live mutation that diverges from Git. With self-heal on, the next reconcile reverts it. Without self-heal the app shows OutOfSync, but the live edit survives only until the next sync, including the automated sync that any new Git commit triggers.
argocd app diff web argocd app set web --sync-policy automated --self-heal --auto-pruneLink to this question
Sync answers 'does live match Git?' — Synced or OutOfSync. Health answers 'is it actually working?' — Healthy, Progressing, Degraded. You can be Synced and Degraded: the YAML is right, the app is crash-looping.
argocd app get web # Sync: Synced | Health: Degraded ← investigate pods/events, not GitLink to this question
Prune deletes live resources that aren't in Git anymore. Without it, GitOps only creates and updates — remove a manifest and the orphan sticks around in the cluster.
syncPolicy:
automated: { prune: true, selfHeal: true }
# Flux: spec.prune: true on KustomizationLink to this questionTags like latest — or even v1.2.3 — can move under you. A digest is an immutable content address. Pinning @sha256:… means desired state is exact, reproducible, and plays nicer with signature verification.
images:
- name: app
digest: sha256:abc123…Link to this questionArgo CD & Flux
Flux is more of a toolkit. You've got a GitRepository or OCIRepository as the source, then Kustomization or HelmRelease objects that reconcile paths or charts. dependsOn handles ordering, and image-automation can bump tags by committing back to Git.
kind: GitRepository
metadata: { name: config }
spec: { url: https://github.com/org/config, ref: { branch: main } }
---
kind: Kustomization
metadata: { name: web }
spec:
path: ./apps/web
prune: true
sourceRef: { kind: GitRepository, name: config }Link to this questionI'd say both are solid CNCF-graduated GitOps engines. Argo is app-centric with a strong UI, Projects, and SSO — great when the platform team wants visibility and multi-tenancy in one product. Flux is lean and CRD-native; teams that live in Git and the CLI often prefer it. I pick on UI needs, tenancy model, and what the ops team already knows — not on 'which is more GitOps.'
Argo's Application CR owns source, destination, and sync policy in one object, and the UI is something non-platform engineers will actually open. Flux splits Source → Kustomization/HelmRelease → ImagePolicy, which maps cleanly to Git and doesn't depend on a central UI — but onboarding usually needs more YAML literacy. ApplicationSets and Flux (per-cluster Kustomization paths with postBuild substitution, or Flux Operator ResourceSets) solve the same fan-out problem differently. Neither replaces progressive delivery; you still layer Rollouts or Flagger. Credential model, RBAC, and how you bootstrap the first controller matter more than feature bingo.
# Prefer Argo when: SSO + UI + AppProjects multi-tenancy matter # Prefer Flux when: Git-native CRDs, OCI sources, image automation compose better # Always: config in Git, pull reconcile, prune + self-heal, no kubectl from CI to prod
Interviewer often follows with: If you had to migrate from one to the other, how would you avoid dual-writing forever?
Link to this questionApp-of-apps is a root Application that syncs child Applications — one entry point to bootstrap a cluster. ApplicationSets generate those children from generators (list, git directories, clusters), so onboarding an app or cluster is data instead of hand-written YAML.
App-of-apps is perfect for bootstrap and a small, stable tree. It gets painful once you have dozens of apps times environments times clusters — every new destination is another Application commit. ApplicationSets template Applications from a generator: a Git folder per app, a cluster list, or a matrix. On Flux, the Flux Operator's ResourceSets play that role. The root still exists, but the explosion of children is generated and pruned when the generator input disappears. I'd still keep clear ownership via AppProjects or Flux namespaces, and the generator inputs themselves live in Git.
kind: ApplicationSet
spec:
generators:
- clusters: {}
template:
metadata: { name: 'web-{{name}}' }
spec:
source: { repoURL: https://github.com/org/config, path: apps/web }
destination: { name: '{{name}}', namespace: web }Interviewer often follows with: How do you keep a bad generator change from wiping every child app?
Link to this questionWaves are annotations that order resources so CRDs and namespaces land before workloads. Hooks — PreSync, PostSync, SyncFail — run Jobs for migrations or smoke tests around the sync. Waves order things; hooks add lifecycle steps.
Without waves, Argo can apply a Deployment before its CRD or operator exists and thrash. Negative waves run first — namespaces, CRDs, operators — zero is the default, positive waves later for apps and NetworkPolicies that select workloads. Hooks are separate Jobs annotated as PreSync for migrate, PostSync for smoke, SyncFail for alert. A failed hook can block the sync, so you want backoff, TTL, and idempotent Jobs. In Flux the equivalent is dependsOn between Kustomizations plus healthChecks — same idea, different API.
metadata:
annotations:
argocd.argoproj.io/sync-wave: "-1"
argocd.argoproj.io/hook: PreSync
argocd.argoproj.io/hook-delete-policy: HookSucceeded
# Job runs migrate; only then wave 0 Deployments syncInterviewer often follows with: What breaks if a PreSync migration isn't backward-compatible with the old pods still running?
Link to this questionAn image updater watches the registry and commits the new tag or digest back to the config repo. The reconciler applies that commit. The cluster still only changes because Git changed — CI never ran kubectl.
kind: ImagePolicy
spec:
imageRepositoryRef: { name: app }
policy: { semver: { range: ">=1.0.0" } }
# image-automation controller opens/commits the bumpLink to this questionAppProjects constrain which source repos, destinations, and cluster-scoped resources an Application can use. They're the tenancy boundary so team A can't sync arbitrary YAML into team B's namespaces.
kind: AppProject
metadata: { name: team-a }
spec:
sourceRepos: ['https://github.com/org/config.git']
destinations:
- { namespace: 'team-a-*', server: https://kubernetes.default.svc }Link to this questionI'd use overlays when I'm patching a shared base with small env deltas. Helm values fit parameterized charts and third-party apps. A lot of GitOps setups use Helm for vendors and Kustomize for first-party apps, or wrap Helm with Kustomize.
# overlays/prod/kustomization.yaml
resources: ['../../base']
images: [{ name: app, digest: sha256:… }]Link to this questionPromotion, secrets & structure
I model each environment as a path or overlay — occasionally a branch, but overlays are usually cleaner. Promotion is a PR that bumps the target env's image digest or chart version. Humans review the Git change; the agent applies it.
cd overlays/prod && kustomize edit set image app=app@sha256:abc… git commit -am "promote app@sha256:abc to prod"Link to this question
Most teams I work with keep app code and cluster config in separate repos so a deploy is a config commit and app CI never needs cluster credentials. Inside the config repo, per-env overlays sharing a base keep things DRY. Separate repos buy stronger RBAC boundaries, but you pay in duplication.
The split that matters is code vs desired-state config. App pipelines build and push images — and preferably sign and attest — then a second change updates the config repo. Mono-repo config with overlays for dev/stg/prod is simplest for shared bases and CODEOWNERS. Repo-per-team or repo-per-env helps when blast radius and write access need hard isolation. Branch-per-env usually turns into merge hell and drift. My default pitch: separate config repo, overlays for envs, promote by digest not mutable tags, CODEOWNERS on prod paths.
config/
apps/web/base/
apps/web/overlays/{dev,stg,prod}/
platform/ # controllers, policies
# CODEOWNERS: overlays/prod → @platform-oncallInterviewer often follows with: How do you stop a developer from merging a prod overlay change without platform review?
Link to this questionI never commit plaintext. Sealed Secrets or SOPS keep ciphertext in Git; External Secrets Operator keeps only a reference in Git and pulls the value from Vault or cloud SM at runtime. If we already have a central secrets platform, I'd lean ESO.
Sealed Secrets encrypts to a cluster-scoped public key and only the controller decrypts — fine for small setups, painful once you're juggling keys across clusters. SOPS encrypts files with age or KMS before commit; it works across tools but key access has to be gated. With ESO, an ExternalSecret CR references Vault or a cloud secrets manager, the Secret gets materialized in-cluster and refreshed, Git stays free of secret material, and rotation lives upstream — though etcd then holds the synced Secret, so encrypt etcd or use ephemeral projection where you can. Pattern I like: GitOps owns the ExternalSecret YAML; Vault owns the value and the audit trail.
# Sealed Secrets
kubeseal < secret.yaml > sealed-secret.yaml && git add sealed-secret.yaml
# ESO — Git holds a reference only
kind: ExternalSecret
spec:
secretStoreRef: { name: vault, kind: ClusterSecretStore }
target: { name: db }
data: [{ secretKey: password, remoteRef: { key: apps/web/db } }]Interviewer often follows with: How do you rotate a secret without a downtime-causing reconcile race?
Link to this questionI treat the config repo like production: branch protection, required reviews, CODEOWNERS on prod overlays, signed commits if we need them, plus Argo AppProjects or Flux RBAC so a team's Application can only sync to their namespaces and allowed sources.
There are two planes — Git RBAC and cluster RBAC. On Git: protect main, require reviews for overlays/prod, restrict who can push to the repo the controller trusts, and prefer deploy keys or GitHub App tokens scoped read-only for the controller. On the cluster: the controller ServiceAccount should be the only principal that can mutate GitOps-managed namespaces; developers get get/list/watch and maybe exec, not edit. Argo AppProjects constrain source repos, destinations, and cluster resources. Without both planes, either anyone can PR-bomb prod or anyone can kubectl-edit around Git.
kind: AppProject
metadata: { name: team-a }
spec:
sourceRepos: ['https://github.com/org/config.git']
destinations:
- { namespace: 'team-a-*', server: https://kubernetes.default.svc }
clusterResourceWhitelist: [] # no cluster-scoped by defaultInterviewer often follows with: How do you allow an emergency hotfix without turning off all the controls permanently?
Link to this questionInstall the GitOps controller and credentials to read the config repo, apply a root app-of-apps or Flux Kustomization pointing at the platform path, then let it reconcile everything else — CRDs, policies, apps. Stateful data still needs its own restore plan.
argocd app create root --repo https://github.com/org/config \ --path clusters/prod --dest-server https://kubernetes.default.svc \ --dest-namespace argocd --sync-policy automatedLink to this question
PR pipelines render and validate manifests — kustomize build, kubeconform, Conftest or Checkov — optionally deploy to an ephemeral or staging cluster with the same overlay pattern, then promote by merging to the prod path. I never point prod at an unreviewed branch.
kustomize build overlays/prod | kubeconform -strict - kustomize build overlays/prod | conftest test - # Argo CD: ApplicationSet Pull Request generator for per-PR preview AppsLink to this question
Least-privilege read access for the controller — deploy key or GitHub App limited to the config repo — tokens in a sealed or ESO-managed secret, rotate on a schedule, and never reuse a human PAT with org-wide scope.
The controller identity is high value: read access to desired state can reveal infra topology; write access for image automation can push malicious commits. I prefer app installations over PATs, contents:read by default, and contents:write only for image-automation bots. If the bot can write, branch protection should still block skipping reviews on prod paths. Audit git access logs. Private Helm or OCI gets separate pull secrets. On compromise: revoke the app or key, rotate, and review recent commits the bot authored.
# Argo repository Secret: type git, url + sshPrivateKey (deploy key read-only) # Image automation: separate bot identity with write to overlays/dev only
Interviewer often follows with: Should image-automation commit directly to main or open PRs?
Link to this questionI'd restrict automation to non-prod paths or PRs, require human or semver gates for prod, and separate the bot identity so it can't skip reviews on production overlays.
Continuous latest or :sha tags into prod create change fatigue and accidental deploys of broken builds. Pattern that works: automate overlays/dev, open PRs for staging, promote to prod via merge with CODEOWNERS and tests. Pin digests in prod. Rate-limit the bot and ignore build-metadata-only noise. If you're doing write-back, scope the GitHub App to contents:write only on allowed paths. Treat image automation as a promotion policy problem, not a convenience toggle.
# ImagePolicy → writes overlays/dev only # prod overlay: digest pinned; promote via PR from staging
Interviewer often follows with: How do you keep the bot from merging its own prod PRs?
Link to this questionProgressive delivery & multi-cluster
I'd keep GitOps owning the desired manifests, but hand traffic shifting to Argo Rollouts or Flagger. They canary or blue-green against Prometheus metrics, auto-promote or roll back, while Git still records the intended version. GitOps deploys the revision; the rollout controller owns the cutover.
A plain Deployment rolling update has no analysis gate: it replaces pods as fast as maxSurge and maxUnavailable allow, checking only readiness. Progressive delivery adds analysis templates — error rate, p99 — at each weight step. Blue-green keeps two full stacks and flips a Service or Ingress; canary shifts a percentage via mesh or Ingress weights, or, without a traffic router, approximates the weight with canary and stable replica counts. On failure the controller reverts traffic without necessarily reverting Git — then you fix Git so desired state stays honest. When I'm explaining blue-green, I sketch the cutover out loud.
# Rollouts AnalysisTemplate queries Prometheus error ratio # Flagger: progressDeadlineSeconds + metric thresholds # Git still pins image digest; controller only shifts weight
Interviewer often follows with: If analysis passes but business KPIs tank, how do you fold those signals in?
Link to this questionI'd use a management cluster — or per-cluster agents — with ApplicationSet cluster generators or Flux multitenancy so one template fans out per registered cluster, with per-cluster overlays or Helm values. Label clusters by env, region, team, and constrain Projects so teams only sync their scope.
Common patterns: a hub Argo registering spoke clusters via cluster secrets — one control plane and a centralized UI, but a big blast radius if the hub is compromised; instance-per-cluster Flux or Argo — better isolation, harder fleet visibility; ApplicationSet with a cluster generator or git files generator for region overlays. Per-cluster differences belong in values files or cluster-labeled overlays, not forked repos. For secrets I'd rather use ESO with per-cluster SecretStores than replicate Sealed Secrets keys. Progressive delivery and policy like Gatekeeper or Kyverno should ship as platform layer on every cluster, not bolted on per app.
generators:
- clusters:
selector:
matchLabels: { env: prod }
template:
spec:
source:
path: 'apps/web/overlays/{{metadata.labels.region}}'
destination: { name: '{{name}}', namespace: web }Interviewer often follows with: How would you do a controlled region-by-region rollout instead of syncing every prod cluster at once?
Link to this questionIf users are hurting, I'd manually abort or promote based on alternate signals, restore metrics permissions or ServiceMonitors, and never leave a half-shifted canary overnight without an owner.
Flagger and Rollouts depend on Prometheus or webhooks. Failures are usually wrong metrics namespace, missing RBAC, histogram vs counter mismatch, or analysis windows that are too short. Break-glass: abort to stable, or promote only with explicit incident-commander approval. Lasting fixes: synthetic checks as backup analysis, alert on AnalysisRun Error, and game-day the failure mode. When I'm aborting I think about traffic cutover the same way I do blue-green. Owning the analysis stack is part of progressive delivery — not an afterthought.
kubectl argo rollouts abort web kubectl argo rollouts get rollouts web # fix ServiceMonitor / RBAC; re-run canary in business hours
Interviewer often follows with: When is automatic promote-on-analysis-timeout dangerous?
Link to this questionI'd pause the ApplicationSet or auto-sync immediately, fix the generator template with a dry-run preview, and add CI that renders generators and asserts destination/path invariants before merge.
ApplicationSets multiply mistakes. Controls that help: an applicationsSync policy of create-only or create-update (controller-wide, or per ApplicationSet once policy override is enabled) so a generator change can't rewrite or delete children, deny sync windows, progressive syncs (beta and opt-in), and PR-rendered previews of generated Applications. Assert that prod destinations only mount overlays/prod, that cluster labels gate generators, and that ApplicationSet changes need platform review. Prefer allow-lists over globbing every cluster. After the incident, audit what synced during the window and revert commits. Pause, generator tests and destination invariants are the three pieces.
kubectl -n argocd scale deploy argocd-applicationset-controller --replicas=0 argocd proj windows add prod --kind deny --schedule "* * * * *" --duration 2h --applications "*" # revert the generator commit, review rendered Apps, then remove the window and scale back # CI: render ApplicationSet and grep -L 'overlays/prod' for prod clusters
Interviewer often follows with: How do you test ApplicationSet changes against a single canary cluster first?
Link to this questionI'd monitor revision skew between clusters, alert when a cluster lags the target SHA, and use mirrored repos or multi-destination sync with an SLO on reconcile latency. Synced to an old commit isn't safe.
Each cluster's agent tracks a Git revision independently. Network partitions, rate limits, or broken repo credentials cause silent lag. I'd dashboard target vs live revision, time-since-reconcile, and sync error rate per cluster. Don't promote traffic until the DR cluster reports the required SHA and healthy apps. For airgapped DR, OCI artifacts with mirrored registries help. 'Synced' is relative to whatever commit the agent sees — make the commit ID a first-class SLI.
# alert if cluster-b revision != cluster-a revision for > 10m # argocd app get web -o json | jq .status.sync.revision
Interviewer often follows with: How do you avoid split-brain if both regions can accept writes to the config repo?
Link to this questionFailure modes & debugging
I'd start with app status and conditions, then diff live vs desired. After that I'm hunting a failing hook, an immutable field fight, a missing CRD, or an RBAC-denied apply. Fix the blocker first — prune or replace only once I understand the diff.
What I see most: PreSync Job crashlooping, someone trying to change an immutable selector, CRD not in an earlier wave, Application pointed at the wrong revision, compareOptions ignoring (or not ignoring) server-side fields, admission webhooks rejecting applies, or two Applications fighting over the same resources. In Flux I'd check Ready on the Kustomization, flux logs, and dependency blockers. I never reach for --force first — I need to know whether Git or live is wrong. If someone kubectl-edited, either commit the intentional change or turn self-heal back on after I've confirmed.
argocd app get web argocd app diff web argocd app history web kubectl -n argocd logs -l app.kubernetes.io/name=argocd-application-controller --tail=100
Interviewer often follows with: When is replace or force justified versus just fixing the manifest?
Link to this questionWorkloads come back by installing the GitOps agent on a fresh cluster and pointing it at the config repo. I still need a plan for persistent data — Velero, snapshots — and for bootstrap credentials: repo access, decryption keys, secret stores. Git restores desired state, not databases.
My DR runbook looks like: recreate the cluster and control plane; restore or recreate the controller's repo credentials and any SOPS/age/KMS access; apply the root app; wait for the platform wave — CRDs, controllers, policies; restore PVCs and DBs from backups before or as apps come up depending on RPO; then verify Sync, Health, and critical SLOs. Multi-cluster GitOps helps if you can fail traffic to a warm region that's already reconciled. Practice it — untested 'Git is our backup' usually fails on secret-zero and stateful restores. The interview signal is separating declarative cluster state from data-plane backups.
1. new cluster + CNI + storage class 2. install Argo/Flux + repo + KMS access 3. sync platform/ (CRDs, OPA, cert-manager) 4. velero restore / DB snapshot 5. sync apps/ and confirm health
Interviewer often follows with: What's your RPO/RTO if the config repo itself is unavailable?
Link to this questionWhen the work is inherently imperative and short-lived — one-off node surgery, interactive debugging — when state can't usefully be declared, or when you need sub-second human control without a commit. I use GitOps for desired cluster state and runbooks or jobs for emergencies, then encode the lasting fix in Git.
GitOps struggles with careful database cutovers, firmware/BIOS, secrets you won't put even as ciphertext in Git without mature ESO, ultra-high-churn job systems where every commit is noise, and environments where Git latency or review exceeds incident needs. Mature teams allow break-glass imperative access with automatic drift alerts and a hard requirement to commit the end state. Saying 'GitOps everywhere' without naming those exceptions is a red flag in a senior interview.
# incident: scale manually kubectl -n web scale deploy/api --replicas=20 # within N minutes: commit replicas (or HPA) so Git matches live
Interviewer often follows with: How do you reconcile GitOps with Helm charts that create random-named resources each release?
Link to this questionI'd make ownership exclusive: one Application or Kustomization per resource set, non-overlapping paths, and prune only on the true owner. Shared platform resources belong in a platform app — product apps consume them, they don't re-declare them.
Dual ownership shows up as perpetual OutOfSync, thrashing applies, and surprise deletes when one side prunes. Fix it by splitting paths cleanly, using Argo's resource tracking annotations, preferring server-side apply with clear field managers, and forbidding wildcard apps that sync overlapping directories. ApplicationSets should generate disjoint destinations. In Flux, dependsOn expresses order without two reconcilers writing the same object. I'd add a CI check that fails if the same GVK/name appears in two synced paths.
# render all apps and assert unique namespace/name/gvk tuples kustomize build apps/web | kubeconform - # Argo: check resource annotations for application ownership
Interviewer often follows with: How should CRDs be owned when both platform and app charts vendor them?
Link to this questionUsually server-side defaults, webhook mutations like sidecars, normalized fields, or ignoreDifferences gaps. I'd diff carefully, ignore known server-populated fields, and make sure CI renders with the same tools and versions Argo uses.
Classic noise: cluster-added caBundle, defaulted clusterIP, Istio or Linkerd injecting containers, HPA fighting replicas, kubectl last-applied annotations. Argo's ignoreDifferences and RespectIgnoreDifferences matter; so do Flux SSA strategies. Align Helm and Kustomize versions between CI and the controller; from Argo CD 3.5 every chart renders with Helm 4 and spec.source.helm.version: v3 is ignored, so a CI job still on Helm 3 can render differently. Either write desired state that expects the mutation, or use policy that forbids surprise mutation in GitOps namespaces.
spec:
ignoreDifferences:
- group: apps
kind: Deployment
jsonPointers: ['/spec/replicas'] # when HPA owns replicasInterviewer often follows with: When is ignoring replicas the wrong fix?
Link to this questionI'd restore the SecretStore from Git first, then put CRDs and SecretStores in a platform app with pruning disabled for those kinds, keep app Applications consuming ExternalSecrets only, and order sync waves so stores exist before app secrets.
Prune is dangerous across shared platform resources. Recover first: re-sync the store from Git history, or restore from etcd or Velero backups. Pattern: platform wave 0 for CRDs, ESO, SecretStore; app wave 1 for ExternalSecret and Deployments, and never let an app prune cluster-scoped shared types. Mark those objects with the Argo CD sync option Prune=false so a bad sync can't remove them. CRDs are the worst case: deleting a CRD deletes every custom object of that kind, and a recreated CRD starts empty, so restoring the CRD alone doesn't bring back the SecretStores or ExternalSecrets that lived under it. Finalizers and deletion policies on ExternalSecret matter — Retain vs Delete. SecretStores aren't app-owned; document that. Add sync-wave annotations and health checks. Deletion safety for secret infrastructure is a senior ops topic, and interviewers know it.
metadata:
annotations:
argocd.argoproj.io/sync-wave: "0" # SecretStore
argocd.argoproj.io/sync-options: Prune=false
# app ExternalSecret: sync-wave "1"Interviewer often follows with: Should ExternalSecret-managed Secrets be pruned by the app Application?
Link to this questionReal-world scenarios
I treat Git as the source of truth. Either commit the change, or use a controlled bypass (a sync window or a pause) with a ticket. Don't leave live-only edits under auto-sync, and make repeats hard: humans get read-only RBAC on managed namespaces, and break-glass is a short-lived, audited RoleBinding.
Out-of-band changes are drift. With self-heal on, the next reconcile reverts them; without it the app sits OutOfSync until the next sync overwrites the edit, so a hotfix that never reached Git is lost either way. Process I'd use: hotfix branch → PR → sync, or a temporary ignore with an expiry. For true emergencies, disable auto-sync, apply, then reverse-commit live state into Git before re-enabling. If an ApplicationSet generates the app, a direct syncPolicy change is overwritten unless the ApplicationSet ignores /spec/syncPolicy through ignoreApplicationDifferences. Prevention: read-only RoleBindings for developers, ValidatingAdmissionPolicy or Kyverno to deny mutations without a GitOps exception label, and alerts when OutOfSync lasts more than N minutes. GitOps discipline under pressure is the interview signal.
argocd app set web --sync-policy none kubectl edit deploy/web # capture live → git commit → argocd app sync web argocd app set web --sync-policy automated
Interviewer often follows with: When is ignoreDifferences the wrong way to keep a kubectl edit?
Link to this questionI'd halt sync or auto-sync for that app immediately, restore prod from the last good Git revision or backup, and enforce destination allow-lists plus CI checks that prod paths only target prod clusters.
Wrong destination is a top GitOps outage class. Controls: AppProject destination restrictions, separate Argo instances per env, CODEOWNERS on Application manifests, and CI that asserts destination name/server plus path invariants. Prefer cluster name labels over raw API URLs people copy wrong. Post-incident, audit what was pruned or applied. Project RBAC and destination allow-lists beat hoping engineers pick the right context.
# AppProject
spec:
destinations:
- name: prod-east
namespace: 'prod-*'
# CI: fail if path overlays/dev && destination prodInterviewer often follows with: Why is a separate Argo CD for prod stronger than one instance with many clusters?
Link to this questionI'd inspect wave dependencies and Job health, delete or fix the stuck Job or hook, adjust waves so optional work can't block forever, and add timeouts on PreSync Jobs.
Waves order resources; hooks can block sync. Deadlocks usually come from a Job without a backoff limit, circular waits, or CRDs that aren't in earlier waves. Break-glass: terminate the hook, sync carefully, or disable auto-sync while fixing Git. Lasting fixes: review sync-wave annotations in CI, hook timeouts, and health customizations. Waves are a concurrency protocol — design them with failure modes in mind.
argocd app get web --show-operation kubectl -n prod get jobs,applications # fix Job or remove blocking hook; sync-wave: "-1" for CRDs
Interviewer often follows with: Should database migrations be a PreSync hook or a separate pipeline job?
Link to this questionI'd bootstrap minimal Project and RBAC with a one-time break-glass apply or a wave-0 platform app outside the loop, then let app-of-apps manage the rest. Never depend on a child to create its own parent prerequisites.
Classic chicken-egg: root Application needs a Project, but the Project only lives under a child path in Git. Pattern: a minimal bootstrap manifest applied once — or via Terraform — or put Project plus root in the same wave-0 path synced by a privileged installer. Document recovery if Argo itself gets deleted. Separate 'control plane install' from 'fleet desired state.'
# 1) kubectl apply -f bootstrap/appproject-root.yaml # 2) apply root Application (points at apps/) # 3) children sync; do not require child to create root Project
Interviewer often follows with: How do you recover if someone deletes the root Application in the cluster?
Link to this questionAnalysis used the wrong SLIs or too-narrow a traffic sample. I'd abort or roll back on real user signal, then fix the queries to include business KPIs and multi-region scrapes before the next promote.
Lying analysis looks like success on infra metrics while checkout fails, scraping only canary pods that skip a bad code path, or windows that are too short. I want dual signals and synthetic transactions. Auto promote-on-green without human review for high-risk changes is dangerous. Abort to stable like you would in blue-green. Ownership of analysis templates belongs on the delivery platform.
kubectl argo rollouts abort web # analysis: add checkout_success + multi-zone Prometheus # require iterations covering peak traffic
Interviewer often follows with: How do you prevent analysis from succeeding when Prometheus returns empty series?
Link to this questionI'd freeze syncs, restore a known-good commit from reflog or a backup remote, point Argo at a verified revision, and block force-push on protected branches going forward.
GitOps assumes immutable history on release branches. Force-push breaks SHAs controllers track and can resurrect vulnerable manifests or drop commits. Recovery: protect main, use merge queues, recover the commit from another remote or CI artifact mirror, hard-refresh apps to a good SHA. Audit what live clusters applied during the incident. Treat a history rewrite as an integrity incident, not a git inconvenience.
git reflog show origin/main argocd app set web --sync-policy none # syncing another revision is refused while auto-sync is on argocd app sync web --revision <known_good_sha> # branch protection: deny force push on main
Interviewer often follows with: Would tagging release SHAs in an OCI registry have helped here? How?
Link to this questionAnyone can ship — or bypass Git gates — to production. I'd split projects by env, map SSO groups to least privilege, deny override in prod, and audit policy.csv changes like IAM.
Wide Argo RBAC is cluster-admin adjacent. AppProjects scope repos, destinations, and cluster resources; RBAC scopes actions like get, sync, override, delete. Since Argo CD 3.0, update and delete on an Application no longer cover its managed resources; those need explicit update/* or delete/* grants. Syncing to an arbitrary revision only counts as override when application.sync.requireOverridePrivilegeForRevisionSync is set to true in argocd-cm. For prod I want sync via automated Git only, or a break-glass group with MFA. Disable exec and override for normal developers. The GitOps UI privilege is part of the security boundary.
p, role:dev, applications, sync, dev-*/*, allow p, role:dev, applications, sync, prod-*/*, deny g, alice@example.com, role:dev
Interviewer often follows with: What does override allow that plain sync does not?
Link to this questionI'd dual-encrypt with old and new keys during the transition, re-encrypt all secrets in Git, update the controllers' keys, then retire the old key. Never rotate keys without a re-encrypt pass.
SOPS/age/KMS rotation needs a transition window. Controllers must hold keys that can decrypt HEAD. Procedure: add the new recipient, sops updatekeys across the repo, sync, remove the old recipient. Keep break-glass decrypt offline. Pair with External Secrets where you can so there's less ciphertext in Git. Secret rotation is a coordinated Git plus controller change.
# .sops.yaml: add new age recipient sops updatekeys -y apps/**/secret*.yaml # update flux decryption keys; sync; remove old recipient
Interviewer often follows with: Why might External Secrets reduce how often you re-encrypt Git?
Link to this questionI'd pin both chart version and values commit in one change — or vendor the chart — add CI that renders that exact pair, and avoid floating 'latest' chart deps in prod.
Multi-source can drift if auto-sync picks a new chart independently of values. Pin versions in Git, use a single lock commit, or package an OCI artifact that binds chart and values. CI should helm template the pair on every PR. Atomicity of desired state across sources is a design choice you have to own.
# same PR: Chart.yaml version bump + values digest helm template web chart-1.2.3 -f values-prod.yaml | kubeconform
Interviewer often follows with: How do ApplicationSets make this pinning problem worse if you're not careful?
Link to this questionI'd read the constraint message, fix the manifest in Git or temporarily widen the constraint with change control, then re-sync. I wouldn't kubectl-delete half-applied objects without understanding dependents.
Admission failures leave partial syncs. Prefer dry-run or server-side apply in CI against the same constraints. Sync waves so namespaces and constraints exist first. Break-glass exceptions need tickets. Policy engines are part of the GitOps path, not an afterthought.
kubectl get k8srequiredlabels.constraints.gatekeeper.sh argocd app sync web --dry-run # commit label fix; sync again
Interviewer often follows with: How do you test Gatekeeper constraints in CI before merge?
Link to this questionI read the upgrade notes for every minor on the path and check four things: RBAC, because update and delete on an Application no longer cover its managed resources and logs access is enforced; resource tracking, which now defaults to annotations; repositories still defined in argocd-cm, which 3.0 no longer supports; and from 3.5, Helm 4 rendering.
Policies that granted update or delete on applications used to apply to child resources; from 3.0 they need update/* or delete/* entries, or server.rbac.disableApplicationFineGrainedRBACInheritance set to false to keep the v2 behavior while you migrate. Without an explicit logs, get grant (or a default role that covers it) users lose the logs tab. Annotation-based tracking is the default; apps using label tracking with ApplyOutOfSyncOnly=true need an explicit sync right after the upgrade. Move repositories, repository.credentials and helm.repositories out of argocd-cm into repository Secrets first. From 3.5, charts render with Helm 4 only and plain-HTTP OCI registries need --insecure-oci-force-http. I rehearse on a staging instance and compare Application sync and health before and after.
kubectl -n argocd get cm argocd-cm -o yaml | grep -E 'resourceTrackingMethod|repositories|repository.credentials' kubectl -n argocd get cm argocd-rbac-cm -o yaml | grep -E 'applications, *(update|delete)' # after upgrade: sync ApplyOutOfSyncOnly apps once; add logs, get grants
Interviewer often follows with: Why would you keep the v2 RBAC inheritance flag only as a temporary bridge?
Link to this questionRelated
Primary references
Found a technical issue on this page? Report it with the tool version you used and the behavior you saw. How resources are maintained.