Blue-green and canary deploys from your pipeline

Ship with zero downtime and instant rollback by shifting traffic between versions instead of replacing in place.

Oct 22, 2024·Updated ·5 min readAdvanced·By SecOpsLog · documentation-verified

A rolling update decides the moment users meet the new version for you: as each pod passes its readiness probe it joins the Service, and by the time an error rate is visible a third of the fleet is already serving it. Blue-green and canary move that moment under your control. Blue-green keeps the old version running, brings the new one up to full size behind a second Service, and switches the selector when you say so; canary sends a slice of traffic first and lets a metric decide. Both are a routing change instead of a rebuild, so rollback takes seconds, and both fail the same way when the database schema is not ready for two versions at once.

Blue-green cutover and rollback

Keep blue idle until green is proven. Rollback is a selector flip.

Blue (live) current production pods labeled env=blue Service / Ingress selector flip = cutover Green (idle) new version, no traffic smoke then switch 1 Cutover point selector at green 100% traffic in one step 2 Watch errors · latency · saturation soak window before teardown 3 Rollback flip selector back to blue instant — no rebuild Keep blue idle until green proves stable. Rollback is routing, not redeploy. blue ←→ Service selector ←→ green
Blue — liveGreen — candidateRollback — flip back

Blue-green with a controller, not a script

rollout.yaml (Argo Rollouts, blueGreen)
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: api
namespace: prod
spec:
replicas: 6
selector: { matchLabels: { app: api } }
template:
metadata: { labels: { app: api } }
spec:
containers:
- name: api
image: registry.acme.dev/shop/api@sha256:9f2a1c…
strategy:
blueGreen:
activeService: api # what users hit; selector rewritten at promotion
previewService: api-preview # the new ReplicaSet, reachable for smoke tests, no production traffic
autoPromotionEnabled: false # hold at the preview stage until promoted
prePromotionAnalysis:
templates: [{ templateName: smoke-and-error-rate }]
args: [{ name: service, value: api-preview }]
scaleDownDelaySeconds: 120 # keep the old ReplicaSet for two minutes after the switch

The kubectl patch service version of blue-green works and is also a runbook step that someone has to remember at 02:00. The Rollout does the same selector rewrite, but sequenced: the new ReplicaSet is created, scaled to full size, pointed at by previewService, held there while prePromotionAnalysis runs against it, and only then swapped into activeService. The old ReplicaSet stays up for scaleDownDelaySeconds, which is the rollback target and also covers a detail the documentation calls out: when a Service selector changes, nodes update their forwarding rules with a delay, and traffic can still reach the old pods for a moment. Scaling them down instantly turns that moment into connection resets.

bash — promote when the preview has earned it, abort when it has not
kubectl argo rollouts get rollout api -n prod
Status: ॥ Paused
Message: BlueGreenPause
Images: registry.acme.dev/shop/api@sha256:3c4d… (stable, active)
registry.acme.dev/shop/api@sha256:9f2a… (preview)
curl -fsS https://api-preview.prod.svc.cluster.local/healthz && kubectl argo rollouts promote api -n prod
rollout 'api' promoted
kubectl get endpointslices -n prod -l kubernetes.io/service-name=api -o jsonpath="{.items[*].endpoints[*].targetRef.name}"
api-7d9c… api-7d9c… api-7d9c… api-7d9c… api-7d9c… api-7d9c…
six endpoints, all from the new ReplicaSet; the old one is scaled down after the delay. Rollback within it: kubectl argo rollouts abort api

Canary when full-size twice is too much

rollout.yaml (canary strategy excerpt)
strategy:
canary:
steps:
- setWeight: 10
- pause: { duration: 5m }
- analysis:
templates: [{ templateName: error-rate }] # Prometheus query: 5xx ratio for the canary pods
- setWeight: 50
- pause: { duration: 10m }
- setWeight: 100
# without a traffic router, weight is approximated by the replica ratio of the two ReplicaSets

Canary trades the second full-size fleet for gradual exposure: ten percent of traffic, a pause, an analysis run that reads the canary's own error rate from Prometheus, then more. A failed analysis aborts the rollout and the stable ReplicaSet is scaled back up, so the worst case is ten percent of requests for five minutes. Without an ingress or mesh integration, Rollouts approximates the weight by replica counts (one canary pod in ten), which is coarse and still far better than all-at-once; with one, the weight is exact. The analysis template should measure what users feel: 5xx ratio and latency, not CPU.

analysistemplate.yaml (the metric that decides)
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
name: error-rate
namespace: prod
spec:
metrics:
- name: http-5xx-ratio
interval: 1m
count: 5
failureLimit: 1 # one bad sample aborts; make this 2 if the metric is noisy
successCondition: result[0] < 0.01
provider:
prometheus:
address: http://prometheus.monitoring:9090
query: |
sum(rate(http_requests_total{app="api",role="canary",code=~"5.."}[2m]))
/ sum(rate(http_requests_total{app="api",role="canary"}[2m]))

The template is where the strategy earns its keep, and it is only as good as the label that separates canary traffic from stable. Argo Rollouts can stamp the canary pods with ephemeral metadata so the query can select them, and a template whose query sums both versions measures nothing about the change. count and failureLimit set how much evidence is needed: five samples a minute apart with one failure allowed is a ten-minute decision, which is the trade against speed.

Choosing a strategy

Rolling (Deployment)Blue-greenCanary
extra capacitya surge of a few podsa second full fleet, brieflya few canary pods
who decides exposurereadiness probesa person or an analysis, before any usera metric, at each step
rollbacka new rollout of the old imageselector flip while the old ReplicaSet existsabort; stable scales back
needsnothinga preview Service and a smoke test worth runninga metric that reflects user impact and a traffic router for exact weights
fitsstateless services with good probesreleases that must be verified in place, or switched back in secondshigh-traffic services where 10% is a real sample
Two versions means two versions of the schema
During a blue-green hold or a canary step, old and new code read and write the same database. A migration that renames a column, changes a type or adds a NOT NULL constraint the old version cannot satisfy breaks one side the instant it runs. Expand first (add the column, backfill, dual-write), ship code that works with both shapes, then contract in a later release. No deployment strategy makes an incompatible migration safe; it only chooses which users see it first.

The switch is only as clean as the pods on both sides of it: the outgoing version has to finish in-flight requests when its endpoints are removed (graceful shutdown done properly), and the incoming one has to be honest about readiness (probe design). The preview Service also gives review apps a production-shaped place to run the smoke test before promotion.

Related posts

Quick reference