Blue-green and canary deploys from your pipeline
Ship with zero downtime and instant rollback by shifting traffic between versions instead of replacing in place.
A rolling update decides the moment users meet the new version for you: as each pod passes its readiness probe it joins the Service, and by the time an error rate is visible a third of the fleet is already serving it. Blue-green and canary move that moment under your control. Blue-green keeps the old version running, brings the new one up to full size behind a second Service, and switches the selector when you say so; canary sends a slice of traffic first and lets a metric decide. Both are a routing change instead of a rebuild, so rollback takes seconds, and both fail the same way when the database schema is not ready for two versions at once.
Keep blue idle until green is proven. Rollback is a selector flip.
Blue-green with a controller, not a script
apiVersion: argoproj.io/v1alpha1kind: Rolloutmetadata:name: apinamespace: prodspec:replicas: 6selector: { matchLabels: { app: api } }template:metadata: { labels: { app: api } }spec:containers:- name: apiimage: registry.acme.dev/shop/api@sha256:9f2a1c…strategy:blueGreen:activeService: api # what users hit; selector rewritten at promotionpreviewService: api-preview # the new ReplicaSet, reachable for smoke tests, no production trafficautoPromotionEnabled: false # hold at the preview stage until promotedprePromotionAnalysis:templates: [{ templateName: smoke-and-error-rate }]args: [{ name: service, value: api-preview }]scaleDownDelaySeconds: 120 # keep the old ReplicaSet for two minutes after the switch
The kubectl patch service version of blue-green works and is also a runbook step that someone has to remember at 02:00. The Rollout does the same selector rewrite, but sequenced: the new ReplicaSet is created, scaled to full size, pointed at by previewService, held there while prePromotionAnalysis runs against it, and only then swapped into activeService. The old ReplicaSet stays up for scaleDownDelaySeconds, which is the rollback target and also covers a detail the documentation calls out: when a Service selector changes, nodes update their forwarding rules with a delay, and traffic can still reach the old pods for a moment. Scaling them down instantly turns that moment into connection resets.
kubectl argo rollouts get rollout api -n prodStatus: ॥ PausedMessage: BlueGreenPauseImages: registry.acme.dev/shop/api@sha256:3c4d… (stable, active) registry.acme.dev/shop/api@sha256:9f2a… (preview)curl -fsS https://api-preview.prod.svc.cluster.local/healthz && kubectl argo rollouts promote api -n prodrollout 'api' promotedkubectl get endpointslices -n prod -l kubernetes.io/service-name=api -o jsonpath="{.items[*].endpoints[*].targetRef.name}"api-7d9c… api-7d9c… api-7d9c… api-7d9c… api-7d9c… api-7d9c…six endpoints, all from the new ReplicaSet; the old one is scaled down after the delay. Rollback within it: kubectl argo rollouts abort apiCanary when full-size twice is too much
strategy:canary:steps:- setWeight: 10- pause: { duration: 5m }- analysis:templates: [{ templateName: error-rate }] # Prometheus query: 5xx ratio for the canary pods- setWeight: 50- pause: { duration: 10m }- setWeight: 100# without a traffic router, weight is approximated by the replica ratio of the two ReplicaSets
Canary trades the second full-size fleet for gradual exposure: ten percent of traffic, a pause, an analysis run that reads the canary's own error rate from Prometheus, then more. A failed analysis aborts the rollout and the stable ReplicaSet is scaled back up, so the worst case is ten percent of requests for five minutes. Without an ingress or mesh integration, Rollouts approximates the weight by replica counts (one canary pod in ten), which is coarse and still far better than all-at-once; with one, the weight is exact. The analysis template should measure what users feel: 5xx ratio and latency, not CPU.
apiVersion: argoproj.io/v1alpha1kind: AnalysisTemplatemetadata:name: error-ratenamespace: prodspec:metrics:- name: http-5xx-ratiointerval: 1mcount: 5failureLimit: 1 # one bad sample aborts; make this 2 if the metric is noisysuccessCondition: result[0] < 0.01provider:prometheus:address: http://prometheus.monitoring:9090query: |sum(rate(http_requests_total{app="api",role="canary",code=~"5.."}[2m]))/ sum(rate(http_requests_total{app="api",role="canary"}[2m]))
The template is where the strategy earns its keep, and it is only as good as the label that separates canary traffic from stable. Argo Rollouts can stamp the canary pods with ephemeral metadata so the query can select them, and a template whose query sums both versions measures nothing about the change. count and failureLimit set how much evidence is needed: five samples a minute apart with one failure allowed is a ten-minute decision, which is the trade against speed.
Choosing a strategy
| Rolling (Deployment) | Blue-green | Canary | |
|---|---|---|---|
| extra capacity | a surge of a few pods | a second full fleet, briefly | a few canary pods |
| who decides exposure | readiness probes | a person or an analysis, before any user | a metric, at each step |
| rollback | a new rollout of the old image | selector flip while the old ReplicaSet exists | abort; stable scales back |
| needs | nothing | a preview Service and a smoke test worth running | a metric that reflects user impact and a traffic router for exact weights |
| fits | stateless services with good probes | releases that must be verified in place, or switched back in seconds | high-traffic services where 10% is a real sample |
The switch is only as clean as the pods on both sides of it: the outgoing version has to finish in-flight requests when its endpoints are removed (graceful shutdown done properly), and the incoming one has to be honest about readiness (probe design). The preview Service also gives review apps a production-shaped place to run the smoke test before promotion.