Skip to content
Infrastructure · 5 min read

Blue-green deployment is a load balancer and a spare

The claim Zero-downtime deployment and instant rollback are frequently treated as advanced capabilities that require a container orchestrator and a platform team. For a single appl...

A Written by Administrator
Blue-green deployment is a load balancer and a spare

The claim

Zero-downtime deployment and instant rollback are frequently treated as advanced capabilities that require a container orchestrator and a platform team. For a single application on a couple of servers, they require neither — they require a load balancer, a second copy of your application, and the discipline to switch traffic between them. Blue-green deployment is one of the highest-value reliability practices a small team can adopt, and the version that fits a small team is genuinely simple.

The idea in one paragraph

You run two identical production environments, called blue and green. At any moment, one of them is live and serving all traffic while the other is idle. To deploy, you release the new version to the idle environment, test it there while it receives no real traffic, and then switch the load balancer to send traffic to it. The formerly-live environment is now idle and holds the previous version, untouched. If the new version misbehaves, you switch back — and because the old environment was never modified, the rollback is instant and total, not a redeploy but a redirect.

Why this beats deploying in place

The common alternative — updating the application on the live server — has two problems blue-green removes. First, there is a window during the in-place update when the application is stopping, restarting, or running half-updated, and during that window requests fail. Blue-green has no such window, because the switch happens between two fully-ready environments. Second, rolling back an in-place deploy means deploying again, in reverse, under pressure, hoping the previous version still builds and starts cleanly — whereas blue-green rollback is flipping a switch to an environment that is already running the old version and known to work.

The whole mechanism

With Nginx as the load balancer, the switch is a single line and a reload:

# /etc/nginx/upstream.conf  -- the only thing that changes on a switch
upstream app {
    server 10.0.1.10:8080;   # blue   (live)
    # server 10.0.1.11:8080; # green  (idle, next release goes here)
}

The deploy sequence becomes a script anyone can run and reason about:

1. Deploy new version to the idle environment (green)
2. Run health checks and a smoke test against green directly
3. Switch: point upstream at green, `nginx -s reload`
4. Watch error rates for a few minutes
5. If bad: switch back to blue, `nginx -s reload`  -- instant
6. If good: blue is now the idle spare for next time

nginx -s reload applies the new upstream without dropping a single in-flight connection — existing requests finish on the old worker while new ones go to the new target — which is what makes the switch itself zero-downtime rather than merely fast.

The step that makes it safe: test before the switch

The reason blue-green is more than just two servers is the testing window it creates. Because the new version is live on the green environment but receiving no public traffic, you can exercise it fully before a single customer touches it — run the health checks, hit the critical paths, confirm the new version actually starts and serves correctly against the real production database and configuration. This catches the class of failure that only appears in the real environment: the missing config value, the migration that did not run, the dependency that behaves differently in production. You find it while the environment is idle, not after you have sent customers to it.

The complication nobody mentions: the database

Two application environments are simple; the database they share is where the care is required, because both blue and green talk to the same database and you cannot keep two copies of the live data in sync. The rule that makes this work is the one from safe schema migrations: every database change must be compatible with both the old and the new application version at once. The old code keeps running against the new schema during the window when both exist, which means schema changes are additive — add the column, deploy code that writes both, switch, then later remove the old column in a separate cycle. A migration that the currently-live version cannot tolerate breaks the whole model, because that version is still serving traffic right up to the switch and must keep working if you roll back.

What it costs

The honest cost is a second environment, which for most of the deployment cycle sits idle. For a small application this is modest — a second small instance, or the same instance running a second copy of the application on a different port if you cannot justify separate hardware. Weigh that against what it buys: deployments with no downtime window, rollback that takes seconds instead of a fraught redeploy, and a place to fully test each release against real production before it serves anyone. For any business where a failed deploy during business hours is expensive, the idle spare pays for itself the first time a bad release is caught in the green environment, or reversed in the ten seconds it takes to switch back — instead of the frantic hour it would otherwise have been.

Where to start

Put a load balancer in front of your application if there is not one already — this single change is the prerequisite and is worth doing on its own. Then stand up the second environment and practise the switch and the rollback on a quiet afternoon, before you depend on them, so that the sequence is familiar rather than improvised. The capability that lets you deploy on a Friday without dread is not a platform. It is a spare and a switch, and it is within reach of a team of two.

#deployment #nginx #reliability #operations

Keep reading