A green deployment pipeline is satisfying.
The build passed. The artifact was produced. Infrastructure accepted it. Health checks are green. The new version is running.
None of that proves the system is ready for production.
Deployment proves that software reached an environment.
Production readiness asks whether the system can operate, be observed, fail safely, recover, be supported and remain owned after release.
That is a much larger standard.
Deployment, Release and Operation Are Different States
Teams often compress three events into one:
Deploy
↓
Release
↓
Operate
A deployment moves code into an environment.
A release exposes capability to users.
Operation begins when the organization becomes responsible for the consequences.
Feature flags make the distinction obvious. Code can be deployed for days before release. A service can be released successfully and still fail operationally hours later when traffic changes, a dependency slows, a secret expires or a background job accumulates memory.
Production readiness therefore needs to reason beyond the deployment event.
Readiness Starts With Ownership
A system that nobody clearly owns is not production ready.
Ownership answers questions such as:
- Who responds when it fails?
- Who understands the critical dependencies?
- Who can roll back?
- Who owns the runbook?
- Who receives alerts?
- Who decides whether degraded behavior is acceptable?
- Who communicates during incidents?
Ownership is not the same as repository access.
A team may have written the code but lack permission to restart production resources. Another team may own infrastructure without understanding application behavior. A vendor may provide a dependency whose failure semantics nobody has mapped.
Readiness makes those boundaries explicit before an incident.
Observability Must Answer Operational Questions
Logs are not observability merely because they exist.
A production-ready system should make important behavior visible.
Can the team tell:
- whether users are succeeding;
- which dependency is failing;
- whether latency is increasing;
- whether memory pressure is accumulating;
- whether a queue is backing up;
- whether a background job stopped;
- whether error rate is localized to one tenant or global;
- whether a new release changed behavior?
Google's SRE launch guidance includes monitoring, capacity, failover, external dependencies, rollout planning, security, backup and recovery because reliability depends on the system around the code.
Observability should be designed from failure questions, not from whatever the framework logs by default.
Health Checks Need Semantics
A process returning HTTP 200 may still be useless.
A meaningful health model distinguishes levels such as:
- process is alive;
- process can serve requests;
- critical dependencies are reachable;
- application can perform an important end-to-end operation.
These signals serve different purposes.
A liveness probe should not restart a service because a remote dependency is temporarily slow. A readiness probe may need to remove an instance from traffic if it cannot serve correctly. A synthetic check may need to test a customer-critical flow through multiple components.
Production readiness means understanding what each signal says and what automation does with it.
Failure Behavior Is Part of the Product
Systems do not only need a success path.
They need designed behavior under failure.
Questions include:
- What happens when the email provider returns errors?
- What happens when the identity provider is unavailable?
- What happens when a database migration partially completes?
- What happens when a queue message is processed twice?
- What happens when a vendor times out after accepting a request?
- What happens when an API token expires overnight?
- What happens when memory usage grows for twelve hours?
These are architecture and product questions because failure changes user outcomes.
Timeouts, retries, idempotency, circuit breakers, dead-letter queues and fallback behavior are mechanisms.
The product decision is what the user and business should experience when dependencies are not healthy.
Capacity Planning Is Not Only for Large Systems
Every production system has capacity limits.
The question is whether the team knows them before users discover them.
Capacity may be constrained by:
- CPU;
- memory;
- database connections;
- vendor rate limits;
- queue throughput;
- storage;
- network;
- license limits;
- concurrency;
- downstream processing windows.
A service can operate normally at average traffic and fail during a predictable batch process because memory consumption is cumulative.
Production readiness should connect expected workload to known limits and monitoring.
Google's historical launch checklist explicitly includes traffic estimates, load tests, capacity, growth and dependency impact. The details are organization-specific, but the principle remains relevant.
Configuration and Secrets Are Production Dependencies
Many incidents do not come from application code.
They come from:
- expired certificates;
- missing environment variables;
- incorrect connection strings;
- stale secrets;
- incompatible feature flags;
- environment drift;
- configuration changed without version history.
A production-ready system needs controlled configuration.
That includes knowing which settings are required, which can change at runtime, how secrets rotate, what happens when rotation fails and how configuration changes are audited.
A deployable application with fragile configuration is not operationally ready.
Data Changes Need Recovery Paths
Application rollback is easy to discuss when code is stateless.
Database changes make rollback harder.
A production-readiness review should ask:
- Is the migration backward compatible?
- Can old and new versions run during rollout?
- What happens if migration stops halfway?
- Can it be resumed?
- Is destructive change delayed until confidence exists?
- How is data validated?
- Is backup restoration tested?
A binary may roll back in seconds while the schema it changed cannot.
Release planning needs to treat data as part of the rollout.
Backup Is Not Recovery
A backup exists.
That does not prove recovery works.
Production readiness requires confidence that the organization can restore what matters within acceptable time and data-loss objectives.
A practical review asks:
- What is backed up?
- How often?
- Where?
- Who can restore it?
- How long does restoration take?
- What data is lost between backups?
- Has restoration been tested?
The same principle applies to rollback, disaster recovery and incident procedures.
Capabilities that exist only on paper may not exist when needed.
Dependencies Need Explicit Failure Semantics
Modern applications rely on external services for identity, communication, payments, analytics, AI, storage and more.
Every dependency adds an operational contract.
A readiness review should make that contract visible:
| Question | Why it matters |
|---|---|
| What timeout applies? | Prevent blocked resources |
| What is retried? | Avoid duplicate side effects |
| What rate limit exists? | Prevent cascading failure |
| What happens when unavailable? | Define user experience |
| Is there fallback? | Determine degradation strategy |
| How is dependency health observed? | Diagnose incidents |
| Who owns vendor escalation? | Reduce response delay |
External dependencies are not outside the system from the user's perspective.
If they are required for the capability, their failure belongs in the product's operational design.
Runbooks Should Encode Decisions, Not Ceremony
A useful runbook helps someone act under pressure.
It should answer:
- What does this alert mean?
- What is the likely impact?
- Which dashboards help?
- What common causes exist?
- What actions are safe?
- When should we roll back?
- Who needs to be involved?
- What should not be done?
A runbook that says "investigate logs" adds little.
Operational documentation is valuable when it reduces decision latency during abnormal conditions.
Operational Acceptance Is a Release Gate
Many organizations have Definition of Done for development but no explicit operational acceptance.
A stronger release process asks whether the capability is production done.
That may include:
- code complete;
- tests passing;
- security concerns reviewed;
- migration rehearsed;
- dashboards available;
- alerts meaningful;
- rollback tested;
- ownership assigned;
- runbook written;
- support informed;
- critical dependencies understood;
- post-release metrics defined.
The exact checklist should be proportional to risk. A small internal feature does not need the same ceremony as a payment workflow.
The principle is that operational responsibility is part of completion.
Production Done Is an Engineering Standard
NILLKAI uses the idea of Production Done to describe this broader completion condition.
The important point is not the label.
It is the refusal to equate "deployed" with "finished."
Software becomes valuable in operation, and operation creates new obligations.
A team that owns outcomes needs to know how the system behaves after the deployment pipeline has stopped being interesting.
Deployment Is Only the Beginning
Production readiness is confidence across the lifecycle of operation.
Can the system serve expected load?
Can the team observe meaningful behavior?
Can failures be contained?
Can data be recovered?
Can a release be reversed?
Can dependencies degrade safely?
Can someone respond at 2 AM without reconstructing the system from source code?
Those questions make readiness more expensive than deployment.
They also make deployment less likely to become an incident.
The release is not complete when the artifact reaches production.
It is complete when the organization is prepared to own what happens next.
References
- Google. Launch Coordination Checklist. Site Reliability Engineering. https://sre.google/sre-book/launch-checklist/
- Google. Reliable Product Launches at Scale. Site Reliability Engineering. https://sre.google/sre-book/reliable-product-launches/
- Google. Creating a Production Launch Plan. https://sre.google/resources/practices-and-processes/production-launch-planning/
- Microsoft. Reliability design principles. Azure Well-Architected Framework. https://learn.microsoft.com/en-us/azure/well-architected/reliability/principles
