503 Service Unavailable
Server Error
The server is currently unable to handle the request due to a temporary overload or scheduled maintenance, which will be alleviated after some delay. The response should include a Retry-After header so clients and crawlers know when to come back. 503 is the most polite of the 5xx codes. It acknowledges the server is healthy enough to respond but cannot serve this specific request right now. Search engines treat 503 with Retry-After as a clear "come back later" signal and avoid de-indexing the URL, which makes 503 the correct status for planned maintenance windows.
When does this happen?
Return 503 during planned maintenance, during deploys, when the server is over capacity, or when a critical dependency (database, cache, queue) is unavailable. The Retry-After header is important. Set it to the expected duration of the issue. For maintenance windows of an hour, send Retry-After: 3600. For overload, a shorter value (60-300 seconds) tells crawlers to come back soon. The response body can include a user-friendly maintenance page. 503 is not the same as 502 or 504: 503 means "I am here but cannot do this now" while 502 means "the upstream is broken" and 504 means "the upstream did not respond." Use 503 when you want graceful degradation and explicit downtime communication. Browsers display whatever HTML the server returns, so a well-designed maintenance page is visible to users during outages.
Common causes
- Planned maintenance. The operator has put the application in maintenance mode while deploying or migrating data.
- Server overload. CPU saturated, request queue full, and the load balancer is shedding traffic with 503.
- Dependent service (database, cache, queue) is down and the application cannot complete most requests.
- Deployment in progress. The new version is starting up and the old version has already drained.
- Auto-scaler is lagging behind a traffic spike and the current capacity cannot keep up.
- Region failover in progress. The active region is being shifted and the gateway returns 503 for inbound traffic during the cutover.
- Circuit breaker tripped. A downstream service has been failing and the application is failing fast with 503 instead of cascading failures.
- Rate-limit-style throttling at the infrastructure level rather than application level.
How to fix it
- Wait and retry. Check Retry-After. 503 is explicitly transient.
- If you operate the service, scale up capacity to handle the traffic level that triggered the overload.
- Verify dependent services (database, cache, queue) are healthy. A downstream outage cascades into 503 upstream.
- Implement a maintenance page with a friendly message and a Retry-After header so users and crawlers know what to do.
- Tune autoscaler thresholds. Chronic 503s during traffic spikes mean the scaler is too slow.
- Add a circuit breaker around fragile downstream calls so the application fails fast with 503 instead of timing out into 502/504.
- Audit deploy procedures. 503 spikes during rollouts often mean the new pods are not ready before old ones drain.
- For client-side, implement retry with exponential backoff and jitter and honor Retry-After to avoid thundering herds.
Real-world examples
- Operator puts the application into maintenance mode at midnight to run a database migration.
- Server returns 503 with Retry-After: 1800 and a maintenance page. Crawlers respect the header; users see the page and come back later.
- Black Friday traffic spike exceeds autoscaler capacity by 30%.
- Load balancer returns 503 to the excess traffic with Retry-After: 60. Autoscaler catches up and 503s fade within minutes.
- Primary database is in failover and writes are paused for 90 seconds.
- Write endpoints return 503 with Retry-After: 90. Reads continue via replicas. Clients with retry logic recover automatically once failover completes.
- Kubernetes rolling deploy where new pods are starting and old pods are being drained.
- Ingress returns 503 for a small window while no pods are ready. Operator tunes maxUnavailable and minReadySeconds to eliminate the gap.
- Circuit breaker tripped after a downstream pricing service failed.
- Application fails fast with 503 on requests that need pricing data. Other endpoints continue to serve normally.
- CDN origin shield is overwhelmed during a viral spike and starts shedding traffic.
- Edge returns 503 for cache misses. Cached assets continue to serve from the edge. Operator scales the origin and adds caching.
Debugging
- Check the Retry-After header. Its value tells you whether this is a short blip or a planned outage.
- Run curl -I https://api.example.com/path during the outage and confirm 503 is coming from your application or your edge layer.
- Inspect application and infrastructure logs for the time window. Look for autoscaler events, database failovers, or deploy markers.
- Compare current request rate to the capacity baseline. Chronic 503s usually correlate with under-provisioning.
- Verify circuit breakers are tripped or untripped as expected. A stuck-open breaker will return 503 indefinitely even after the downstream recovers.
How it differs from related codes
HTTP 500
500 is an unexpected error. Something broke unintentionally. 503 is a deliberate "unavailable". Maintenance, overload, or graceful degradation. 503 is recoverable on its own; 500 requires investigation.
HTTP 502
502 means the gateway got a bad response from the upstream. 503 means the upstream itself is saying it cannot serve right now. 503 is more polite and includes Retry-After; 502 is a symptom of upstream brokenness.
HTTP 429
429 is per-client rate limiting. You are over your quota. 503 is server-wide unavailability. Everyone is affected. Retry-After applies to both, but the cause differs entirely.
Related status codes
See HTTP 503 in your redirect chains?