504 Gateway Timeout
Server Error
The server, while acting as a gateway or proxy, did not receive a timely response from an upstream server it needed to access in order to complete the request. 504 differs from 502 in that the upstream did not produce a malformed response. It failed to respond within the gateway's timeout window. The gateway terminates the upstream request and emits 504 to the client. Common emitters include Nginx, HAProxy, Cloudflare, AWS ALB, and Kubernetes ingress.
When does this happen?
504 surfaces when the upstream takes too long. Long-running database queries, blocked locks, slow third-party API calls, and synchronous workloads that should have been async all produce 504s. Cloudflare's 524 is a related but distinct code (cf-specific timeout); ALB has a configurable target timeout that defaults to 60 seconds; Nginx's proxy_read_timeout default is also 60 seconds. Browsers display whatever HTML the gateway returns. The fix can be on either side: optimize the upstream to respond faster, or extend the gateway timeout to give the upstream more room. Persistent 504s on indexed pages cause de-indexing, similar to 502 and 500. Watch for 504s that mask deadlocks: if a query is held by another transaction, every other waiter will time out before the original transaction completes.
Common causes
- Backend application is too slow. Slow query, blocking lock, or CPU saturated.
- Database query exceeded the gateway's read timeout. Usually a missing index or a query plan regression.
- Network latency between proxy and backend is too high. Common across regions or VPCs.
- Proxy timeout configured too aggressively for the actual workload. Defaults are 60s but legitimate work may take longer.
- Backend is waiting on a third-party API that is itself slow or down.
- Deadlock in the application. Request handler is waiting on a mutex that will never release.
- Synchronous handler doing work that should be async. Generating reports, processing large files, etc.
- Backend's response buffering exceeds the proxy's allowed body size and the proxy gives up.
How to fix it
- Optimize the slow upstream operation. Profile, add indexes, or split the work.
- Increase the proxy's read timeout if the work is legitimately slow. Proxy_read_timeout in Nginx, response timeout on ALB target groups.
- Move long-running work to a background job queue and return 202 with a polling URL instead of blocking the request.
- Add caching in front of slow endpoints. Redis, Varnish, or a CDN cache reduces backend load and hides slow responses.
- Check for deadlocks. A request that holds a lock and waits on the same lock will hang until killed.
- Audit network paths. A 504 on cross-region calls often improves with a closer replica or a VPC peering fix.
- Set explicit timeouts on third-party calls inside your application so they fail fast instead of cascading into 504.
- Use observability (traces, slow query logs) to find the specific operation that exceeded the budget.
Real-world examples
- Reporting endpoint generates a 500MB CSV synchronously by querying a table without an index on the date column.
- Query takes 90 seconds. Proxy times out at 60s and returns 504. Engineer moves report generation to a background job and serves the result via a polling endpoint.
- Backend calls a third-party payment provider that is having an incident.
- Payment provider's API hangs for 120 seconds. Application's proxy times out at 60s and returns 504. Operator adds a 5-second client-side timeout and a circuit breaker for the payment provider.
- Database deadlock blocks every request that touches a hot row.
- Affected requests time out at the gateway and return 504. DBA kills the long-running transaction and traffic recovers.
- Application running in us-east tries to call a backend in eu-west across the public internet.
- Network latency plus a slow query exceed the 60s default. Gateway returns 504. Operator moves the dependency to the same region.
- Cloudflare cannot get a response from the origin within its default 120-second Proxy Read Timeout.
- Cloudflare returns 524 (its variant of 504). Operator checks origin response times and optimizes the slow path.
- Migration job runs an UPDATE on a 10-million-row table during peak traffic.
- Table is locked; requests to that table time out at the gateway. Application returns 504 until the migration completes.
Debugging
- Check Nginx (or your gateway's) error log for "upstream timed out (110: Connection timed out) while reading response header from upstream."
- Run curl -w '%{time_total}' --max-time 120 https://api.example.com/slow-endpoint to measure how long the upstream actually takes.
- Inspect slow-query logs in the database (pg_stat_activity, slow_query_log) for queries that exceed the gateway timeout.
- Capture distributed traces. A 504 at the edge often points at a specific span buried two services deep.
- Check the proxy's read timeout (proxy_read_timeout, target group response timeout) and compare against the upstream's p99.
- Use APM (Datadog, Sentry, New Relic) to identify slow endpoints and prioritize the worst offenders.
How it differs from related codes
HTTP 502
502 is the upstream returning a bad or no response. 504 is the upstream not responding in time. Both surface at the gateway, but the cause differs. 502 is brokenness, 504 is slowness.
HTTP 503
503 is the server saying "I am unavailable" with a Retry-After. 504 is the gateway giving up on a backend that did not say anything. 503 is more cooperative.
HTTP 500
500 is an application-level error that returned promptly. 504 is the absence of a response within the budget. If your application returns 500 to the gateway, the client sees 500, not 504.
Related status codes
See HTTP 504 in your redirect chains?