Repository navigation
orchestrator-kubernetes: Fix replica metrics fetches failing on idle connections - #39696
Conversation
…connections The Kubernetes client's 60 s read timeout fires on pooled connections that idle for close to 60 s, because hyper-timeout runs the read timer while a connection sits idle and does not reset it when a request is written. The per-minute replica metrics fetch reuses such connections for processes 1+ and fails with `client error (SendRequest)`, leaving their metrics NULL. Raise the read timeout to 120 s, above the pool's 90 s idle expiry. Count failed per-process metrics fetches in `mz_orchestrator_kubernetes_process_metrics_fetch_failures_total`, labeled by fetch step and error kind, and log the full error chain. Closes: CPU-306 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ggevay
left a comment
There was a problem hiding this comment.
LGTM, thanks! Minor comments.
The branch name contains cpu-306, so merging auto-closes CPU-306. Its upstream half (getting back to 60 s) stays open, though, and TODO(CPU-306) would then point at a closed issue. Maybe keep CPU-306 open, or point the TODO at the kube-rs PR instead.
| // therefore times out before the response arrives and fails with | ||
| // `client error (SendRequest)`. | ||
| // | ||
| // TODO(CPU-306): Return to 60 s once kube-client enables hyper-timeout's |
There was a problem hiding this comment.
kube-client won't enable it by default: the upstream change (kube-rs#2111) exposes it as an opt-in Config::reset_reader_on_write, and it can only ship in a kube-client release after 4.2.0 (we're on 3.1.0). Suggest: "TODO: set reset_reader_on_write and return to 60 s once we're on a kube-client release with kube-rs#2111."
| // The read timeout bounds how long a hung Kubernetes call can block an orchestrator worker, | ||
| // which handles one command at a time. | ||
| // | ||
| // NOTE: The read timeout must exceed the connection pool's 90 s idle expiry. hyper-timeout |
There was a problem hiding this comment.
Nit: exceeding 90 s isn't quite sufficient. Pool idle time counts against the next response for any read timeout, so at 120 s a request on a connection that idled ~89 s has ~31 s left. That's harmless for these millisecond-scale requests, but the actual constraint is "90 s plus the slowest response".
|
I re-opened CPU-306 after merging the PR. Thanks for the reviews! |
For multi-process replicas, the per-minute metrics fetch fails for processes 1+ with
client error (SendRequest), leaving their rows inmz_cluster_replica_metrics_historyNULL. hyper-timeout runs the kube client's read timer while a pooled connection sits idle and does not reset it on write. The extra connections for processes 1+ come back after about 60 s idle, so the 60 s read timeout fires before the response arrives. This raises the read timeout to 120 s, above the pool's 90 s idle expiry, with a TODO to return to 60 s once kube-client enablesreset_reader_on_write.It also adds
mz_orchestrator_kubernetes_process_metrics_fetch_failures_total, labeled by fetch step and error kind, and logs the full error chain, which previously stopped atSendRequest. Unit tests cover the error classification insrc/orchestrator-kubernetes/src/metrics/tests.rs.Raising the timeout doubles how long a hung Kubernetes call can block an orchestrator worker. CPU-306 found no orchestrator call hitting the timeout in recent production logs or Kubernetes-based CI.
Closes: CPU-306
🤖 Generated with Claude Code