Skip to content

Implement exponential backoff circuit breaker for failing producer peers in tpu_sync. - #887

Open
copybara-service[bot] wants to merge 1 commit into
mainfrom
test_978163735
Open

copybara-service[bot] wants to merge 1 commit into
mainfrom
test_978163735

Conversation

@copybara-service

@copybara-service copybara-service Bot commented Sep 8, 2026

Copy link
Copy Markdown

Implement exponential backoff circuit breaker for failing producer peers in tpu_sync.

When a prefill node experiences failures (e.g. frozen process, unresponsive socket, or crash), incoming requests targeting that producer can consume worker threads in push_pool_, potentially starving healthy producers.

  • Add a 2-state exponential backoff Circuit Breaker (CLOSED <-> OPEN) per peer endpoint:
    • Trips from CLOSED to OPEN after 2 consecutive connection/handshake failures.
    • Initial ban duration is 2 minutes, doubling on repeated failures up to 16 minutes (2m -> 4m -> 8m -> 16m).
    • Automatically transitions back to CLOSED upon expiry of the ban deadline.
    • Fast-fails incoming reads for banned peers in <1us without occupying threads in push_pool_.
    • Straggler protection: in-flight requests that fail while the circuit is already OPEN do not trigger additional backoff escalation.
  • Guard configuration behind opt-in environment variable (disabled by default):
    • Environment variable: TPU_RAIDEN_ENABLE_CIRCUIT_BREAKER ("1" / "true").
    • Programmatic setter/getter: set_circuit_breaker_enabled(bool) / circuit_breaker_enabled().
  • Add unit tests verifying fast-failover, non-starvation of healthy peers when enabled, disabled-by-default behavior, and straggler protection.

@google-cla

google-cla Bot commented Sep 8, 2026

Copy link
Copy Markdown

Thanks for your pull request! It looks like this may be your first contribution to a Google open source project. Before we can look at your pull request, you'll need to sign a Contributor License Agreement (CLA).

View this failed invocation of the CLA check for more information.

For the most up to date status, view the checks section at the bottom of the pull request.

…ers in tpu_sync.

When a prefill node experiences failures (e.g. frozen process, unresponsive socket, or crash), incoming requests targeting that producer can consume worker threads in push_pool_, potentially starving healthy producers.

* Add a 2-state exponential backoff Circuit Breaker (CLOSED <-> OPEN) per peer endpoint:
  - Trips from CLOSED to OPEN after 2 consecutive connection/handshake failures.
  - Initial ban duration is 2 minutes, doubling on repeated failures up to 16 minutes (2m -> 4m -> 8m -> 16m).
  - Automatically transitions back to CLOSED upon expiry of the ban deadline.
  - Fast-fails incoming reads for banned peers in <1us without occupying threads in push_pool_.
  - Straggler protection: in-flight requests that fail while the circuit is already OPEN do not trigger additional backoff escalation.
* Guard configuration behind opt-in environment variable (disabled by default):
  - Environment variable: `TPU_RAIDEN_ENABLE_CIRCUIT_BREAKER` ("1" / "true").
  - Programmatic setter/getter: `set_circuit_breaker_enabled(bool)` / `circuit_breaker_enabled()`.
* Add unit tests verifying fast-failover, non-starvation of healthy peers when enabled, disabled-by-default behavior, and straggler protection.

PiperOrigin-RevId: 978163735
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants