[release-3.27] iam: separate backoffs, add jitter, and increase for conflicts #2

stbenjam · 2021-12-22T19:17:26Z

Cherry-pick of my proposed fix upstream to 3.27 (what's used in the 4.10 openshift installer)

We can hold off on merging this for now, to give the upstreams folks a chance to review.

Within our project, we can have multiple instances of terraform
modifying iam policies, and in many cases these instances are kicked off
at exactly the same time. We're running into errors where we exceed the
backoff max (which in reality is 16 seconds, not 30). Also, Google
reccomends that backoffs contain jitter [1] to prevent clients from
retrying all at once in synchronized waves.

This change (1) separates the 3 distinct backoffs used in the iam policy
read-modify-write cycle, (2) introduces jitter on each retry, and (3)
increases the conflict max backoff to 5 minutes.

[1] https://cloud.google.com/iot/docs/how-tos/exponential-backoff

Within our project, we can have multiple instances of terraform modifying iam policies, and in many cases these instances are kicked off at exactly the same time. We're running into errors where we exceed the backoff max (which in reality is 16 seconds, not 30). Also, Google reccomends that backoffs contain jitter [1] to prevent clients from retrying all at once in synchronized waves. This change (1) separates the 3 distinct backoffs used in the iam policy read-modify-write cycle, (2) introduces jitter on each retry, and (3) increases the conflict max backoff to 5 minutes. [1] https://cloud.google.com/iot/docs/how-tos/exponential-backoff

openshift-bot · 2022-03-22T23:29:38Z

Issues go stale after 90d of inactivity.

Mark the issue as fresh by commenting /remove-lifecycle stale.
Stale issues rot after an additional 30d of inactivity and eventually close.
Exclude this issue from closing by commenting /lifecycle frozen.

If this issue is safe to close now please do so with /close.

/lifecycle stale

openshift-bot · 2022-04-21T23:58:12Z

Stale issues rot after 30d of inactivity.

Mark the issue as fresh by commenting /remove-lifecycle rotten.
Rotten issues close after an additional 30d of inactivity.
Exclude this issue from closing by commenting /lifecycle frozen.

If this issue is safe to close now please do so with /close.

/lifecycle rotten
/remove-lifecycle stale

openshift-ci bot added the lifecycle/stale Denotes an issue or PR has remained open with no activity and has become stale. label Mar 22, 2022

openshift-ci bot added lifecycle/rotten Denotes an issue or PR that has aged beyond stale and will be auto-closed. and removed lifecycle/stale Denotes an issue or PR has remained open with no activity and has become stale. labels Apr 21, 2022

stbenjam closed this Apr 22, 2022

stbenjam deleted the fix-retries-ocp branch April 22, 2022 00:39

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[release-3.27] iam: separate backoffs, add jitter, and increase for conflicts #2

[release-3.27] iam: separate backoffs, add jitter, and increase for conflicts #2

stbenjam commented Dec 22, 2021 •

edited

Loading

openshift-bot commented Mar 22, 2022

openshift-bot commented Apr 21, 2022

[release-3.27] iam: separate backoffs, add jitter, and increase for conflicts #2

[release-3.27] iam: separate backoffs, add jitter, and increase for conflicts #2

Conversation

stbenjam commented Dec 22, 2021 • edited Loading

openshift-bot commented Mar 22, 2022

openshift-bot commented Apr 21, 2022

stbenjam commented Dec 22, 2021 •

edited

Loading