Resilience Demo: naive vs. resilient integration
Two clients calling the same intentionally-flaky fake third-party dependency — one plainHttpClientcall, one wrapped in a Polly v8 resilience pipeline. Same scripted outage hits both at the same moment, in the same k6 run.
One small, self-contained repo — no database, clone and run locally in minutes.
What the resilient client does differently
The naive client is a plain HttpClient call — no retry, no timeout override, no circuit breaker. The resilient client wraps the same call in a Polly v8 pipeline via Microsoft.Extensions.Http.Resilience:
- Retry with backoff + jitter — Transient failures (a dropped connection, a momentary 500) get a few automatic retries with exponential backoff, instead of surfacing to the caller on the first hiccup.
- Circuit breaker — Once failures cross a threshold, the breaker opens and fails fast for a cooldown window — no point hammering a dependency that's already down. It half-opens automatically to test recovery, then fully closes once healthy again.
- Per-attempt timeout — Each individual attempt is bounded, not the whole retry budget — a hung call doesn't block the pipeline from moving on to the next attempt or, once the breaker opens, failing fast.
- Graceful fallback — When the pipeline ultimately can't get a real answer, the endpoint returns a clear "temporarily unavailable" response instead of an unhandled 500 — the caller gets something predictable to handle.
Load test results
During the 30-second total outage, the resilient client handled 2.75x the requests of the naive client (440 vs 160) by failing fast instead of hanging on every call — see exactly where below.
Timeline: watch it happen
Every 5 seconds of the same 90-second run, both clients hit by the identical outage window.
Requests handled per 5s
p95 latency
Error rate
View as table
| t (s) | Naive requests | Naive p95 | Naive error rate | Resilient requests | Resilient p95 | Resilient error rate |
|---|---|---|---|---|---|---|
| 0 | 84 | 728ms | 0.00% | 86 | 728ms | 0.00% |
| 5 | 96 | 89ms | 0.00% | 94 | 85ms | 0.00% |
| 10 | 100 | 30ms | 0.00% | 100 | 36ms | 0.00% |
| 15 | 100 | 30ms | 0.00% | 100 | 27ms | 0.00% |
| 20 | 40 | 3.0s | 100.0% | 20 | 8.8s | 100.0% |
| 25 | 20 | 3.0s | 100.0% | 31 | 27ms | 100.0% |
| 30 | 20 | 3.0s | 100.0% | 97 | 28ms | 100.0% |
| 35 | 20 | 3.1s | 100.0% | 100 | 21ms | 100.0% |
| 40 | 34 | 3.1s | 100.0% | 97 | 25ms | 100.0% |
| 45 | 26 | 3.1s | 100.0% | 95 | 20ms | 100.0% |
| 50 | 47 | 41ms | 0.00% | 99 | 31ms | 65.7% |
| 55 | 100 | 60ms | 0.00% | 94 | 46ms | 0.00% |
| 60 | 94 | 66ms | 0.00% | 95 | 75ms | 0.00% |
| 65 | 99 | 65ms | 0.00% | 96 | 67ms | 0.00% |
| 70 | 87 | 70ms | 0.00% | 97 | 56ms | 0.00% |
| 75 | 99 | 32ms | 0.00% | 99 | 30ms | 0.00% |
| 80 | 98 | 42ms | 0.00% | 96 | 36ms | 0.00% |
| 85 | 100 | 23ms | 0.00% | 95 | 23ms | 0.00% |
Phase summary
Healthy
Dependency behaving normally.
Requests handled (same window)
380vs380
Error rate
0.00%vs0.00%
p95 latency
704msvs713ms
Baseline — with nothing wrong, both clients behave identically. The resilience pipeline adds no meaningful overhead when it isn't needed.
Outage
Dependency dropping every connection.
Requests handled (same 30s)
160vs440
Error rate
100.0%vs100.0%
p95 latency
3.1svs2.2s
Same 100% error rate is expected — no amount of resilience code invents a response a fully-down dependency won't give. The real difference is speed and throughput: once its circuit breaker trips, the resilient client stops making real network calls and fails instantly, so it processes 440 requests to naive's 160 in the identical 30-second window (2.75x the throughput), at roughly 29% lower p95 latency per failed request. The timeline below shows the trip itself: a brief spike past 8s while the breaker is still deciding, then a hard drop to ~20-30ms once it opens.
Recovery
Dependency back up — does each client notice?
Requests handled (same window)
724vs771
Error rate
0.00%vs8.4%
p95 latency
58msvs55ms
Naive recovers instantly — it's stateless, with no breaker to reset, so it just starts succeeding the moment the dependency does. The resilient client's ~8% error rate here is concentrated almost entirely in the first 5 seconds after recovery (65.7% in that single bucket) — its circuit breaker's half-open verification window, a handful of trial calls that have to succeed before it trusts the dependency again. By the next bucket it's back to 0%. A genuine, worthwhile trade-off of circuit breakers — not a flaw.
Last captured 7/29/2026.