Making microservices resilient

Timeouts, retries with backoff, circuit breakers, idempotency and the outbox: how services survive each other's failures.

Making microservices resilient

In a microservice system, some call always fails somewhere: a service restarts, the network hiccups, a database is slow. Design for these partial failures instead of hoping they don’t happen.

The toolbox

  • Timeout: never wait forever for another service. A slow dependency must not hold your threads and requests.
  • Retry with backoff: retry short, temporary failures, waiting longer each time (for example 200 ms, 400 ms, 800 ms) and adding a little randomness (jitter), so many clients don’t retry at the same moment.
  • Circuit breaker: after many failures, stop calling the broken service for a while and fail fast. Then try again carefully.
  • Bulkhead: limit how many calls can go to one dependency, so one slow service can’t use up everything.
  • Fallback: return a cached or default answer when a call fails.

In .NET, Microsoft.Extensions.Http.Resilience (built on Polly) adds these to HttpClient in a few lines.

A retry, by hand

int calls = 0;
Console.WriteLine(await RetryAsync(() => ++calls < 3 ? throw new HttpRequestException("busy") : Task.FromResult("ok")));
async Task<string> RetryAsync(Func<Task<string>> action)
{
    for (int attempt = 1; ; attempt++)
    {
        try { return await action(); }
        catch (HttpRequestException) when (attempt < 5) { await Task.Delay(10 * (1 << attempt)); } // wait longer each time
    }
}

It prints ok: the first two calls fail, the third succeeds, with a growing wait between tries.

Data across services

  • Idempotency: retries mean the same request can arrive twice. Make handlers safe to repeat, for example with a request ID.
  • Outbox: save the business change and the message to send in the same database transaction; a background job publishes the messages. No lost events, but a message may be published twice, so consumers must be idempotent.
  • Saga: a business process across services is a series of local steps, each with a compensating step to undo it if a later step fails.

At the edge

An API gateway gives clients one entry point and handles routing, authentication and rate limiting for the services behind it.

Read more: https://learn.microsoft.com/en-us/dotnet/architecture/microservices/implement-resilient-applications/

#tip · TIP-099


Write a comment