Posts

Architecture Patterns for Resilience Distributed Systems

In today’s cloud-native landscape, distributed systems have become the backbone of most services. However, with distribution comes the inevitability of failures — network issues, service outages, or resource contention. The good news is that resilience patterns exist to help systems handle these failures gracefully, ensuring stability and reliability.

In this post, we’ll explore common resilience patterns that can be employed to strengthen distributed systems and mitigate the impact of failure.

1. Retry Pattern

The Retry Pattern is a simple yet effective method to handle transient failures. When an operation fails, instead of giving up immediately, the system retries the operation after a delay. This is particularly useful for temporary issues, like network timeouts or service unavailability.

When to use it?

This pattern is ideal for operations where failures are likely to be temporary, such as external API calls or database queries.

+---------+       Request        +----------+
|  Client | -------------------> | Service  |
+---------+        Retry          +----------+
     | <---(Service fails)---|         |
     | Retry (after delay)   |         |

In the diagram above, when the service fails, the client waits for a short period before trying again. This cycle repeats a limited number of times, ensuring the issue isn’t permanent before abandoning the request.

2. Circuit Breaker Pattern

The Circuit Breaker Pattern is designed to stop a system from making calls to a service if failures reach a certain threshold. Once the threshold is met, the circuit “opens,” blocking further calls for a period. After the timeout, the circuit goes to a “half-open” state, allowing a few calls to check if the service has recovered.

When to use it?

This pattern is essential for preventing cascading failures. It protects a system from repeatedly calling a failing service, which could result in system-wide downtime.

+---------+      Request      +----------+
|  Client | ----------------> | Service  |
+---------+   (Closed Circuit) +----------+
    |      Circuit Breaker (opens on failure)
    |
    |<--------------------------|
  (Circuit Open - Requests blocked)

The system monitors the success rate of service requests. If failures exceed a defined limit, the circuit breaker stops further calls until the service shows signs of recovery.

3. Backpressure Pattern

The Backpressure Pattern ensures that a system can handle incoming traffic by slowing down or rejecting requests when downstream systems are overloaded. Instead of letting a system be overwhelmed, it communicates back to the source of the requests, signaling it to reduce the traffic.

When to use it?

This pattern is used in streaming or message-based systems where the producer can slow down the rate of requests based on the consumer’s processing capacity.

+---------+    Data Stream    +----------+
| Producer| --------------->  | Consumer |
+---------+                   +----------+
      |       Backpressure signal (slow down)      |
      |<-------------------------------------------|

When the consumer is overloaded, it sends a backpressure signal to the producer, requesting it to slow down or pause sending more data until it can catch up.

4. Deadline Pattern:

Sets a maximum time for an operation or series of operations. If the deadline is reached, the operation is aborted, regardless of whether all individual tasks within it have completed or not.

When to use it?

Often used in distributed systems to enforce limits on operation durations, helping ensure system responsiveness by preventing long-running operations from continuing indefinitely. While a timeout is applied to a single operation, a deadline applies to an entire chain of dependent operations.

Consider a system making multiple downstream service calls (like microservices). A client requests some data, and the system has multiple services processing the request. The deadline ensures that all operations complete within the specified time frame. If they don’t, the whole process is aborted.

+---------+     Request     +----------+     +----------+
|  Client | --------------> | Service A |--->| Service B |
+---------+                 +----------+     +----------+
        |                (Deadline for the entire process)
      (Aborted if not completed within the deadline)

5. Bulkhead Pattern

Inspired by ship design, the Bulkhead Pattern isolates parts of a system so that a failure in one part doesn’t cascade to others. Each component of the system has its own resources (e.g., thread pools, memory) to prevent resource exhaustion.

When to use it?

Use this pattern when you want to prevent one overloaded service from affecting the performance or availability of others.

+----------+         +-----------+
| Service A | <-----> | Resource A |
+----------+         +-----------+
       |
       |
+----------+         +-----------+
| Service B | <-----> | Resource B |
+----------+         +-----------+

By isolating resources for different services, you can ensure that a problem in one service (like Service A) doesn’t affect the resources available to other services (like Service B).

6. Timeout Pattern

The Timeout Pattern ensures that an operation doesn’t wait indefinitely for a response. If the response takes longer than the predefined timeout, the operation is terminated. This avoids situations where unresponsive services cause cascading delays across the system.

When to use it?

This is useful for any service communication where the response time is critical. It prevents system deadlocks caused by hanging requests.

+---------+     Request     +----------+
|  Client | --------------> | Service  |
+---------+    Timeout       +----------+
      |<-----(Timeout error)----|

The client sends a request, and if the service doesn’t respond within the timeout window, the operation fails gracefully, preventing the system from hanging.

7. Failover Pattern

The Failover Pattern switches to a backup service or resource when the primary one becomes unavailable. Failover can happen automatically or manually, ensuring continued service availability.

When to use it?

This pattern is ideal for mission-critical services that cannot tolerate downtime, such as database clusters or replicated services.

+---------+     Request     +----------+
|  Client | --------------> | Service A|
+---------+                 +----------+
     |    Failover            |
     |----------------------->| Backup |
                              +--------+

If the primary service (Service A) fails, the system automatically redirects traffic to the backup service, ensuring continuity.

8. Failover & Redundancy Pattern

The Failover & Redundancy Pattern ensures that if a primary component or service fails, the system automatically switches to a backup or redundant instance, maintaining continuity. Redundancy refers to having multiple instances of a component available, while failover is the process of switching to the backup when the primary one fails. Together, they enhance system availability and reliability.

When to use it?

This pattern is essential in high-availability systems where downtime is unacceptable, such as in mission-critical applications, cloud environments, or systems with strict service-level agreements (SLAs). It ensures minimal disruption during failures by maintaining backup services ready to take over immediately.

+---------+     Primary Request    +----------+
|  Client | ---------------------> | Service A|
+---------+                        +----------+
      |                             |
      |    Primary fails, switch    |
      |----------------------------> +--------+
                                     | Backup |
                                     +--------+

In this pattern, multiple instances of a service (Service A and Backup) are kept ready. When the primary service fails, the system immediately fails over to the backup service, ensuring uninterrupted functionality.

9. Auto-Scaling & Self-Healing Pattern

The Auto-Scaling & Self-Healing Pattern enables systems to automatically adjust their resources based on demand and recover from failures without manual intervention. Auto-scaling dynamically adds or removes instances of services in response to traffic or resource usage, ensuring optimal performance and cost-efficiency. Self-healing detects failures and automatically replaces or restarts unhealthy components, maintaining system stability.

When to use it?

This pattern is critical in cloud environments and distributed systems where workloads can fluctuate, and high availability is required. It ensures that systems efficiently handle spikes in demand and recover from component failures without affecting the user experience.

+---------+     Monitor Usage     +----------+
| Service | <-------------------> | Auto-Scaler|
+---------+                        +----------+
      |                             |
      | Scale up/down as needed     |
      |----------------------------> 

+---------+   Detect Failure    +----------+
| Service | ------------------> | Self-Healing|
+---------+                     +-----------+
     |                              |
     | Restart/Replace Instance      |
     |-------------------------------|

In this pattern, auto-scaling dynamically adjusts the number of instances based on resource usage, and self-healing mechanisms detect failures and automatically restart or replace failing instances, keeping the system healthy and responsive.

10. Fallback Pattern

The Fallback Pattern provides an alternative action or response if a service fails or is unavailable. Instead of failing completely, the system provides a degraded but functional response.

When to use it?

This pattern is useful in non-critical operations where a partial or default response is better than none.

+---------+     Request     +----------+
|  Client | --------------> | Service  |
+---------+                 +----------+
      |<--- (Service fails)---|
      |       Fallback         |
      |<---------------------->|

In the event of a failure, the system returns a pre-configured fallback response, keeping the user experience as smooth as possible.

11. Throttling Pattern

The Throttling Pattern limits the number of requests a service can handle in a given time period. By throttling incoming requests, you ensure that a system isn’t overwhelmed and can continue operating within its capacity.

When to use it?

This pattern is especially useful in high-traffic scenarios, protecting your services from being overwhelmed.

+---------+      Request     +----------+
|  Client | ---------------> | Service  |
+---------+      Throttling    +----------+
   |      (Some requests dropped or delayed)
   |<---------------------------|

Here, the service allows a limited number of requests per unit time. Any additional requests are either delayed or dropped.

12. Queue-Based Load Leveling

In the Queue-Based Load Leveling Pattern, incoming requests are placed in a queue and processed sequentially. This allows the system to handle traffic spikes by processing requests at a steady rate.

When to use it?

This pattern is essential for services that experience periodic surges in requests, such as e-commerce systems during sales.

+---------+   Request  +----------+    Process    +----------+
|  Client | ---------> |   Queue   | -----------> |  Service  |
+---------+            +----------+              +----------+

Requests are queued and processed one by one, preventing the service from being overwhelmed by a sudden influx of traffic.

13. Cache-Aside Pattern

The Cache-Aside Pattern loads data into a cache from a data store on demand. When an application requests data, it first checks the cache. If the data is not present, it retrieves the data from the database, stores it in the cache, and then returns it to the client. This helps reduce load on the database and improves performance.

When to use it?

This pattern is particularly useful when you need to improve response times for frequently requested data without constantly querying a database.

+---------+    Request     +---------+
|  Client | -------------> |  Cache  |
+---------+                +---------+
       |      Cache miss         |
       |------------------------>| 
       |                       +---------+
       |      Data loaded       |Database |
       |<-----------------------|         |
+---------+     Data returned    +---------+

The client checks the cache first. If there’s a cache miss, the system retrieves data from the database and stores it in the cache for future requests.

14. Compensating Transaction Pattern

The Compensating Transaction Pattern is used to undo the effects of a previously committed transaction in a distributed system. When something goes wrong after one or more steps of a process, a compensating action can be taken to revert the system to a consistent state.

When to use it?

Use this pattern when managing distributed transactions where rolling back a complete transaction isn’t possible. For example, in complex workflows involving multiple services, like payment processing or booking systems.

+---------+    Transaction    +----------+
| Service | ----------------> |  Step 1  |
+---------+                   +----------+
      |   |   Step 2 completes, Step 3 fails   |
      |<----------------------------------------|
  Compensating action to undo Step 1 and Step 2

When a transaction fails at a later stage, compensating transactions are triggered to revert previous successful steps.

15. Event Sourcing Pattern

The Event Sourcing Pattern records changes to an application’s state as a sequence of events. Instead of storing the current state of the system, the system stores a log of all changes (events), which can be replayed to derive the current state.

When to use it?

Use this pattern when you need an audit log or to rebuild the system’s state at any point in time, such as in financial services or event-driven architectures.

+---------+    Event log    +------------+
|  Client | --------------> |  Event Log |
+---------+                 +------------+
       |        Replay events to rebuild state     |
       |------------------------------------------>|
+---------+     Rebuild state     +------------+
| Service | --------------------> |  Service   |
+---------+                        +------------+

All system changes are captured as events, allowing the system to reconstruct its state by replaying these events.

16. Priority Queue Pattern

The Priority Queue Pattern ensures that critical requests are processed before less important ones. It assigns priority to requests and processes them based on their importance, allowing time-sensitive or high-priority requests to be handled first.

When to use it?

This pattern is ideal for systems where certain tasks or users require preferential treatment, like healthcare systems or real-time services.

+---------+      Request 1 (high priority)      +----------+
|  Client | ----------------------------------> | Service  |
+---------+                                      +----------+
       |
       |---> Request 2 (low priority) waits until Request 1 is processed

High-priority tasks are processed first, ensuring time-sensitive operations aren’t delayed by less critical ones.

17. Idempotency Pattern

The Idempotency Pattern ensures that performing the same operation multiple times will produce the same result. This is crucial in distributed systems where requests may be duplicated due to network issues or retry mechanisms.

When to use it?

This pattern is essential when an operation may be executed more than once but should only have an effect as if it were executed once — such as payment processing or order creation.

+---------+    Request 1     +---------+
|  Client | ---------------> | Service |
+---------+                  +---------+
      |           |
      |         Retry Request 1 (same effect as initial)
      |---------------------->

Here, the client sends the same request multiple times, but the service ensures the operation’s result remains consistent and no duplicate actions (like charging a customer twice) occur.

18. Saga Pattern

The Saga Pattern is a coordination pattern for managing distributed transactions. It breaks a long-running transaction into a series of smaller steps, each managed as a local transaction. If one of these steps fails, compensating transactions are invoked to undo the previous steps, similar to the compensating transaction pattern but on a larger scale.

When to use it?

This pattern is ideal when managing complex workflows that span multiple services, such as booking systems (e.g., hotel and flight bookings).

+---------+   Step 1 (Success)   +---------+
|  Service| -------------------> | Service |
+---------+   Step 2 (Success)   +---------+
       |          |
       |        Step 3 (Fails)
       |-----------------------> Trigger compensating steps for Step 1 and 2

When a step fails, previously successful steps are rolled back using compensating actions, ensuring data consistency across distributed services.

19. Load Shedding Pattern

The Load Shedding Pattern is a strategy used to reject excess requests when a system is under heavy load. The idea is to drop less important or non-essential requests to ensure that critical ones can still be processed.

When to use it?

Use this pattern when your system experiences high load and you need to prioritize critical traffic, like real-time services where certain operations (e.g., health checks) are more important than others (e.g., reporting).

+---------+    Incoming Requests (overload)    +----------+
|   Load  | ---------------------------------> | Service  |
| Balancer|                                    +----------+
   |         Reject non-essential requests
   |
   |-----> Drop some requests based on priority or load

Here, the system drops lower-priority requests to maintain performance for higher-priority tasks when under stress.

20. Rate Limiting Pattern

The Rate Limiting Pattern limits the number of requests a system or service can handle over a specified period. This ensures that a service doesn’t get overwhelmed by a flood of requests, protecting it from resource exhaustion.

When to use it?

This pattern is commonly used in APIs, especially public-facing ones, where you need to prevent abuse or excessive traffic from overwhelming your service.

+---------+    Requests     +----------+
|  Client | --------------> | API Gate |
+---------+  Rate Limiting   +----------+
   |
   |<--- "429 Too Many Requests" (after limit reached)

After a certain limit is reached, the API gate responds with an error code (like HTTP 429), preventing further requests from reaching the backend.

21. Service Mesh Pattern

The Service Mesh Pattern uses a dedicated infrastructure layer to manage service-to-service communication. This layer provides features like traffic management, retries, circuit breaking, load balancing, and observability without requiring changes to the application code.

When to use it?

Use this pattern when you need to control and monitor service-to-service communication in a microservices architecture. Service meshes are often deployed in Kubernetes environments, using tools like Istio or Linkerd.

+-----------+    Request    +----------+
|   Client  |  -----------> |  Proxy   |
+-----------+               +----------+
                              |  Service-to-service communication
                              |  managed by the service mesh

The service mesh adds an additional layer of control and resilience to your system, handling failures and retries transparently for your services.

22. Graceful Degradation Pattern

The Graceful Degradation Pattern allows a system to reduce its functionality in a controlled way when certain services or resources become unavailable. Rather than failing completely, the system continues operating at a reduced level of service.

When to use it?

This pattern is useful in systems where partial availability is preferable to complete unavailability, such as an e-commerce site showing limited product information if the product catalog service fails.

+---------+    Request    +---------+
|  Client | ----------->  | Service |
+---------+               +---------+
                               |
                    Service degraded due to partial failure

Instead of fully failing, the system reduces its functionality gracefully, such as showing cached or partial data.

23. Shadow Traffic Pattern

The Shadow Traffic Pattern involves routing a portion of real traffic (or production-like traffic) to a new service or instance, without affecting the live system. This allows you to test the resilience and performance of a new service or version without putting actual customers at risk. It’s an advanced pattern for verifying that the new system can handle real-world loads and errors before being fully deployed.

When to use it?

This pattern is useful when you want to test new versions of services or systems under realistic load conditions but are not ready to switch all traffic to the new service immediately.

+---------+   Request    +-----------+
|  Client | -----------> |  Primary  |   
+---------+               | Service  |
                          +-----------+
                               |
                        Shadow Traffic  (Copy of real traffic)
                               |
                        +-----------+
                        |  Shadow   |
                        |  Service  |
                        +-----------+

Shadow traffic duplicates real production traffic to a new instance (shadow) of the service to verify its performance and resilience.

24. Steady-State Pattern

The Steady-State Pattern focuses on maintaining the stability of a system over time. In distributed systems, it’s crucial that systems can handle regular operations without degradation or unexpected failures. Steady-state monitoring involves continuously observing key indicators (e.g., CPU usage, memory consumption, network traffic) to ensure the system remains healthy and performs optimally.

When to use it?

This pattern is applicable to production environments where maintaining high availability and performance is critical. Continuous monitoring ensures that minor issues are caught before they become critical.

+---------+   Monitor   +---------+
| System  | -----------> | Health  |    
| Metrics |             |   Check |
+---------+             +---------+
         |
      Continuous Health Monitoring

The system constantly monitors metrics and health checks to ensure it stays in a steady, healthy state.

25. Graceful Shutdown Pattern

What is it?

The Graceful Shutdown Pattern ensures that a service or system can shut down without disrupting ongoing operations. It allows a service to finish processing in-progress requests before terminating, preventing data corruption or incomplete transactions.

When to use it?

This pattern is important when deploying updates, scaling down services, or restarting systems to ensure in-flight requests are properly completed.

+---------+  Shutdown Signal   +---------+
| Service | -----------------> | Request |
|  Node   |                    | Handler |
+---------+                    +---------+
          |     Finish Active Requests    |
          |----------------------------->|
          |  Safe Shutdown after Completion

The system receives a shutdown signal but waits to terminate until all active requests have been completed gracefully.

26. Redundancy Pattern

The Redundancy Pattern is about deploying multiple instances of services or resources so that if one instance fails, another instance can take over without service disruption. This increases system availability and resilience, especially in critical services where downtime is unacceptable.

When to use it?

This pattern is useful for high-availability systems where maintaining uptime is essential, such as in financial systems, healthcare platforms, or any service where downtime can lead to significant impact.

+---------+   Request    +---------+
|  Client | -----------> | Service |
+---------+              +---------+
      |                       |
  Failover if failure occurs   |
      |----------------------->|
+---------+  Redundant Instance +---------+

When the primary service fails, a redundant instance automatically takes over, ensuring continuous availability.

27. Service Discovery Pattern

The Service Discovery Pattern automatically detects services in a system and routes traffic to healthy instances. As distributed systems scale, services often move between nodes or containers, and their locations can change frequently. Service discovery ensures that clients always connect to available, healthy service instances without needing hardcoded addresses.

When to use it?

This pattern is essential in large-scale distributed systems or microservice architectures where services dynamically register themselves or are replaced during scaling, deployments, or failover.

+---------+  Lookup Service  +---------+
|  Client | ------------->   |  Service|
+---------+                  +---------+
      |                       |
+---------+  Discover New     +---------+
| Registry |  Service Instance|
+---------+  When Needed      +---------+

Services dynamically register with a discovery system, and clients query the registry to find available services.

28. Chaos Engineering Pattern

The Chaos Engineering Pattern introduces controlled failures into a system to test its resilience and identify weaknesses before real-world issues cause outages. By simulating crashes, latency, or resource exhaustion in production environments, chaos engineering helps ensure that systems can withstand and recover from disruptions. The goal is to make failures a routine part of the system lifecycle and validate that systems degrade gracefully under stress.

When to use it?

This pattern is crucial for large-scale distributed systems, cloud-native environments, or microservices architectures where failures are inevitable. Chaos engineering is especially beneficial in environments where high availability, reliability, and fault tolerance are critical. It helps uncover weaknesses that might go unnoticed during normal operations.

+---------+   Inject Failure    +-----------+
|  System | ------------------> | Chaos Tool |
+---------+                     +-----------+
     |                              |
     | Monitor System Response       |
     |-------------------------------|

In chaos engineering, failures like server crashes, network delays, or database failures are intentionally introduced, and the system’s behavior is observed to ensure it remains resilient and stable.

29. Strangler Fig Pattern

The Strangler Fig Pattern is used to incrementally replace legacy systems with new implementations. As new functionality is developed, traffic is routed away from the legacy system to the new one. Over time, the legacy system is “strangled” as it is replaced with new functionality.

When to use it?

This pattern is effective when you need to modernize a monolithic application or gradually replace legacy systems without performing a large-scale migration in one go.

+---------+   Requests  +-----------+
|  Client | --------->  |  Legacy   |   
+---------+             |  System   |
      |                 +-----------+
      |  Traffic gradually shifted to new service
      |----------------------------->|
+---------+      New Service         +---------+

As the new service grows, more traffic is routed to it, while the legacy system eventually becomes obsolete.

30. Pessimistic and Optimistic Locking Patterns

  • Pessimistic Locking locks data when it is read, preventing other users from accessing it until the lock is released. This ensures no one else can modify the data while it’s being worked on.
  • Optimistic Locking allows multiple users to access data simultaneously. Conflicts are checked when the data is updated, and if a conflict is detected, the update fails.

When to use them?

Use Pessimistic Locking in environments where data consistency is critical and conflicts must be avoided at all costs (e.g., financial transactions). Use Optimistic Locking in scenarios where conflicts are rare, and it’s more efficient to handle them after they occur.

For Pessimistic Locking:

+---------+  Lock Data   +---------+
| Service | -----------> |   Data  |
+---------+              +---------+
       |    Data locked until changes are complete

For Optimistic Locking:

+---------+  Modify Data  +---------+
| Service |  -----------> |  Data   |
+---------+               +---------+
       |    Conflict detected, update fails 
       |<-----------------------------|

31. Event-Driven Architecture Pattern

The Event-Driven Architecture Pattern decouples services by triggering actions based on events rather than direct service calls. Each service or component publishes events to an event broker, and other services subscribe to these events and respond accordingly. This promotes loose coupling between services and ensures resilience by allowing systems to continue functioning independently of one another.

When to use it?

Use this pattern in systems where loose coupling, scalability, and responsiveness to real-time events are required. It’s particularly helpful in environments where services must react to a variety of asynchronous events or need to process large volumes of real-time data, like IoT systems or microservice-based architectures.

+---------+   Event        +---------+
| Service |  Publisher  -> |  Event  |
+---------+                |  Broker |
                           +---------+
                             |   |
                 +-----------+   +-----------+
                 |  Event Consumer           |
                 +---------------------------+

The event broker receives events from services, and event consumers react asynchronously, without needing to know the source of the event.

32. Health Endpoint Monitoring Pattern

The Health Endpoint Monitoring Pattern involves exposing an endpoint in a service that external systems (like monitoring tools or load balancers) can ping to determine whether the service is healthy and operating as expected. The service can run checks (such as checking database connectivity or external API availability) and return a status code or detailed information about the health of the service.

When to use it?

This pattern is useful for monitoring critical components in your architecture, ensuring that health checks are automated and can inform load balancers to route traffic only to healthy instances. It’s also helpful for alerting or automatically restarting services that fail health checks.

+---------+  GET /health   +---------+
| Monitor | -------------> | Service|
+---------+                +---------+
                             |
                       Health Check Passed 
                           (200 OK)

Monitoring tools periodically ping the /health endpoint of services to ensure they are healthy and able to handle traffic.

33. Retry with Exponential Backoff Pattern

The Retry with Exponential Backoff Pattern involves retrying failed operations but increasing the delay between each retry attempt exponentially. This allows systems to handle transient failures (like temporary network issues) without overwhelming a service or causing a cascading failure across the system. It also prevents immediate retries that might overload a downstream service.

When to use it?

This pattern is valuable when working with external services, databases, or APIs that may occasionally become unavailable due to transient issues. By implementing exponential backoff, you reduce the risk of overwhelming the system or external services with too many retry attempts.

+---------+  Attempt 1  +---------+
| Service | ----------> | Service |
+---------+   (Fail)    +---------+
      |     Wait 1s         |
+---------+  Attempt 2  +---------+
| Service | ----------> | Service |
+---------+   (Fail)    +---------+
      |     Wait 2s         |
+---------+  Attempt 3  +---------+
| Service | ----------> | Service |
+---------+   (Success) +---------+

Retries are attempted after increasing intervals, providing the failing system time to recover before further attempts.

34. Circuit Breaker with Bulkhead Pattern

This is a combination of the Circuit Breaker Pattern and Bulkhead Pattern. The Bulkhead Pattern isolates failures within different parts of a system so that one component failure doesn’t bring down the entire system. Coupled with the Circuit Breaker, you can isolate services or components into different groups (bulkheads), and when a failure happens, the circuit breaker kicks in to stop further damage.

When to use it?

This pattern is ideal for large, complex systems where failures in one part could propagate and bring down other parts of the system. It helps to contain issues and protect core functionalities while failures in isolated bulkheads are handled.

+---------+  Isolated          +---------+
| Client  |    Bulkheads       | Service |
+---------+    (A, B, C)       +---------+
      |     +---+    +---+    +---+
      |---->| A |----| B |----| C |
            +---+    +---+    +---+
                      |      Circuit Breaker triggers
                      +------>  Circuit Open

Each service component is isolated in bulkheads, and when failures occur, only affected bulkheads are stopped, while others continue to function.

35. Caching with TTL (Time to Live) Pattern

The Caching with Time-to-Live (TTL) Pattern involves caching data but assigning a TTL to cache entries. After the TTL expires, the cache entry is refreshed by retrieving the latest data from the source. This helps in balancing load and avoiding stale data while ensuring that cached responses remain fresh.

When to use it?

This pattern is useful in systems with frequent reads but relatively infrequent updates, such as user profile data, product catalogs, or API response caching. It can also be used when systems need to reduce load on a database or service but want to maintain data freshness.

+---------+  Get from Cache    +---------+
|  Client |  ----------------> |  Cache  |
+---------+                    +---------+
      |                        TTL Expired
+---------+  Fetch from Source +---------+
|  Source |  ----------------> |  Cache  |
+---------+                    +---------+

Cache entries are valid for a specified time, and once the TTL expires, new data is fetched from the source.

36. Static Content Hosting Pattern

The Static Content Hosting Pattern involves serving static resources (e.g., HTML, CSS, JS, images) directly from highly resilient and scalable storage or content delivery networks (CDNs) rather than dynamic servers. This makes the system more resilient by reducing load on application servers and preventing downtime due to spikes in traffic or failures in the backend.

When to use it?

This pattern is useful for web applications or systems with a lot of static content, such as websites, blogs, or single-page applications (SPAs), where delivering static content from a fast, highly available source can offload pressure from application servers.

+---------+  Request  +---------+
|  Client | --------> |  CDN    |
+---------+           +---------+
      |  Static content delivered from edge servers
      +--------------------------------------------->

The CDN delivers static content while dynamic content is handled by the application servers.

37. Transactional Outbox Pattern

The Transactional Outbox Pattern ensures that database updates and messages sent to a message broker (or another service) happen atomically. This prevents the risk of a database update succeeding but the event notification (e.g., Kafka message) failing, or vice versa, which can lead to inconsistent system states.

When to use it?

This pattern is especially useful when integrating services that rely on both data consistency and messaging systems (like event-driven architectures), particularly in microservices.

+---------+  Update Database  +---------+
| Service | -----------------> |  DB    |
+---------+                    +---------+
     |   Store message in Outbox table
     +----------------------------->|
+---------+    Read from Outbox    +---------+
|   Poll  | -------------------->  |  Broker |
+---------+                        +---------+

Messages and database changes are stored transactionally, ensuring both are either completed or rolled back.

38. Shadowing Pattern (Canary Releasing)

The Shadowing (Canary Releasing) Pattern involves deploying a new version of a service to a small subset of users or requests to observe its behavior before fully rolling it out. This minimizes the risk of widespread failure by only exposing a limited set of users or traffic to the new changes.

When to use it?

This pattern is useful when rolling out new versions of a service or feature in production environments and reducing the risk of a full-scale outage. It allows for incremental testing of new releases with minimal impact.

+---------+    Old Version     +---------+
|  Users  | ----------------> | Service |
+---------+                    +---------+
      |           Canary Release Active
      |-----> Small % of Traffic Sent to New Version
+---------+                    +---------+
|  Users  | ---------------->  | Canary  |
+---------+                    +---------+

A small portion of traffic is routed to the new version of the service, allowing gradual observation and control over changes.

39. Dead Letter Queue Pattern

The Dead Letter Queue Pattern handles messages that cannot be processed successfully by moving them into a separate queue for further analysis. Rather than discarding unprocessed messages, they are sent to the “dead letter queue” for inspection, alerting, or manual intervention.

When to use it?

This pattern is helpful in message-driven systems (e.g., message brokers or event-driven architectures) where you want to handle failed or problematic messages without losing them.

+---------+  Process Failed Message  +---------+
| Service | -----------------------> | DLQ     |
+---------+                          +---------+
         |
+---------+   Inspect & Analyze    +---------+
|  Admin  | <---------------------- | DLQ     |
+---------+                          +---------+

Messages that cannot be processed are moved into the dead letter queue for later investigation.

40. Idempotent Consumer Pattern

The Idempotent Consumer Pattern ensures that even if a message is processed multiple times (due to retries, duplicates, etc.), the result will always be the same as if it had been processed once. This prevents issues arising from duplicate message handling, ensuring that side effects (like database writes or API calls) are not repeated.

When to use it?

This pattern is essential in message-driven or event-driven architectures where systems may receive the same message multiple times (e.g., network issues, message retries). It’s commonly used in financial systems to prevent duplicate transactions or other critical side effects.

+---------+   Duplicate Message  +---------+
| Service | --------------------> | Consumer|
+---------+    Idempotent Check   +---------+
     |                            Result same for duplicates
+---------+   Process Once        +---------+
| Database| <-------------------- | Consumer|
+---------+                        +---------+

The consumer checks if the message has already been processed before applying any state changes, ensuring idempotency.

41. Fail-Fast Pattern

The Fail-Fast Pattern involves immediately halting operations when a failure is detected, rather than allowing it to propagate or attempting to recover from it. This helps to avoid wasting resources and prevents further complications from cascading failures.

When to use it?

This pattern is useful in systems where it’s better to stop and alert immediately on error rather than attempting retries or workarounds. It is especially helpful in environments where quick failure detection and resolution are prioritized, such as in real-time or critical systems.

+---------+  Detect Error  +---------+
| Service |  ------------->|   Fail  |
+---------+                |   Fast  |
                          +---------+
                             |
                       Immediate Stop & Alert

The system immediately stops processing when a failure is detected, reducing the risk of further issues.

42. CQRS (Command Query Responsibility Segregation) Pattern

The CQRS Pattern separates the responsibility of handling commands (updates to data) from handling queries (reads from data). This allows the system to optimize for reads and writes separately, often resulting in better scalability and performance.

When to use it?

This pattern is helpful in systems with high write and read workloads, where separating these responsibilities can reduce load on databases and optimize for each operation. It’s also useful in event-driven architectures and systems that need to scale read and write independently.

+---------+     Commands      +---------+
| Command | ----------------> | Command |
|  Side   |                    |Handler |
+---------+                    +---------+
      |
+---------+    Queries       +---------+
| Query   | ----------------> | Query   |
|  Side   |                    |Handler |
+---------+                    +---------+

Commands (writes) are handled separately from queries (reads), allowing for optimized processing of each.

43. Feature Toggle Pattern

The Feature Toggle Pattern allows enabling or disabling features in a system dynamically without deploying new code. Feature toggles (or flags) control which features are visible or active based on configuration settings.

When to use it?

This pattern is helpful for continuous deployment, rolling out new features gradually, and performing A/B testing or canary releases. It enables safe experimentation and iterative development.

+---------+    Toggle Status    +---------+
|  Client | ------------------>| Feature |
+---------+   Feature Enabled   | Toggle  |
                          |         +---------+
+---------+     Feature     +---------+
|   Service| <---------------- | Client  |
+---------+                     +---------+

Features are toggled on or off based on configuration settings, allowing for dynamic control over feature availability.

44. Resource Quota Pattern

The Resource Quota Pattern allocates a set amount of resources (CPU, memory, storage) to different services or users. It ensures that no single entity can consume more than its allotted share, preventing resource starvation and ensuring fair distribution.

When to use it?

This pattern is useful in multi-tenant environments or shared systems where you need to manage resource allocation and prevent any single user or service from overwhelming the system.

+---------+  Allocate Quota  +---------+
|  Service| ----------------> | Resource|
+---------+   Resource Limits  | Manager |
      |                         +---------+
+---------+  Enforce Limits    +---------+
|   Client| <------------------|  Service|
+---------+                      +---------+

Resources are allocated and enforced according to pre-defined quotas to ensure fair use and prevent resource exhaustion.

45. Request Collapsing Pattern

The Request Collapsing Pattern involves combining multiple similar requests into a single request to reduce the load on backend systems. This pattern prevents redundant processing and improves efficiency.

When to use it?

This pattern is useful when multiple requests for similar data or operations are received simultaneously. It reduces redundant work and minimizes the impact on backend services.

+---------+  Multiple Requests  +---------+
|  Client | ------------------> | Collap- |
+---------+     Combine         | ser     |
                             +---------+
+---------+   Single Request   +---------+
| Backend | <----------------- | Backend |
+---------+                    +---------+

Similar requests are combined into one, reducing redundant load on backend systems.

46. State Machine Pattern

The State Machine Pattern uses finite state machines to manage complex state transitions and behavior in a system. It defines a set of states, transitions, and actions to handle different conditions and responses.

When to use it?

This pattern is useful for managing complex workflows, ensuring predictable behavior, and handling various states and transitions explicitly. It’s often used in process automation, user interfaces, and protocol handling.

+---------+    Event A   +---------+
|   State | ----------> | State   |
|    1    |   Action    |    2    |
+---------+              +---------+
     ^                        |
     |                        |
     +---- Event B  +---------+
          Action    | State   |
                    |    3    |
                    +---------+

State transitions are explicitly defined and managed, ensuring predictable and controlled behavior.

47. Circuit Breaker with Failover Pattern

The Circuit Breaker with Failover Pattern extends the Circuit Breaker Pattern by incorporating failover mechanisms. If a circuit breaker triggers, the system automatically switches to a backup or failover service to maintain availability.

When to use it?

This pattern is useful when you need both fault isolation and redundancy. It helps ensure that service disruptions do not completely impact users by providing an alternative path for operations.

+---------+   Failures Detected   +---------+
| Service | -------------------> | Circuit |
|   A     |   Break/Failover    | Breaker |
+---------+                      +---------+
      |                           |
+---------+    Failover Service   +---------+
| Backup  | <-------------------> | Service |
| Service |     Alternative       |   B     |
+---------+                      +---------+

If the primary service fails, the system switches to a backup service to maintain functionality.

48. Adaptive Load Balancing Pattern

The Adaptive Load Balancing Pattern dynamically adjusts the distribution of requests or load among multiple servers or services based on real-time performance metrics. It helps ensure that no single server is overwhelmed.

When to use it?

This pattern is useful in systems with fluctuating workloads or where server performance can vary. It helps optimize resource utilization and maintain performance by adapting to changing conditions.

+---------+   Performance Metrics   +---------+
|  Load   | ----------------------> | Adaptive|
| Balancer|   Adjust Distribution  |  Load   |
+---------+                       +---------+
      |                               |
      |    Request Distribution       |
+---------+   Based on Metrics   +---------+
| Server  | <------------------- | Server  |
|    A    |                      |    B    |
+---------+                      +---------+

The load balancer adjusts request distribution based on real-time performance metrics to optimize resource utilization.

49. Health Check Pattern

The Health Check Pattern involves regularly monitoring the health and status of services or components to detect failures and take corrective actions, such as restarting or redirecting traffic.

When to use it?

This pattern is valuable for ensuring system availability and reliability. It helps detect and address issues before they impact users, improving overall system robustness.

+---------+   Health Check   +---------+
|  Service| <-------------- | Monitor |
|    A    |   Status/Health |  Tool   |
+---------+                  +---------+
      |                           |
      |    Health Status         |
+---------+   Action (Restart)   +---------+
|  Service| -------------------> | Service |
|    A    |                      |   B     |
+---------+                      +---------+

Health checks are used to monitor services and take corrective actions if issues are detected.

50. Hot Standby Pattern

The Hot Standby Pattern involves having a backup service or system that is fully operational and kept synchronized with the primary service. In case of a failure, the backup can take over with minimal disruption.

When to use it?

This pattern is useful for systems requiring high availability and minimal downtime. It ensures that a backup is ready to take over immediately if the primary service fails.

+---------+   Active Service    +---------+
| Primary | -----------------> | Hot    |
| Service |  Continuous Sync   | Standby|
+---------+                    +---------+
      |                            |
      |   Failover Triggered       |
+---------+   Backup Service      +---------+
| Backup  | <------------------- |  Primary|
| Service |   Take Over          |  Service|
+---------+                    +---------+

A hot standby service is kept in sync with the primary and takes over immediately if the primary fails.

51. Data Replication Pattern

The Data Replication Pattern involves creating and maintaining copies of data across multiple locations or systems to ensure data availability and durability.

When to use it?

This pattern is useful for systems where data durability and availability are critical. It helps protect against data loss and improves access to data by distributing it across multiple locations.

+---------+    Write Data    +---------+
|  Client | ----------------> | Primary |
|         |    Replication   | Database|
+---------+                    +---------+
      |                           |
      |   Data Replication        |
+---------+   +---------+        +---------+
|  Replica|   | Replica|        | Replica|
|  Database|   | Database|        | Database|
+---------+   +---------+        +---------+

Data is replicated across multiple databases to ensure availability and durability.

52. Contextual Resilience Pattern

The Contextual Resilience Pattern involves designing systems to be resilient based on specific contexts or usage scenarios. It adapts the resilience strategies to fit the particular needs and requirements of different contexts.

When to use it?

This pattern is useful when a system has varying requirements based on its use cases or environments. It ensures that resilience strategies are tailored to the specific needs of different contexts.

+---------+   Contextual Data    +---------+
|  Service| -------------------> | Resilience|
|   A     |  Adapt Strategies   |  Manager  |
+---------+                      +---------+
      |                             |
      |  Context-Specific Resilience|
+---------+   Adjusted Resilience  +---------+
|  Context| <-------------------- | Strategy|
|  A      |                      |    1    |
+---------+                      +---------+

Resilience strategies are adapted based on specific contexts or scenarios to meet varying requirements.

53. Distributed Consensus Pattern

The Distributed Consensus Pattern ensures that multiple distributed nodes or systems agree on a single value or state, even in the presence of failures or network partitions.

When to use it?

This pattern is critical in distributed systems that require agreement or coordination among multiple nodes, such as in distributed databases or coordination services.

+---------+   Consensus Request  +---------+
| Node 1  | -------------------> | Consensus|
+---------+                    |  Service |
      |                         +---------+
      |     Consensus Response   |
+---------+   +---------+        +---------+
| Node 2  | <---| Consensus|------| Node 3 |
|         |     |  Response|      |         |
+---------+   +---------+        +---------+

Nodes participate in a consensus process to agree on a single value or state, ensuring consistency across the distributed system.

54. Compensating Transactions Pattern

The Compensating Transactions Pattern involves performing compensating actions to reverse or mitigate the effects of a previous transaction when a distributed transaction fails.

When to use it?

This pattern is useful in distributed systems where transactions span multiple services or components. It helps ensure data consistency and reliability by handling failures gracefully.

+---------+    Main Transaction  +---------+
| Service | ------------------>  | Service |
|    A    |   Perform Actions   |    B    |
+---------+                     +---------+
      |                            |
      |    Compensating Actions    |
+---------+  Compensate Failure  +---------+
| Service | <------------------  | Service |
|    A    |   Undo Actions      |    B    |
+---------+                     +---------+

Compensating actions are taken to undo or mitigate the effects of a failed transaction, maintaining system consistency.

55. Service Degradation Pattern

The Service Degradation Pattern involves intentionally reducing the level of service to maintain overall system availability when some components fail. It allows a system to continue functioning, albeit with reduced capabilities.

When to use it?

This pattern is useful when a system needs to remain operational even if some features or components are not fully functional. It helps to prioritize critical functionality while degrading non-essential features gracefully.

+---------+   Service   +---------+
|  Client | ---------->|  Service|
|         |    (Full)   |   A     |
+---------+             +---------+
       |
       |   Partial Service
       |
+---------+   Degraded   +---------+
|  Client | <------------| Service |
|         |    (Partial) |   B     |
+---------+             +---------+

The system provides partial functionality when some components fail, ensuring critical features remain available.

56. Decoupling Pattern

The Decoupling Pattern separates different components or services to reduce dependencies and increase system flexibility. It involves using asynchronous communication or message brokers to manage interactions between components.

When to use it?

This pattern is useful for increasing system resilience by reducing inter-component dependencies. It allows components to operate independently and handle failures more gracefully.

+---------+   Asynchronous   +---------+
|  Client | ---------------->| Service |
|         |   Messages/Events |   A     |
+---------+                    +---------+
     |                             |
     |   Decoupled Components       |
+---------+   +---------+         +---------+
| Service |   |  Message|         | Service |
|    B    |   |  Broker |         |   C     |
+---------+   +---------+         +---------+

Components communicate through asynchronous messages or events, reducing direct dependencies and increasing resilience.

57. Data Sharding Pattern

The Data Sharding Pattern involves dividing a large dataset into smaller, more manageable pieces (shards) to improve performance, scalability, and fault tolerance. Each shard is managed separately, allowing for parallel processing and distributed load.

When to use it?

This pattern is useful for systems with large datasets or high load, where distributing data across multiple nodes can improve performance and scalability. It helps to manage large volumes of data more effectively.

+---------+   Data Shards    +---------+
|  Client | ---------------> | Shard   |
|         |    Request       | Manager |
+---------+                  +---------+
      |                           |
      |   Data Distribution       |
+---------+   +---------+         +---------+
|  Shard 1 |   | Shard 2 |         | Shard 3 |
|         |   |         |         |         |
+---------+   +---------+         +---------+

Data is divided into shards, with each shard managed separately to improve performance and scalability.

58. Distributed Snapshot Pattern

The Distributed Snapshot Pattern involves taking consistent snapshots of a distributed system’s state at a given point in time. It helps to capture and restore the system’s state for recovery or analysis.

When to use it?

This pattern is useful for systems that require consistent backups or state captures. It helps to ensure that the entire system’s state can be restored to a known good point in case of failure.

+---------+    Snapshot         +---------+
|  Client |  ----------------> | Snapshot |
|         |  Request            |  Service |
+---------+                      +---------+
      |                            |
      |   Consistent Snapshots     |
+---------+   +---------+         +---------+
|  Data   |   |  Data   |         |  Data   |
| Store 1 |   | Store 2|         | Store 3 |
+---------+   +---------+         +---------+

Snapshots of the distributed system’s state are captured to enable consistent recovery or analysis.

59. Compartmentalization Pattern

The Compartmentalization Pattern involves dividing a system into isolated compartments or segments to limit the impact of failures and improve security. Each compartment operates independently and interacts with others through well-defined interfaces.

When to use it?

This pattern is useful for improving fault isolation, security, and manageability. It helps to prevent failures in one compartment from affecting others and enhances system resilience.

+---------+   Isolated    +---------+
|  Client |  ------------> | Compartment|
|         |   Request      |    A      |
+---------+                   +---------+
      |                            |
      |   Inter-Compartment        |
+---------+   +---------+         +---------+
| Compartment|   | Compartment|   | Compartment|
|    B       |   |    C       |   |    D       |
+---------+   +---------+         +---------+

The system is divided into isolated compartments that operate independently and communicate through defined interfaces.

60. State Transfer Pattern

The State Transfer Pattern involves transferring the state of a system from one component to another, usually to ensure continuity or enable failover. It ensures that the new component or service can pick up where the previous one left off.

When to use it?

This pattern is useful when migrating or upgrading systems, or during failover scenarios. It helps to maintain continuity and minimize disruption by preserving the system’s state.

+---------+   State Transfer   +---------+
|  Client |  -----------------> | Service |
|         |     State Data       |   A     |
+---------+                    +---------+
      |                             |
      |   New Service                |
+---------+   +---------+           +---------+
|  Client |   |  Service|           | Service |
|         |   |   B     |           |   A     |
+---------+   +---------+           +---------+

State data is transferred from one service to another to ensure continuity during migration or failover.

61. Eventual Consistency Pattern

The Eventual Consistency Pattern allows a system to be in an inconsistent state temporarily, with the guarantee that it will eventually reach a consistent state. It is often used in distributed systems where strict consistency is not feasible.

When to use it?

This pattern is useful in distributed systems where immediate consistency is challenging. It helps to improve availability and scalability by allowing temporary inconsistencies.

+---------+   Eventual Consistency  +---------+
|  Client |  --------------------> |  Service|
|         |    (Temporary State)    |   A     |
+---------+                      +---------+
      |                            |
      |  Consistency Updates        |
+---------+   +---------+          +---------+
|  Client |   | Service |          | Service |
|         |   |   B     |          |   A     |
+---------+   +---------+          +---------+

The system allows temporary inconsistencies, with the assurance that it will eventually become consistent.

62. Immutable Data Pattern

The Immutable Data Pattern involves using immutable data structures that cannot be modified once created. Instead of updating existing data, new versions are created. This pattern simplifies concurrency control and improves system reliability.

When to use it?

This pattern is useful for systems that require strong consistency and concurrency control. It helps to prevent data corruption and simplifies reasoning about data changes.

+---------+   Immutable Data  +---------+
|  Client |  ----------------|  Service|
|         |   (No Updates)    |   A     |
+---------+                    +---------+
      |
      |   New Data Version
      |
+---------+   +---------+    +---------+
|  Data   |   |  Data   |    |  Data   |
|  V1     |   |  V2     |    |  V3     |
+---------+   +---------+    +---------+

Data is immutable, and new versions are created instead of modifying existing data.

63. Health Monitoring Pattern

The Health Monitoring Pattern involves continuously monitoring the health of components and services within a system. It provides real-time insights into the system’s status and helps detect issues before they impact users.

When to use it?

This pattern is useful for proactive failure detection and management. It helps to identify and address issues early, improving overall system reliability and performance.

+---------+   Health Checks    +---------+
|  Client |  -----------------> |  Service|
|         |   (Monitoring)      |   A     |
+---------+                    +---------+
      |                            |
      |  Health Data               |
+---------+   +---------+          +---------+
| Monitoring|  | Service |          | Service |
|   System  |  |  B      |          |   A     |
+---------+   +---------+          +---------+

Continuous health monitoring provides insights into the status of services and helps in proactive issue management.

64. Dependency Injection Pattern

The Dependency Injection Pattern involves injecting dependencies into a component rather than having the component create or manage its dependencies. It promotes loose coupling and enhances testability and flexibility.

When to use it?

This pattern is useful for improving modularity and making components easier to test and maintain. It helps to manage dependencies more effectively and supports better separation of concerns.

+---------+   Dependency    +---------+
|  Client |  ----------->  |  Service|
|         |    Injection     |   A     |
+---------+                  +---------+
      |
      |   Injected Dependency
      |
+---------+   +---------+   +---------+
|  Service |   |  Service|   |  Service|
|   B      |   |   C     |   |   D     |
+---------+   +---------+   +---------+

Dependencies are injected into components, promoting loose coupling and improving flexibility.

65. Snapshot Isolation Pattern

The Snapshot Isolation Pattern involves providing a consistent view of data by using snapshots. This ensures that transactions see a consistent state of the data, even if other transactions are modifying it.

When to use it?

This pattern is useful for systems that require consistent reads while allowing concurrent writes. It helps to maintain data integrity and prevent anomalies in transactional systems.

+---------+   Snapshot View   +---------+
|  Client |  ---------------->|  Database|
|         |   (Consistent)    |   A     |
+---------+                    +---------+
      |
      |   Snapshot Reads
      |
+---------+   +---------+   +---------+
|  Snapshot|   |  Snapshot|   | Snapshot|
|   1      |   |   2      |   |   3     |
+---------+   +---------+   +---------+

Transactions see a consistent snapshot of the data, even if other transactions are in progress.

66. Retry-After Pattern

The Retry-After Pattern involves responding to a request with a Retry-After header indicating when the client should attempt the request again. This helps clients manage retry attempts and avoid overwhelming the server.

When to use it?

This pattern is useful for managing client-side retry behavior in response to rate limiting or temporary service unavailability. It helps to provide guidance on when to retry a request.

+---------+    Initial Request    +---------+
|  Client |  -------------------->| Service |
|         |    (Rate Limited)     |   A     |
+---------+                      +---------+
      |                             |
      |   Retry-After Header        |
+---------+   +---------+           +---------+
|  Client |   |  Delay  |           | Service |
|         |   |  Period |           |   A     |
+---------+   +---------+           +---------+
      |
      |    Retry Request
+---------+   +---------+
|  Client |   |  Service|
|         |   |   A     |
+---------+   +---------+

The server provides a Retry-After header to guide the client on when to retry the request.

67. Partitioning Pattern

The Partitioning Pattern involves dividing a system into smaller, manageable parts or partitions. This can help in scaling, isolating failures, and improving performance by distributing the load.

When to use it?

This pattern is useful for scaling systems and managing large datasets. It helps to distribute load, improve performance, and isolate failures within partitions.

+---------+   Partition 1   +---------+
|  Client |  -------------> |  Service|
|         |    (Partition)   |   A     |
+---------+                   +---------+
      |                             |
      |   Partition 2               |
+---------+   +---------+           +---------+
| Service |   | Service|           | Service |
|   B     |   |   C    |           |   D     |
+---------+   +---------+           +---------+

The system is divided into partitions, each handling a portion of the load or data.

68. Defer Pattern

The Defer Pattern involves postponing the execution of a task or operation until a more appropriate time. This can help in managing load, handling transient failures, or improving user experience.

When to use it?

This pattern is useful for deferring non-critical operations, reducing immediate load, or managing resources more efficiently.

+---------+   Initial Request   +---------+
|  Client |  -----------------> | Service |
|         |    (Defer)          |   A     |
+---------+                    +---------+
      |                            |
      |  Deferred Operation        |
+---------+   +---------+         +---------+
| Service |   |  Task   |         | Service |
|   B     |   |  Queue  |         |   A     |
+---------+   +---------+         +---------+

Tasks or operations are deferred and managed through a queue or scheduler.

69. Compensating Action Pattern

The Compensating Action Pattern involves taking corrective actions to undo or mitigate the effects of a failed operation. It helps to ensure consistency and recover from failures.

When to use it?

This pattern is useful for managing failures in distributed transactions or processes. It helps to ensure that the system remains consistent by compensating for errors or failures.

+---------+   Initial Action    +---------+
|  Client |  ----------------> |  Service|
|         |     (Success)       |   A     |
+---------+                    +---------+
      |                             |
      |  Compensating Action        |
+---------+   +---------+          +---------+
| Service |   |  Rollback |         | Service |
|   B     |   |   Action  |         |   A     |
+---------+   +---------+          +---------+

If the initial action fails, compensating actions are taken to restore consistency.

70. Queue-Based Load Leveling Pattern

The Queue-Based Load Leveling Pattern involves using a queue to buffer and manage requests or tasks. This helps to smooth out load spikes and ensure that the system can handle bursts of traffic.

When to use it?

This pattern is useful for managing variable workloads and smoothing out spikes in traffic. It helps to prevent overload and ensure system stability.

+---------+   Request Queue    +---------+
|  Client |  ---------------->|  Queue  |
|         |     (Buffer)       |         |
+---------+                    +---------+
      |                             |
      |   Process Requests           |
+---------+   +---------+          +---------+
|  Service |   | Service|          | Service |
|   A      |   |   B    |          |   C     |
+---------+   +---------+          +---------+

Requests are buffered in a queue and processed at a controlled rate.

71. Message Broker Pattern

The Message Broker Pattern involves using a message broker to handle communication between services. It decouples components, enabling asynchronous communication and improving system flexibility and resilience.

When to use it?

This pattern is useful for decoupling services, enabling asynchronous communication, and improving scalability and fault tolerance.

+---------+   Publish Message   +---------+
|  Client |  -----------------> | Broker  |
|         |   (Message)         |         |
+---------+                    +---------+
      |                            |
      |  Distribute Messages        |
+---------+   +---------+          +---------+
| Service |   | Service|          | Service |
|   A     |   |   B    |          |   C     |
+---------+   +---------+          +---------+

A message broker manages and distributes messages between services, enabling decoupled and asynchronous communication.

72. Service Proxy Pattern

The Service Proxy Pattern involves using a proxy to handle requests and responses between clients and services. The proxy can add functionality such as caching, load balancing, or security.

When to use it?

This pattern is useful for adding additional functionality or abstraction to service communication, such as caching or load balancing, without modifying the underlying services.

+---------+   Proxy Service    +---------+
|  Client |  ----------------> | Proxy   |
|         |   (Request)        | Service |
+---------+                    +---------+
      |                            |
      |  Forward Request           |
+---------+   +---------+          +---------+
|  Service |   |  Service|          | Service |
|   A      |   |   B    |          |   C     |
+---------+   +---------+          +---------+

A proxy service manages requests and responses, adding additional functionality such as caching or load balancing.

Conclusion:

Resilience patterns are key to building fault-tolerant systems that can recover from and gracefully handle failures. By incorporating patterns like retries, circuit breakers, and failover mechanisms, you ensure your system remains robust, reliable, and available under all circumstances.

Understanding these patterns — and knowing when to apply them — will help you design systems that can stand the test of time and failure. If you’re working with distributed systems, mastering these patterns is a must.

Leave a comment

Your email address will not be published. Required fields are marked *.