Software ArchitectureJune 13, 2026·Skaitymo trukmė – 14 min.·Miracle KaluMiracle Kalu

Event-Driven Architecture in Production: Lessons from a High-Throughput Order Pipeline

Abstract visualization of flowing data through a distributed event-driven pipeline

The Event-Driven Promise

Event-driven architecture is easy to sell and hard to operate. The pitch is seductive: decouple your services, scale independently, react to changes in real time, and build systems that feel alive. The reality is a stack of trade-offs around ordering, idempotency, observability, schema evolution, and failure modes that most tutorials gloss over.

Two years ago, my team rewrote a monolithic order processor into an event-driven pipeline. The system now handles several million events per day across checkout, payment, inventory, fulfillment, and notifications. The migration drew heavily on established patterns from Martin Fowler's taxonomy of event-driven systems and on operational lessons from companies like LinkedIn, Uber, Netflix, and Slack that run event-driven platforms at internet scale. This article is the guide I wish we had before we started.

Why We Chose Events

The old system was a synchronous chain of API calls. When a customer placed an order, the web handler called the inventory service, which called the payment service, which called the fulfillment service, which called the notification service. Every hop added latency and multiplied the blast radius of a failure.

If the notification service was slow, checkout timed out. If the inventory service had a brief spike, payments failed. If fulfillment was down, the whole order was rejected even though payment had already been collected. The system was tightly coupled in ways that made operational incidents predictable and recovery slow.

Events solved the immediate coupling problem. The web handler publishes an OrderPlaced event and returns immediately. Inventory, payment, fulfillment, and notification each consume the event independently. If notification is slow, checkout does not care. If fulfillment is temporarily down, orders can still be accepted and processed later.

That was the theory. The practice required us to answer a long list of questions we had not fully considered.

The Four Patterns of Event-Driven Architecture

Martin Fowler identifies four distinct patterns that people often lump together under the label "event-driven." Understanding the differences is essential because each pattern solves a different problem and introduces a different cost structure.

Event notification is the lightest pattern. A service publishes a small message saying that something happened, and downstream services react. The event usually contains only an identifier. Consumers that need more information must query the producer. This keeps payloads small but preserves some coupling because consumers still need to know how to talk to the producer.

Event-carried state transfer includes enough data in the event for consumers to act without calling back. A CustomerAddressChanged event carries the new address. A OrderPlaced event carries the line items. This reduces coupling and improves resilience because consumers can keep working even when the producer is unavailable. The cost is larger payloads, replicated state, and eventual consistency.

Event sourcing makes the event log the system of record. Instead of storing current state in a database and publishing events as a side effect, you append events to an immutable log and derive current state by replaying them. Git is the canonical example: the working tree is derived from the commit log. Event sourcing gives you strong audit trails, temporal queries, and the ability to rebuild state, but it adds significant complexity around schema evolution, external system integration, and snapshotting.

CQRS separates the model used for writes from the model used for reads. It is not strictly about events, but it pairs naturally with event sourcing and event-carried state transfer because events can populate read-optimized projections. The write model handles commands and emits events; the read model consumes events and maintains projections tailored to specific queries.

Our pipeline uses event-carried state transfer for the core order lifecycle, event notification for lightweight triggers, and a limited form of event sourcing for an audit log. We deliberately avoided full CQRS for most services because the read and write models were not different enough to justify the split.

Event Design Is the Most Important Decision

The biggest mistake teams make is treating events as thin wrappers around database changes. They emit OrderCreated, OrderUpdated, and OrderDeleted events that mirror a CRUD table. This leaks internal state and forces consumers to reconstruct intent from a stream of diffs.

We learned to design events around business intent, not database rows. Instead of OrderUpdated, we use OrderPlaced, PaymentConfirmed, InventoryReserved, OrderShipped, and OrderCancelled. Each event name describes something that happened in the business, not something that changed in a table.

A well-designed event answers three questions:

  1. What happened? Use a past-tense verb in the event name.
  2. To what? Include the aggregate identifier and enough context to understand the subject.
  3. For what purpose? Include the data a consumer needs to act, without including everything you might someday need.
{
  "eventId": "evt_01J8XQ...",
  "eventType": "OrderPlaced",
  "aggregateId": "order_48291",
  "timestamp": "2026-06-10T14:23:11Z",
  "payload": {
    "orderId": "order_48291",
    "customerId": "cust_7721",
    "currency": "EUR",
    "total": 149.95,
    "lineItems": [
      { "sku": "SHOE-42-BLK", "quantity": 1, "unitPrice": 89.95 },
      { "sku": "SOCK-3PK", "quantity": 2, "unitPrice": 30.00 }
    ],
    "shippingAddress": {
      "country": "DE",
      "postalCode": "10115"
    }
  }
}

Notice what is not in the event. There is no internal database primary key, no version field, no status enum, and no data that only one consumer needs. The event is a contract, and every field in it is a commitment.

Choreography versus Orchestration

Once you have events, you must decide who coordinates the workflow. There are two broad approaches.

Choreography means each service reacts to events independently and emits new events when it finishes. There is no central coordinator. The order service emits OrderPlaced; inventory consumes it and emits InventoryReserved; payment consumes OrderPlaced and emits PaymentConfirmed; fulfillment waits for both InventoryReserved and PaymentConfirmed before shipping. This is highly decoupled and scales well, but the global workflow is implicit. You cannot read one piece of code and see the whole process.

Orchestration introduces a central coordinator, often called a workflow engine or saga orchestrator, that explicitly sequences the steps. The orchestrator sends commands to services and waits for responses. This makes the workflow visible and easier to reason about, but it reintroduces a central dependency and can become a bottleneck.

We use choreography for the standard order flow because the steps are well understood and the services are stable. We use orchestration for complex refund and exchange workflows because they involve branching logic, compensating actions, and timeouts that are hard to express as pure choreography. Many production systems, including those at Uber and Netflix, end up using both.

Ordering and Partitioning Are Not Optional

One of the first surprises was that event consumers do not automatically process events in the order a human would expect. If a customer places an order and then immediately cancels it, the OrderCancelled event might arrive before the OrderPlaced event in a different partition or consumer group.

We had to decide what ordering guarantees we actually needed. For payment and inventory, order matters a great deal. A cancellation must not be processed before the original placement. For notifications, order matters less. A shipping confirmation arriving before a payment confirmation is confusing but not catastrophic.

Our solution was to partition events by aggregate ID. Every event related to the same order goes to the same partition, which guarantees ordering within that order's stream. Consumers that need strict ordering subscribe to those partitions. Consumers that do not need it can process events from any partition in parallel.

function partitionKeyFor(event: OrderEvent): string {
  if (event.aggregateId.startsWith('order_')) {
    return event.aggregateId;
  }
  return event.eventType;
}

This is not free. Partitioning by order means that a single hot order cannot be parallelized. During flash sales, a few SKUs can create partition hotspots and slow down the whole pipeline. We mitigate this by separating reservation events from general order lifecycle events, so inventory reservation can scale independently.

Consumer group rebalancing is another ordering trap. When a consumer joins or leaves the group, partitions are reassigned. If a consumer has already processed some events from a partition but has not committed the offset, the newly assigned consumer may reprocess those events. This is why idempotency and at-least-once delivery semantics are non-negotiable.

Idempotency Saves Your Sanity

Events are delivered at least once. Network hiccups, consumer restarts, and retries all produce duplicates. If your consumer is not idempotent, you will ship the same order twice, charge the same card twice, or reserve the same inventory twice.

Every consumer must be idempotent by design. The simplest pattern is to track processed event IDs in a deduplication store with a TTL that matches your retention window. A more robust pattern is to make the downstream operation itself idempotent.

For example, our payment processor accepts an idempotencyKey on every charge request. We use the event ID as the key. If the same PaymentConfirmed event is processed twice, the second call returns the previously created charge instead of creating a new one.

async function handlePaymentConfirmed(event: PaymentConfirmedEvent): Promise<void> {
  const charge = await payments.charge({
    amount: event.payload.amount,
    currency: event.payload.currency,
    source: event.payload.paymentMethod,
    idempotencyKey: event.eventId,
  });

  await orderProjection.update(event.aggregateId, {
    paymentStatus: 'confirmed',
    chargeId: charge.id,
    paidAt: charge.createdAt,
  });
}

The idempotency key is not a convenience. It is the contract that makes the event consumer safe under retry. Without it, you are building a system that works most of the time and fails catastrophically under load.

Exactly-Once Semantics: Worth It or Overkill?

A growing trend in event-driven systems is the push for exactly-once processing semantics. Kafka supports idempotent producers and transactions. Flink and Kafka Streams provide exactly-once guarantees for stream processing. For financial, healthcare, and e-commerce systems, exactly-once can eliminate the need for manual reconciliation.

We considered exactly-once for payment processing but ultimately stayed with at-least-once plus idempotent consumers. The reason was operational simplicity. Exactly-once semantics require transactional producers, careful consumer configuration, and a deep understanding of how commits interact with side effects. At our scale and team size, idempotent consumers plus clear metrics were easier to reason about and debug.

The decision depends on your tolerance for duplicates and your team's operational maturity. If a duplicate event would cause a regulatory or financial problem that cannot be fixed by idempotency, exactly-once is worth the complexity. Otherwise, at-least-once with idempotent operations is usually the pragmatic choice.

Error Handling and Dead Letter Queues

In a synchronous system, a failed call usually returns an error to the caller. In an event-driven system, a failed consumer does not have a caller. The event just sits there, retrying, blocking the partition, and potentially poisoning downstream consumers.

We implemented a tiered error-handling strategy:

  1. Transient failures are retried with exponential backoff and jitter. Network timeouts, database contention, and third-party rate limits fall into this bucket.
  2. Business rule violations are parked in a quarantine queue for manual review. These are not bugs; they are cases the code explicitly does not know how to handle.
  3. Permanent failures go to a dead letter queue after a small number of retries. An alert fires, and an engineer investigates.

The retry count and backoff strategy depend on the consumer. Payment processing retries three times over thirty seconds because a card network might be briefly unavailable. Inventory reservation retries more aggressively because a sale window is short. Notification sending retries for hours because a temporary provider outage should not lose the event.

Crucially, retries must not break ordering. If a consumer fails on event N and keeps retrying, event N+1 on the same partition is blocked. We use a retry topic with a delay mechanism so that failed events are reinserted while the partition continues.

Backpressure and Flow Control

Not every consumer can keep up with the producer. During a marketing campaign, the checkout service might publish ten thousand events per second while the analytics consumer can only process two thousand. Without backpressure, the analytics consumer falls behind, memory grows, and eventually the process crashes.

The first defense is scaling. We run multiple consumer instances per consumer group, limited by partition count. If we need more throughput, we increase partitions, which requires planning because repartitioning a live topic is disruptive.

The second defense is flow control. Some consumers can skip non-critical events during overload. Our analytics consumer, for example, can sample high-frequency events when lag exceeds a threshold. Core order processing cannot sample, so it scales out instead.

The third defense is circuit breaking. If a downstream dependency like the payment processor is failing, we stop hammering it and let the retry topic buffer events. This protects the dependency and prevents the consumer from burning CPU on requests that are guaranteed to fail.

Observability Changes Everything

Debugging event-driven systems is harder than debugging synchronous ones. You cannot follow a single request through a call stack. You have to reconstruct a story from scattered logs, metrics, and traces.

We built three observability layers:

  • Distributed tracing with a correlation ID that propagates from the initial HTTP request through every event produced and consumed. A single trace shows the full lifecycle of an order.
  • Event metrics including publish rate, consume rate, lag per partition, retry rate, and dead letter rate. These are the vital signs of the pipeline.
  • Audit log of every event published and consumed, stored in an object store with a long retention window. This is invaluable for incident investigation and compliance.

The metric that saved us most often was consumer lag. A sudden spike in lag tells you a consumer is falling behind before it fully fails. We alert on lag percentile thresholds, not just on errors.

Our standard incident runbook starts with three questions: Is the producer publishing? Is the consumer keeping up? Is the dead letter queue growing? Those three metrics usually tell us which part of the pipeline is sick.

Serialization and Schema Evolution

Events live longer than services. A year after you publish an event, three new teams might consume it, and two original producers might have changed shape. If you treat event schemas as an afterthought, you create a minefield of implicit dependencies.

We started with JSON for readability and debugging, then moved performance-critical topics to Avro. A study by Confluent found that switching from JSON to Avro can reduce payload size by around seventy percent and cut latency by tens of milliseconds per event in high-frequency environments. For a pipeline processing millions of events daily, that difference matters.

We use a schema registry with forward and backward compatibility checks. Every event has a versioned schema. Producers validate events against the schema before publishing. Consumers declare which schema versions they support.

The rule we enforce is simple: additive changes only, unless every consumer has been migrated. You can add optional fields. You cannot remove fields or change the type of existing fields without a formal deprecation cycle.

# schemas/order-placed/v1.yaml
name: OrderPlaced
version: 1
type: record
fields:
  - name: orderId
    type: string
  - name: customerId
    type: string
  - name: total
    type: decimal
    scale: 2
  - name: lineItems
    type: array
    items:
      type: record
      fields:
        - name: sku
          type: string
        - name: quantity
          type: int
        - name: unitPrice
          type: decimal
          scale: 2

When we needed to add a giftMessage field, we created v2 with the new optional field, updated producers to publish v2, and gave consumers two sprints to migrate. Because the field was optional, older consumers could still read v2 events without changes. This discipline feels bureaucratic until your first breaking change. Then it feels like the only reason the system still works.

When Event-Driven Is the Wrong Choice

Event-driven architecture is not a universal upgrade. We learned to avoid it in three situations.

First, when strong consistency is required and the business cannot tolerate eventual consistency. If an action must immediately reflect in every read model, synchronous coordination is often simpler and safer.

Second, when the workflow is fundamentally sequential and every step must complete before the next begins. Forcing this into events adds complexity without decoupling anything meaningful.

Third, when the team does not have the operational maturity to run distributed systems. Event-driven systems fail in subtle ways. If you do not have good observability, runbooks, and on-call rotation, you will spend more time fighting the architecture than benefiting from it.

What We Would Do Differently

If we started again, we would design the event catalog before writing any consumer. We wrote code first and named events second, which produced inconsistent vocabulary across services. Renaming events is painful because consumers depend on the names.

We would also invest earlier in consumer lag alerting. The first time a consumer fell behind during a sale, we discovered the problem from customer complaints, not metrics. Lag alerting is now one of our most important operational signals.

Finally, we would start with a schema registry from day one. Retrofitting schema validation onto a live event stream meant coordinating deploys across teams and temporarily accepting risk we could have avoided. We also would have chosen Avro for high-throughput topics earlier instead of carrying JSON overhead until the lag became visible.

Takeaways

Event-driven architecture is a powerful tool for decoupling and scaling, but it shifts complexity from the call graph to the operational model. Understand which pattern you are using: notification, event-carried state transfer, event sourcing, or CQRS. Design events around business intent, not database rows. Choose choreography or orchestration based on workflow complexity. Enforce ordering only where it is truly needed. Make every consumer idempotent. Treat retries, dead letter queues, backpressure, and observability as first-class concerns. Govern schemas formally. And be honest about whether your problem actually needs events or whether a simpler synchronous design will serve you better.

The goal is not to use more events. The goal is to build systems that fail less often, recover faster, and remain understandable as they grow.

Dalintis:

XLinkedIn
Miracle Kalu

Autorius

Miracle Kalu

Senior Full Stack Engineer

Patiko tai, ką perskaitėte?

Svarstau vyresniojo inžinieriaus pozicijas ir techninio konsultavimo projektus. Susisiekime.

Susisiekti →

Paskelbta 2026 m. birželio 13 d. · Skaitymo trukmė – 14 min.