Distributed Transactions in Practice: Reliable Orders with Outbox, TCC, and Saga

A practical Java order workflow comparing local transactions, the transactional outbox, TCC, and Saga for inventory, payments, rewards, and refunds.

An order can fail long after the customer clicks “Pay.”

Inventory may have been reserved while the payment request times out. The payment provider may have charged the customer while the order service is restarting. A rewards service may receive the same payment event twice. A refund may finish successfully even though the customer has already spent the points earned from the original purchase.

Distributed transactions are therefore a business design problem as much as a technical one. A coordinator can move a workflow forward, but it cannot decide whether a paid order should be canceled, whether shipped inventory should be restored, or how to account for rewards that have already been spent.

The reference workflow below keeps four services separate: orders, inventory, payments, and rewards. A third-party payment provider sits outside the system. I use the workflow to answer four implementation questions: where a local transaction ends, how a retry is recognized, how a committed change becomes an event, and how a refund remains recoverable.

All identifiers, amounts, and timestamps are illustrative. They make the failure paths testable without presenting invented examples as production measurements.

Start with the promises the order must keep

Before selecting a transaction framework, define what customers and operators should be able to rely on.

For this order workflow, the important promises are:

  • An order becomes ready for fulfillment only after payment is confirmed and inventory is committed.
  • Repeating a request must not create another charge, inventory reservation, or rewards grant.
  • Points and promotional coupons may arrive after payment, but unfinished work must remain visible and recoverable.
  • A refund remains in progress until the payment provider confirms its outcome.
  • Compensation follows the actual business state, including shipment, returns, and rewards already used.

These promises require different consistency boundaries.

Creating an inventory reservation and reducing available stock belong in one inventory database transaction. Recording a confirmed payment and the event announcing it belong in one payment database transaction. Sending that event to other services happens afterward.

Across services, the order needs explicit intermediate states. A useful starting point is:

PAYMENT_CONFIRMED and READY_FOR_FULFILLMENT represent different business facts. If the customer has paid but inventory confirmation is still being retried, the system needs to show that situation accurately.

A single SUCCESS flag would hide the difference between “money received” and “goods can be shipped.” Keeping those states separate also gives support staff something useful to investigate when an order stops progressing.

The same principle applies to failure. A timeout is an observation about communication, not a reliable statement about the remote business outcome. Recovery should use a stable operation identifier to discover or complete that outcome.

Reserve inventory before sending the customer to payment

Consider a product with one unit remaining. Two customers can both read an available quantity of one before either request updates the row. A local transaction around each request does not make an application-level “read, check, write” sequence safe by itself.

For a simple reservation, the inventory service can combine the availability condition and the update:

UPDATE inventory
SET available = available - :quantity,
    reserved  = reserved + :quantity
WHERE sku_id = :sku_id
  AND available >= :quantity;

The service must validate that quantity is positive and check the affected-row count. If no row was updated, it must reject the reservation and roll back the surrounding transaction. An update affecting zero rows is not automatically a database error.

The reservation record belongs in that same transaction. It identifies the order, product, quantity, and reservation status, with a unique business key such as (order_id, sku_id).

That record matters when the first request succeeds but its response is lost. A retry with the same order identifier should find the existing reservation and return its status. It should not deduct available stock again. The unique constraint must enforce this under concurrent retries, with a conflicting attempt rolling back its inventory changes.

For orders containing several products, the service should either reserve all required lines or roll back the attempt. Accessing inventory rows in a consistent order can also reduce deadlock opportunities.

Once the local transaction commits, database locks can be released. The reservation remains as business data while the customer completes payment.

This reservation model follows the business-level idea behind TCC: Try reserves the resource, Confirm consumes the reservation, and Cancel releases it. In a typical XA-backed two-phase commit, a prepared database transaction may retain locks while awaiting the coordinator’s decision. Here, the reservation is committed as ordinary business data, so database locks do not need to remain held throughout checkout.

That moves responsibility into application design. A complete TCC implementation needs durable branch states, repeatable Confirm and Cancel operations, and protection against a delayed Try arriving after cancellation. The workflow described here borrows that resource model; three endpoints alone would not make it a complete TCC protocol. These failure cases are also addressed by mechanisms such as Seata’s TCC transaction fence.

The difficult case is a payment confirmation arriving at the same time as reservation expiry.

A cleanup job should not release stock merely because a timer fired. It needs to coordinate cancellation with the order and payment state. Inside the inventory service, confirmation and release should compete through a conditional transition from RESERVED, with the inventory adjustment performed in the same local transaction.

If release has already won when payment is confirmed, the order cannot simply become ready for fulfillment. It needs an explicit recovery decision: attempt another reservation or initiate a refund.

That decision is part of the checkout contract. A lock alone cannot supply it.

Commit the event with the business change

After payment is confirmed, several actions may follow: commit the inventory reservation, update the order, grant points, issue a promotional coupon, and send a notification.

Putting all of them inside one synchronous request makes payment completion depend on every downstream service. A rewards outage could then make a successful purchase appear unsuccessful.

Instead, the payment service can commit two things together:

An independent relay reads committed outbox records and publishes them to the message broker. If publishing fails, the event remains available for retry.

The important boundary is that the outbox record is written inside the local transaction; message delivery happens after commit. This closes the gap where the database commits but the application crashes before sending the event. It is the central idea of the transactional outbox pattern.

A payment event might contain:

{
  "eventId": "evt-payment-4821",
  "eventType": "PaymentConfirmed",
  "schemaVersion": 1,
  "paymentId": "pay-4821",
  "orderId": "order-7308",
  "userId": "user-204",
  "amountMinor": 12900,
  "currency": "USD",
  "occurredAt": "2026-09-08T09:30:00Z"
}

The event identifier supports delivery deduplication. The payment and order identifiers support business reconciliation. They serve different purposes.

The relay can still publish twice. For example, the broker accepts the message, but the relay crashes before marking the outbox record as sent. The next attempt sends the same event again.

Consumers therefore need both delivery-level and business-level protection.

The points service should record message consumption and its points-ledger change in one local transaction. A unique key on the rewards action—such as (order_id, reward_type)—also prevents a second grant if the same business action is accidentally announced under a different event identifier.

A preliminary “does this message exist?” query is only an optimization. Concurrent consumers can both pass that check. Database uniqueness constraints and transaction rollback must enforce the actual guarantee.

For promotional coupons, the same idea becomes a unique issuance key such as (order_id, campaign_id). Retrying delivery should return or rediscover the existing coupon.

At this point, “eventually consistent” has a testable business meaning. The payment service can already show a confirmed payment while the rewards service still shows points as pending. That temporary difference is acceptable because rewards are not a prerequisite for fulfillment.

Convergence depends on the recovery mechanisms continuing to work: events must remain available, retries must resume, and consumers must process them safely. A permanently invalid event will not become valid simply because it is retried. Such failures need an alert and a repair path.

For this workflow, “eventually consistent” means that unfinished rewards remain traceable and recoverable. It should not mean that the team assumes everything will eventually fix itself.

There is another distinction worth making: a coupon issued after purchase can be asynchronous. A coupon used to reduce the checkout price affects what the customer owes and needs validation and appropriate reservation before payment. Both are called “coupons,” but they belong to different consistency boundaries.

If the platform already uses RocketMQ, transaction messages offer another way to coordinate local transaction outcomes with message publication. They require a reliable local transaction status check and still leave downstream processing, retries, and idempotency to consumers. They do not turn the entire order workflow into one atomic transaction.

A local transaction should protect one business fact

Each service should commit the fact it owns before asking another service to react. For payment confirmation, that usually means updating the payment record and writing the outbox row in the same database transaction:

@Transactional
public void confirmPayment(Payment payment) {
    payment.confirm();
    paymentRepository.save(payment);
    outboxRepository.insert(
        OutboxMessage.forAggregate(
            payment.orderId(),
            "PaymentConfirmed",
            payment.version()));
}

This is illustrative Spring-style code, not a complete library listing. The transaction commits the payment fact and the message to be relayed; the relay still runs after commit and may publish more than once. A unique key such as (aggregate_id, aggregate_version) prevents the same business revision from creating two outbox rows. See the Spring transaction documentation for the annotation contract.

Treat refunds as a durable workflow

Refunds reveal why distributed recovery cannot be reduced to reversing SQL statements.

A refund can involve the payment provider, inventory, points, coupons, and the order record. Some actions are reversible immediately; others depend on what happened after purchase.

A practical refund workflow starts by creating a durable refund request with a stable refund_id. The payment adapter uses that identifier as an idempotency key where the provider supports it.

If the provider call times out, the request remains unresolved. The system queries the existing refund or retries using the same identifier. Creating a new identifier on every retry risks creating a second financial operation.

A minimal state model could be:

A communication timeout normally leaves the request in PROCESSING until its outcome is established. A definitive provider rejection may move it to FAILED.

Once the monetary refund is confirmed, the payment service records that outcome and an outbox event in the same local transaction. Other services then apply their own adjustments.

Inventory restoration depends on fulfillment state. An unshipped order may allow committed inventory to be returned to availability. A shipped product should generally pass through a return and inspection process before becoming sellable again. A refund event alone does not prove that the warehouse has usable stock.

Points also need a business rule. If a purchase earned 100 points and the customer already spent 80, deleting the original grant would erase useful accounting history. A ledger should preserve the grant and record an adjustment. The policy might permit a negative balance, offset future earnings, or route the case for review.

Promotional coupons present a similar problem. An unused coupon may be revoked. A coupon already redeemed on another order cannot be handled by pretending the redemption never happened.

Partial refunds add another layer: the system must track how much has already been refunded and adjusted, rather than treating every refund event as a request to reverse the whole order.

An orchestrated Saga can coordinate this workflow. The coordinator records progress, requests actions, and tracks unfinished work. Each participant commits a local transaction, while compensation follows explicit business rules.

A Saga also changes the isolation model. Each completed step commits locally, so other requests may observe its result before the whole workflow finishes. A customer could spend newly granted points while a refund is being processed, or a fulfillment worker could attempt to ship an order whose cancellation has started. Compensation does not automatically prevent these interactions.

The services need business-level guards around conflicting transitions. For example, fulfillment can reject a shipment transition once the order enters an agreed cancellation state, while rewards can apply a defined adjustment policy instead of deleting an earlier grant. These guards must be enforced atomically within the service that owns the relevant state.

The Saga coordinates recovery across services; local transactions and state-transition rules protect each participant while that recovery is in progress.

The monetary outcome should remain visible even if a later rewards adjustment fails. An operator should be able to see “refund succeeded; points adjustment pending” instead of a misleading generic “refund failed.”

Design the failure paths before choosing the framework

With those boundaries written down, the framework choice becomes a smaller decision.

Short transactions across a small number of controlled, XA-capable resources may justify XA coordination. Inventory with a clear reservation model may fit TCC. Refunds and other long-running processes often fit a Saga. Post-payment rewards and notifications are natural candidates for outbox-based events.

Seata can provide infrastructure for several of these approaches, but selecting a mode does not resolve the business questions. A framework cannot determine whether a used coupon should be revoked or whether a late payment should trigger a new inventory reservation.

Before implementation is considered complete, the following failure cases deserve deliberate testing:

FailureExpected behavior
Inventory reservation succeeds, but the response is lostRetry discovers the same reservation without reserving again.
Payment is confirmed, but event publication is delayedThe committed outbox event remains pending and can be retried.
The same payment event reaches rewards twiceOnly one points grant and one campaign coupon are created.
Payment confirmation races with reservation releaseOne reservation transition wins; the order enters an explicit recovery path if necessary.
A refund request times outThe system checks the existing refund using its stable identifier.
A refund succeeds, but rewards adjustment failsThe refund remains successful while the adjustment is retried or reviewed.

Operational visibility should follow the same model. Useful signals include the age of the oldest unpublished event, orders stuck between payment and fulfillment, unresolved refund requests, and compensation steps exceeding their retry window.

Logging order_id, payment_id, refund_id, and event_id consistently helps connect the workflow across services. An error count tells the team that something broke; a durable state record tells them what remains to be done.

Review the workflow with these five questions

  • What single business fact does each local transaction commit?
  • Which stable identifier makes a retry find the original operation?
  • Can a consumer apply the same event twice without a second charge or grant?
  • Which state transition wins when timeout, cancellation, and confirmation race?
  • Where can an operator see work that is still pending?

For this order system, reliability comes from a small set of concrete decisions: reserve inventory with an atomic local update, preserve events alongside committed business changes, make repeated operations safe, and represent refunds as recoverable workflows.

Each decision answers a failure the system can actually experience. Together, they provide a more useful design than choosing one distributed transaction mechanism and expecting every business process to fit inside it.

Leave a Reply

Your email address will not be published. Required fields are marked *