Reliable Messaging in Java Services: From Payment Events to Correct Reward Balances

A customer pays for an order, but the promised reward points never appear. Another customer receives the same reward twice. Both problems can happen even when the message broker is available and persistence is enabled.

The first failure might occur before a message reaches the broker: the payment transaction commits, and the application stops before publishing its event. The second might occur after successful processing: the points service commits its database transaction, then crashes before acknowledging the message.

Reliable messaging has to cover both ends of that journey. Broker durability matters, but it cannot decide whether a payment event should exist or whether a reward has already been granted.

This article follows a Java service workflow in which a confirmed payment earns points. The examples describe a concrete design rather than measured production incidents. The goal is to make each interrupted operation recoverable while keeping the business result correct across retries, delayed delivery, and historical replay.

Commit the payment fact and its event together

Consider a payment handler that updates MySQL and then publishes OrderPaid. If the application stops between those operations, the database contains a paid order with no durable record that an event still needs to be sent.

Publishing first creates the opposite problem: a consumer may grant points for a payment transition that later rolls back.

A transactional outbox closes this gap by recording the accepted business change and its event within one local database transaction. However, simply placing an UPDATE and an INSERT between BEGIN and COMMIT is not enough. The application must verify that the intended state transition actually occurred.

For example:

UPDATE payment_order
SET status = 'PAID',
    version = version + 1
WHERE payment_id = :paymentId
  AND status = 'UNPAID';

An affected-row count of zero is not a SQL error. It might mean the payment was already processed, the record does not exist, or the current state does not allow this transition.

The handler therefore needs to distinguish those outcomes before inserting an event. A repeated notification for the same accepted payment should resolve to the existing result. An incompatible state should enter an explicit investigation or reconciliation path. Only an accepted transition should produce its corresponding event.

Before applying the transition, the handler validates the notification against the expected business record, including the provider reference, amount, currency, and payment identity. Delivery of a callback is not itself proof that the application should mark an arbitrary order as paid.

The event should have a stable application identifier. A uniqueness constraint representing the payment transition can prevent two callback handlers from creating separate events for the same fact. If either the transition or event insertion fails, the transaction must roll back.

The outbox must also belong to the same transaction-capable database boundary as the payment record. Placing it in a separate central database would reintroduce the cross-system write problem.

After commit, delivery becomes recoverable. A polling relay claims pending records, sends them, and updates their status after the required broker confirmation. A CDC-based relay instead follows committed database changes and maintains its own progress. These approaches have different operational details and should not be described as the same scanning mechanism.

Neither eliminates duplicates. If publication succeeds but recording SENT fails, the relay may publish again. Its job is to preserve the opportunity to deliver; consumers must tolerate the uncertainty.

Figure 1. Two separate local transactions are connected by recoverable message delivery. The broker sits between them; it does not make both databases part of one atomic transaction.

Know what the broker’s confirmation actually proves

A producer call can finish before the broker has accepted the message, depending on the client API and send mode. Reliable publishing therefore needs to inspect the asynchronous result, callback, or confirmation rather than treating successful submission to a client buffer as successful delivery.

Even then, a timeout can leave the outcome unknown. The broker may have accepted the message while its response was lost. A retry should retain the same application event identity.

Broker features provide different guarantees:

MechanismWhat it helps establishWhat it does not establish
Producer confirmationThe configured broker acceptance condition was metThe reward transaction committed
Replication and persistenceAccepted messages survive specified failuresEvery business event was generated correctly
Consumer acknowledgment or offset commitProcessing progress is recordedExternal business effects cannot repeat
Broker transactionOperations within its supported transaction boundary are coordinatedArbitrary MySQL updates join that boundary automatically

For Kafka, acks=all refers to acknowledgment by the current in-sync replicas. With a topic configured for three replicas and min.insync.replicas=2, writes using this acknowledgment policy require a sufficient ISR; the producer does not simply wait for a fixed number of arbitrary replicas. Producer idempotence addresses supported producer retry behavior, not every new application-level send of the same business event.

For RabbitMQ, publisher confirms and consumer acknowledgments cover separate interactions. A publisher also needs to handle routing failures: confirmation alone does not prove delivery to the intended queue. Depending on the publishing contract, mandatory and returned-message handling can expose unroutable messages. A durable exchange preserves routing infrastructure; the exchange itself is not message storage.

RocketMQ transaction messages coordinate producer-side local transaction outcomes with message visibility through transaction processing and status checks. They do not mean that the downstream points database has already been updated.

This is also the right way to discuss exactly-once processing. A guarantee can be strong within a defined system boundary while requiring additional coordination outside it. The points service still needs a durable rule for deciding whether a reward has already taken effect.

Make the reward transaction idempotent at the business level

Suppose event E90001 represents a confirmed payment that earns 100 points. The consumer can record the event identifier and update the points database in one local transaction.

That protects against redelivery of E90001. It does not protect against a producer defect or a repair tool generating E90002 for the same payment.

The business needs two identities:

IdentityPurpose
Event identifierRecognize repeated delivery of one event
Reward operation keyPrevent repeated execution of the same reward entitlement

A reward key might identify a payment and a reward program, provided that combination matches the business model. If one payment can legitimately earn several rewards, the key must distinguish those entitlements without changing across retries.

For this design, the points ledger enforces uniqueness on that stable reward key. The consumer inbox separately records event processing. Both records and the balance change belong to the same local transaction.

The consumer first inserts its inbox record, then inserts the unique reward ledger entry and updates the balance. It verifies the expected write results before committing. Only after commit does it acknowledge delivery.

The affected-row checks matter. If the points account does not exist, an UPDATE can affect zero rows without raising a database error. The service must either initialize the account through a safe, defined path or fail the transaction. It must not commit a “processed” record while granting no points.

Duplicate handling also needs precision. Only the expected uniqueness conflict should be interpreted as evidence of an existing operation. A connection failure, unrelated constraint violation, or malformed record is not proof that the reward succeeded.

If the transaction encounters a duplicate, roll back the attempted work and inspect the existing committed record. A matching reward key with the same intended effect can resolve to success. The same key carrying a different customer or point amount is a conflict that needs investigation, not a duplicate to ignore.

Acknowledgment happens after the business transaction commits. If the consumer stops before acknowledgment, redelivery finds the durable operation and avoids granting again.

Figure 2. A crash after commit can cause another delivery without causing another reward. The duplicate path resolves against durable records; the original grant is protected by a database uniqueness constraint.

The ledger provides a stronger long-term boundary than relying only on a consumer-group name. Renaming a consumer, rebuilding a projection, or republishing an event should not authorize another reward for the same entitlement.

Preserve the reward decision across retries and refunds

A reliable transport can still produce the wrong result if the consumer recalculates business rules differently on each attempt.

Imagine that a payment earns 100 points on Monday. Processing fails, and the message is replayed on Thursday after a promotion changes the reward to 200 points. A consumer that always applies today’s rule may grant a different amount for the same historical payment.

The workflow needs a reproducible decision. It can record the determined reward amount, or retain the inputs and rule version needed to calculate it. If the points service owns that decision, its first accepted processing should persist the result so later attempts resolve to the same operation.

A reward record might include:

{
  "rewardKey": "payment-P1042-base-reward",
  "paymentId": "P1042",
  "customerId": 10086,
  "points": 100,
  "ruleVersion": "base-reward-v3"
}

The exact fields depend on ownership. The important requirement is that retrying delivery does not silently renegotiate the reward.

Refunds introduce another operation. Reversing a reward should create a separate ledger entry linked to the original grant, with its own stable idempotency key. Deleting the original inbox record or grant to “allow reprocessing” destroys the history needed to reason about the account.

Partial refunds need an allocation rule: the reversed amount may depend on refunded items, eligible spend, rounding, and previous reversals. A unique refund-operation key prevents repeated delivery from applying the same reversal twice, while business constraints prevent cumulative reversal beyond the allowed amount.

The account policy must also address points already spent. A negative balance, future deductions, or manual resolution may be appropriate depending on the product. Messaging infrastructure cannot select that policy.

These choices make the ledger useful for both recovery and explanation. Support staff can see which payment granted points, which rule applied, and which refund reversed part of the grant.

Separate latest-state projections from operations that must execute

Ordering requirements depend on what the consumer does.

An order-summary projection may accept a newer complete snapshot and ignore an older one:

UPDATE order_snapshot
SET status = :status,
    version = :incomingVersion
WHERE order_id = :orderId
  AND version < :incomingVersion;

Under a suitable snapshot contract, this prevents an old PAID state from overwriting a newer SHIPPED state. It does not prove that every intermediate event was processed. The projection also needs a defined insertion path when no row exists.

That distinction is critical for points. If version 3 arrives before version 2 and the consumer simply keeps the largest version, it may skip a reward operation represented by version 2.

A projection that displays current state can have different recovery rules from a consumer that must apply every grant and reversal. Version-based replacement is appropriate only when the newer representation contains everything that consumer needs.

Routing related events to one partition or ordered queue can help, but producer behavior, consumer concurrency, and retry handling must preserve the intended execution order. Sending failed records to another retry stream can also allow later events to proceed.

For example, a refund reversal might arrive before the original grant is available. The consumer can retain a pending operation and resolve the dependency from authoritative records, or use another explicitly defined policy. It should not silently discard the reversal because the prerequisite is temporarily missing.

Likewise, a numerical version gap means a missing event only when the consumer’s contract includes every consecutive revision. A subscriber interested in selected event types cannot assume that all skipped numbers represent delivery failures.

Recover with bounded retries and business reconciliation

Transient database errors can justify retries. Invalid payloads and unsupported rule versions usually require correction. A third-party timeout may require a status lookup using a stable operation identifier because the external effect could already have occurred.

Retry policies should distinguish these cases and bound attempts, elapsed time, and concurrent work. Backoff with jitter reduces synchronized retries, but the service still needs an overall downstream capacity budget.

A dead-letter queue marks an unresolved operation. It should retain enough context to investigate the event, failure, and attempted processing, with access controls appropriate to the payload. Replay should follow root-cause correction and begin in small batches. It must preserve the business operation key even if the replay mechanism uses a new transport identifier.

The replay horizon also affects deduplication retention. If inbox records are deleted after a week but operators can replay events from six months ago, event-level deduplication no longer protects that history. Business ledger uniqueness and retention policies need to cover the recovery procedures the team actually supports.

Backlog recovery requires spare capacity. If producers generate 20,000 messages per second and consumers complete 5,000, the backlog grows by 15,000 per second. These are illustrative rates, but they show why expanding consumers helps only when partitions, processing logic, and downstream services can support more work.

Even matching the production rate stops growth without clearing the backlog. To drain it, sustainable consumption must exceed continuing arrivals. Recovery planning should therefore include database headroom, retry traffic, and replay work rather than treating consumer count as the only control.

Finally, reconcile business records independently of transport success. For the reward workflow, compare eligible confirmed payments with their expected grants, then compare refund obligations with reversal entries. Match stable business keys and verify customer, amount, and relevant rule information.

The comparison needs a defined cutoff and allowance for normal processing delay. A payment committed a moment ago may legitimately have no reward entry yet. Persistent gaps beyond the expected processing window deserve investigation and controlled repair.

Message counts cannot establish business equality. User balances alone cannot either: balances also reflect spending, expiration, adjustments, and refunds.

A useful failure exercise is to stop the consumer immediately after its database commit but before acknowledgment. Restart it and verify that the same reward resolves to one ledger effect. Repeat the exercise with two different event identifiers for the same reward key, a missing account, a changed reward rule, and a reversal arriving before its grant.

These checks examine the promise the customer cares about: the correct points are granted and reversed, with an understandable history.

For this workflow, reliable messaging means that a committed payment has a recoverable event, repeated delivery resolves to the same reward operation, and unresolved differences remain visible until repaired. Broker features support that design, but the durable business records determine whether it has actually succeeded.

Leave a Reply

Your email address will not be published. Required fields are marked *