Distributed Locks in Java: Redis, Redlock, and ZooKeeper Through an Order Workflow

Two checkout requests reach different Java service instances while a product has one unit left. Each instance enters its own synchronized block, reads the same stock value, and creates an order. Both eventually write zero back to the database. The final inventory looks plausible, but two customers have been promised the same item.

A distributed lock can coordinate those instances, but acquiring a lock is only the beginning of the design. The client may pause, its lease may expire, or the database transaction may commit after the application has released ownership. The difficult question is what prevents an outdated request from changing business state when coordination fails.

This article uses an order workflow as a concrete design example: reserve inventory, confirm an order, grant rewards, and request a refund. The examples illustrate failure cases and implementation choices rather than benchmark results. Redis, Redlock, and ZooKeeper each have a role, but the choice starts with the business rule that must survive a retry or an outage.

Protect the inventory rule before choosing a lock

For a single inventory row in MySQL, a conditional update is often the most direct protection. Instead of reading stock into Java and writing back a calculated value, ask the database to decrement it only when enough stock remains:

UPDATE inventory
SET available = available - :quantity
WHERE product_id = :productId
  AND available >= :quantity;

Validate that quantity is positive and check the affected-row count. One affected row means the condition passed; zero means there was no matching row with sufficient availability. The inventory change and its reservation record should commit in the same local transaction. A unique reservation key, such as the order-line identifier, prevents a retried request from deducting inventory again; a duplicate must resolve to the existing outcome without committing another decrement.

This does not turn an entire order workflow into one database transaction. It gives the inventory service a clear local invariant: available stock cannot become negative through this write path. Every inventory writer must respect that rule. A Redis lock may reduce overlapping attempts, but stock correctness should not depend on every caller remembering to acquire it.

The same distinction helps elsewhere in checkout. A unique reward-ledger entry can prevent duplicate points for one payment. A conditional coupon update can allow redemption only from an available state. A refund request needs a stable provider idempotency key because a process can crash after the provider accepts the request. Serializing calls with a lock does not prevent a later retry from repeating the external effect.

A dedicated database lock table is a different mechanism. Deleting an old lock row after a timeout does not stop its previous owner from running. For data already owned by the database, transactions, uniqueness constraints, and conditional writes are usually easier to reason about than building a separate lease protocol in another table.

A Redis lock is a lease with an owner

A basic Redis acquisition sets a unique ownership value and an expiry atomically:

SET lock:order:1042 unique-acquisition-token NX PX 30000

The token identifies this acquisition. Release must atomically compare it before deleting the key; renewal needs the same ownership check. Separate SETNX and expiry commands leave a crash window. T

Consider a worker that starts generating an order document and then pauses for longer than its lease. A second worker acquires the expired lock and finishes a newer document. When the first worker resumes, it can still send a stale write. An ownership-safe unlock prevents it from deleting the second worker’s lock, but does not retract that stale write.

Lease expiry does not stop a paused worker
A acquires lease
A pauses; lease expires
B acquires and writes
A resumes with stale work
The protected resource needs its own rule for rejecting stale effects.

Renewal reduces accidental expiry during ordinary long-running work. It cannot guarantee progress during an unbounded JVM pause or a network partition. Increasing the timeout changes how long another worker waits after a crash; it does not remove the failure scenario.

For a resource that supports it, fencing provides an additional boundary. Each ownership epoch receives a monotonically increasing token. The resource atomically records the highest accepted token alongside the protected mutation and rejects lower tokens. Once it has accepted token 42 from worker B, a delayed write carrying token 41 from worker A must fail. Token ordering must remain valid across coordinator recovery; a random ownership UUID is not a fencing token.

This check belongs at the resource that performs the effect, not in a separate check-then-write call. If the destination cannot enforce fencing, use its own conditional-write or idempotency facility, or redesign the operation.

Keep the database commit inside the coordinated operation

Java code introduces a second boundary: the transaction may outlive the method’s apparent critical section. With Spring’s default proxy-based transaction management, an @Transactional method normally returns to the proxy before the proxy completes the transaction. Releasing a distributed lock in that method’s finally block can therefore happen before commit.

Moving the annotation to a private method on the same object does not fix this under the default proxy model. Self-invocation bypasses the transactional proxy. Use an externally invoked transactional service or an explicit transaction boundary.

Normal execution: acquire before the transaction, release after completion
Acquire order lock
Begin transaction
Update and commit
Release order lock
On failure, finish rollback before cleanup. Lease loss still requires database-level protection.

The following application fragment uses a dedicated TransactionTemplate and a Redisson lock. The template uses PROPAGATION_REQUIRES_NEW, so it completes its own transaction before returning. The entry point is intended to run without an outer transaction; introducing nested transactions into a larger workflow requires a separate atomicity decision.

// Initialization: configure once before serving requests.
TransactionTemplate orderTx = new TransactionTemplate(transactionManager);
orderTx.setPropagationBehavior(
    TransactionDefinition.PROPAGATION_REQUIRES_NEW);

// Application method fragment; repositories and logger are injected.
RLock lock = redisson.getLock("order:" + orderId);
boolean acquired = false;
try {
    acquired = lock.tryLock(200, TimeUnit.MILLISECONDS);
    if (!acquired) {
        throw new OrderBusyException(orderId);
    }

    orderTx.executeWithoutResult(status -> {
        // Uses a conditional state transition and durable deduplication.
        orderRepository.applyCommandOnce(orderId, commandId);
        // Written in the same database transaction.
        outboxRepository.appendIfAbsent(commandId, orderId);
    });
} catch (InterruptedException ex) {
    Thread.currentThread().interrupt();
    throw new OrderInterruptedException(orderId, ex);
} finally {
    if (acquired) {
        try {
            lock.unlock();
        } catch (RuntimeException releaseFailure) {
            // Cleanup failure must not hide the business outcome.
            log.error("Order lock release failed: {}",
                      orderId, releaseFailure);
        }
    }
}

The repository methods above represent application contracts, not library APIs. Their implementations must enforce legal state transitions and command uniqueness within the transaction. A transaction error must propagate so Spring can roll back; the example does not turn an unsuccessful operation into a success.

The acquisition overload shown specifies a wait time without a fixed lease time, allowing Redisson’s watchdog behavior. Supplying an explicit leaseTime instead establishes automatic expiry after that interval. Redisson also checks ownership when unlocking; an expired owner does not simply delete another owner’s lock.

The 200-millisecond wait is illustrative, not a universal setting. Choose it from the request budget and observed contention. An unlock failure should be visible in monitoring, but a client retry must use the same command identifier because the database transaction may already have committed. A connection error during commit can also leave the outcome uncertain; durable command records make that uncertainty recoverable.

Scheduling asyncUpdateStock() and then unlocking would protect only task submission. The later database write would execute outside the coordinated section. Either complete the protected write before releasing ownership or submit a durable job whose consumer enforces its own transaction and concurrency rules. An in-memory queue is not a durable acceptance record.

What Redlock and ZooKeeper change

Redlock acquires leases on a majority of independent Redis masters within a bounded acquisition interval. Remaining validity must account for acquisition time and clock-drift assumptions; a failed attempt needs cleanup of partial acquisitions. It is not simply one Redis primary with several replicas, and a majority response alone is not the complete algorithm.

This changes the coordination failure model, but does not make downstream operations transactional or automatically fence old clients. For the order example, adding more lock servers cannot replace the inventory condition, reward uniqueness, or refund idempotency key. Redisson’s current documentation marks its RedLock object as deprecated, so copying an older RedissonRedLock example is not a good starting point for new code.

ZooKeeper uses a different coordination model. In its lock recipe, contenders create ephemeral sequential nodes under a shared path. The contender with the lowest sequence owns the lock; others watch their immediate predecessor and recheck when it disappears. Watching the predecessor avoids waking every waiting client on each release.

An ephemeral node is removed when its session ends or expires, not immediately after every transient disconnect. This distinction matters during partitions. A client that cannot confirm its session state should suspend protected work; after session loss it must treat its old ownership as invalid.

ZooKeeper is useful when the platform already needs ordered coordination or leader ownership and the team can operate the ensemble. Its stronger coordination semantics do not physically stop a paused worker from contacting another database or service. Resource-side conditional writes or fencing remain relevant, especially for jobs that publish results after long computation.

For example, a report-generation worker could write an immutable output artifact and then conditionally update a “current report” pointer using its accepted generation token. Losing ownership may waste some computation, but the old worker cannot replace the newer published report. That is a clearer correctness contract than assuming the worker will always notice session loss in time.

Make contention and failure visible in the order workflow

Lock granularity should follow the shared resource. An order-level lock coordinates competing transitions for one order, but does not serialize two different orders that purchase the same product. Inventory needs its own database protection. Conversely, a single global order lock would make unrelated customers wait for one another. If several locks are unavoidable, acquire them in a consistent order and keep the protected work bounded.

Moving coupon redemption and reward calculation under separate locks also does not make the overall workflow atomic. Each service needs a durable state transition, and failures between services need retry or compensation. For post-payment rewards, a committed outbox event and an idempotent ledger entry provide a more useful recovery path than holding an order lock while waiting for every downstream call.

When lock acquisition times out, the API should return a retryable busy result or durably accept a pending command. It should not continue the same critical mutation without protection, and “accepted” should not be presented as “completed.” For a refund, the pending record should retain the same provider request key across retries, including retries after a timeout with an unknown external outcome.

A useful verification exercise deliberately pauses the first worker until another can take over, then resumes the first. Check the business state rather than only the lock key. The following cases are design tests, not claims about measured results:

Injected failureBusiness result to verify
Two orders compete for the last unitAt most one inventory reservation commits; a retried command cannot deduct twice.
A worker resumes after losing its leaseA stale state transition or fenced publication is rejected at the resource.
The client retries after a commit timeoutThe command identifier resolves to one durable outcome.
A refund call times out after provider acceptanceThe same provider idempotency key is reused; no second refund is created.
Redis or ZooKeeper becomes unavailableWork pauses, fails retryably, or enters a durable pending state without bypassing the invariant.

Monitor acquisition latency, contention by resource, ownership-loss and release errors, transaction duration, and the age of pending commands. High contention may reveal an oversized critical section or a hot product; it does not automatically justify increasing every timeout. Track invariant violations and duplicate business effects as well, because a healthy lock-service dashboard cannot prove the order workflow is correct.

For this order design, I would begin with database-enforced reservations and idempotent transitions, then add coordination where overlapping work has a meaningful cost. Redis can be useful for that coordination; ZooKeeper can fit broader ownership requirements. The decision is successful when inventory, rewards, and refunds remain correct after ownership changes, not merely when the happy-path lock example runs.

Leave a Reply

Your email address will not be published. Required fields are marked *