Microservice Overengineering: Lessons from Simplifying a Small Java Platform

A practical reflection on microservice overengineering, deployment coupling, cross-service consistency, and a checklist for choosing module boundaries in a small Java platform.

In one of my previous projects, neither the team nor the application was particularly large. Yet we introduced an architecture that required us to operate as though both were.

We separated users, live classrooms, orders, and products into independent services. We added message queues, rate-limiting rules, and the configuration needed to make those components work together. Part of the motivation was preparing for future growth. Another part, looking back, was the attraction of using a more “advanced” architecture.

The cost became visible in everyday work. Development and maintenance grew more complicated. Deployments took longer. Network problems and message-queue failures left business data inconsistent across services, making failures difficult to understand and recover from.

We eventually reduced the number of independent services. User management and authentication remained separate, while the other business capabilities were consolidated. We retained the microservice framework, but stepped back from the original level of fragmentation.

There was a related problem inside the code: excessive abstraction had scattered business logic across multiple layers. Even when no network was involved, understanding a complete operation could require too much navigation.

Both problems taught me the same lesson. Every boundary needs to earn its place.

Write down the pressure before extracting a service

A business noun is not, by itself, a deployment argument. Before moving a module across a network, record the pressure that the boundary is meant to relieve and the new work the team is willing to own.

Possible pressureEvidence to collectCost to accept
Different capacity profileLoad or resource measurements show that one capability needs a different scaling policy.Separate deployment, alerts, pools, and capacity planning.
Independent release ownershipA team can own the contract, compatibility tests, rollback, and on-call work.Versioned APIs, contract tests, and release coordination.
Security or failure isolationAn access, compliance, or availability requirement cannot be met inside the existing process.New authentication, observability, and partial-failure paths.
Work can finish laterThe user can see a pending state and the operation has a stable idempotency key.Retries, duplicate delivery handling, and recovery monitoring.

The table is a decision template, not a set of measurements from the project. Its purpose is to make “we may need to scale later” more specific before the team pays the cost of a remote boundary.

For a small platform, I would also record a baseline: deployment duration, number of cross-service calls in a common operation, time to repair a partial failure, and the age of the oldest pending message. Without a baseline, a later architecture change can sound successful while the original problem remains unmeasured.

We Treated Business Categories as Deployment Boundaries

Users, products, orders, and live classrooms are distinct business concepts. Giving them clear responsibilities in the code made sense. We took the additional step of giving them separate deployments before the business had demonstrated a need for that separation.

Those are different decisions.

An order module can own order rules without running in a separate process. A product module can expose a clear interface without requiring an HTTP call. Separating responsibilities helps organize an application; distributing those responsibilities adds communication and operational requirements.

At the time, we gave considerable weight to future possibilities. Independent services would let us expand the system later. The framework appeared to provide a foundation for a larger platform.

What we underestimated was how much work that foundation required immediately.

Every deployed service needed configuration, health checks, logs, and a release process. Every remote interaction needed a decision about timeouts and failure handling. Every workflow spanning multiple data owners needed a way to recover when only part of it succeeded.

Our small team had to maintain those decisions alongside the product itself. We had introduced capabilities for a future operating model while paying for them with our current development capacity.

One Business Operation Became Several Separate Outcomes

The consistency problems were especially frustrating because the business intent was often straightforward, even when the implementation was distributed.

Consider a simplified course-purchase workflow: a confirmed purchase should result in the student receiving access to a course. This example illustrates the coordination problem; it is not a reconstruction of a particular incident from the project.

If order state and course access are maintained by separate services, completing one does not automatically complete the other. The order service might commit its database change and then publish an event. Another service receives that event and grants access.

The gap between those steps matters.

A publish attempt can fail after the order transaction commits. The broker can accept the event while the consumer remains unavailable. A consumer can update its database and then fail before acknowledging the message, causing redelivery.

Illustrative course purchase workflow across service boundaries

An illustrative workflow: committing an order, delivering an event, and granting access are separate outcomes. Each boundary needs a recovery strategy.

Synchronous calls have their own uncertainty. If a caller times out, the downstream operation may have failed, may still be running, or may have completed without its response reaching the caller. Retrying a state-changing request safely requires more than increasing the timeout.

In our project, network and MQ problems did lead to inconsistencies between services. Investigating them meant determining which parts of a business operation had completed and which had not.

A queue does not remove that responsibility. It provides a mechanism for moving work between components. Reliable business processing still depends on how events are recorded, retried, consumed, and reconciled.

For example, an outbox can record an event in the same local transaction as a business update, reducing the risk of committing the update without retaining the event for delivery. It does not make the entire cross-service workflow atomic. Consumers still need to handle duplicate delivery, and operations still need a way to detect work that has stopped progressing.

These mechanisms are useful. They also require implementation, testing, and ongoing attention. We had made ourselves responsible for those concerns earlier than the business required.

The Maintenance Cost Extended Beyond Failures

Incidents made the problem obvious, but routine work was already carrying the cost.

Deploying several connected services requires more coordination than deploying one application. A change to a contract raises questions about compatibility and release order. A configuration mismatch can produce symptoms that look like application bugs. Troubleshooting requires following the same business operation across multiple processes.

These concerns are manageable when service independence produces a meaningful benefit. In our case, development and maintenance had become expensive relative to the size of the project.

The distinction I would make now is between separate deployment and useful independence. Two services may have different deployment packages but still need to be understood, changed, and released together. In that situation, the network boundary adds work while providing limited freedom.

Rate-limiting configuration contributed another kind of complexity. Rate limiting itself can be valuable in a small application: a database connection pool or an expensive endpoint may need protection regardless of system size.

But a useful rule needs a reason. What resource does it protect? What workload informed the threshold? What should the user experience when the limit is reached? Can retries from upstream increase the pressure?

Adding rules without clear answers creates more behavior that the team must remember during development and incidents. Configuration becomes part of the application’s logic, even when it lives outside the code.

The problem was the combined maintenance burden. We had to keep the business working while also managing a growing collection of interactions and settings.

Over-Abstraction Made the Business Harder to Read

Inside the application, we encountered a similar problem at a smaller scale.

Business behavior was distributed across too many layers. To understand an operation, a developer had to follow the execution path through several abstractions and assemble the actual rule from pieces in different locations.

Abstraction is valuable when it captures a stable concept or isolates a real source of variation. It becomes expensive when the reader must pass through layers that contribute little information.

Imagine tracing why an order receives a particular status. Ideally, the application should make the important decisions visible: validate the current state, apply the business rule, record the result, and trigger any necessary follow-up work.

Those responsibilities can still use separate components. The difficulty arises when the sequence becomes hidden behind generic handlers and indirect calls, so the reader must inspect many implementations simply to discover what happens.

That kind of structure also complicates changes. Before modifying one rule, the developer must work out whether an abstraction is shared intentionally, whether other callers depend on its behavior, and where the business decision actually belongs.

I no longer judge an abstraction mainly by how many future implementations it could support. I look at what it makes easier today. Does it keep a rule in one place? Does it isolate something that changes independently? Does it make the next modification easier to understand and test?

An interface with one implementation is not automatically wrong. It might provide a useful integration boundary. But an additional layer should have a purpose beyond making the design appear extensible.

We Consolidated the Business Services

Our eventual adjustment was to reduce the scope of distribution.

We retained user management and authentication as a separate service and consolidated the other business services. The microservice framework remained, but fewer business boundaries also had to be deployment boundaries.

Illustrative consolidation of business modules

A simplified view of the consolidation direction. Connections illustrate service boundaries rather than exact historical call paths; databases and messaging infrastructure are omitted.

This was a pragmatic reduction in complexity. It also involved a compromise worth acknowledging: under the business requirements at the time, even the remaining user and authentication separation was not strictly necessary. We kept that boundary largely to leave room for future expansion.

I would not present that decision as proof that every small platform needs an independent identity service. It was the boundary we chose to retain.

Consolidation also does not require abandoning clear responsibilities. Products, orders, and classroom operations can remain distinct modules within an application. Their interfaces can express business intent without making every interaction cross a network.

Nor does combining deployments automatically solve consistency. Operations can share a local transaction only when their storage and transaction arrangement allows it. Calls to external systems remain outside that transaction, and asynchronous work still needs failure handling.

The practical value of consolidation is that it reduces the places where distribution is mandatory. The team can decide where a remote boundary is justified, instead of inheriting one for every business category.

That distinction better reflects what we needed. The project needed understandable business organization. It did not need every responsibility to operate as a separate service.

Consolidation Still Needs Visible Module Contracts

Putting capabilities back into one application should reduce network failure modes without turning the code into one undifferentiated package. A simple module shape can keep the business boundaries visible:

app/
  orders/
    OrderService.java
    OrderRepository.java
  classrooms/
    ClassroomService.java
    ClassroomRepository.java
  identity/
    UserService.java

This tree is illustrative. The useful property is that an order rule has a clear owner and a small interface; it does not require an HTTP call merely because the capability has a different name. If a module later needs independent scaling or release ownership, the interface and measurements provide a starting point for extracting it.

How I Evaluate Complexity Now

If I were making the initial decision again, I would begin by separating business responsibilities in the code and evaluating deployment boundaries independently.

Before extracting a service, I would want to identify a concrete pressure. Perhaps a workload needs substantially different capacity. Perhaps a team needs independent release ownership. Perhaps security or availability requirements demand isolation.

I would then include the operational consequences in the decision. Who will diagnose partial failures? How will incompatible changes be released? What happens when the dependency is unavailable? How will inconsistent business state be detected and repaired?

Those questions are part of the architecture, not tasks to postpone until after the services exist.

I would apply the same reasoning to messaging. Background notifications, exports, and other work that can complete later may justify asynchronous processing. A queue should support those requirements with explicit retry and recovery behavior. Its presence on an architecture diagram is not evidence that the workflow is reliable.

For code structure, I would prioritize making important business operations readable from a clear entry point. Shared behavior should be extracted when the common rule is understood, and extension points should reflect meaningful variation. Anticipating every possible future change can make the changes we actually need more difficult.

A review checklist for a small platform

  • What concrete pressure requires a process boundary?
  • Which team owns compatibility, deployment, and recovery for that boundary?
  • Which business facts may be temporarily inconsistent, and how are they reconciled?
  • Can a failed or duplicated message be retried without a second business effect?
  • What measurement will show that the boundary solved the original problem?

The experience did not make me reject microservices. It made me account for their costs earlier.

We had spent too much effort preparing for a system larger than the one we were operating. Consolidating services was an acknowledgment that the team needed a structure it could develop, deploy, and troubleshoot effectively.

Today, I am more comfortable leaving a capability inside an application until there is a specific reason to move it out. Clear module boundaries preserve options. Independent services introduce obligations. The decision to cross that line deserves evidence from the business and the team that will maintain it.

Leave a Reply

Your email address will not be published. Required fields are marked *