When we added an AI customer support feature to an existing Spring Boot application, the initial requirement sounded straightforward: let customers ask questions and have Qwen answer them.
Calling the model was the easy part. The requirement became more interesting when customers asked questions such as:
“Where are the headphones I bought last week?”
“If they still haven’t arrived tomorrow, can I request a refund?”
The first question needs access to orders and logistics. The second refers to the same purchase, but also introduces company policy and potentially a financial operation. A model that writes fluent answers cannot resolve these questions from its general knowledge alone.
We chose Spring AI Alibaba because the application already used Spring Boot. Orders, logistics, and customer accounts remained behind existing Java services. The model became a natural-language interface to those capabilities.
The important design work was deciding what information the model should receive, which operations it could request, and which decisions had to remain inside the application.
Keep the Model Behind the Application Boundary
Spring AI Alibaba made it possible to integrate model interaction into the existing Spring application rather than scatter provider-specific HTTP calls throughout the business code.
ChatClient provides the application-facing interaction API. Tools connect selected business capabilities, memory supplies relevant conversation context, and retrieval supplies supporting documents. These components help organize the implementation, but they do not automatically establish authorization or business correctness.

Figure 1. The application controls identity, access, and execution. The model receives selected context and can request operations through explicitly exposed tools.
The distinction between requesting and executing a tool matters. The model returns a proposed tool call and arguments. Application-side code resolves and executes that request, then supplies the result for further model processing. The model does not receive unrestricted access to the underlying Java services. Spring AI tool-calling documentation
For a working implementation, I would first pin a compatible Spring Boot, Spring AI, and Spring AI Alibaba dependency set. Examples from different release lines should not be combined casually. The current Alibaba quick-start documentation uses spring-ai-alibaba-starter-dashscope and the spring.ai.dashscope.api-key configuration property; older examples may use different dependencies. Spring AI Alibaba quick start
The examples below describe application boundaries rather than provide a complete, version-tested starter project. Tool registration, memory configuration, and streaming integration should follow the documentation for the chosen dependency set.
A system instruction asking the assistant not to invent information is useful. However, it cannot replace access checks, structured tool results, or a defined response when information is unavailable.
Find the Right Order Before Answering About It
The question “Where are the headphones I bought last week?” contains several useful search clues. It does not necessarily identify one order.
Returning the customer’s most recent order would be a convenient shortcut, but it could select a different purchase. The customer may also have bought two similar items during that period.
A better sequence is to search authorized orders using the product description and an appropriate date range. Relative dates need a known reference date and timezone. If several candidates remain, the assistant should present a small, privacy-conscious summary and ask the customer to choose.
If no candidate is found, the response should say so and ask for more information. If the order service times out, that is a different outcome: the assistant cannot conclude that the customer has no matching order.
Before any of these steps, the application must know who the customer is.
User identity comes from the authenticated request, not from the conversation or model-generated arguments.
A tool that accepts an arbitrary userId from the model and immediately queries that account creates an unnecessary access-control risk. A safer boundary supplies the authenticated principal separately and treats model arguments as untrusted search inputs.
The following Java-style pseudocode illustrates that separation:
OrderSearchResult searchOrders(
OrderSearchCriteria modelCriteria,
TrustedRequestContext context) {
var principal = context.requireAuthenticatedPrincipal();
var criteria = validateAndLimit(modelCriteria);
return orderService.searchVisibleOrders(
principal.tenantId(),
principal.userId(),
criteria);
}
Here, TrustedRequestContext and the other types represent application-owned abstractions. The context is constructed by the backend; it is not another object the model is allowed to populate.
Spring AI’s ToolContext offers a way to pass application context separately from model-generated tool arguments. The application still needs to populate it from trusted sources and enforce authorization inside its services. Spring AI Tool Context
An order identifier is also not proof of ownership. Every order lookup must check that the authenticated customer can access the requested record, including when the identifier came from a previous conversation turn.
The result returned to the model should be a narrow support view. Shipment status and relevant timestamps may be sufficient; full addresses, payment details, and internal account fields usually are not.
Read-only tools avoid modifying state, but they can still disclose sensitive information. Their permissions and output deserve the same care as an ordinary customer-facing API.
Remember the Conversation, Refresh the Facts, Retrieve the Right Policy
Once the customer has identified an order, the follow-up question becomes easier to interpret:
“If it still hasn’t arrived tomorrow, can I request a refund?”
Conversation memory helps resolve “it.” It should not be treated as a fresh source of logistics data.
A useful separation is to keep the selected order reference as validated conversation state and retrieve changing facts again when the answer depends on them. Yesterday’s delivery estimate may no longer be correct. A previous assistant answer may also have been incomplete.
The conversation itself needs an ownership boundary. A client-supplied conversation identifier is not sufficient authorization to load its history. The backend must verify that the conversation belongs to the current user and tenant.
Creating a ChatClient does not automatically provide persistent, isolated memory. The appropriate memory implementation and advisor need to be configured, together with conversation identification and retention behavior. Spring AI chat-memory documentation
RAG addresses another part of the question: the refund policy.
Instead of placing every support document in the system prompt, the application retrieves relevant passages and provides them with the customer’s question and current order facts. Document ingestion typically involves preparing content, splitting it into useful sections, generating embeddings, and storing searchable representations.
The important question is not simply whether retrieval returned something. It is whether the returned policy applies to this order.
A policy may differ by region, seller, product category, or purchase date. The newest document is not automatically the correct one for an older purchase. Useful metadata includes the policy identifier, version, effective dates, scope, and an accessible source reference.

Figure 2. The diagram shows the path when candidate orders are available. Multiple matches require clarification; no matches or failed lookups require their own responses. Policy retrieval supports an explanation, not automatic refund execution.
When investigating an incorrect policy answer, I would inspect the retrieved evidence before changing models. Did retrieval preserve the exception attached to the main rule? Did it select an expired policy? Did it find a shipping explanation rather than a refund condition?
Chunk boundaries, metadata filters, keyword matching, and reranking can help, but each change should be evaluated against representative questions. Increasing the number of retrieved passages can also add conflicting material and consume more context.
The answer should make its basis visible. It can summarize the shipment status, explain the relevant condition, and cite the approved policy source. A citation should resolve to evidence the application actually retrieved, rather than a reference invented during generation.
If the policy is missing or ambiguous, a useful answer can say what is known and offer a human-support route. It should not turn an uncertain interpretation into a guaranteed refund promise.
Retrieved text must also remain evidence, not authority over the application. A passage telling the assistant to ignore permissions or invoke a refund should not alter which tools are available or bypass their checks.
Explaining a Refund Is Different from Executing One
The same conversational interface can make a read operation and a write operation look deceptively similar.
“Can I request a refund?” asks for information.
“Please refund this order” asks the system to change business state.
In our design, query tools and modifying tools needed different treatment. A refund operation belongs behind the existing refund workflow, including authorization, eligibility checks, amount calculation, and idempotency.
The model can help collect the customer’s intent and explain the next step. It should not decide the refundable amount solely from retrieved prose or a remembered order total.
Before a modifying action proceeds, the application should establish the exact order, the proposed action, and any required customer confirmation or human approval. The backend must validate the current state again at execution time: an order may have been canceled or refunded since the conversation began.
Confirmation is strongest when it refers to a specific proposed operation. A vague “yes” should not authorize a different order or a changed amount after the conversation has moved on.
Execution also needs to survive an interrupted chat response. Suppose the refund request is accepted, but the model connection fails before the final message reaches the customer. Retrying the conversation must not create a second refund.
The backend should retain an operation identifier and provide status lookup. Repeated attempts should follow the existing idempotency rules. An unknown outcome remains pending until the business system resolves it.
This is the same boundary discussed in designing reliable orders and refunds with distributed transactions: a conversational interface does not remove the need for explicit state transitions and recovery.
A successful chat response is not evidence that a refund succeeded. Conversely, a failed chat response does not prove that no refund was initiated.
Make Incomplete Information and Service Failures Visible
After tools, memory, and retrieval are connected, the remaining work looks familiar to a backend engineer: latency, concurrency, timeouts, retries, and operational visibility.
Streaming can improve the experience by displaying output as it becomes available. It does not necessarily reduce total generation time, and a request may still wait for several tools before useful answer text can appear.
A backend implementation also needs more than a method returning Flux<String>. The HTTP endpoint, event format, frontend handling, and error behavior must support streaming consistently. When a customer disconnects, unnecessary generation should be canceled where supported, while already accepted business operations continue to have an independently queryable status.
I would avoid streaming a confident refund conclusion before the required checks complete. A progress message such as “Checking the shipment status” is different from announcing an outcome.
Tool execution needs budgets as well. One customer turn can trigger multiple model calls, several business queries, and retrieval. The application should bound total duration, tool rounds, result sizes, and concurrent requests rather than rely only on a timeout for each individual dependency.
Retries require particular care after a write tool might have run. A model timeout should not trigger an automatic replay of the entire workflow.
Context is another finite resource. System instructions, conversation history, retrieved passages, and tool results all compete for space. A bounded history window, compact tool responses, and selective retrieval are more sustainable than continually appending everything.
Summaries can reduce history size, but they should not become an authoritative store for financial state. A summarized conversation saying “refund approved” cannot replace the refund service’s record.
For fallback behavior, the response must reflect the failure that actually occurred. A logistics timeout should produce “I cannot retrieve the current shipment status,” not “Your order has not shipped.” If the model is unavailable, the application may still offer the normal order page or a human-support channel.
The service-level protections are similar to those used elsewhere in Java systems. The difference is that retries and partial failures must account for both generated text and any business operations already executed.
I would evaluate the assistant with cases such as these before expanding its tool access:
| Scenario | Expected behavior |
|---|---|
| Two purchases match “last week’s headphones” | Ask the customer to identify the intended order. |
| The customer supplies another account’s order number | Deny access without disclosing order details. |
| Logistics lookup times out | Report unavailable information rather than inventing a shipment state. |
| Retrieval finds an inapplicable policy version | Exclude it or acknowledge that the answer cannot be confirmed. |
| A retrieved passage asks the model to bypass permissions | Keep application permissions and tool boundaries unchanged. |
| A refund is accepted but the chat stream disconnects | Recover the existing operation status without repeating the refund. |
These are acceptance cases, not a claim of measured production accuracy.
Operational traces should connect model calls, tool execution, retrieved document identifiers, and resulting business operations. Useful metrics include time to first useful output, total latency, tool failures, token usage, and unresolved operations.
Logging every raw prompt and tool result is not necessary to achieve that visibility. Order data and conversation content should be minimized or redacted, with appropriate access and retention controls.
What made Spring AI Alibaba useful in this project was the fit with the existing application. The model could interpret a customer’s question, while established Java services remained responsible for authorized access and business execution.
The next step was not automatically a larger agent architecture. It was making the current assistant reliable: selecting the correct order, refreshing facts, retrieving applicable policy, and refusing to turn uncertainty into a completed business action.
That is the boundary I would keep as the system grows. Let the model help customers express what they need and understand the result. Let the application determine what it knows, what the customer may access, and what has actually been done.
