Why does Spring AI chat memory appear to save a message while the next model request still contains only the latest user input? I hit this while tracing a support flow: the repository held the member’s previous order number, but the ChatModel did not receive it.
The useful question is not whether the conversation exists. It is which messages reached the model on this request.
Why are stored messages missing from the Spring AI model request?
Saving a message and including it in the next request are separate operations. A conversation can exist in storage without being used as model context.
This article traces that boundary through Spring AI 2.0.1. It uses a recording ChatModel instead of a remote LLM, so the tests can inspect exact messages rather than judge an unpredictable generated answer.
Scope and Version
The examples use:
| Component | Version or scope |
|---|---|
| Spring AI | 2.0.1 |
| Java | 17 |
| API | Synchronous ChatClient calls |
| Memory | MessageWindowChatMemory with an in-memory repository |
| Model | A local recording test double |
| External services | None; no API key required |
The source links are pinned to the v2.0.1 tag. The examples do not cover streaming, tool-call loops, database repository integration, or model answer quality.
This matters because memory behavior can change between framework versions. A statement about a missing conversation ID or eviction policy should not be assumed to apply unchanged to an older application.
Separate History, Retained Memory, and the Request
Three different objects are often called “chat history”:
| Layer | Purpose | Debugging question |
|---|---|---|
| Full application transcript | Preserve the conversation for display or audit | Was the message recorded? |
| Retained chat memory | Keep context under a memory policy | Is the message still in the retained window? |
| Model request | Provide context for one invocation | Was the message actually included? |
Spring AI separates memory management from repository storage. Its documentation also distinguishes chat memory from a complete historical record. Chat memory reference
In this example:
- ChatMemoryRepository provides conversation storage operations.
- MessageWindowChatMemory manages the retained message window.
- MessageChatMemoryAdvisor connects that memory to a ChatClient interaction.
A repository bean or a table full of messages does not, by itself, establish that a particular ChatClient reads those messages.

Simplified synchronous flow without tools. Conversation authorization belongs to the application before this flow begins.
Capture the Request Instead of Asking Whether the Model Remembers
The recording model stores each Prompt it receives and returns a fixed assistant response:
static final class RecordingModel implements ChatModel {
final List<List<Message>> requests = new ArrayList<>();
@Override
public ChatResponse call(Prompt prompt) {
requests.add(List.copyOf(prompt.getInstructions()));
return new ChatResponse(List.of(
new Generation(new AssistantMessage("ACK"))));
}
}
This is a single-threaded test double, not an LLM simulator. It does not infer an order number or evaluate an answer.
Its purpose is narrower: expose the messages crossing the ChatModel boundary. The supplied test class includes imports, helper methods, and an optional failure switch.
A captured Prompt is also not a capture of a provider’s final HTTP payload. Provider-specific serialization and server-side behavior require a separate integration test.
Reproduce the Problem: Saved Memory, Plain Client
First, save a message:
ChatMemory memory = MessageWindowChatMemory.builder().build();
memory.add("A", new UserMessage("My order is ORD-42."));
Then call a client that has no memory advisor:
ChatClient plain = ChatClient.builder(model).build();
plain.prompt()
.user("Which order?")
.advisors(a -> a.param(ChatMemory.CONVERSATION_ID, "A"))
.call()
.content();
The recording model receives:
USER: Which order?
The saved message remains in memory, but it is absent from the request.
Supplying a conversation ID does not automatically install a component that reads the memory. The plain client has no connection to the memory object we populated.
This is a useful diagnostic distinction: storage can be working while request enrichment is missing.
Connect the Memory Advisor
Build the client with the memory advisor:
ChatMemory memory = MessageWindowChatMemory.builder()
.maxMessages(20)
.build();
ChatClient client = ChatClient.builder(model)
.defaultAdvisors(
MessageChatMemoryAdvisor.builder(memory).build())
.build();
Supply the conversation ID on each call:
client.prompt()
.user("My order is ORD-42.")
.advisors(a -> a.param(ChatMemory.CONVERSATION_ID, "A"))
.call()
.content();
client.prompt()
.user("Which order?")
.advisors(a -> a.param(ChatMemory.CONVERSATION_ID, "A"))
.call()
.content();
The second recorded request now contains:
USER: My order is ORD-42.
ASSISTANT: ACK
USER: Which order?
After the second call succeeds, the retained conversation contains:
USER: My order is ORD-42.
ASSISTANT: ACK
USER: Which order?
ASSISTANT: ACK
The test asserts message order and roles as well as text. Flattening everything into one string would conceal useful information about how the history was assembled.
This result proves that the advisor passed the earlier exchange to the model boundary. It does not prove that a real model will answer correctly.
Read the Source Along the Request Path
There is no need to read the entire repository to explain this behavior.
Start with these version-pinned entry points:
| Class | Relevant method | What to inspect |
|---|---|---|
| BaseChatMemoryAdvisor | getConversationId | How the context key is resolved |
| MessageChatMemoryAdvisor | before | How retained and current messages are combined |
| MessageWindowChatMemory | add, get, process | How memory is retrieved and retained |
| ChatMemoryRepository | saveAll | Whether a write appends or replaces |
| BaseAdvisor | adviseCall | Where before, downstream invocation, and after occur |
In v2.0.1, the memory advisor reads retained messages, combines them with the current prompt, and avoids prepending an already-present matching history sequence. It positions the first system message at the front. It then adds the current user or tool-response message to memory before returning the enriched request. Its after method records assistant outputs from the response. MessageChatMemoryAdvisor source
For synchronous calls, BaseAdvisor runs before, invokes the next component, then runs after. This ordering explains why a downstream failure does not necessarily undo the earlier memory write. BaseAdvisor source
These are the framework-specific steps. The following experiments verify their observable consequences with the published dependency.
Conversation IDs Select Memory; They Do Not Authorize Access
The tests compare three cases:
| Calls | History visible to the second call |
|---|---|
| A followed by A | Earlier A exchange |
| A followed by B | No A exchange |
| Two simulated users both using shared | Earlier exchange under shared |
The framework is doing exactly what the key requests. It does not know that the second caller represents another authenticated user.
In this version, omitting the conversation ID fails before the recording model is invoked. The base advisor explicitly requires the context parameter. Conversation ID resolution source
In an application, validate ownership before using a client-supplied conversation identifier. Include the appropriate tenant and conversation namespace, and apply the same authorization to reading, clearing, and listing conversations.
An unguessable ID can reduce accidental collisions. It is not a substitute for access control.
Changing IDs on every request creates the opposite problem: apparent forgetting caused by looking up a different conversation each time.
A Four-Message Window Can Produce a Five-Message Request
Configure maxMessages to four, then make three sequential calls:
U1 -> ACK
U2 -> ACK
U3 -> ACK
The third request captured by the model contains:
U1, ACK, U2, ACK, U3
That is five messages.
After the response, the retained memory contains:
U2, ACK, U3, ACK
On the fourth call, the captured request is:
U2, ACK, U3, ACK, U4
U1 has now disappeared.
The important distinction is between the already-assembled request and the subsequently updated memory window. Limiting retained memory is not the same as enforcing a limit on all messages in the current request.
Nor is a message count a token budget. Four short messages and four long documents can have very different context costs.
For applications with strict input limits, budget the assembled request, including current input, system instructions, retrieved documents, and any other added context.
Eviction Can Leave Fewer Messages Than the Configured Maximum
A second test adds these four messages to a window configured for three:
U1, A1, U2, A2
The retained result is:
U2, A2
It contains two messages rather than three. Retaining the last three would start with A1, an assistant reply separated from its initiating user message.
The v2.0.1 implementation advances an eviction boundary to a user message, with special handling for system messages. Its get method delegates to the repository; its add path reads, processes, and saves the retained list. MessageWindowChatMemory source
Do not describe this policy as “always keep exactly the last N messages.” It also does not mean N conversation turns, and it does not summarize forgotten information.
This is one reason to keep a full business transcript separately when complete history is required.
A Failed Call Can Leave a User Message Behind
The recording model can be configured to throw an exception.
The test submits “Will fail.” The call raises an error, but retained memory contains:
USER: Will fail
There is no assistant response.
The test then switches the model back to success and resubmits the same text. In this exact one-message case, the second captured request contains only one copy: the history-matching guard recognizes the retained message in the current prompt.
However, the successful retry leaves this stored sequence:
USER: Will fail
USER: Will fail
ASSISTANT: ACK
A subsequent different question receives both stored user messages plus the reply. Prompt assembly and memory writes therefore need separate assertions; the guard is not a general request-idempotency mechanism.
This is an application-level retry of the entire ChatClient call. It is not a claim about every provider’s internal retry strategy.
A failed response does not automatically mean no state changed.
If the product needs retry-safe conversation handling, define request identity and failure status at the application layer. Do not blindly clear the entire conversation on failure; that can remove valid history or another concurrent request’s work.
Identical text also does not necessarily mean a duplicate request. A user can legitimately repeat a question. Use request identity rather than text equality when distinguishing retries.
Persistence and Retention Solve Different Problems
One test creates a second MessageWindowChatMemory using the same in-memory repository instance. It can read the message stored through the first object.
Creating a new memory object with a fresh repository yields an empty conversation.
This checks object and storage wiring. It is not a process-restart or database durability test.
An in-memory repository does not provide durable recovery across process restarts. A persistent repository can address storage lifetime, but does not automatically preserve every message excluded by the memory policy.
The repository contract describes saveAll as replacing the conversation’s existing messages with the supplied list. ChatMemoryRepository source
When evaluating persistent storage, test restart behavior, ordering, supported message types, schema migration, and retention explicitly. This article does not validate a JDBC deployment.
Treat Concurrent Turns as a Separate Design Problem
The source-level read–process–save sequence is a reason to investigate concurrent behavior. It is not evidence of a global transaction spanning model invocation and memory updates.
Two requests for one conversation may begin before either assistant response is available. Even if individual storage operations are thread-safe, that alone does not establish the conversational ordering your application expects.
For a product that requires sequential turns, consider serializing work per conversation. Multiple application instances need coordination beyond an in-process lock.
Alternatively, model turns as durable events and construct context from explicitly committed turns. The suitable design depends on latency, throughput, and consistency requirements.
The supplied tests are sequential. They do not claim to reproduce a concurrent lost update or prove a production-safe locking strategy.
Run the Example Without an API Key
The package contains a small Maven project and nine JUnit tests.
Its core dependency is:
<dependency>
<groupId>org.springframework.ai</groupId>
<artifactId>spring-ai-client-chat</artifactId>
<version>2.0.1</version>
</dependency>
There is no provider starter, Spring Boot application, external database, or remote model.
From the example directory, run:
mvn test
The included POM fixes the compiler and test plugin versions. The optional central-settings.xml can be used when a local Maven mirror cannot resolve public dependencies.
All nine tests passed with Spring AI 2.0.1, Java 17.0.17, and Maven 3.8.6. The suite covers:
| Test | Expected observation |
|---|---|
| Stored memory without an advisor | Current message only |
| Advisor with the same conversation ID | Prior user and assistant messages included |
| Different conversation ID | Earlier conversation not included |
| Shared conversation key | Earlier messages are reused |
| Missing conversation ID | Failure before the model call |
| Four-message retained window | Five-message request in the demonstrated sequence |
| Turn-boundary eviction | A two-message retained result with maxMessages set to three |
| Shared versus fresh repository | Storage identity determines retained state |
| Failed invocation followed by retry | Failed user message remains and is included again |
These are framework integration tests against the recording model. They verify request composition and local memory effects, not an LLM’s reasoning, provider HTTP serialization, persistent storage, streaming cancellation, or tool execution.
Debug the Boundary That Is Actually Failing
When a user reports that the assistant forgot something, check the chain in order:
- Resolve the authorized conversation ID.
- Check whether the relevant message remains in the selected memory.
- Confirm that the actual ChatClient uses the intended memory advisor and instance.
- Inspect message roles and order at the model boundary.
- If context is present, investigate provider limits and model behavior separately.
In production, prefer message counts, roles, request IDs, and redacted diagnostics over logging full personal conversations.
Related Spring AI Guides
- Spring AI Advisor Chain: Request and Response Order Explained
- Spring AI RAG Debugging: When Retrieved Documents Do Not Reach the Model
The useful distinction is concrete: a stored message is available data; an injected message is model context. Trace the path between them before treating a memory symptom as a model problem.

[…] The fixture verifies context construction, not embedding quality or answer quality. For the preceding memory boundary, see Spring AI Chat Memory: Why Stored Messages May Not Reach the Model. […]
[…] If history appears to be missing, compare the advisor order with the captured prompt; the related chat memory troubleshooting example uses the same evidence-first […]