Spring AI Chat Memory: Stored Messages Missing from Model Requests

Trace how Spring AI loads chat memory into model requests, and reproduce common issues with conversation IDs, message windows, and failed-call retries.

Why does Spring AI chat memory appear to save a message while the next model request still contains only the latest user input? I hit this while tracing a support flow: the repository held the member’s previous order number, but the ChatModel did not receive it.

The useful question is not whether the conversation exists. It is which messages reached the model on this request.

Why are stored messages missing from the Spring AI model request?

Saving a message and including it in the next request are separate operations. A conversation can exist in storage without being used as model context.

This article traces that boundary through Spring AI 2.0.1. It uses a recording ChatModel instead of a remote LLM, so the tests can inspect exact messages rather than judge an unpredictable generated answer.

Scope and Version

The examples use:

ComponentVersion or scope
Spring AI2.0.1
Java17
APISynchronous ChatClient calls
MemoryMessageWindowChatMemory with an in-memory repository
ModelA local recording test double
External servicesNone; no API key required

The source links are pinned to the v2.0.1 tag. The examples do not cover streaming, tool-call loops, database repository integration, or model answer quality.

This matters because memory behavior can change between framework versions. A statement about a missing conversation ID or eviction policy should not be assumed to apply unchanged to an older application.

Separate History, Retained Memory, and the Request

Three different objects are often called “chat history”:

LayerPurposeDebugging question
Full application transcriptPreserve the conversation for display or auditWas the message recorded?
Retained chat memoryKeep context under a memory policyIs the message still in the retained window?
Model requestProvide context for one invocationWas the message actually included?

Spring AI separates memory management from repository storage. Its documentation also distinguishes chat memory from a complete historical record. Chat memory reference

In this example:

  • ChatMemoryRepository provides conversation storage operations.
  • MessageWindowChatMemory manages the retained message window.
  • MessageChatMemoryAdvisor connects that memory to a ChatClient interaction.

A repository bean or a table full of messages does not, by itself, establish that a particular ChatClient reads those messages.

A request passes through MessageChatMemoryAdvisor before reaching ChatModel; the advisor reads retained history, saves the current user message, and saves the assistant reply after a successful call.

Simplified synchronous flow without tools. Conversation authorization belongs to the application before this flow begins.

Capture the Request Instead of Asking Whether the Model Remembers

The recording model stores each Prompt it receives and returns a fixed assistant response:

static final class RecordingModel implements ChatModel {
    final List<List<Message>> requests = new ArrayList<>();

    @Override
    public ChatResponse call(Prompt prompt) {
        requests.add(List.copyOf(prompt.getInstructions()));
        return new ChatResponse(List.of(
                new Generation(new AssistantMessage("ACK"))));
    }
}

This is a single-threaded test double, not an LLM simulator. It does not infer an order number or evaluate an answer.

Its purpose is narrower: expose the messages crossing the ChatModel boundary. The supplied test class includes imports, helper methods, and an optional failure switch.

A captured Prompt is also not a capture of a provider’s final HTTP payload. Provider-specific serialization and server-side behavior require a separate integration test.

Reproduce the Problem: Saved Memory, Plain Client

First, save a message:

ChatMemory memory = MessageWindowChatMemory.builder().build();
memory.add("A", new UserMessage("My order is ORD-42."));

Then call a client that has no memory advisor:

ChatClient plain = ChatClient.builder(model).build();

plain.prompt()
        .user("Which order?")
        .advisors(a -> a.param(ChatMemory.CONVERSATION_ID, "A"))
        .call()
        .content();

The recording model receives:

USER: Which order?

The saved message remains in memory, but it is absent from the request.

Supplying a conversation ID does not automatically install a component that reads the memory. The plain client has no connection to the memory object we populated.

This is a useful diagnostic distinction: storage can be working while request enrichment is missing.

Connect the Memory Advisor

Build the client with the memory advisor:

ChatMemory memory = MessageWindowChatMemory.builder()
        .maxMessages(20)
        .build();

ChatClient client = ChatClient.builder(model)
        .defaultAdvisors(
                MessageChatMemoryAdvisor.builder(memory).build())
        .build();

Supply the conversation ID on each call:

client.prompt()
        .user("My order is ORD-42.")
        .advisors(a -> a.param(ChatMemory.CONVERSATION_ID, "A"))
        .call()
        .content();

client.prompt()
        .user("Which order?")
        .advisors(a -> a.param(ChatMemory.CONVERSATION_ID, "A"))
        .call()
        .content();

The second recorded request now contains:

USER: My order is ORD-42.
ASSISTANT: ACK
USER: Which order?

After the second call succeeds, the retained conversation contains:

USER: My order is ORD-42.
ASSISTANT: ACK
USER: Which order?
ASSISTANT: ACK

The test asserts message order and roles as well as text. Flattening everything into one string would conceal useful information about how the history was assembled.

This result proves that the advisor passed the earlier exchange to the model boundary. It does not prove that a real model will answer correctly.

Read the Source Along the Request Path

There is no need to read the entire repository to explain this behavior.

Start with these version-pinned entry points:

ClassRelevant methodWhat to inspect
BaseChatMemoryAdvisorgetConversationIdHow the context key is resolved
MessageChatMemoryAdvisorbeforeHow retained and current messages are combined
MessageWindowChatMemoryadd, get, processHow memory is retrieved and retained
ChatMemoryRepositorysaveAllWhether a write appends or replaces
BaseAdvisoradviseCallWhere before, downstream invocation, and after occur

In v2.0.1, the memory advisor reads retained messages, combines them with the current prompt, and avoids prepending an already-present matching history sequence. It positions the first system message at the front. It then adds the current user or tool-response message to memory before returning the enriched request. Its after method records assistant outputs from the response. MessageChatMemoryAdvisor source

For synchronous calls, BaseAdvisor runs before, invokes the next component, then runs after. This ordering explains why a downstream failure does not necessarily undo the earlier memory write. BaseAdvisor source

These are the framework-specific steps. The following experiments verify their observable consequences with the published dependency.

Conversation IDs Select Memory; They Do Not Authorize Access

The tests compare three cases:

CallsHistory visible to the second call
A followed by AEarlier A exchange
A followed by BNo A exchange
Two simulated users both using sharedEarlier exchange under shared

The framework is doing exactly what the key requests. It does not know that the second caller represents another authenticated user.

In this version, omitting the conversation ID fails before the recording model is invoked. The base advisor explicitly requires the context parameter. Conversation ID resolution source

In an application, validate ownership before using a client-supplied conversation identifier. Include the appropriate tenant and conversation namespace, and apply the same authorization to reading, clearing, and listing conversations.

An unguessable ID can reduce accidental collisions. It is not a substitute for access control.

Changing IDs on every request creates the opposite problem: apparent forgetting caused by looking up a different conversation each time.

A Four-Message Window Can Produce a Five-Message Request

Configure maxMessages to four, then make three sequential calls:

U1 -> ACK
U2 -> ACK
U3 -> ACK

The third request captured by the model contains:

U1, ACK, U2, ACK, U3

That is five messages.

After the response, the retained memory contains:

U2, ACK, U3, ACK

On the fourth call, the captured request is:

U2, ACK, U3, ACK, U4

U1 has now disappeared.

The important distinction is between the already-assembled request and the subsequently updated memory window. Limiting retained memory is not the same as enforcing a limit on all messages in the current request.

Nor is a message count a token budget. Four short messages and four long documents can have very different context costs.

For applications with strict input limits, budget the assembled request, including current input, system instructions, retrieved documents, and any other added context.

Eviction Can Leave Fewer Messages Than the Configured Maximum

A second test adds these four messages to a window configured for three:

U1, A1, U2, A2

The retained result is:

U2, A2

It contains two messages rather than three. Retaining the last three would start with A1, an assistant reply separated from its initiating user message.

The v2.0.1 implementation advances an eviction boundary to a user message, with special handling for system messages. Its get method delegates to the repository; its add path reads, processes, and saves the retained list. MessageWindowChatMemory source

Do not describe this policy as “always keep exactly the last N messages.” It also does not mean N conversation turns, and it does not summarize forgotten information.

This is one reason to keep a full business transcript separately when complete history is required.

A Failed Call Can Leave a User Message Behind

The recording model can be configured to throw an exception.

The test submits “Will fail.” The call raises an error, but retained memory contains:

USER: Will fail

There is no assistant response.

The test then switches the model back to success and resubmits the same text. In this exact one-message case, the second captured request contains only one copy: the history-matching guard recognizes the retained message in the current prompt.

However, the successful retry leaves this stored sequence:

USER: Will fail
USER: Will fail
ASSISTANT: ACK

A subsequent different question receives both stored user messages plus the reply. Prompt assembly and memory writes therefore need separate assertions; the guard is not a general request-idempotency mechanism.

This is an application-level retry of the entire ChatClient call. It is not a claim about every provider’s internal retry strategy.

A failed response does not automatically mean no state changed.

If the product needs retry-safe conversation handling, define request identity and failure status at the application layer. Do not blindly clear the entire conversation on failure; that can remove valid history or another concurrent request’s work.

Identical text also does not necessarily mean a duplicate request. A user can legitimately repeat a question. Use request identity rather than text equality when distinguishing retries.

Persistence and Retention Solve Different Problems

One test creates a second MessageWindowChatMemory using the same in-memory repository instance. It can read the message stored through the first object.

Creating a new memory object with a fresh repository yields an empty conversation.

This checks object and storage wiring. It is not a process-restart or database durability test.

An in-memory repository does not provide durable recovery across process restarts. A persistent repository can address storage lifetime, but does not automatically preserve every message excluded by the memory policy.

The repository contract describes saveAll as replacing the conversation’s existing messages with the supplied list. ChatMemoryRepository source

When evaluating persistent storage, test restart behavior, ordering, supported message types, schema migration, and retention explicitly. This article does not validate a JDBC deployment.

Treat Concurrent Turns as a Separate Design Problem

The source-level read–process–save sequence is a reason to investigate concurrent behavior. It is not evidence of a global transaction spanning model invocation and memory updates.

Two requests for one conversation may begin before either assistant response is available. Even if individual storage operations are thread-safe, that alone does not establish the conversational ordering your application expects.

For a product that requires sequential turns, consider serializing work per conversation. Multiple application instances need coordination beyond an in-process lock.

Alternatively, model turns as durable events and construct context from explicitly committed turns. The suitable design depends on latency, throughput, and consistency requirements.

The supplied tests are sequential. They do not claim to reproduce a concurrent lost update or prove a production-safe locking strategy.

Run the Example Without an API Key

The package contains a small Maven project and nine JUnit tests.

Its core dependency is:

<dependency>
    <groupId>org.springframework.ai</groupId>
    <artifactId>spring-ai-client-chat</artifactId>
    <version>2.0.1</version>
</dependency>

There is no provider starter, Spring Boot application, external database, or remote model.

From the example directory, run:

mvn test

The included POM fixes the compiler and test plugin versions. The optional central-settings.xml can be used when a local Maven mirror cannot resolve public dependencies.

All nine tests passed with Spring AI 2.0.1, Java 17.0.17, and Maven 3.8.6. The suite covers:

TestExpected observation
Stored memory without an advisorCurrent message only
Advisor with the same conversation IDPrior user and assistant messages included
Different conversation IDEarlier conversation not included
Shared conversation keyEarlier messages are reused
Missing conversation IDFailure before the model call
Four-message retained windowFive-message request in the demonstrated sequence
Turn-boundary evictionA two-message retained result with maxMessages set to three
Shared versus fresh repositoryStorage identity determines retained state
Failed invocation followed by retryFailed user message remains and is included again

These are framework integration tests against the recording model. They verify request composition and local memory effects, not an LLM’s reasoning, provider HTTP serialization, persistent storage, streaming cancellation, or tool execution.

Debug the Boundary That Is Actually Failing

When a user reports that the assistant forgot something, check the chain in order:

  1. Resolve the authorized conversation ID.
  2. Check whether the relevant message remains in the selected memory.
  3. Confirm that the actual ChatClient uses the intended memory advisor and instance.
  4. Inspect message roles and order at the model boundary.
  5. If context is present, investigate provider limits and model behavior separately.

In production, prefer message counts, roles, request IDs, and redacted diagnostics over logging full personal conversations.

Related Spring AI Guides

The useful distinction is concrete: a stored message is available data; an injected message is model context. Trace the path between them before treating a memory symptom as a model problem.

2 Comments

Leave a Reply

Your email address will not be published. Required fields are marked *