Text Embeddings for RAG: From Token Lookup to Retrieval Quality

How do tokens become useful retrieval vectors? This article explains embedding lookup, contextual encoding, and retrieval training, then uses a small experiment to explore ranking quality, encoding pitfalls, and model evaluation.

A retrieval system can generate vectors successfully and still return the wrong passage.

In a customer support knowledge base, a question about refunds after settlement might retrieve a paragraph about settlement schedules. The wording overlaps, the topic is related, and the similarity score may look convincing. Yet the passage does not answer the question.

Understanding this gap requires looking beyond the statement that “embeddings turn text into numbers.”

In the previous article, Making RAG Reliable: Document Quality, Retrieval, and Evidence in a Customer Support Assistant, I discussed the evidence pipeline around retrieval. This article focuses on the representation itself: how token vectors become text vectors, why training matters, and how to investigate retrieval quality before changing the rest of the system.

The discussion uses a conventional single-vector dense retrieval setup. Other approaches, including sparse and multi-vector retrieval, use different representations and scoring methods.

Token Lookup Is Only the Beginning

The word embedding appears at several levels of a language-model application.

Inside a typical Transformer, an input embedding layer maps token IDs to learned vectors. In a retrieval API, an embedding usually refers to a representation of an entire query or passage.

Those representations are connected, but they are not interchangeable.

A tokenizer first converts text into tokens and their corresponding IDs. Tokens are not necessarily words. Depending on the tokenizer, a word may become several subword pieces, and multilingual text may be segmented in ways that do not resemble human word boundaries.

The IDs identify entries in a vocabulary. Their numerical order has no semantic meaning: token 8,000 is not “more meaningful” than token 800.

Suppose an illustrative model has a vocabulary of V tokens and an embedding width of H. Its input embedding matrix has shape:

Embedding matrix: V × H

For a token ID i, the lookup selects row i:

initial_token_vector = embedding_matrix[i]

For T input tokens, the resulting sequence has shape T × H, ignoring the batch dimension.

With fixed weights, the same token ID selects the same initial lookup vector. Context-dependent meaning develops later, as the representations pass through the Transformer and interact with other positions.

For example, if “apple” is represented by the same token ID in two inputs, its initial lookup vector is the same whether the sentence concerns fruit or a technology company. Its later contextual representation can differ.

An illustrative single-vector encoder. Token IDs are examples, and the output dimension D may equal the hidden width H.

A text embedding model then derives an overall representation from the contextual token representations.

Depending on the model, this may involve a designated token, masked mean pooling, another aggregation method, or a learned projection. Mean pooling normally excludes padding positions so that the representation does not depend on how much padding was added to a batch.

Pooling does not necessarily reduce the feature dimension. Aggregating T × H token representations can produce one H-dimensional vector. A separate projection may change that width to D, but D is not inherently much smaller than H.

The pooling and normalization configuration belongs to the model. Arbitrarily replacing its pooling method changes the representation and can damage retrieval quality. Sentence Transformers’ model-construction documentation explains how these components fit together.

Training Determines What “Close” Means

A Transformer followed by pooling produces a vector. That alone does not establish that distances between those vectors are useful for retrieval.

A language model trained to predict tokens learns representations that support its training objective. A retrieval model needs representations that help distinguish relevant passages from competing candidates.

Contrastive learning is one common way to develop that behavior. Training examples associate a query with relevant passages and negative examples. The objective encourages the relevant pair to receive a higher score than negatives.

For an illustrative query such as “How are refunds handled after settlement?”, a passage explaining post-settlement refunds is relevant. A passage about settlement schedules shares vocabulary but does not answer the question. That makes it a useful potential hard negative.

The training objective uses labeled relationships to shape retrieval scores. It does not guarantee correct ordering for every future query.

Hard-negative selection needs care. A passage that contains an alternative valid answer should not automatically be treated as irrelevant. Poor labels can teach the model to separate evidence that the application actually needs together.

Retrieval models may use shared query and document encoder weights or a jointly trained compatible pair. The important property is that their outputs are designed to be compared.

There are also training objectives beyond a simple query-positive-negative example. The Sentence Transformers loss overview describes different objectives for different training data.

This distinction helps explain why averaging arbitrary language-model outputs is not a substitute for evaluating a retrieval-trained model. The shape of the output tells us how many numbers were produced. Training and evaluation tell us whether those numbers are useful.

It also explains a limitation: two statements with opposite conclusions may remain close because they discuss the same topic. “Refunds are supported” and “Refunds are not supported” both concern refund eligibility. Retrieval identifies candidate evidence; the application still needs to interpret conditions, negation, and policy applicability.

Similarity Is a Ranking Signal, Not a Confidence Percentage

Common vector comparisons include cosine similarity, dot product, and Euclidean distance.

Cosine similarity compares direction. Dot product depends on both direction and magnitude. Euclidean distance measures separation in the vector space.

For nonzero vectors that have both been L2-normalized to unit length:

cosine(q, d) = q · d

squared_euclidean_distance(q, d) = 2 − 2(q · d)

Consequently, exact comparisons over the same normalized candidates produce equivalent rankings when cosine or dot product is maximized, or Euclidean distance is minimized. Approximate indexes, quantization, numerical precision, and tie handling can introduce differences in actual systems.

Normalization should still follow the model’s intended use. It is not automatically appropriate for every representation or scoring method.

A similarity score of 0.85 does not mean an answer is 85% likely to be correct. Score distributions depend on the model, training objective, inputs, and scoring method. A threshold calibrated for one model should not be copied to another without evaluation.

Even a high similarity score cannot establish that a passage is current, authorized, or applicable to a particular order. Those checks belong to the evidence pipeline described in the previous article.

A Small Experiment Before Adding a Vector Database

For an initial investigation, I would remove the approximate index from the equation.

A small collection can be encoded and searched by directly comparing every passage vector with the query vector. This makes it easier to examine the encoder and scoring behavior without introducing index parameters.

The following example uses BGE-M3 through Sentence Transformers. It exercises its dense single-vector representation; it does not demonstrate the model’s sparse or multi-vector retrieval capabilities. The BGE-M3 model card documents those distinctions and the supported usage.

The policy statements below are fictional. Each describes a deliberately narrow condition so that the expected evidence can be inspected.

Install the required package in an isolated Python environment:

python -m pip install sentence-transformers

Save this as embedding_demo.py:

import os
from importlib.metadata import version

import numpy as np
from sentence_transformers import SentenceTransformer

MODEL_ID = "BAAI/bge-m3"

# Set this to an immutable model commit for repeatable comparisons.
revision = os.getenv("EMBEDDING_MODEL_REVISION")
model_options = {"revision": revision} if revision else {}

model = SentenceTransformer(MODEL_ID, **model_options)

documents = [
    {
        "id": "before-settlement",
        "text": (
            "Before settlement, eligible course orders can be "
            "refunded to the original payment method."
        ),
    },
    {
        "id": "after-settlement",
        "text": (
            "After settlement, course refunds require a manual "
            "application and finance approval. Automatic refunds "
            "to the original payment method are unavailable."
        ),
    },
    {
        "id": "settlement-timing",
        "text": (
            "Settlement runs at the end of each business day. "
            "Payout timing depends on the payment provider."
        ),
    },
    {
        "id": "promo-package",
        "text": (
            "For package PROMO-2026-A, cancellation requests must "
            "be reviewed under the promotional terms accepted "
            "when the package was purchased."
        ),
    },
    {
        "id": "password-reset",
        "text": (
            "Users who forget their password can request a reset "
            "link from the sign-in page."
        ),
    },
]

cases = [
    (
        "How can I request a refund after settlement?",
        "after-settlement",
    ),
    (
        "The order has already settled. How do I get my money back?",
        "after-settlement",
    ),
    (
        "Can a settled order be refunded automatically "
        "to the original payment method?",
        "after-settlement",
    ),
    (
        "What cancellation rules apply to PROMO-2026-A?",
        "promo-package",
    ),
]

passage_vectors = model.encode(
    [document["text"] for document in documents],
    normalize_embeddings=True,
    convert_to_numpy=True,
)

query_vectors = model.encode(
    [query for query, _ in cases],
    normalize_embeddings=True,
    convert_to_numpy=True,
)

# Dot product equals cosine similarity for these unit vectors.
scores = query_vectors @ passage_vectors.T

print("Model:", MODEL_ID)
print("Requested revision:", revision or "default; not pinned")
print("sentence-transformers:", version("sentence-transformers"))
print("transformers:", version("transformers"))
print("torch:", version("torch"))
print("Configured maximum sequence length:", model.max_seq_length)
print("Passage vector shape:", passage_vectors.shape)

top1_hits = 0

for row, (query, expected_id) in enumerate(cases):
    ranking = np.argsort(-scores[row], kind="stable")
    expected_position = next(
        position
        for position, index in enumerate(ranking, start=1)
        if documents[int(index)]["id"] == expected_id
    )

    top1_hits += int(expected_position == 1)

    print(f"\nQuery: {query}")
    print(f"Expected evidence: {expected_id}")
    print(f"Expected evidence rank: {expected_position}")

    for index in ranking[:3]:
        index = int(index)
        print(
            f"  {documents[index]['id']}: "
            f"{float(scores[row, index]):.4f}"
        )

print(f"\nTop-1 evidence hits: {top1_hits}/{len(cases)}")

Run it with:

python embedding_demo.py

The first run may download the model. This is not a minimal-size model, so allow for its storage and runtime memory requirements.

No similarity scores are supplied here as measured results. The script prints the values produced in your environment. Before comparing runs or publishing a result, pin the model revision and dependencies; an installation command without fixed versions does not fully specify an experiment.

The useful observation is whether the expected passage appears where it should, and which competing passage outranks it when it does not.

For example, if the settlement-timing passage ranks above the refund rule, inspect whether shared vocabulary is dominating answer relevance. If a paraphrased query performs differently from the direct question, keep both in the evaluation set.

The four cases are a diagnostic starting point. Even perfect results would not establish production readiness.

Encoding Details Can Change the Result

A model name is only part of the encoding configuration.

Some models require query and passage prefixes or task instructions. E5, for example, documents the use of query: and passage: for asymmetric retrieval. BGE-M3 uses different guidance. Prefix conventions should come from the selected model’s instructions, not from a generic RAG template.

Input truncation is another important source of failure. A passage can be accepted by an encoding API while its ending is omitted because it exceeds the configured input limit. If the exception to a refund rule appears at the end, the vector may represent a materially incomplete passage.

Measure token length with the actual tokenizer, including any instructions and special tokens. A character count does not reliably predict model input length.

A useful truncation experiment places the decisive policy condition near the beginning and then near the end of an otherwise similar long passage. Record whether either input exceeds the encoding limit before interpreting the ranking difference. Position effects and actual truncation are different issues.

Pooling, normalization, output dimension, and numeric precision also belong in the recorded configuration. The same output dimension does not make two models’ vectors compatible.

When changing models, build a separate version of the document embeddings and route queries to the matching encoder configuration. Evaluate that version before switching traffic, and preserve a rollback path. Merely changing the query model leaves the stored document vectors in the old space.

The same caution applies to manually shortening vectors. Truncation of dimensions is appropriate only when the model or API supports that representation and the resulting quality has been evaluated.

Choose a Model by Its Failures and Operating Cost

A public benchmark can help identify candidates, but it does not reproduce a support knowledge base’s product names, policy exceptions, or language mix.

I would compare models using a fixed set of queries and evidence labels that includes paraphrases, exact identifiers, negation, multilingual questions where relevant, and passages with conditions that change the answer.

Keep a portion of the questions separate from tuning. Otherwise, repeatedly adjusting the system around visible examples can create an improvement that does not generalize.

Start with exact similarity search over the same passages to examine encoder behavior. Then evaluate the approximate index separately. A passage missing from an ANN result does not necessarily mean the embedding model failed to rank it well under exact comparison.

Measure evidence coverage and ranking alongside query-encoding latency, ingestion throughput, memory demand, and index size. A larger vector is not automatically a better representation.

For a concrete capacity estimate, one million 1,024-dimensional float32 vectors contain about 4.096 GB of raw vector values. That arithmetic excludes document text, metadata, index structures, replicas, and other overhead. It is a storage baseline, not a deployment memory estimate.

Keep tools and models distinct when making the choice:

LayerWhat it determines
Model checkpointLearned representations and supported encoding conventions
Encoding frameworkHow tokenization, pooling, and inference are executed
Serving runtimeBatching, concurrency, hardware use, and operational behavior
Search indexCandidate-search speed, filtering behavior, and approximation tradeoffs

A framework such as Sentence Transformers is not restricted to English simply because one frequently used model is English-focused. Language suitability belongs primarily to the chosen model and its demonstrated behavior on the application’s data.

Finally, avoid expecting embeddings to replace other controls. Exact order identifiers may be better served by structured lookup. Document permissions require enforcement. A relevant refund paragraph does not establish that a specific customer owns the order or qualifies for approval.

Embeddings make text comparable under a learned representation. Understanding the lookup layer explains where the initial numbers come from. Understanding contextual encoding and training explains why those numbers may help retrieval. Testing realistic failures shows whether they help enough for the application being built.

Leave a Reply

Your email address will not be published. Required fields are marked *