Skip to content

Retrieval (RAG)

Answer questions from your own documents, find them by meaning, and let the user see where each answer came from.

Retrieval-augmented generation (RAG) answers from your documents and not from whatever the model happens to remember. Each passage is turned into an embedding, a vector that captures what it is about, and stored in a vector database. A question gets the same treatment, the database returns the passages nearest in meaning, and the model answers from those and lists their ids as its sources. Because the match is on meaning, a customer's wording doesn't have to match the document's. A check on the way out rejects any source the model did not actually retrieve, so every citation points at something it saw.

Use it when

  • Answers must come from your documents and not from the model's memory.
  • Users ask in their own words, not the documents'.
  • Users need to see where an answer came from.
  • "I couldn't find that" is better than a guess.

Look elsewhere when

  • The documents are small enough to put in the prompt: do that and skip the database.
  • The answer needs a calculation or a query over structured data: use tools (tool_calling, code_mode).
documents ─ embed ─▶ Chroma (its own service)
                          ▲
question ─ embed ─▶ search_docs ─▶ nearest passages ─▶ answer + sources ─▶ citation check

What it shows

  • Search by meaning, in a real vector database. Each passage is embedded (a vector that captures what it is about) and stored in Chroma, which runs as its own Docker service. search_docs embeds the question the same way and asks Chroma for the nearest passages. "How long can I send my hiking footwear back for my money?" shares no word with the returns policy, and finds it first
  • It is measured, not assumed. test_live.py runs 15 questions against a keyword baseline with real embeddings: the right passage came first for 15 of 15 by vector search and 12 of 15 by keyword search (which found nothing for the paraphrase above, and the wrong passage for another)
  • An index you do not manage by hand. ensure_index names the collection after the embedding model and a fingerprint of the documents, and fills it only if it is not full already. Edit a document or change the model and a new collection is built; the next run does nothing if it is already there
  • Nearest is not relevant. A vector search always returns something. MAX_DISTANCE drops the clearly unrelated passages (the right one was never farther than 0.31 in the 15 test questions; unrelated questions were 0.43 or more away; MAX_DISTANCE is 0.37, between the two), and the prompt tells the model to answer only from a passage that really says it. "Do you sell tents?" returns the nearest passages and the model still says it could not find it
  • Trustworthy citations: an output validator rejects any source the model didn't actually retrieve (RagDeps.retrieved records every id a search returned), so each id the user sees was in a real search result
  • The address is a dependency. RagDeps.chroma_url (from CHROMA_URL) points the agent at the local service or a staging database; ChromaUnavailable says where it looked

See it run: sample_run.md is a recorded run against a real model: what each agent was asked, which tools it called, and what it returned.

Why vectors and not both

The obvious next idea is a hybrid: run keyword and vector search and merge the rankings. We measured it (reciprocal-rank fusion, with the vector side weighted 1x, 2x and 3x) and it did worse than vectors alone on this set: 14 of 15 first, not 15. A single spurious keyword hit (a question about "a charge to get a package delivered quickly" matched the Trail Club passage) outranked the right vector answer however heavily the vector side was weighted. With a strong embedding model and a corpus like this one, keyword search added noise. That can change with a weaker embedding model or a corpus full of exact identifiers, so measure on yours; the keyword baseline in test_live.py is the starting point for that comparison.

Running it

docker compose -f examples/rag/service/docker-compose.yml up -d --wait
export CHROMA_URL=http://$(docker compose -f examples/rag/service/docker-compose.yml port chroma 8000)
uv run --with "chromadb-client>=1.5,<1.6" python -m examples.rag.agent
docker compose -f examples/rag/service/docker-compose.yml down

The embedding model follows your LLM's provider: Google or OpenAI work out of the box. Anthropic has no embedding model, so with an Anthropic LLM set AGENT_EMBEDDING_MODEL in .env (for example google:gemini-embedding-001 or openai:text-embedding-3-small; the key for that provider must be set too). The first run embeds the documents; later runs reuse them.

To use it in your project, add_agent.py copies the agent into agent/agents/, the service into services/<name>/, and installs chromadb-client:

uv run python scripts/add_agent.py rag --name support_docs

To adapt it, replace KNOWLEDGE_BASE with your own documents (long ones should be split into passages first) and rewrite prompts/rag.txt for your domain. Re-measure MAX_DISTANCE for your embedding model: distances are not comparable between models.

Notes

  • The tests use the real service. Chroma's slim client and its full engine cannot be installed together (both provide the chromadb module), so there is no in-process Chroma for the offline tests. They run against the Docker service with a scripted embedder (each text maps to a vector the test chose, so distances are known exactly) and are skipped without CHROMA_URL. The live tests then use real embeddings
  • Old collections stay. A new embedding model or edited documents create a new collection; the old one is left in the database. Delete it when you no longer need it
  • Pin the server to the client's version. service/docker-compose.yml names chromadb/chroma:1.5.9 to match chromadb-client 1.5; latest was a minor version behind the client when this was written

Source

All of it is in examples/rag/.

"""Retrieval-augmented generation: answer from documents by meaning, and cite only what was retrieved.

Use this pattern when:
- Answers must come from your own documents, not the model's memory
- Users ask in their own words, not the documents' (so matching on keywords misses)
- Users need to see where an answer came from
- The model should say "I couldn't find that" rather than guess

How it works:
    1. Each passage is turned into an *embedding* (a vector that captures its meaning) and stored in
       a vector database, Chroma, which runs as its own service. This happens once per set of
       documents and embedding model (`ensure_index`)
    2. A retrieval tool (`search_docs`) embeds the question the same way and asks Chroma for the
       passages whose vectors are nearest, dropping any that are not near enough
    3. The model answers from those passages and lists their ids as `sources`
    4. An output validator rejects any source the model did not actually retrieve, so a citation
       can be trusted: it points at something the model really saw

"How long can I send my hiking footwear back for my money?" shares no word with the returns policy
("Unworn items can be returned within 45 days…"), so a keyword search finds nothing and a vector
search finds it first. The tests measure that against a keyword baseline.

Two things to know about vector search. It always returns the *nearest* passages, even when none
answers the question, so `MAX_DISTANCE` drops the clearly unrelated ones and the model is told to
answer only if a passage really contains the answer. And the embedding model is part of the index: a
different model makes vectors of a different size and meaning, so the index is named after both the
model and the documents, and a changed set of either is re-indexed into a new collection.

The policies are fictional on purpose: a model cannot answer from memory, so a correct answer proves
retrieval worked. Replace `KNOWLEDGE_BASE` with your own documents.
"""

from __future__ import annotations

import asyncio
import hashlib
import os
from dataclasses import dataclass, field
from typing import Any
from urllib.parse import urlparse

import chromadb
from pydantic import BaseModel
from pydantic_ai import Agent, Embedder, ModelRetry, RunContext
from pydantic_ai.capabilities import RaiseContentFilterError
from pydantic_ai.embeddings import TestEmbeddingModel
from pydantic_ai.usage import UsageLimits

from agent.config import settings
from agent.logging import agent_label, configure_logging, get_logger
from agent.prompts.templates import load_prompt
from agent.runs import Flow, RunResult

logger = get_logger(__name__)
LABEL = agent_label(__name__)  # names this agent's run spans in Logfire traces

# Guardrail against runaway agentic loops (see examples/single for the details). A multi-part
# question may search several times, so the request limit has headroom.
USAGE_LIMITS = UsageLimits(
    request_limit=12, total_tokens_limit=100_000, cost_limit=settings.cost_limit
)

MAX_RESULTS = 3  # passages returned per search; keep context small and relevant
# A passage farther than this (cosine distance: 0 is identical, 1 is unrelated) is not returned. With
# Gemini embeddings the right passage was never farther than 0.31 across the 15 test questions, and
# questions with no connection to the documents were 0.43 or more away; this sits between the two.
# Distances are not comparable between embedding models, so re-measure when you change yours.
MAX_DISTANCE = 0.37

# Chroma's own port, for a server you start yourself on it. The Docker service publishes on a random free
# host port instead, so set CHROMA_URL to the address `docker compose port` shows (see service/).
DEFAULT_CHROMA_URL = "http://127.0.0.1:8000"
# The embedding model to use with each LLM provider's key, when AGENT_EMBEDDING_MODEL is not set.
# (Anthropic has no embedding model, so with an Anthropic LLM you must choose one yourself.)
DEFAULT_EMBEDDING_MODELS = {
    "google": "google:gemini-embedding-001",
    "google-gla": "google:gemini-embedding-001",
    "openai": "openai:text-embedding-3-small",
}


# --- The knowledge base ---
@dataclass(frozen=True)
class Passage:
    id: str
    title: str
    text: str


# Replace with your own documents. Fictional here, so correct answers can only come from retrieval.
KNOWLEDGE_BASE: tuple[Passage, ...] = (
    Passage(
        "returns-policy",
        "Returns",
        "Unworn items can be returned within 45 days of delivery for a full refund. Boots must be "
        "returned in their original box. Sale items are final sale and cannot be returned.",
    ),
    Passage(
        "shipping-rates",
        "Standard shipping",
        "Standard shipping is free on orders over $75 and costs $6.95 on smaller orders. Orders "
        "ship within 2 business days.",
    ),
    Passage(
        "express-shipping",
        "Express shipping",
        "Express shipping (2-day) costs $18. It is not available to Alaska, Hawaii or PO boxes.",
    ),
    Passage(
        "warranty-boots",
        "Boot warranty",
        "Birchwood boots carry a 3-year warranty against defects in materials and workmanship. "
        "Normal wear is not covered.",
    ),
    Passage(
        "warranty-packs",
        "Pack warranty",
        "Backpacks and duffels carry a lifetime warranty on zippers and stitching.",
    ),
    Passage(
        "store-hours",
        "Store hours",
        "The Portland flagship store is open 9am to 7pm Monday to Saturday and 11am to 5pm on "
        "Sunday.",
    ),
    Passage(
        "gift-cards",
        "Gift cards",
        "Gift cards never expire and can be used online or in store. They cannot be redeemed for "
        "cash.",
    ),
    Passage(
        "price-match",
        "Price matching",
        "We match the price of any authorized retailer on identical in-stock items within 14 days "
        "of purchase.",
    ),
    Passage(
        "loyalty",
        "Trail Club",
        "Trail Club members earn 1 point per dollar and get free shipping on every order, "
        "whatever the amount.",
    ),
    Passage(
        "repairs",
        "Repairs",
        "We repair boots and packs (soles, zippers, straps) for a flat $25 fee at any store or "
        "by mail.",
    ),
    Passage(
        "recall-rn-2291",
        "Recall notice RN-2291",
        "Recall notice RN-2291: Birchwood Ridgeline boots from production lot 7B-114 may have a "
        "loose heel counter. Stop wearing them and contact us for a free replacement pair. Other "
        "lots are not affected.",
    ),
    Passage(
        "recall-rn-2290",
        "Recall notice RN-2290",
        "Recall notice RN-2290: Birchwood Summit gloves from production lot 4C-087 may shed lining "
        "fibres. Stop using them and contact us for a refund. Other lots are not affected.",
    ),
)


# --- Embeddings and the vector database ---
class EmbeddingModelNotConfigured(Exception):
    """No embedding model is configured and none can be inferred from the LLM provider."""


class ChromaUnavailable(Exception):
    """The Chroma server could not be reached."""


def chroma_url_from_env() -> str:
    return os.environ.get("CHROMA_URL", DEFAULT_CHROMA_URL)


def embedding_model() -> Any:
    """The embedding model to use: AGENT_EMBEDDING_MODEL, else the one that goes with the LLM's provider.

    Under the offline test model (`AGENT_MODEL=test`) this is Pydantic AI's `TestEmbeddingModel`, so
    nothing calls a real provider, just as the agent itself does not.
    """
    configured = os.environ.get("AGENT_EMBEDDING_MODEL")
    if configured:
        return configured
    provider = settings.model.split(":")[0]
    if provider == "test":
        return TestEmbeddingModel()
    if provider in DEFAULT_EMBEDDING_MODELS:
        return DEFAULT_EMBEDDING_MODELS[provider]
    raise EmbeddingModelNotConfigured(
        f"No embedding model for the {provider!r} provider. Set AGENT_EMBEDDING_MODEL in .env, "
        "for example google:gemini-embedding-001 or openai:text-embedding-3-small."
    )


def connect(url: str) -> Any:
    """A client for the Chroma server at `url`, checked with a heartbeat."""
    parts = urlparse(url)
    secure = parts.scheme == "https"
    try:
        client = chromadb.HttpClient(
            host=parts.hostname or "127.0.0.1",
            port=parts.port or (443 if secure else 8000),
            ssl=secure,
        )
        client.heartbeat()
    except ValueError as exc:  # Chroma reports a refused connection as a ValueError
        if "connect" not in str(exc).lower():
            raise
        raise ChromaUnavailable(
            f"No Chroma server at {url}. Start the service (`docker compose up -d --wait` in the "
            "directory with its docker-compose.yml) and set CHROMA_URL to the address Docker "
            "published, or point CHROMA_URL at your own Chroma server."
        ) from exc
    return client


def index_name(embedding_model_id: str, knowledge_base: tuple[Passage, ...]) -> str:
    """The collection's name: a fingerprint of the embedding model and every passage.

    Changing either gives a new name, so a new collection is built; the old one is left alone.
    """
    digest = hashlib.sha256(embedding_model_id.encode())
    for p in knowledge_base:
        digest.update(f"\0{p.id}\0{p.title}\0{p.text}".encode())
    return f"kb-{digest.hexdigest()[:16]}"


# --- Dependencies ---
@dataclass
class RagDeps:
    """Runtime dependencies for the retrieval agent."""

    knowledge_base: tuple[Passage, ...] = KNOWLEDGE_BASE
    chroma_url: str = field(default_factory=chroma_url_from_env)
    # Ids of the passages retrieved during this run; citations are checked against it. A fresh
    # RagDeps per run (the default in run_rag) keeps one question's sources from leaking into the next.
    retrieved: set[str] = field(default_factory=set)
    # Normally left unset: built from the settings above on first use. Set them to use another
    # Chroma client or embedding model, such as an in-process client in tests.
    client: Any = None
    embedder: Embedder | None = None
    collection: Any = None  # the index; filled in by ensure_index


async def ensure_index(deps: RagDeps) -> Any:
    """The collection holding `deps.knowledge_base`'s embeddings, building it if it is not there yet.

    Safe to call on every run: a collection with all the passages is used as it is.
    """
    if deps.collection is not None:
        return deps.collection
    if deps.embedder is None:
        deps.embedder = Embedder(embedding_model())
    if deps.client is None:
        deps.client = await asyncio.to_thread(connect, deps.chroma_url)
    model = deps.embedder.model  # the name it was given, or the model instance
    model_id = model if isinstance(model, str) else f"{model.system}:{model.model_name}"
    name = index_name(model_id, deps.knowledge_base)
    collection = await asyncio.to_thread(
        deps.client.get_or_create_collection, name, configuration={"hnsw": {"space": "cosine"}}
    )
    if await asyncio.to_thread(collection.count) != len(deps.knowledge_base):
        # A passage is embedded with its title, which is part of what it is about.
        embedded = await deps.embedder.embed_documents(
            [f"{p.title}. {p.text}" for p in deps.knowledge_base]
        )
        await asyncio.to_thread(
            collection.upsert,
            ids=[p.id for p in deps.knowledge_base],
            documents=[p.text for p in deps.knowledge_base],
            metadatas=[{"title": p.title} for p in deps.knowledge_base],
            embeddings=[list(vector) for vector in embedded.embeddings],
        )
        logger.info(
            "Indexed passages", extra={"collection": name, "count": len(deps.knowledge_base)}
        )
    deps.collection = collection
    return collection


@dataclass(frozen=True)
class Hit:
    """A passage the search returned, and how far its vector is from the question's."""

    id: str
    title: str
    text: str
    distance: float


async def vector_search(deps: RagDeps, query: str, limit: int = MAX_RESULTS) -> list[Hit]:
    """The passages nearest in meaning to `query`, nearest first, none farther than MAX_DISTANCE."""
    collection = await ensure_index(deps)
    assert deps.embedder is not None  # ensure_index sets it
    embedded = await deps.embedder.embed_query(query)
    found = await asyncio.to_thread(
        collection.query,
        query_embeddings=[list(embedded.embeddings[0])],
        n_results=limit,
        include=["documents", "metadatas", "distances"],
    )
    return [
        Hit(id=pid, title=meta["title"], text=text, distance=distance)
        for pid, text, meta, distance in zip(
            found["ids"][0],
            found["documents"][0],
            found["metadatas"][0],
            found["distances"][0],
            strict=True,
        )
        if distance <= MAX_DISTANCE
    ]


# --- Output type ---
class Answer(BaseModel):
    # `result` is the conventional output field in these examples; the generated
    # eval starter reads it when present (see evals/helpers.py).
    result: str
    sources: list[str] = []


# --- Agent ---
rag_agent: Agent[RagDeps, Answer] = Agent(
    settings.model,
    name=LABEL,
    output_type=Answer,
    deps_type=RagDeps,
    capabilities=[RaiseContentFilterError()],
    instructions=load_prompt("rag"),  # prompts/rag.txt; copied to agent/prompts/<name>.txt
)


# --- Tools ---
@rag_agent.tool
async def search_docs(ctx: RunContext[RagDeps], query: str) -> str:
    """Search the company documentation for passages relevant to a question or topic.

    Args:
        query: What to look for: the question itself or a short description of the topic.

    Returns:
        Up to three passages, each as `[id] Title: text`, nearest in meaning first; or a note that
        nothing relevant was found.

    Raises:
        ModelRetry: When the query is empty.
    """
    if not query.strip():
        raise ModelRetry(
            "The query is empty. Search for the question or a description of the topic."
        )
    hits = await vector_search(ctx.deps, query)
    logger.info(
        "Search", extra={"query": query, "hits": [(h.id, round(h.distance, 3)) for h in hits]}
    )
    if not hits:
        return "No relevant passages found. Rephrase the query, or tell the user you could not find it."
    ctx.deps.retrieved.update(hit.id for hit in hits)
    return "\n\n".join(f"[{hit.id}] {hit.title}: {hit.text}" for hit in hits)


# --- Grounding ---
@rag_agent.output_validator
def check_citations(ctx: RunContext[RagDeps], answer: Answer) -> Answer:
    """Reject a citation the model never retrieved; the message goes back to the model.

    This is what makes `sources` trustworthy: every id the user sees was in a search result.
    """
    invented = [source for source in answer.sources if source not in ctx.deps.retrieved]
    if invented:
        raise ModelRetry(
            f"You cited {invented}, which no search returned. Cite only passages returned by "
            "search_docs, or leave sources empty if nothing relevant was found."
        )
    return answer


async def run_rag(user_input: str, deps: RagDeps | None = None) -> RunResult[Answer]:
    """Answer `user_input` from the knowledge base.

    Returns:
        A RunResult: `.output` is the `Answer` (text and the ids it relies on); the searches the
        model made are in `.all_messages()`.

    Raises:
        ChromaUnavailable: When the Chroma server can't be reached.
        EmbeddingModelNotConfigured: When no embedding model is set and none fits the LLM provider.
    """
    if deps is None:
        deps = RagDeps()
    logger.info("Running retrieval agent", extra={"user_input": user_input})
    await ensure_index(
        deps
    )  # fail before any model call if the database or embeddings are not usable
    flow = Flow(USAGE_LIMITS)
    result = await flow.run(rag_agent, user_input, deps=deps)
    return flow.finish(result.output)


if __name__ == "__main__":
    configure_logging()
    print(asyncio.run(run_rag("How long can I send my hiking footwear back for my money?")).output)
You answer customer questions about Birchwood Outfitters using only the company's documentation.

- Always call search_docs before answering, and answer only from the passages it returns. Never
  use outside knowledge about the company.
- Search by meaning: pass the customer's question, or a short description of the topic. You may
  search more than once, for example one search per topic in a multi-part question.
- The search returns the passages nearest in meaning to the question, and nearest is not the same as
  relevant: a passage can be returned without containing the answer. Answer only from a passage that
  actually says it. If none of them does, say you could not find it in the documentation.
- In `sources`, list the id of every passage your answer relies on, exactly as shown in brackets, and
  nothing else. Leave `sources` empty when you could not find the answer.
title = "Retrieval (RAG)"
pattern = "rag"
summary = "Answer from your own documents by meaning, with embeddings in a Chroma vector database running as a service, and cite only passages the model really retrieved."
# Worded to share no key word with the returns policy, so only a search by meaning finds it.
smoke_input = "How long can I send my hiking footwear back for my money?"
expected_tools = ["search_docs"]

# The agent is only a client of the database: the slim HTTP client, not the whole engine. (The full
# `chromadb` package cannot be installed beside it: both provide the `chromadb` module.) The tests
# need the real server too, so they are skipped without CHROMA_URL, like the service below.
dependencies = ["chromadb-client>=1.5,<1.6"]

# Optional: the embedding model. By default it follows your LLM's provider (Google or OpenAI).
env = ["AGENT_EMBEDDING_MODEL"]

# The database runs as a docker-compose service (service/). The release check starts it, finds the
# port Docker chose, and passes the address to the example in CHROMA_URL.
services = ["chroma"]

[service.chroma]
port = 8000
env = "CHROMA_URL"
url = "http://{address}"

[smoke.rag_agent]
output = { result = "Returns are accepted within 45 days.", sources = [] }

[entrypoint]
deps = "RagDeps"
run = "run_rag"
"""The index, the vector search, the tool and the citation check, against a real Chroma server.

The tests that need the database run against the example's Docker service: the release check starts
it and passes its address in CHROMA_URL (see example.toml), and without it they are skipped. They
use a *scripted* embedder, which maps each text to a vector chosen by the test, so what is checked
is Chroma's real behaviour (cosine distance, ranking, counts, upserts) with exactly known inputs and
no embedding provider. Meaning-level quality of real embeddings is what test_live.py measures.
Needs chromadb-client (declared in example.toml), so the module is skipped without it.
"""

import math
import os
import uuid
from collections.abc import Sequence

import pytest

pytest.importorskip("chromadb", reason="needs chromadb-client")

from pydantic_ai import Embedder, ModelRetry, RunContext  # noqa: E402
from pydantic_ai.embeddings import EmbeddingModel, EmbeddingResult, TestEmbeddingModel  # noqa: E402
from pydantic_ai.messages import ModelResponse, RetryPromptPart, ToolCallPart  # noqa: E402
from pydantic_ai.models.function import AgentInfo, FunctionModel  # noqa: E402
from pydantic_ai.models.test import TestModel  # noqa: E402
from pydantic_ai.usage import RequestUsage, RunUsage  # noqa: E402

from examples.rag import agent as module  # noqa: E402
from examples.rag.agent import (  # noqa: E402
    DEFAULT_EMBEDDING_MODELS,
    KNOWLEDGE_BASE,
    MAX_DISTANCE,
    MAX_RESULTS,
    Answer,
    ChromaUnavailable,
    EmbeddingModelNotConfigured,
    Passage,
    RagDeps,
    check_citations,
    chroma_url_from_env,
    connect,
    embedding_model,
    ensure_index,
    index_name,
    rag_agent,
    run_rag,
    search_docs,
    vector_search,
)

# --- A scripted embedder ---


def at(degrees: float) -> list[float]:
    """A unit vector at an angle: two such vectors are as far apart (cosine) as their angle says."""
    return [math.cos(math.radians(degrees)), math.sin(math.radians(degrees))]


def cosine_distance(a_degrees: float, b_degrees: float) -> float:
    return 1 - math.cos(math.radians(a_degrees - b_degrees))


class ScriptedEmbeddings(EmbeddingModel):
    """Embeds each text as the vector the test gave it, and counts how many texts it was asked for."""

    def __init__(self, vectors: dict[str, list[float]], name: str | None = None):
        super().__init__()
        self.vectors = vectors
        self.name = name or f"scripted-{uuid.uuid4().hex[:8]}"  # a fresh index per test
        self.texts_embedded = 0

    @property
    def model_name(self) -> str:
        return self.name

    @property
    def system(self) -> str:
        return "scripted"

    async def embed(
        self, inputs: str | Sequence[str], *, input_type, settings=None
    ) -> EmbeddingResult:
        texts, _ = self.prepare_embed(inputs, settings)
        self.texts_embedded += len(texts)
        return EmbeddingResult(
            embeddings=[self.vectors[text] for text in texts],
            inputs=texts,
            input_type=input_type,
            usage=RequestUsage(input_tokens=len(texts)),
            model_name=self.name,
            provider_name=self.system,
            provider_response_id=str(uuid.uuid4()),
        )

    async def max_input_tokens(self) -> int | None:
        return 1024

    async def count_tokens(self, text: str) -> int:
        return len(text.split())


def text_of(p: Passage) -> str:
    return f"{p.title}. {p.text}"  # what the agent embeds for a passage


# Three passages at 0, 30 and 90 degrees. A question at 10 degrees is 0.015 from the first, 0.060
# from the second and 0.826 from the third, which is beyond MAX_DISTANCE.
ALPHA = Passage("alpha", "Alpha", "The first passage.")
BETA = Passage("beta", "Beta", "The second passage.")
GAMMA = Passage("gamma", "Gamma", "The third passage.")
SMALL_KB = (ALPHA, BETA, GAMMA)
SMALL_VECTORS = {text_of(ALPHA): at(0), text_of(BETA): at(30), text_of(GAMMA): at(90)}


@pytest.fixture
def address() -> str:
    server = os.environ.get("CHROMA_URL")
    if not server:
        pytest.skip("needs the Chroma service running (CHROMA_URL)")
    return server


def make_deps(
    address: str, kb=SMALL_KB, vectors=None, questions=None
) -> tuple[RagDeps, ScriptedEmbeddings]:
    """Deps for a real Chroma and an embedder that knows the passages and the given questions."""
    model = ScriptedEmbeddings({**(vectors or SMALL_VECTORS), **(questions or {})})
    return RagDeps(knowledge_base=kb, chroma_url=address, embedder=Embedder(model)), model


def ctx(deps: RagDeps) -> RunContext[RagDeps]:
    return RunContext(deps=deps, model=TestModel(), usage=RunUsage())


# --- Choosing and naming things, without the database ---


def test_the_index_is_named_after_the_documents_and_the_model():
    base = index_name("scripted:one", SMALL_KB)
    assert base == index_name("scripted:one", SMALL_KB)  # stable
    assert base.startswith("kb-") and 3 <= len(base) <= 512
    assert base != index_name("scripted:two", SMALL_KB)  # another model: vectors of another meaning
    changed = (ALPHA, Passage("beta", "Beta", "The second passage, edited."), GAMMA)
    assert base != index_name("scripted:one", changed)  # a passage was edited
    assert base != index_name("scripted:one", SMALL_KB[:2])  # a passage was removed


def test_an_embedding_model_can_be_chosen_in_the_environment(monkeypatch):
    monkeypatch.setenv("AGENT_EMBEDDING_MODEL", "openai:text-embedding-3-large")
    assert embedding_model() == "openai:text-embedding-3-large"


@pytest.mark.parametrize(
    ("llm", "expected"),
    [
        ("google:gemini-3.1-flash-lite", "google:gemini-embedding-001"),
        ("google-gla:gemini-3-pro", "google:gemini-embedding-001"),
        ("openai:gpt-5.2", "openai:text-embedding-3-small"),
    ],
)
def test_the_embedding_model_follows_the_llm_provider(monkeypatch, llm, expected):
    monkeypatch.delenv("AGENT_EMBEDDING_MODEL", raising=False)
    monkeypatch.setattr(module.settings, "model", llm)
    assert embedding_model() == expected
    assert expected in DEFAULT_EMBEDDING_MODELS.values()


def test_the_offline_test_model_gets_the_offline_embedding_model(monkeypatch):
    monkeypatch.delenv("AGENT_EMBEDDING_MODEL", raising=False)
    monkeypatch.setattr(module.settings, "model", "test")
    assert isinstance(embedding_model(), TestEmbeddingModel)


def test_a_provider_without_embeddings_asks_you_to_choose_one(monkeypatch):
    monkeypatch.delenv("AGENT_EMBEDDING_MODEL", raising=False)
    monkeypatch.setattr(module.settings, "model", "anthropic:claude-sonnet-5-5")
    with pytest.raises(EmbeddingModelNotConfigured, match="AGENT_EMBEDDING_MODEL"):
        embedding_model()


def test_the_database_address_comes_from_the_environment(monkeypatch):
    monkeypatch.setenv("CHROMA_URL", "http://chroma.internal:9000")
    assert chroma_url_from_env() == "http://chroma.internal:9000"
    assert RagDeps().chroma_url == "http://chroma.internal:9000"
    monkeypatch.delenv("CHROMA_URL")
    assert chroma_url_from_env() == module.DEFAULT_CHROMA_URL


def test_an_unreachable_database_is_reported_with_how_to_start_it():
    with pytest.raises(ChromaUnavailable, match="docker compose"):
        connect("http://127.0.0.1:1")


def test_a_secure_address_uses_tls_and_the_default_port(monkeypatch):
    seen = {}

    class Recording:
        def __init__(self, **kwargs):
            seen.update(kwargs)

        def heartbeat(self):
            return 1

    monkeypatch.setattr(module.chromadb, "HttpClient", Recording)
    connect("https://chroma.example.com")
    assert seen == {"host": "chroma.example.com", "port": 443, "ssl": True}
    connect("http://chroma.example.com")
    assert seen == {"host": "chroma.example.com", "port": 8000, "ssl": False}


def test_a_failure_that_is_not_a_connection_problem_is_not_disguised(monkeypatch):
    def broken(**kwargs):
        raise ValueError("the tenant does not exist")

    monkeypatch.setattr(module.chromadb, "HttpClient", broken)
    with pytest.raises(ValueError, match="tenant"):
        connect("http://127.0.0.1:8000")


async def test_a_run_fails_before_any_model_call_when_the_database_is_down():
    def must_not_run(messages, info):
        raise AssertionError("the model was called although there is no database")

    deps = RagDeps(chroma_url="http://127.0.0.1:1", embedder=Embedder(ScriptedEmbeddings({})))
    with rag_agent.override(model=FunctionModel(must_not_run)), pytest.raises(ChromaUnavailable):
        await run_rag("anything", deps)


# --- The index, in a real Chroma ---


async def test_the_index_holds_every_passage_with_its_text_and_title(address):
    deps, model = make_deps(address)
    collection = await ensure_index(deps)

    stored = collection.get(ids=["alpha", "beta", "gamma"], include=["documents", "metadatas"])
    assert dict(zip(stored["ids"], stored["documents"], strict=True)) == {
        "alpha": "The first passage.",
        "beta": "The second passage.",
        "gamma": "The third passage.",
    }
    assert {m["title"] for m in stored["metadatas"]} == {"Alpha", "Beta", "Gamma"}
    assert collection.count() == 3
    assert model.texts_embedded == 3
    assert collection.configuration_json["hnsw"]["space"] == "cosine"


async def test_an_index_that_is_already_complete_is_used_not_rebuilt(address):
    first, model = make_deps(address)
    await ensure_index(first)
    assert model.texts_embedded == 3

    # A second run with the same documents and model finds the collection full: no embedding calls.
    second = RagDeps(knowledge_base=SMALL_KB, chroma_url=address, embedder=Embedder(model))
    await ensure_index(second)
    assert model.texts_embedded == 3
    assert second.collection.name == first.collection.name


async def test_the_index_is_looked_up_once_per_run(address):
    deps, model = make_deps(address)
    first = await ensure_index(deps)
    assert (
        await ensure_index(deps) is first
    )  # cached on deps: no second lookup, no second embedding
    assert model.texts_embedded == 3


async def test_a_partly_built_index_is_completed(address):
    deps, model = make_deps(address)
    collection = await ensure_index(deps)
    collection.delete(ids=["gamma"])  # e.g. an earlier run that stopped part way
    assert collection.count() == 2

    again = RagDeps(knowledge_base=SMALL_KB, chroma_url=address, embedder=Embedder(model))
    await ensure_index(again)
    assert again.collection.count() == 3


async def test_a_different_embedding_model_gets_its_own_index(address):
    first, model_one = make_deps(address)
    second, model_two = make_deps(address)  # a different model name, so a different collection
    assert (await ensure_index(first)).name != (await ensure_index(second)).name
    assert model_one.texts_embedded == model_two.texts_embedded == 3


async def test_a_connection_is_made_from_the_address_when_no_client_is_given(address):
    deps, _ = make_deps(address)
    assert deps.client is None
    await ensure_index(deps)
    assert deps.client is not None


# --- Searching the index ---


async def test_the_nearest_passages_come_first_with_their_distances(address):
    deps, _ = make_deps(address, questions={"q": at(10)})
    hits = await vector_search(deps, "q")

    assert [h.id for h in hits] == ["alpha", "beta"]  # nearest first; gamma is too far to return
    assert hits[0].distance == pytest.approx(cosine_distance(10, 0), abs=1e-4)
    assert hits[1].distance == pytest.approx(cosine_distance(10, 30), abs=1e-4)
    assert (hits[0].title, hits[0].text) == ("Alpha", "The first passage.")


async def test_a_passage_beyond_the_maximum_distance_is_dropped(address):
    assert (
        cosine_distance(10, 90) > MAX_DISTANCE > cosine_distance(10, 30)
    )  # the setup is as described
    deps, _ = make_deps(address, questions={"far": at(180), "near": at(10)})
    assert await vector_search(deps, "far") == []  # opposite direction: nothing is near enough
    assert "gamma" not in [h.id for h in await vector_search(deps, "near", limit=3)]


async def test_the_number_of_results_is_limited(address):
    deps, _ = make_deps(address, questions={"q": at(10)})
    assert [h.id for h in await vector_search(deps, "q", limit=1)] == ["alpha"]
    assert MAX_RESULTS == 3


# --- The tool ---


async def test_a_search_returns_labelled_passages_and_records_what_was_retrieved(address):
    deps, _ = make_deps(address, questions={"q": at(10)})
    text = await search_docs(ctx(deps), "q")

    assert text == "[alpha] Alpha: The first passage.\n\n[beta] Beta: The second passage."
    assert deps.retrieved == {"alpha", "beta"}


async def test_a_search_with_nothing_near_says_so_and_records_nothing(address):
    deps, _ = make_deps(address, questions={"q": at(180)})
    text = await search_docs(ctx(deps), "q")
    assert text.startswith("No relevant passages found")
    assert deps.retrieved == set()


async def test_an_empty_query_asks_the_model_to_retry():
    with pytest.raises(ModelRetry, match="query is empty"):
        await search_docs(ctx(RagDeps()), "   ")


# --- The citation check ---


def test_citing_what_was_retrieved_is_accepted():
    deps = RagDeps(retrieved={"returns-policy"})
    answer = Answer(result="45 days", sources=["returns-policy"])
    assert check_citations(ctx(deps), answer) is answer


def test_no_sources_is_accepted():
    answer = Answer(result="I could not find that.", sources=[])
    assert check_citations(ctx(RagDeps()), answer) is answer


def test_citing_something_never_retrieved_is_rejected_with_the_offending_id():
    deps = RagDeps(retrieved={"returns-policy"})
    with pytest.raises(ModelRetry, match="invented-id"):
        check_citations(ctx(deps), Answer(result="x", sources=["returns-policy", "invented-id"]))


# --- Through the agent loop, with a scripted model ---


def scripted(*steps: dict):
    """A model that follows `steps`: a tool call ({"tool": ..., "args": ...}) or the final output."""
    remaining = list(steps)

    def model_fn(messages, info: AgentInfo) -> ModelResponse:
        step = remaining.pop(0)
        if "tool" in step:
            return ModelResponse(parts=[ToolCallPart(step["tool"], step["args"])])
        return ModelResponse(parts=[ToolCallPart(info.output_tools[0].name, step["output"])])

    return FunctionModel(model_fn)


async def test_the_model_searches_then_answers_with_a_grounded_citation(address):
    deps, _ = make_deps(address, questions={"q": at(10)})
    model = scripted(
        {"tool": "search_docs", "args": {"query": "q"}},
        {"output": {"result": "It is the first.", "sources": ["alpha"]}},
    )
    with rag_agent.override(model=model):
        result = await run_rag("which?", deps)
    assert result.output.sources == ["alpha"]
    assert [step.agent for step in result.steps] == ["rag"]


async def test_an_ungrounded_citation_is_sent_back_and_the_model_corrects_it(address):
    deps, _ = make_deps(address, questions={"q": at(10)})
    model = scripted(
        {"tool": "search_docs", "args": {"query": "q"}},
        {"output": {"result": "It is the first.", "sources": ["made-up"]}},
        {"output": {"result": "It is the first.", "sources": ["alpha"]}},
    )
    with rag_agent.override(model=model):
        result = await run_rag("which?", deps)

    retries = [p for m in result.all_messages() for p in m.parts if isinstance(p, RetryPromptPart)]
    assert len(retries) == 1 and "made-up" in str(retries[0].content)
    assert result.output.sources == ["alpha"]


async def test_each_run_starts_with_nothing_retrieved(address):
    """Sources from one question must not make a citation valid in the next."""
    first_deps, model = make_deps(address, questions={"q": at(10)})
    first = scripted(
        {"tool": "search_docs", "args": {"query": "q"}},
        {"output": {"result": "a", "sources": ["alpha"]}},
    )
    with rag_agent.override(model=first):
        await run_rag("one", first_deps)

    second_deps = RagDeps(knowledge_base=SMALL_KB, chroma_url=address, embedder=Embedder(model))
    cites_without_searching = scripted(
        {"output": {"result": "b", "sources": ["alpha"]}},
        {"output": {"result": "b", "sources": []}},
    )
    with rag_agent.override(model=cites_without_searching):
        result = await run_rag("two", second_deps)
    assert result.output.sources == []  # the first attempt was rejected: nothing was retrieved yet


def test_the_real_knowledge_base_has_unique_ids_and_the_two_notices_differ_only_in_their_codes():
    ids = [p.id for p in KNOWLEDGE_BASE]
    assert len(ids) == len(set(ids))
    by_id = {p.id: p for p in KNOWLEDGE_BASE}
    assert "RN-2291" in by_id["recall-rn-2291"].text and "RN-2290" in by_id["recall-rn-2290"].text
"""Live check: real embeddings in a real Chroma service, and a real model answering from what it finds.

`-m eval`. The release check (scripts/release_check.py) builds and starts the Chroma service, then runs
this with its address in CHROMA_URL. To run it by hand: start the service, then set the variable:

    docker compose -f examples/rag/service/docker-compose.yml up -d --wait
    export CHROMA_URL=http://$(docker compose -f examples/rag/service/docker-compose.yml port chroma 8000)

The first half measures *retrieval* alone against a keyword baseline: that is why this example uses
vectors. The second half checks the whole agent.
"""

import os
import re

import pytest

pytest.importorskip("chromadb", reason="needs chromadb-client")

from evals.trace import traced_run  # noqa: E402
from examples.live_support import assert_every_agent_ran, run_as_script  # noqa: E402
from examples.rag import agent as module  # noqa: E402
from examples.rag.agent import KNOWLEDGE_BASE, MAX_DISTANCE, RagDeps, vector_search  # noqa: E402

pytestmark = pytest.mark.eval

if not os.environ.get("CHROMA_URL"):
    pytest.skip("needs the Chroma service running (CHROMA_URL)", allow_module_level=True)


# --- The baseline: what this example used before it had a vector database ---

STOPWORDS = frozenset(
    [
        "a",
        "an",
        "and",
        "are",
        "as",
        "at",
        "be",
        "by",
        "can",
        "do",
        "does",
        "for",
        "from",
        "how",
        "i",
        "if",
        "in",
        "is",
        "it",
        "me",
        "my",
        "of",
        "on",
        "or",
        "our",
        "the",
        "to",
        "we",
        "what",
        "when",
        "where",
        "which",
        "who",
        "will",
        "with",
        "you",
        "your",
    ]
)


def tokens(text: str) -> set[str]:
    words = re.findall(r"[a-z0-9]+", text.lower())
    return {w[:-1] if len(w) > 3 and w.endswith("s") else w for w in words if w not in STOPWORDS}


def keyword_ids(query: str) -> list[str]:
    """Passage ids by keyword overlap (a title word counts double), best first; none if no word is shared."""
    wanted = tokens(query)
    scored = []
    for passage in KNOWLEDGE_BASE:
        in_title = tokens(passage.title)
        in_text = in_title | tokens(passage.text)
        score = sum((word in in_text) + (word in in_title) for word in wanted)
        if score:
            scored.append((-score, passage.id))
    return [pid for _, pid in sorted(scored)]


# Questions in a customer's words, each with the passage that answers it. A mix: paraphrases that share
# no word with the answer, ones that do, and ones built on codes.
CASES = [
    ("How long can I send my hiking footwear back for my money?", "returns-policy"),
    ("Is there a charge to get a package delivered quickly?", "express-shipping"),
    ("Do you fix broken zippers?", "repairs"),
    ("Can I pay with a gift card instead of cash?", "gift-cards"),
    ("My boots fell apart after two years, am I covered?", "warranty-boots"),
    ("When can I visit the Portland shop on weekends?", "store-hours"),
    ("Will you match a lower price from another shop?", "price-match"),
    ("What perks do rewards members get?", "loyalty"),
    ("Does my backpack zipper have a guarantee?", "warranty-packs"),
    ("Can I send back something I bought on clearance?", "returns-policy"),
    ("What does recall notice RN-2291 cover?", "recall-rn-2291"),
    ("Is lot 4C-087 affected?", "recall-rn-2290"),
    ("RN-2290", "recall-rn-2290"),
    ("Which gloves were recalled?", "recall-rn-2290"),
    ("How long do I have to return boots?", "returns-policy"),
]
UNRELATED = [
    "What is the capital of France?",
    "How do I bake sourdough bread?",
    "Who won the world cup in 2018?",
]


@pytest.fixture(scope="module")
async def retrieval():
    """For every question, what the vector search returned (in order) and what the baseline did."""
    deps = RagDeps()
    return {
        question: ([h.id for h in await vector_search(deps, question)], keyword_ids(question))
        for question, _ in CASES
    }


async def test_the_right_passage_is_nearly_always_first(retrieval):
    first = sum(1 for question, want in CASES if retrieval[question][0][:1] == [want])
    assert first >= len(CASES) - 1  # measured: all 15 with Gemini embeddings; one miss of slack


async def test_the_right_passage_is_never_lost_to_the_distance_cutoff(retrieval):
    """MAX_DISTANCE must keep what is relevant: the right passage is among the results for every question."""
    missing = [q for q, want in CASES if want not in retrieval[q][0]]
    assert missing == []


async def test_vector_search_finds_what_the_keyword_baseline_misses(retrieval):
    vector_first = sum(1 for q, want in CASES if retrieval[q][0][:1] == [want])
    keyword_first = sum(1 for q, want in CASES if retrieval[q][1][:1] == [want])
    assert vector_first >= keyword_first + 3  # measured: 15 against 12

    paraphrase = "How long can I send my hiking footwear back for my money?"
    assert (
        retrieval[paraphrase][1] == []
    )  # no word in common with the policy: the baseline finds nothing
    assert retrieval[paraphrase][0][0] == "returns-policy"  # the vector search finds it first


async def test_a_question_that_is_one_passages_word_for_word_still_works(retrieval):
    """Meaning is not at the cost of an exact code: the two near-identical notices are told apart."""
    assert retrieval["What does recall notice RN-2291 cover?"][0][0] == "recall-rn-2291"
    assert retrieval["Is lot 4C-087 affected?"][0][0] == "recall-rn-2290"


async def distances(question: str) -> dict[str, float]:
    """The distance from the question to every passage in the index."""
    deps = RagDeps()
    collection = await module.ensure_index(deps)
    embedded = await deps.embedder.embed_query(question)
    found = collection.query(
        query_embeddings=[list(embedded.embeddings[0])], n_results=len(KNOWLEDGE_BASE)
    )
    return dict(zip(found["ids"][0], found["distances"][0], strict=True))


async def test_questions_with_no_connection_to_the_documents_return_nothing():
    deps = RagDeps()
    for question in UNRELATED:
        assert await vector_search(deps, question) == [], question


async def test_the_cutoff_sits_clear_of_both_the_right_passages_and_the_unrelated_questions():
    """If this fails after a change of embedding model, MAX_DISTANCE needs re-measuring."""
    right = max([(await distances(q))[want] for q, want in CASES])
    unrelated = min([min((await distances(q)).values()) for q in UNRELATED])
    margin = 0.03
    assert right + margin < MAX_DISTANCE < unrelated - margin, (right, MAX_DISTANCE, unrelated)


# --- The agent, end to end ---


async def ask(question: str):
    async def helper(text: str):
        return await module.run_rag(text)

    return await traced_run(helper, question)


@pytest.fixture(scope="module")
async def returns():
    return await ask("How long can I send my hiking footwear back for my money?")


@pytest.fixture(scope="module")
async def two_topics():
    return await ask("Is shipping free on a $60 order, and how long is the boot warranty?")


@pytest.fixture(scope="module")
async def recall():
    return await ask("What does recall notice RN-2291 cover, and what should I do?")


@pytest.fixture(scope="module")
async def not_covered():
    return await ask("Do you sell tents?")


async def test_a_question_in_the_customers_own_words_is_answered_with_its_source(returns):
    # "45 days" exists only in the invented policy, so it can't come from the model's memory, and the
    # question shares no word with the policy, so a keyword search could not have found it.
    assert "search_docs" in returns.tools_called
    output = returns.result.output
    assert "45" in output.result
    assert "returns-policy" in output.sources


async def test_a_two_part_question_draws_on_both_passages(two_topics):
    output = two_topics.result.output
    assert {"shipping-rates", "warranty-boots"} <= set(output.sources)
    answer = output.result.lower()
    # $60 is under the $75 free-shipping threshold, so it pays the $6.95 standard rate: a right answer
    # gives the rate or the threshold, whichever way it is worded.
    assert "6.95" in answer or "75" in answer
    assert "3" in answer and "year" in answer  # the 3-year boot warranty


async def test_a_question_about_one_code_is_not_answered_with_the_other_notice(recall):
    output = recall.result.output
    assert "recall-rn-2291" in output.sources
    assert "7B-114" in output.result
    assert "4C-087" not in output.result  # that is the other notice's lot


async def test_every_citation_is_a_passage_the_model_really_retrieved(returns, two_topics, recall):
    """The output validator enforces this; the message history shows the same thing independently."""
    for traced in (returns, two_topics, recall):
        seen = " ".join(
            str(part.content)
            for message in traced.result.all_messages()
            for part in message.parts
            if type(part).__name__ == "ToolReturnPart" and part.tool_name == "search_docs"
        )
        for source in traced.result.output.sources:
            assert f"[{source}]" in seen


async def test_a_question_the_documents_do_not_cover_cites_nothing(not_covered):
    """A vector search always returns the nearest passages, so the model must judge that none answers."""
    assert "search_docs" in not_covered.tools_called  # it looked
    assert not_covered.result.output.sources == []
    assert not_covered.result.output.result.strip()  # and it said something rather than nothing


async def test_the_agent_talks_to_the_service_docker_published():
    assert RagDeps().chroma_url == os.environ["CHROMA_URL"]
    assert RagDeps().chroma_url.startswith("http://127.0.0.1:")


async def test_the_agent_ran(returns, two_topics, recall, not_covered):
    ran = returns.agents_ran | two_topics.agents_ran | recall.agents_ran | not_covered.agents_ran
    assert_every_agent_ran(module, ran)


async def test_the_demo_script_runs():
    assert "returns-policy" in await run_as_script("examples.rag.agent")
# The Chroma vector database, as a service. Start it by hand with:
#     docker compose up -d --wait
# scripts/release_check.py starts and stops it for you when it checks the example. It keeps its data
# in the container only, so `docker compose down` forgets it and the next run re-indexes.
services:
  chroma:
    # Pinned to the version of the chromadb-client package the agent uses: a client and a server a
    # minor version apart can disagree about the wire format.
    image: chromadb/chroma:1.5.9
    ports:
      # Publish on 127.0.0.1 only, on a free host port (the first field is left empty so Docker
      # picks one); `docker compose port chroma 8000` says which.
      - "127.0.0.1::8000"
    healthcheck:
      # Healthy once the heartbeat endpoint answers, so `up --wait` returns when it is ready.
      test:
        - CMD
        - bash
        - -c
        - exec 3<>/dev/tcp/127.0.0.1/8000 && printf 'GET /api/v2/heartbeat HTTP/1.0\r\n\r\n' >&3 && grep -q nanosecond <&3
      interval: 1s
      timeout: 3s
      retries: 40

Recorded run · gemini-3.1-flash-lite · 1 step · $0.0004

Recorded 2026-10-07 with google:gemini-3.1-flash-lite · 1 step · 1,027 tokens · $0.0004 · 2.0 s. Model output varies between runs. Regenerate with uv run python scripts/record_example.py rag.

Input

How long can I send my hiking footwear back for my money?

Steps

1. rag 1,027 tokens · $0.0004

Prompt

How long can I send my hiking footwear back for my money?

What happened - called search_docs({"query": "return policy for hiking footwear refund time limit"}) - search_docs returned: [returns-policy] Returns: Unworn items can be returned within 45 days of delivery for a full refund. Boots must be returned in their original box. Sale items are final sale and cannot be returned. [re …

Output

{
  "result": "You can return unworn hiking footwear within 45 days of delivery for a full refund. Please note that the boots must be returned in their original box and that sale items are final sale.",
  "sources": [
    "returns-policy"
  ]
}

Result

run_rag(...).output

{
  "result": "You can return unworn hiking footwear within 45 days of delivery for a full refund. Please note that the boots must be returned in their original box and that sale items are final sale.",
  "sources": [
    "returns-policy"
  ]
}