Conversation¶
An agent that remembers what was said, keeps its context bounded, and streams its replies.
A model has no memory of its own: it sees only what you send with each request. A conversation is the agent sending its earlier messages along with the new one, so it can refer back to them, and taking the new messages from each result to send next time. That history grows with every turn, and so do the context and the bill, so a history processor keeps only the most recent turns, cutting at turn boundaries so a tool call is never separated from its result. Replies can also be streamed, appearing as they are written. This is memory within a conversation, not memory across sessions.
Use it when
- Users talk to the agent over several turns and expect it to remember.
- A long conversation shouldn't grow the context, and the bill, without limit.
- Replies should appear as they are written.
Look elsewhere when
- Each request stands alone:
single. - The agent must remember across sessions: store the facts yourself and put them in the instructions.
turn 1 ─┐
turn 2 ─┼→ history ─→ (keep the last N turns) ─→ model ─→ reply ─→ history for the next turn
turn 3 ─┘
What it shows
- Memory is message history: each turn passes the earlier messages in, and the result's messages
become the next turn's history.
Conversationholds that for you - A bounded window: a history processor (
ProcessHistory) keeps the lastmax_turnsuser turns before every model request, reading the size from deps so each conversation can choose its own. It cuts only at turn boundaries, so a tool call is never separated from its result - Streaming:
Conversation.stream()yields the reply as text deltas and updates the history when it finishes. The agent returns plain text (nooutput_type), which is what streams; a structured output has no text to stream until it is complete - What it doesn't do: once a turn leaves the window the model no longer sees it. That is context management, not long-term memory; to remember across sessions, store facts yourself and put them in the instructions
See it run: sample_run.md is a recorded run against a real model: what each agent was asked, which tools it called, and what it returned.
run_chat(text, deps, history=...) is one turn and returns a RunResult like every other example;
Conversation is the convenient way to hold a dialogue:
conversation = Conversation() # or Conversation(ChatDeps(max_turns=6))
await conversation.say("My name is Priya.")
reply = await conversation.say("What is my name?") # reply.output: "Your name is Priya."
async for delta in conversation.stream("Tell me a joke."):
print(delta, end="") # the reply, as it is generated
Source¶
All of it is in examples/conversation/.
"""A conversation: memory across turns, a bounded context window, and streaming.
Use this pattern when:
- Users talk to the agent over several turns and expect it to remember
- The conversation can grow longer than you want to send to the model every time
- The reply should appear as it is generated, not all at once
How it works:
1. Each turn passes the earlier messages in (`message_history`); the result's messages become the
next turn's history, so the agent remembers what was said
2. A history processor keeps only the most recent `max_turns` turns before each model request, so
cost and context stay bounded. It cuts at turn boundaries, so a tool call is never separated
from its result
3. `Conversation` holds the history for you, with `say()` (a full reply and its RunResult) and
`stream()` (the reply as text deltas)
What this is not: long-term memory. Once a turn falls out of the window the model no longer sees it.
To remember across sessions, store facts yourself and put them in the instructions.
"""
from __future__ import annotations
from collections.abc import AsyncIterator
from dataclasses import dataclass, field
from pydantic_ai import Agent, RunContext
from pydantic_ai.capabilities import ProcessHistory, RaiseContentFilterError
from pydantic_ai.messages import ModelMessage, ModelRequest, UserPromptPart
from pydantic_ai.usage import UsageLimits
from agent.config import settings
from agent.logging import agent_label, configure_logging, get_logger
from agent.prompts.templates import load_prompt
from agent.runs import Flow, RunResult
logger = get_logger(__name__)
LABEL = agent_label(__name__) # names this agent's run spans in Logfire traces
# Per turn, not per conversation: each call to say() or stream() is bounded by this.
USAGE_LIMITS = UsageLimits(
request_limit=10, total_tokens_limit=100_000, cost_limit=settings.cost_limit
)
# --- Dependencies ---
@dataclass
class ChatDeps:
"""Runtime dependencies for the chat agent."""
max_turns: int = 20 # how many of the user's most recent turns the model gets to see
def __post_init__(self) -> None:
if self.max_turns < 1:
raise ValueError("max_turns must be at least 1")
# --- The context window ---
def trim_to_recent_turns(messages: list[ModelMessage], max_turns: int) -> list[ModelMessage]:
"""The last `max_turns` turns of `messages`.
A turn starts at a request carrying the user's prompt and includes everything after it (the
reply, any tool calls and their results). Cutting only at turn starts keeps every tool call
together with its result, which providers require.
"""
starts = [
i
for i, message in enumerate(messages)
if isinstance(message, ModelRequest)
and any(isinstance(part, UserPromptPart) for part in message.parts)
]
if len(starts) <= max_turns:
return messages
return messages[starts[-max_turns] :]
async def keep_recent_turns(
ctx: RunContext[ChatDeps], messages: list[ModelMessage]
) -> list[ModelMessage]:
"""The history processor: runs before every model request, using the window from deps."""
return trim_to_recent_turns(messages, ctx.deps.max_turns)
# --- Agent ---
chat_agent: Agent[ChatDeps, str] = Agent(
settings.model,
name=LABEL,
deps_type=ChatDeps, # no output_type: plain text, which is also what streams
capabilities=[RaiseContentFilterError(), ProcessHistory(keep_recent_turns)],
instructions=load_prompt("conversation"), # prompts/…; copied to agent/prompts/<name>.txt
)
async def run_chat(
user_input: str,
deps: ChatDeps | None = None,
*,
history: list[ModelMessage] | None = None,
) -> RunResult[str]:
"""One turn of conversation: reply to `user_input`, given the earlier `history`.
Returns:
A RunResult: `.output` is the reply. For the next turn, pass
`result.steps[0].result.all_messages()` as `history` (`Conversation` does this for you).
"""
if deps is None:
deps = ChatDeps()
logger.info("Chat turn", extra={"user_input": user_input, "history": len(history or [])})
flow = Flow(USAGE_LIMITS)
result = await flow.run(chat_agent, user_input, deps=deps, message_history=history)
return flow.finish(result.output)
@dataclass
class Conversation:
"""A conversation that remembers: holds the history and feeds it back each turn."""
deps: ChatDeps = field(default_factory=ChatDeps)
history: list[ModelMessage] = field(default_factory=list)
async def say(self, text: str) -> RunResult[str]:
"""Send a message and wait for the whole reply."""
result = await run_chat(text, self.deps, history=self.history)
self.history = result.steps[0].result.all_messages()
return result
async def stream(self, text: str) -> AsyncIterator[str]:
"""Send a message and yield the reply as it is generated, a few words at a time.
Streaming returns text, not a RunResult; the conversation history is updated once the
reply is complete, exactly as after `say()`.
"""
async with chat_agent.run_stream(
text, deps=self.deps, message_history=self.history, usage_limits=USAGE_LIMITS
) as response:
# debounce_by=None delivers every chunk the provider sends, as it arrives. The default
# (0.1 s) merges chunks that arrive close together, which suits a UI that repaints;
# pass a number of seconds here if yours does.
async for delta in response.stream_text(delta=True, debounce_by=None):
yield delta
self.history = response.all_messages()
if __name__ == "__main__":
import asyncio
async def demo() -> None:
conversation = Conversation()
print(
(
await conversation.say("Hi! My name is Priya and I'm planning a trip to Lisbon.")
).output
)
print((await conversation.say("Which city did I say I was visiting?")).output)
print("streaming: ", end="")
async for delta in conversation.stream("And what is my name?"):
print(delta, end="", flush=True)
print()
configure_logging()
asyncio.run(demo())
You are a friendly assistant having an ongoing conversation.
- Keep answers short: one to three sentences unless asked for more.
- Use what the user has told you earlier in the conversation when it is relevant.
- If they ask about something they never told you, say you don't know it rather than guessing.
title = "Conversation"
pattern = "conversation"
summary = "Remember earlier turns with message history, bound the context window by turns, and stream replies as they are generated."
smoke_input = "Hi! My name is Priya and I'm planning a trip to Lisbon."
[entrypoint]
deps = "ChatDeps"
run = "run_chat"
"""The window, the history it keeps, and streaming — offline, with scripted models."""
import pytest
from pydantic_ai import RunContext
from pydantic_ai.messages import (
ModelMessage,
ModelRequest,
ModelResponse,
TextPart,
ToolCallPart,
ToolReturnPart,
UserPromptPart,
)
from pydantic_ai.models.function import AgentInfo, FunctionModel
from pydantic_ai.models.test import TestModel
from pydantic_ai.usage import RunUsage
from examples.conversation.agent import (
ChatDeps,
Conversation,
chat_agent,
keep_recent_turns,
run_chat,
trim_to_recent_turns,
)
# --- The window ---
def turn(n: int) -> list[ModelMessage]:
return [
ModelRequest(parts=[UserPromptPart(f"q{n}")]),
ModelResponse(parts=[TextPart(f"a{n}")]),
]
def conversation_of(count: int) -> list[ModelMessage]:
return [message for n in range(1, count + 1) for message in turn(n)]
def prompts(messages: list[ModelMessage]) -> list[str]:
return [
str(part.content)
for message in messages
if isinstance(message, ModelRequest)
for part in message.parts
if isinstance(part, UserPromptPart)
]
def test_a_conversation_within_the_window_is_untouched():
messages = conversation_of(3)
assert trim_to_recent_turns(messages, 3) is messages
assert trim_to_recent_turns(messages, 10) is messages
def test_only_the_most_recent_turns_are_kept():
assert prompts(trim_to_recent_turns(conversation_of(5), 2)) == ["q4", "q5"]
assert prompts(trim_to_recent_turns(conversation_of(5), 1)) == ["q5"]
def test_the_window_starts_at_a_users_prompt_never_mid_turn():
kept = trim_to_recent_turns(conversation_of(4), 2)
assert isinstance(kept[0], ModelRequest) and prompts(kept)[0] == "q3"
assert len(kept) == 4 # two whole turns: a request and a reply each
def test_a_tool_call_stays_with_its_result_because_the_whole_turn_is_one_unit():
with_tool = [
ModelRequest(parts=[UserPromptPart("look it up")]),
ModelResponse(parts=[ToolCallPart("lookup", {"q": "x"}, tool_call_id="c1")]),
ModelRequest(
parts=[ToolReturnPart("lookup", "found", tool_call_id="c1")]
), # no user prompt
ModelResponse(parts=[TextPart("here it is")]),
]
messages = [*conversation_of(2), *with_tool]
kept = trim_to_recent_turns(messages, 1)
assert kept == with_tool # the tool-return request did not start a new "turn"
def test_an_empty_history_is_fine():
assert trim_to_recent_turns([], 3) == []
@pytest.mark.parametrize("bad", [0, -1])
def test_a_window_smaller_than_one_turn_is_rejected(bad):
with pytest.raises(ValueError, match="at least 1"):
ChatDeps(max_turns=bad)
def test_the_default_window_is_twenty_turns():
assert ChatDeps().max_turns == 20
async def test_the_processor_reads_the_window_from_deps():
ctx = RunContext(deps=ChatDeps(max_turns=1), model=TestModel(), usage=RunUsage())
assert prompts(await keep_recent_turns(ctx, conversation_of(3))) == ["q3"]
# --- Remembering, with a scripted model ---
def echo_model(seen: list[list[ModelMessage]]):
"""Replies `reply N` and records the messages each request carried."""
def model_fn(messages, info: AgentInfo) -> ModelResponse:
seen.append(list(messages))
return ModelResponse(parts=[TextPart(f"reply {len(seen)}")])
return FunctionModel(model_fn)
async def test_a_turn_with_no_history_sees_only_itself():
seen: list = []
with chat_agent.override(model=echo_model(seen)):
result = await run_chat("hello")
assert prompts(seen[0]) == ["hello"] and result.output == "reply 1"
async def test_history_passed_in_is_what_the_model_sees():
seen: list = []
with chat_agent.override(model=echo_model(seen)):
first = await run_chat("my name is Priya")
await run_chat("what is my name?", history=first.steps[0].result.all_messages())
assert prompts(seen[1]) == ["my name is Priya", "what is my name?"]
async def test_a_conversation_remembers_every_turn():
seen: list = []
conversation = Conversation()
with chat_agent.override(model=echo_model(seen)):
for text in ("one", "two", "three"):
await conversation.say(text)
assert prompts(seen[2]) == ["one", "two", "three"]
assert prompts(conversation.history) == ["one", "two", "three"]
async def test_the_window_limits_what_the_model_sees_and_what_is_kept():
seen: list = []
conversation = Conversation(ChatDeps(max_turns=2))
with chat_agent.override(model=echo_model(seen)):
for text in ("one", "two", "three", "four"):
await conversation.say(text)
assert prompts(seen[3]) == ["three", "four"] # "one" and "two" fell out of the window
assert "one" not in prompts(conversation.history)
async def test_say_returns_a_run_result_for_the_turn():
with chat_agent.override(model=echo_model([])):
result = await Conversation().say("hi")
assert [step.agent for step in result.steps] == ["conversation"] and result.output == "reply 1"
# --- Streaming ---
def streaming_model(chunks: list[str], seen: list[list[ModelMessage]]):
async def stream_fn(messages, info: AgentInfo):
seen.append(list(messages))
for chunk in chunks:
yield chunk
def model_fn(messages, info: AgentInfo) -> ModelResponse:
seen.append(list(messages))
return ModelResponse(parts=[TextPart("".join(chunks))])
return FunctionModel(model_fn, stream_function=stream_fn)
async def test_a_reply_streams_as_chunks_that_add_up_to_the_whole_reply():
chunks = ["Hello, ", "Priya", "!"]
conversation = Conversation()
with chat_agent.override(model=streaming_model(chunks, [])):
received = [delta async for delta in conversation.stream("hi")]
assert received == chunks
final = conversation.history[-1]
assert isinstance(final, ModelResponse) and final.parts[0].content == "Hello, Priya!"
async def test_a_streamed_turn_becomes_part_of_the_history_for_the_next_turn():
seen: list = []
conversation = Conversation()
with chat_agent.override(model=streaming_model(["Nice to meet you."], seen)):
_ = [delta async for delta in conversation.stream("I am Priya")]
await conversation.say("who am I?")
assert prompts(seen[-1]) == ["I am Priya", "who am I?"]
"""Live check: the model remembers across turns, forgets what leaves the window, and streams. `-m eval`."""
import pytest
from evals.trace import traced_run
from examples.conversation import agent as module
from examples.live_support import assert_every_agent_ran, run_as_script
pytestmark = pytest.mark.eval
async def turn(conversation: module.Conversation, text: str):
async def helper(message: str):
return await conversation.say(message)
return await traced_run(helper, text)
@pytest.fixture(scope="module")
async def remembering():
conversation = module.Conversation()
first = await turn(
conversation, "Hi! My name is Priya and I'm planning a trip to Lisbon in May."
)
city = await turn(conversation, "Which city did I say I was visiting?")
name = await turn(conversation, "And what is my name?")
return conversation, first, city, name
@pytest.fixture(scope="module")
async def forgetting():
"""A one-turn window: by the third message, the first has left what the model can see."""
conversation = module.Conversation(module.ChatDeps(max_turns=1))
await turn(conversation, "My name is Priya.")
await turn(conversation, "What is the capital of France?")
name = await turn(conversation, "What is my name?")
return conversation, name
@pytest.fixture(scope="module")
async def streamed():
conversation = module.Conversation()
chunks = [
delta async for delta in conversation.stream("Write three short sentences about rivers.")
]
follow_up = await turn(conversation, "What did I just ask you to write about?")
return conversation, chunks, follow_up
async def test_the_model_uses_what_it_was_told_earlier(remembering):
_, _, city, name = remembering
assert "lisbon" in city.result.output.lower()
assert "priya" in name.result.output.lower()
async def test_each_turn_is_one_step_and_the_history_grows(remembering):
conversation, first, city, name = remembering
for traced in (first, city, name):
assert [step.agent for step in traced.result.steps] == ["conversation"]
users = [
p for m in conversation.history for p in m.parts if type(p).__name__ == "UserPromptPart"
]
assert len(users) == 3 # nothing was dropped inside the default 20-turn window
async def test_a_turn_outside_the_window_is_really_forgotten(forgetting):
conversation, name = forgetting
# The model never saw "Priya" in this request, so it cannot know it. (It may say so or ask.)
assert "priya" not in name.result.output.lower()
users = [
p for m in conversation.history for p in m.parts if type(p).__name__ == "UserPromptPart"
]
assert "My name is Priya." not in [
str(p.content) for p in users
] # dropped from the history too
async def test_a_streamed_reply_arrives_in_pieces_that_add_up_to_what_was_said(streamed):
conversation, chunks, _ = streamed
assert chunks and all(isinstance(c, str) for c in chunks)
final = conversation.history[-3] # the streamed reply, before the follow-up turn
assert isinstance(final.parts[0].content, str)
assert "".join(chunks) == final.parts[0].content
async def test_a_streamed_turn_is_remembered_like_any_other(streamed):
_, _, follow_up = streamed
assert "river" in follow_up.result.output.lower()
async def test_the_agent_ran(remembering, forgetting, streamed):
ran = remembering[1].agents_ran | forgetting[1].agents_ran | streamed[2].agents_ran
assert_every_agent_ran(module, ran)
async def test_the_demo_script_runs():
out = await run_as_script("examples.conversation.agent")
assert "lisbon" in out.lower() and "streaming:" in out
Recorded run · gemini-3.1-flash-lite · 1 step · $0.0001
Recorded 2026-10-07 with google:gemini-3.1-flash-lite · 1 step · 115 tokens · $0.0001 · 1.2 s.
Model output varies between runs. Regenerate with uv run python scripts/record_example.py conversation.
Input
Hi! My name is Priya and I'm planning a trip to Lisbon.
Steps
1. conversation
115 tokens · $0.0001
Prompt
Hi! My name is Priya and I'm planning a trip to Lisbon.
Output
Hi Priya! It's lovely to meet you. Lisbon is a fantastic choice for a trip—are you looking for any specific recommendations for your visit?
Result
run_chat(...).output