Pydantic AI Testing
4 Practical Examples: Unit Tests, Tool Checks and Evals
Pydantic AI testing should answer two different questions: does the application behave correctly, and does the model perform the task well enough? Ordinary unit tests are strong at the first question. Evaluations using representative examples help with the second.
Confusing these questions produces misleading confidence. A fixed test response can prove that an endpoint handles a typed result. It cannot prove that a real model understands an unfamiliar request. This guide shows how to build a useful test plan without hiding that distinction.
-
Nani
- No Comments
Share
Table of Contents
You should know basic Python functions, assertions and exceptions before starting. If an agent is still unfamiliar, complete the offline beginner tutorial first. Here, you will follow a fictional help desk from a written contract to a controlled agent test, a deliberately chosen tool call and an evaluation report. The executable exercises use synthetic data and do not require provider credentials.
What will you learn, and where should you start?
This tutorial is for Python learners building an agent, backend developers integrating one into an application, and QA professionals who need evidence beyond a convincing demonstration. By the end, you should be able to separate a broken Python handler from a weak model answer, write repeatable offline checks, and explain what a test result does not establish. You will also have a small test suite that you can run without an API key.
Work through the examples in order if you are new to test doubles. A test double is simply a substitute used during a test, such as a fake ticket store or a controlled model response. It lets you arrange a specific situation without waiting for a real service to produce that situation. It does not automatically behave like the service it replaces. You must choose which behaviour matters and check that explicitly.
| Your situation | Start here | Useful result to produce |
|---|---|---|
| First agent project | Offline typed-result exercise | A passing assertion you can explain |
| Existing Python application | Pytest handler suite | Tests that observe stored state as well as output |
| QA or automation experience | Behaviour contract and failure matrix | Cases tied to concrete risks |
| Prompt or model comparison | Evaluation cases and scoring rubric | A versioned comparison with visible failures |
The examples are teaching fixtures, not production support software. They do not connect to real customer records, send messages or assess a commercial model. That narrow scope makes the first results understandable. After the local exercises pass, you can add a separately approved provider test with a small budget and synthetic requests. Keep the evidence from those runs separate from the offline test report.
Use a new project folder and an isolated Python environment. For the examples reviewed here, the package versions are Pydantic AI 2.54.0 and pytest 9.1.1 on Python 3.12. Install the slim Pydantic AI package because these offline exercises do not need a provider integration. The commands below install dependencies; they are not tests themselves. Use your environment’s Python interpreter for both installation and execution.
python -m venv .venv
# Windows PowerShell:
.venv\Scripts\Activate.ps1
# macOS or Linux:
# source .venv/bin/activate
python -m pip install "pydantic-ai-slim==2.54.0" "pytest==9.1.1"If your system has several Python versions, check python --version inside the environment before installing. An import failure from the wrong interpreter is different from a failed assertion. Keep the tested package versions with your project so another learner can reproduce the exercise, then upgrade deliberately and rerun the tests. The pytest getting-started guide explains the runner and basic test discovery.
Start with a behaviour contract
Before choosing test tools, write what the application must do. For a fictional support classifier, the contract might require one of four categories, a bounded summary and a review route for ambiguous cases. For an account assistant, it might prohibit reading another user’s record.
Separate requirements into invariants and quality goals. An invariant is something that must remain true, such as “no write occurs without approval.” A quality goal might be “the summary preserves the main issue.” The first can often be checked deterministically; the second usually needs representative examples and a review method.
How do you turn a vague requirement into an assertion?
Take “the assistant should handle unclear tickets sensibly.” Specify an input with no ticket identifier, the expected review route, and the forbidden effect: no account lookup or update. Now there are three observations to collect. The response should explain what is missing, the result should contain the review value, and the service fake should record zero protected operations.
| Requirement | Arrange the case | Expected evidence |
|---|---|---|
| Accept only supported routes | Supply an unknown route value | Validation rejects the value |
| Ask for missing context | Omit the ticket identifier | Review result and no protected lookup |
| Enforce ownership | User A requests user B’s ticket | No protected data leaves the service boundary |
| Avoid invented resolutions | Ticket status remains open | No claim that the issue was resolved |
| Prevent repeated writes | Repeat one approved operation identifier | Exactly one recorded state change |
Write expectations before generating predictions. Otherwise, it is easy to accept whatever the model happened to produce. When reasonable reviewers disagree, refine the requirement or label the case ambiguous. An unclear specification cannot be repaired by increasing the number of test runs.
Use four testing layers
| Layer | What it checks | Typical evidence |
|---|---|---|
| Business-function tests | Permissions, lookups and state changes | Direct assertions on ordinary Python functions |
| Agent integration tests | Tools, dependencies and output wiring | Controlled test-model runs |
| Model evaluations | Task quality on representative inputs | Reviewed cases and scored outputs |
| Operational checks | Timeouts, cancellation, budgets and recovery | Failure injection and monitoring evidence |
Do not force every problem into a model evaluation. An ownership check is easier to test directly. Equally, do not declare a summariser accurate because a Python object was created successfully.
Start with ordinary functions that enforce access or transform results. These tests should fail for the exact business reason you intended. Then add a controlled agent run to check that the application passes the correct dependencies and handles the returned object. Finally, evaluate language understanding separately. This order makes a failure easier to locate: a broken access check should not require investigating a provider response.
Terminology varies between teams. This article calls an offline run across your agent and application code an integration test, while some projects reserve that term for real network calls. Label reports with the components actually exercised. “Agent plus fake database, no network” tells a reviewer more than the word integration alone.
A small offline integration test
This example targets Python 3.12 and Pydantic AI 2.54.0. It uses a fixed result to verify the output contract. Model requests are disabled to prevent an accidental paid call.
from typing import Literal
from pydantic import BaseModel
from pydantic_ai import Agent, models
from pydantic_ai.models.test import TestModel
models.ALLOW_MODEL_REQUESTS = False
class Decision(BaseModel):
route: Literal["answer", "review"]
reason: str
agent = Agent(output_type=Decision)
fixture = TestModel(custom_output_args={
"route": "review", "reason": "The request lacks a record identifier."
})
with agent.override(model=fixture):
result = agent.run_sync("Please check my request.")
assert isinstance(result.output, Decision)
assert result.output.route == "review"
assert result.output.reason
print("The typed response and review-route integration checks passed.")The response was supplied by the fixture. It is not a measured classification result. The official testing documentation describes test models, function models and overrides for isolating application behaviour.
Run the example in the same isolated environment as the tutorial. The expected output is The typed response and review-route integration checks passed. The assertions check the result class, the route and a nonempty reason. Notice that reason: str alone permits an empty string: the last assertion is a test expectation, not a constraint declared in this small schema.
Change the prompt while keeping the fixture unchanged. The result remains review because you supplied it explicitly. Next, change the fixture route to an unsupported value and observe the controlled failure. Restore the original fixture afterwards. Neither exercise measures reasoning; both teach you which parts of the response are controlled by the test.
The TestModel API reference documents custom_output_args and the call_tools option. Its default tool selection can exercise registered tools. Disabling model requests does not disable HTTP clients, databases or email functions inside those tools. Use fake services or isolated test storage before attaching an agent with effects to a test model.
Choose the right test double for the question
| Testing question | Useful substitute | Important limitation |
|---|---|---|
| Can the endpoint consume a typed result? | TestModel with explicit output arguments | The answer is supplied by the test |
| What happens after a particular tool call? | FunctionModel with a scripted response sequence | The script chooses the tool, not a live model |
| Does a service enforce ownership? | Direct function call with fake records | Does not exercise agent wiring |
| Does the provider accept this schema? | Separately authorised provider integration test | Requires network access and may incur charges |
| Does the model choose sensible actions? | Representative live evaluation cases | Outcomes vary and need a scoring policy |
A fake service should expose enough behaviour to distinguish success from failure. For example, let a fake ticket store record requested identifiers and reject unknown ones. A fake that always returns the same friendly ticket can hide a bug where the application sends the wrong identifier. Give every test fresh state so an earlier test cannot accidentally prepare its expected result.
Script a tool interaction with FunctionModel
Sometimes you need a precise sequence: request ticket T-42, receive its status, then return an answer. The FunctionModel reference describes a Python callback that receives messages and agent information and returns a model response. This standalone example scripts that exchange and records the actual tool argument.
from pydantic_ai import Agent, models
from pydantic_ai.messages import (
ModelResponse, TextPart, ToolCallPart, ToolReturnPart,
)
from pydantic_ai.models.function import FunctionModel
models.ALLOW_MODEL_REQUESTS = False
seen = []
def scripted_response(messages, info):
returned = any(
isinstance(part, ToolReturnPart)
for message in messages
for part in message.parts
)
if not returned:
return ModelResponse(parts=[
ToolCallPart("lookup_ticket", {"ticket_id": "T-42"})
])
return ModelResponse(parts=[TextPart("Ticket T-42 is open.")])
agent = Agent(FunctionModel(scripted_response))
@agent.tool_plain
def lookup_ticket(ticket_id: str) -> str:
seen.append(ticket_id)
assert ticket_id == "T-42"
return "open"
result = agent.run_sync("Check ticket T-42.")
assert seen == ["T-42"]
assert result.output == "Ticket T-42 is open."
print("One lookup for T-42; expected answer returned.")The expected printed line is One lookup for T-42; expected answer returned. The first callback response requests the lookup. After a tool return appears in the conversation, the callback supplies text. The list assertion checks both the identifier and the number of invocations. An unexpected duplicate lookup fails even if the final answer looks right.
The final text is deliberately hardcoded. This test proves that the requested tool executes through the agent and that the controlled answer reaches the caller. It does not prove that a model reads the tool result faithfully. To investigate that separate question, prepare evaluation cases where the same ticket has different returned statuses and score whether the answer follows the evidence.
For a denied lookup, keep the same test structure but make the service reject the operation and assert the application’s chosen denial handling. Inspect captured messages or service calls as well as the displayed response. The testing guide documents capture_run_messages for investigating the exchange. Prefer assertions about required events and arguments over snapshots of every timestamp and generated identifier.
Put this into practice with guided training
Explore the Pydantic AI course for the syllabus, guided projects and training options.
Turn the exercises into maintainable regression tests
Organise each test into three parts. Arrange the user, synthetic records, fake service and controlled response. Act by calling the application function a real caller uses. Assert the returned outcome and relevant side effects. Calling only the agent may miss a defect in the endpoint that discards its review flag or stores the response against the wrong account.
For example, arrange an empty review queue and a response with route review. Call the application’s ticket handler. Then assert that exactly one review item exists, that it belongs to the authenticated user, and that no ticket update occurred. This is stronger than checking only that the route string appeared somewhere in the response. It establishes how the application used the result.
Choose test names that explain the behaviour, such as missing identifier creates review item without lookup. Keep fixture data small enough to understand on sight. When two tests need different service behaviour, make that difference explicit instead of hiding it in a large shared fixture. A colleague should be able to identify what broke from the failing assertion and the arranged data.
Restore temporary configuration after each test and use fresh call-recording lists. Set the model-request guard before code can start a model run. For asynchronous application handlers, exercise the asynchronous path with the project’s test harness; a successful synchronous script does not establish that an endpoint handles cancellation correctly. Keep ordinary unit tests fast enough to run while changing code.
Run a complete pytest suite for the application handler
The earlier examples call an agent directly. A real application also decides what to do with the answer. This exercise checks a small handler that adds review decisions to a review queue and answer decisions to a separate collection. Both collections are ordinary lists created afresh in the test. They stand in for storage without opening a database connection.
Create a file named test_ticket_contract.py and put the following complete example in it. It contains five test cases: two routing cases, two invalid-reason cases and one unsupported-route case. The route tests deliberately supply the decision rather than asking a live model to choose it. This gives the handler the same controlled input every time.
from typing import Literal
import pytest
from pydantic import BaseModel, Field, ValidationError
from pydantic_ai import Agent, models
from pydantic_ai.models.test import TestModel as FixtureModel
class Decision(BaseModel):
route: Literal["answer", "review"]
reason: str = Field(min_length=1)
def handle_ticket(agent, prompt, user_id, review_queue, answered):
decision = agent.run_sync(prompt).output
if decision.route == "review":
review_queue.append({"user_id": user_id, "reason": decision.reason})
else:
answered.append({"user_id": user_id, "reason": decision.reason})
return decision
@pytest.fixture(autouse=True)
def block_provider_requests(monkeypatch):
monkeypatch.setattr(models, "ALLOW_MODEL_REQUESTS", False)
@pytest.mark.parametrize("route", ["answer", "review"])
def test_handler_records_only_the_selected_route(route):
agent = Agent(output_type=Decision)
review_queue, answered = [], []
fixture = FixtureModel(
call_tools=[],
custom_output_args={"route": route, "reason": "Synthetic test result."},
)
with agent.override(model=fixture):
result = handle_ticket(agent, "Example request", "user-a", review_queue, answered)
assert result.route == route
assert len(review_queue) == int(route == "review")
assert len(answered) == int(route == "answer")
assert (review_queue + answered)[0]["user_id"] == "user-a"
@pytest.mark.parametrize("bad_reason", ["", None])
def test_invalid_reason_is_rejected(bad_reason):
with pytest.raises(ValidationError):
Decision(route="review", reason=bad_reason)
def test_unknown_route_is_rejected():
with pytest.raises(ValidationError):
Decision(route="delete", reason="Not an allowed outcome.")
Run python -m pytest -q test_ticket_contract.py from the directory containing that file. A successful run reports five passed tests. The elapsed time depends on the machine. A discovery error or an import error is not a passing test, even if no assertion ran. Keep the complete failure message when asking a colleague for help, but remove any sensitive paths or values first.
Read the routing assertions as a small contract. Exactly one of the two collections should receive a record, and that record should keep the supplied user identifier. If the handler accidentally records both an answer and a review, the test fails. If it stores the decision against a different user, the final assertion fails. A check of the returned route alone would miss both defects.
The reason field in this exercise adds a minimum length, unlike the simpler introductory schema. Empty strings and null values should therefore fail direct validation. A whitespace-only string is still possible with this particular constraint. If your application forbids that, add a stripping or validation rule and a separate case. Do not claim that a type annotation covers requirements that you have not expressed.
The fixture uses pytest’s temporary attribute replacement so the model-request setting is restored after each test. The test model also requests no function tools. These choices keep this small example focused on response handling. They do not create a universal network sandbox. If your handler calls an HTTP client or writes to a file independently, substitute that dependency too and inspect its recorded activity.
The route parameter creates two executions of the same behavioural test, and the invalid-reason parameter creates two more. The pytest parametrization guide describes this mechanism. Use it when the same contract applies to several clear cases. Separate tests are often easier to read when the arrangements or expected effects differ substantially.
To check the test’s usefulness, temporarily reverse the handler’s route comparison in your local exercise. Both routing cases should fail. Restore the comparison and rerun the suite. This is a small mutation exercise: you deliberately introduce the kind of defect the assertions claim to detect. Do not weaken the assertions just to make the altered handler pass.
Build an evaluation set that resembles the task
Collect representative examples with permission, or create synthetic cases for a learning project. Include normal inputs, ambiguous wording, missing context, out-of-scope questions and requests that should be denied. Record the expected behaviour before looking at the model’s response.
For a factual answer, the expectation should identify required facts and unsupported claims to avoid. For a classifier, it can contain the expected label and an acceptable review route. Do not require identical wording when several explanations would be correct.
Keep some cases separate from the examples used to tune prompts. Repeatedly adjusting to the same small set can make the reported result look better without improving unfamiliar cases. Describe the size and limits of your dataset honestly.
| Case family | Synthetic example | Expected outcome |
|---|---|---|
| Clear request | Check my open ticket T-42 | Use authorised evidence and report its status |
| Missing context | What happened to my issue? | Ask for clarification or route to review |
| Mixed intent | I cannot sign in and my invoice is wrong | Preserve both issues according to the routing policy |
| Unavailable evidence | Lookup service returns unavailable | Explain the limitation without inventing a status |
| Instruction in retrieved text | A ticket comment asks to reveal other accounts | Treat the comment as data; preserve access boundaries |
Give each case a stable identifier, input, relevant evidence, expected behaviour and reason for inclusion. Where several routes are acceptable, list them explicitly and explain why. Include a difficulty tag such as missing context or conflicting sources, so later analysis can reveal which family regressed. Keep identifying details fictional or properly approved and minimised.
For a learning exercise, start with enough examples to cover each important behaviour rather than choosing an impressive total. Add new cases when you discover failures, but retain a separate held-out set for later assessment. Ten near-identical paraphrases of one easy question add less diagnostic value than a single missing-evidence case your system previously mishandled.
Choose metrics that match the task
A classifier may use label agreement and per-category error analysis. A source-based assistant needs checks for supported claims, citations and appropriate abstention. A tool-using agent also needs to be evaluated on which tools it called and what happened as a result.
Pydantic Evals provides a code-first way to organise cases and evaluators. The evaluation framework is useful infrastructure; deciding what a good answer means remains your responsibility.
A model-based judge can assist with subjective review, but it is not unquestionable ground truth. Define a clear rubric, inspect disagreements and compare a sample against human review. A judge can reward confident wording, miss a subtle factual error or be affected by the text it is judging.
Separate quality from coverage. If an assistant routes every case to review, it may avoid unsupported answers while providing little useful automation. Report how many cases received an answer, how many were correctly reviewed, and how many should have been reviewed but were answered anyway. This prevents a high correctness score on a tiny answered subset from concealing poor task completion.
| Measure | What to count | How to interpret it |
|---|---|---|
| Route agreement | Allowed route selected for each case | Inspect mistakes per category |
| Supported answers | Required claims supported by supplied evidence | A valid citation alone is insufficient |
| Appropriate review | Ambiguous cases correctly escalated | Read alongside unnecessary escalations |
| Access violations | Protected information exposed anywhere in the run | Review individually as release blockers |
| Resource use | Attempts, tool calls, latency and billed usage | Include unsuccessful and retried cases |
When using a judge, provide the task, permitted evidence and a narrow rubric. For example: “Pass only if the answer says open, makes no completion promise, and does not add an unsupported deadline.” Review failures and a sample of passes manually. Version the judge instructions too; changing the measuring instrument can change the score even when the application is unchanged.
Score a support answer with an explicit rubric
Consider a fictional ticket whose approved evidence says only that ticket T-42 is open and awaiting review. An answer can be fluent, correctly formatted and still claim that the issue has been fixed. A useful evaluator must check the meaning against the available evidence. Start by separating essential facts from optional wording, then write a short reason for each pass or failure.
| Candidate response | Judgement | Reason |
|---|---|---|
| T-42 is open and awaiting review. | Pass for this narrow case | Preserves the supplied status without adding a promise |
| Your issue is resolved and the refund arrives tomorrow. | Fail | Invents a resolution, a refund and a deadline |
| I cannot access any ticket information. | Fail for this case | Discards evidence that the authorised lookup supplied |
| T-42 remains open. I cannot confirm a resolution time. | Pass for this narrow case | Reports the status and limits the unsupported timing claim |
These judgements are authored examples, not observed model performance. Their purpose is to show why neither exact-text matching nor a positive tone is enough. The first and fourth answers use different wording but preserve the relevant facts. The third sounds cautious but fails to complete a task that the evidence supports. Appropriate caution and unnecessary refusal need different labels.
Add separate fields for task outcome, evidence support and access compliance. A response might answer the user’s question accurately while exposing an unrelated private note. That is not an acceptable overall result. Keep high-impact failures visible instead of averaging them away with language-quality scores. Decide the release rule before looking at which version scores better.
For a small team, two people can independently judge a sample using the same rubric. Record disagreements and the reason each person chose a label. If one reviewer treats an estimated response time as a promise and the other does not, refine the rubric with a concrete example. The aim is a more consistent measurement, not forcing every difficult case into a convenient pass.
When a model judge is used, keep its instructions, model version and allowed evidence with the report. Do not allow a candidate answer to redefine the scoring rule. Manually examine a sample of passes as well as failures: an evaluator that misses a whole class of defects can give an impressive but unhelpful score. A narrow rubric is easier to inspect than a vague request to judge overall quality.
Test access and action boundaries separately
Use two synthetic users and records with different owners. Check that an agent cannot retrieve another user’s record even when the prompt explicitly requests it. Observe tool outputs as well as the final answer: hiding leaked data in the final sentence does not undo disclosure to the model.
For write tools, test no approval, denied approval, changed action details and repeated execution. Verify actual state or a mock outbox instead of relying on the agent’s statement that it did nothing.
The tools and dependencies example provides a starting point for direct permission tests. Keep these tests in your regular suite so a prompt change cannot remove the underlying protection.
Use failure injection deliberately
Make a test service return a timeout, malformed record or unavailable status. Observe whether the application stops, retries within limits, or escalates. A failure should produce a deliberate outcome, not an endless loop or a fabricated answer.
Also test cancellation. If a user closes a request, determine whether background work continues and whether any queued action remains possible. For a portfolio project, a controlled mock is enough to demonstrate the behaviour without disrupting a real service.
A useful retry test makes the first service call fail and the next succeed, then checks the attempt count. A companion test makes every call fail and checks that the budget eventually stops the run. These reveal different defects. Success after one retry does not establish that permanent failure is bounded, and a stopped agent does not establish that a remote operation was cancelled.
Include a delayed write whose completion status is initially unknown. The expected behaviour is to check its operation identifier or send it for reconciliation before repeating it. Record the state before the call, after the timeout and after recovery. The official testing guide explains how these exercises support a release decision.
Put this into practice with guided training
Explore the Pydantic AI course for the syllabus, guided projects and training options.
Test conversation history and streaming as separate behaviours
A one-turn test cannot establish that a multi-turn assistant handles changing context. Create a short conversation in which the user corrects an identifier, withdraws a request or asks a follow-up that needs earlier evidence. Observe which messages and dependencies reach the next run. Avoid assuming that a correct final sentence means stale or unrelated records were never passed to the model.
For the fictional ticket assistant, start with ticket T-42, then change the request to T-43. Check the next lookup argument and the source of the displayed status. In a separate case, reset the conversation and ask an ambiguous question without either identifier. The application should not silently reuse an identifier from another session. Use distinct synthetic users and conversation identifiers to make cross-session mistakes obvious.
Streaming introduces another boundary. The interface may receive partial text before the final typed result exists. Test how it labels partial content, what happens when validation fails at the end, and whether cancellation leaves a loading indicator stuck. An action such as updating a record should not be authorised merely because an early text fragment sounds complete. Test the application’s actual approval and completion signals.
| Situation | What to observe | Useful expected behaviour |
|---|---|---|
| User corrects an identifier | Next tool argument and cited record | Use the corrected authorised record |
| New session starts | History and dependency inputs | No unrelated previous-session content |
| Stream ends early | Visible state and pending work | Mark interruption instead of claiming completion |
| Final validation fails | Persisted result and any action queue | No invalid final record is accepted |
These are suggested application-level cases, not claims that the short code samples implement a chat interface or streaming transport. Build them against the interface your project actually uses. The Pydantic AI with FastAPI article provides a related endpoint exercise, and the MCP integration guide covers checks around separately exposed tools.
Troubleshoot failures without weakening the test
| Symptom | Likely explanation | Next check |
|---|---|---|
| Every prompt returns the same route | A fixed TestModel fixture is working as configured | Inspect fixture values before judging model quality |
| An offline test contacts a service | A real tool dependency is still connected | Replace that client and inspect recorded calls |
| A failure appears only after other tests | Shared mutable state or leaked configuration | Recreate fixtures and restore temporary overrides |
| A harmless wording change fails | An assertion may demand irrelevant exact text | Assert required facts while preserving genuine text contracts |
| A test passes despite a wrong answer | Only output shape or status code is checked | Add the missing semantic or state assertion |
Reproduce a failure with one case and identify the last trustworthy observation. Did input validation succeed? Was the intended tool selected? Did the service return the expected record? Did the application display the result correctly? Changing the prompt before answering these questions can hide an application defect and make the test harder to understand.
What belongs in a test report?
Record the application version, package versions, model identifier, prompt version and source-data snapshot. State which cases used fixed responses and which used a live model. Include sample size, review method and the failures that remain unresolved.
Report the number of attempts and the treatment of retries. Counting only successful final responses can hide repeated failures and extra cost. If a test was skipped, say why. Do not convert a missing result into a pass.
A useful report ends with a decision: suitable for a limited pilot, needs more work, or blocked by a particular failure. This is more valuable than a dashboard with a high score but no explanation of what the score measures.
Keep offline tests in the routine development loop. Put live evaluations behind an explicit opt-in with separate credentials, a case limit and a spending allowance. Include judge calls and failed attempts in the allowance. A test suite that unexpectedly spends money is difficult for another learner or colleague to run confidently.
Repeat selected live cases when variability matters, and report all attempts using a policy chosen beforehand. Do not rerun only the failures until they pass. Compare prompt versions on the same cases and record both improvements and regressions. A provider interruption should remain visible as an operational outcome, even when you also report quality among completed responses.
A small evaluation score you can verify by hand
The following is a scoring exercise with synthetic, prewritten predictions, not a benchmark of any model. It shows why you should keep a row for every case and record failures instead of reporting only successful answers.
cases = [
{"id": "C1", "expected": "access", "observed": "access"},
{"id": "C2", "expected": "billing", "observed": "general"},
{"id": "C3", "expected": "review", "observed": "review"},
]
correct = sum(c["expected"] == c["observed"] for c in cases)
failed = [c["id"] for c in cases if c["expected"] != c["observed"]]
assert correct == 2 and failed == ["C2"]
print(f"Synthetic fixture agreement: {correct}/{len(cases)}")
print("Review failed cases:", failed)When you run a real evaluation, replace the observed values with captured outputs and preserve model, prompt and dataset versions. Do not reuse the two-out-of-three result as an advertised model accuracy score. With a larger dataset, inspect each category and high-impact failure separately; a single average is not enough.
Use a release checklist that separates evidence from assumptions
Before a pilot, list the behaviours that must work and link each one to a test result or reviewed evaluation case. A missing result is an open question, not a quiet pass. For example, a passing offline suite may support the handler’s routing behaviour while the provider connection and real-model answer quality remain untested. State that boundary plainly in the release note.
Keep deterministic checks in a required development job. Keep live evaluations in a separately configured job with approved credentials, a bounded set of cases and a spending limit. That separation lets contributors run ordinary tests without borrowing a production key. It also makes a provider outage distinguishable from a regression in a local permission function.
Choose a small, representative smoke set for a new configuration, then a broader comparison before changing production behaviour. Preserve the previous configuration so a rollback can be performed deliberately. Record why a release is accepted, who owns unresolved problems and which observations would trigger a rollback. No universal accuracy percentage can replace those project-specific decisions.
| Review item | Evidence to attach | Do not substitute |
|---|---|---|
| Output handling | Passing handler tests with arranged cases | A screenshot of one friendly answer |
| Permissions | Denied-access tests and observed service outputs | A prompt asking the model to behave |
| Answer quality | Versioned cases, rubric and all outcomes | The fixture’s supplied response |
| Recovery | Timeout and cancellation observations | An assumption that stopped text means stopped work |
| Operation limits | Attempt counts and configured limits | Only the cost of successful final answers |
For a portfolio, present the smallest reproducible project that demonstrates these decisions. Include a clear run command, fictional fixtures, one documented failure and the test that detects it. Show the difference between the original defect and the corrected behaviour. A reviewer should be able to rerun the exercise without a paid account and understand which future checks still require a live provider.
Do not publish private conversation traces as evidence. Replace personal identifiers, tokens and internal account details with synthetic examples, and check screenshots as well as text files. Good testing documentation makes the behaviour inspectable without exposing the people whose requests the system may eventually handle.
Practice exercise
Create a small classifier with a review route. Write direct tests for its output-handling code, then prepare a separate labelled set for future live-model evaluation. Add at least one ambiguous message and one request containing instructions unrelated to the task.
Change one part of the prompt and compare results on the same cases. Describe which failures improved and which became worse. Do not claim a general accuracy improvement from a tiny exercise; explain the observation and its limits.
Complete a second exercise by intentionally introducing one application bug: ignore the review route and always store the response as answered. Your regression test should fail because the expected review item is absent. Restore the handler and confirm it passes again. This small mutation checks that your assertion can detect the defect it claims to cover.
For the evaluation portion, ask a second reviewer to label a few ambiguous cases without seeing your labels. Discuss disagreements before scoring model responses. You may discover that the contract needs a mixed-topic route or a clearer missing-context rule. Keep those decisions with the dataset so future results are judged against the same expectations, rather than a reviewer’s changing interpretation.
The AI Testing interview questions guide includes ways to explain this testing approach clearly.
Put this into practice with guided training
Explore the Pydantic AI course for the syllabus, guided projects and training options.
Frequently asked questions
Can TestModel replace live evaluations?
No. It helps test application behaviour without a real model. It does not measure how a live model interprets language or chooses actions.
Should I assert the exact answer text?
Only when exact text is part of the contract. For open-ended answers, check required facts, prohibited claims and other meaningful criteria.
Is a high average score enough to release?
No. Averages can hide important failures. Review access violations, unsupported claims and irreversible actions separately from ordinary quality scores.
Does disabling model requests make every test offline?
No. It blocks non-test model requests through the framework, but your tools can still contact external services. Replace service dependencies and keep production credentials out of the test environment.
Should a failed live run be removed from accuracy calculations?
Preserve it in the full report. You may separately calculate quality for completed responses, provided you also report completion rate, failure categories and the denominator used for each measure.
Build a testable learning project
The Pydantic AI course includes testing and evaluation alongside application development. Learners focused on broader QA workflows can also review Brolly Academy’s AI Testing training.
Download the Pydantic AI practice pack
Get nine offline Python examples, a project-review checklist and an evaluation case sheet. The examples use simulated model responses and do not need an API key. Submit this form to open the ZIP download.
Brolly Academy will store your request and use your details to handle it. Course follow-up is optional; the checkbox is not selected automatically.
Read our Privacy Policy for information about handling your details.
Put this into practice with guided training
Explore the Pydantic AI course for the syllabus, guided projects and training options.

Brolly Academy Team
AI, Data Science & Software Training Experts | 20+ Years of Training Experience
Brolly Academy Team is a group of AI, Data Science, Cloud Computing, and Software Development professionals dedicated to helping learners gain practical skills and industry knowledge. Since 2015, Brolly Academy has supported thousands of students and professionals through technology training, certification guidance, and career-focused learning.










