Install the package, point a probe at a target, bind evaluators to a scenario, and print the results. This run uses a dummy target so you can confirm the loop before connecting a real model.
Step 1: Target
Define the system under test.
Step 1: Target
Define the system under test.
DummyTarget ships with TrustTest so you can exercise the loop without a live model.Step 2: Probe
Generate the test cases.
Step 2: Probe
Generate the test cases.
DatasetProbe turns a small Q&A dataset into test cases against the target.test_set has two test cases: the question, the model response, and the context used to score it.Step 3: Scenario
Choose evaluators and failure criteria.
Step 3: Scenario
Choose evaluators and failure criteria.
Bind the test set to the metrics that decide pass or fail.
any_fail fails the scenario if either evaluator fails.Step 4: Run
Evaluate the test set and print results.
Step 4: Run
Evaluate the test set and print results.
Complete example
Full Python script for the quickstart.
Complete example
Full Python script for the quickstart.