Evaluate your agents
Astropods can automatically grade an agent’s production traces against a set of evaluators, surfacing issues like exposed personal data, leaked credentials, or negative user sentiment.
What evals are for
Use evals to:
- Catch regressions in an agent’s output, such as exposed personal data, leaked credentials, or responses that reveal system instructions.
- Score behavior that is hard to assert on directly, like whether a tool call was unnecessary or whether the user’s tone turned negative.
Evaluators don’t block a deploy or change agent behavior; they grade traces after the fact. Verified evaluator outputs can also be curated into a dataset for comparing future runs.
How it works
In the dashboard, go to an agent’s Traces & Evals page to run and review evals. You run the agent’s active evaluation set against its recent production traces. Each evaluator grades a trace independently and doesn’t change agent behavior.
The default evaluation set
Every agent evaluates against the Astropods default evaluation set until you activate a custom one. See the preset registry in the Evaluator Spec for the full list of default evaluators and what each one checks.
Set custom evaluators
To run different checks, add an EVALUATION.yaml file beside astropods.yml at your project root, listing the presets to keep and any custom evaluators you define. Validate it locally, then activate it against the server:
ast eval push requires a blueprint that’s already been pushed; it only activates the evaluation set and doesn’t build or push a container image. Activating a custom set replaces the default set entirely, it isn’t merged with it. An agent with no activated set falls back to the Astropods default.
See the Evaluator Spec for the full EVALUATION.yaml format: preset references, custom evaluator fields, additional trace context, and output schemas.
Next steps
- Evaluator Spec — the full
EVALUATION.yamlformat and the default preset registry - Monitor your agents — the trace collection that evals run against