Skip to main content

Accuracy Evaluation

Judge agent responses against expected answers:

Batch Evaluation

Run multiple cases and get aggregated metrics:

Performance Evaluation

Measure runtime and memory usage:

Reliability Evaluation

Verify the agent calls expected tools:

Agent-as-Judge Evaluation

Custom criteria evaluation:

Team Evaluation

All eval types support team evaluation: