What Ashr does
Ashr provides automated, multi-modal testing and evaluation infrastructure for AI agents. It lets teams run comprehensive tests against agents in real environments before deploying to production, catching failures and regressions early.
Key capabilities
Ashr runs evaluations against agents with actual integrations, tracks performance across runs with datasets and metrics, and helps debug failures by surfacing full conversation traces, tool calls, and step-by-step behavior. It supports prompt version control with inline diffs showing how changes affect pass rates, and integrates via Python or TypeScript SDKs so tests can run from code. Example agent scenarios include claims handling, flight rebooking, refund disputes, scheduling, and underwriting.
Who it's for
Ashr is built for AI development teams shipping production agents that interact with real systems, including users at universities and startups, who need reliable testing before agents reach end users.