About Kitaru
Replay-based evals for AI agents and LLM workflows: your production traces, re-run against your next change, whether the agent is a tool-using loop or one LLM call inside a product flow. Built for agents that write into a system of record, where testing in production would create phantom bookings and duplicate claims. Import the runs your agent already made, turn what your team notices into an evaluator, then compare two experiment runs to see what a new model or prompt would have done. Open source, self-hosted, built by the ZenML team.