Record
Agent Behavior Evals Lab
A public lab for agent behavior evaluation, and a blind audit result that is worse than the one on its own cases.
- 21.8Percent caught on the blind independent-author audit
- 98.0Percent caught on the lab's own adversarial corpus
What it is
A public lab for agent behavior evaluation. It holds 320 real-agent evaluation records across 8 framework and model combinations, a public leaderboard of severity-weighted pass rates with 95 percent confidence intervals for 6 local models, and a GitHub Marketplace Action called Agent Behavior Safety Gate.
What it proves
The pre-registered blind red-team audit is the part worth reading. It ran on an independent-author corpus with the manifest hash committed before any fixes were made, which is what separates an audit from a demonstration. The catch rate was 14.5 percent before those fixes and 21.8 percent after them.
What it does not prove
The blind number is low. It sits next to 98.0 percent on the lab's own adversarial corpus, and the distance between those two numbers is the honest description of the work: cases written by the author are far easier to catch than cases written by someone else. The blind number stays published on purpose.
Where it ships
The lab ships publicly as Senthira, the agent behavior safety gate, at senthira.com. It is an open lab at design-partner stage, with no paid engagements yet. An earlier and unrelated direction called Senthira Flow is paused and is no longer listed as work.
Takeaway
What to take from it.
A blind number you dislike is worth more than a good number from cases you wrote yourself.