Suggestion for the Big Benchmarks Collection: animal welfare benchmarks (TAC, ANIMA, MORU)

#1167
by sparrow8i8 - opened

Hi! I maintain the animal welfare and moral reasoning benchmarks built by CaML (Compassion Aligned Machine Learning) and would like to suggest them for the Big Benchmarks Collection, which currently has no coverage of moral reasoning about animals:

Live multi-model leaderboard with uncertainty quantification: https://compassionbench.com (public read API at https://compassionbench.com/api). Happy to provide anything else needed to evaluate the suggestion!

A focusing update from our side: if you consider adding just one of these, make it TAC. It is the agentic one (tool-calling booking agents, implicit welfare, no LLM judge), it is nowhere near saturation (top frontier welfare rate 64.7% vs a 30.5% frontier mean, against a 65% documented chance baseline), and it is contamination-protected (gated dataset, BIG-bench canary GUID, published probe script). Paper: https://arxiv.org/abs/2606.18142

Sign up or log in to comment