Spaces:
Running on CPU Upgrade
Suggestion for the Big Benchmarks Collection: animal welfare benchmarks (TAC, ANIMA, MORU)
Hi! I maintain the animal welfare and moral reasoning benchmarks built by CaML (Compassion Aligned Machine Learning) and would like to suggest them for the Big Benchmarks Collection, which currently has no coverage of moral reasoning about animals:
- TAC (Travel Agent Compassion): agentic benchmark measuring implicit animal welfare awareness in tool-calling ticket-booking agents. Paper: https://arxiv.org/abs/2606.18142. Dataset: https://huggingface.co/datasets/sentientfutures/tac (gated to prevent training contamination). Runnable via
inspect_evals/tac. - ANIMA: 115 questions scored across 13 moral reasoning dimensions: https://huggingface.co/datasets/sentientfutures/anima https://arxiv.org/abs/2604.13076 https://inspect.aisi.org.uk/evals/#/eval/anima
Live multi-model leaderboard with uncertainty quantification: https://compassionbench.com (public read API at https://compassionbench.com/api). Happy to provide anything else needed to evaluate the suggestion!
A focusing update from our side: if you consider adding just one of these, make it TAC. It is the agentic one (tool-calling booking agents, implicit welfare, no LLM judge), it is nowhere near saturation (top frontier welfare rate 64.7% vs a 30.5% frontier mean, against a 65% documented chance baseline), and it is contamination-protected (gated dataset, BIG-bench canary GUID, published probe script). Paper: https://arxiv.org/abs/2606.18142