Submitted by Xiangyi Li 25 ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces BenchFlow 33 2
Submitted by Xiangyi Li 65 SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks BenchFlow 1.67k 4