-
MMLU-Pro Leaderboard
π₯256More advanced and challenging multi-task evaluation
-
Stick To Your Role! Leaderboard
π62Benchmarking LLMs on the stability of simulated populations
-
ZeroEval Leaderboard
π53Explore ZeroEval embedding benchmark online
-
Open Medical-LLM Leaderboard
π₯437Explore and submit models for benchmarking
Hristo Panev
hppdqdq
AI & ML interests
None yet
Recent Activity
liked a model 3 days ago
outsourc-e/Qwen3.8-27B-Unleashed-GGUF liked a model 20 days ago
MiniMaxAI/MiniMax-Music3 liked a model 28 days ago
ReadyArt/gemma-4-31B-it-scotoma-GGUFOrganizations
None yet