Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
Edit Datasets filters
Main
Tasks
Libraries
Languages
Licenses
Other
1
Reset Other
preprints
Synthetic
art
code
medical
finance
biology
legal
chemistry
agent
climate
music
Apply filters
Datasets
5,270
Full-text search
Edit filters
Sort: Trending
Active filters:
evaluation
Clear all
atroposhealth/precision-evidence-bench
Viewer
•
Updated
4 days ago
•
209
•
110
•
5
Qwen/RecreationBench
Viewer
•
Updated
about 5 hours ago
•
500
•
100
•
5
Alibaba-Aone/aacr-bench
Viewer
•
Updated
Feb 2
•
2.15k
•
561
•
9
Qwen/AgentWorldBench
Viewer
•
Updated
Jul 4
•
2.17k
•
1.18k
•
108
TIGER-Lab/MMLU-Pro
Benchmark
•
Updated
May 2
•
12.1k
•
246k
•
514
nvidia/compute-eval
Viewer
•
Updated
7 days ago
•
3.1k
•
720
•
30
MiniMaxAI/OctoCodingBench
Viewer
•
Updated
Jan 13
•
72
•
429
•
362
treadon/abliteration-eval
Viewer
•
Updated
Apr 14
•
283
•
268
•
3
treadon/disinhibition-eval
Viewer
•
Updated
Apr 30
•
248
•
48
•
2
llamaindex/ExtractBench
Benchmark
•
Updated
Aug 19
•
370
•
20.2k
•
32
harborframework/terminal-bench-2.1
Benchmark
•
Updated
7 days ago
•
158k
•
11
alexshpunt/explicit-edit-benchmark
Viewer
•
Updated
1 day ago
•
86
•
7.81k
•
2
dnhkng/arc-agi-3-teaching-suite
Viewer
•
Updated
about 17 hours ago
•
100
•
75
•
2
LocalLLaMA/terminal-bench-mini
Viewer
•
Updated
1 day ago
•
14
•
305
•
2
openthaigpt/thai-ocr-evaluation
Viewer
•
Updated
Sep 30, 2024
•
104
•
192
•
9
Salesforce/CRMArenaPro
Viewer
•
Updated
Jul 9, 2025
•
8.61k
•
2.04k
•
18
Naholav/claude_4_math_evaluation_500
Preview
•
Updated
Jul 7, 2025
•
79
•
1
DAGroup-PKU/RoVid-X
Preview
•
Updated
May 21
•
17.6k
•
72
GSMA/ot-full
Viewer
•
Updated
Mar 26
•
20.6k
•
869
•
7
harborframework/terminal-bench-2.0
Benchmark
•
Updated
Apr 24
•
89.6k
•
51
KRAFTON/ArtiBench
Viewer
•
Updated
Feb 25
•
1k
•
963
•
6
rl-rag/hle-gpt-oss-120b-no-python-260222
Viewer
•
Updated
Feb 25
•
9.71k
•
4.83k
•
1
PhysionLabs/Physion-Eval
Preview
•
Updated
Jun 14
•
732
•
19
internlm/WildClawBench
Benchmark
•
Updated
Aug 15
•
8.99k
•
71
kensho/WILD-raw
Viewer
•
Updated
May 7
•
7.24M
•
418
•
3
llamaindex/ParseBench
Benchmark
•
Updated
Apr 19
•
169k
•
22.4k
•
129
lthn/MMLU-Pro
Viewer
•
Updated
Apr 10
•
12.1k
•
88
•
1
surgeai/GDP.pdf
Viewer
•
Updated
9 days ago
•
100
•
61k
•
21
besimple-ai/voice-code-bench
Viewer
•
Updated
9 days ago
•
300
•
729
•
13
PaintBenchAnonymousNeurIPS26/PaintBench
Viewer
•
Updated
May 7
•
2.73k
•
50
•
1
Previous
1
2
3
...
100
Next