Personal Assistant Benchmark
Scores a personal assistant by what it did on the device
None defined yet.
SEABED: SouthEast Asian Benchmark for Evaluating Audio Reasoning
GVD: Governed Versioning and Deduplication for Document Repositories
CentificAIResearch is the official Hugging Face organization for Centific Applied AI Research (CAIR).
Centific works with frontier AI labs and enterprises to build production-ready AI systems. We bring together 1.8 million vetted domain experts, 1K+ PhDs, and platforms for data collection, annotation, model fine-tuning, safety evaluation, and localization across 230 languages and locales.
Centific Applied AI Research (CAIR) is focused on one question: what kind of data and evaluation does it take to make AI work reliably in the real world?
We are a team of researchers and engineers working across healthcare AI, physical AI, vision AI, audio AI, AI safety, agentic systems, and multilingual AI.
š See all research publications
We maintain the PRISM Evaluation Suite, covering 7 domains, 12 benchmarks, 25K+ eval tasks, and 50+ models evaluated.
Scores a personal assistant by what it did on the device
Explore AI benchmark scores with trust verification details
A PDF-grounding benchmark for healthcare document work
Tiered, gated evaluation of finance agent tasks
Evaluate AI models on journal entry audit tasks