Hermes: Learning Contextual Reasoning Unlocks Test-Time Scaling Paper • 2609.38332 • Published 4 days ago • 9
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution Paper • 2609.38349 • Published 4 days ago • 16
AQuA: A Benchmarking Tool for Label Quality Assessment Paper • 2306.09467 • Published Jun 15, 2023 • 1
Investigating Compositional Reasoning in Time Series Foundation Models Paper • 2502.06037 • Published Feb 9, 2025 • 2
ARFBench: Benchmarking Time Series Question Answering Ability for Software Incident Response Paper • 2604.21199 • Published Apr 23 • 1
SpIDER: Spatially Informed Dense Embedding Retrieval for Software Issue Localization Paper • 2512.16956 • Published Feb 5
STAMP: Spatial-Temporal Adapter with Multi-Head Pooling Paper • 2511.10848 • Published Nov 13, 2025 • 2
MOMENT: A Family of Open Time-series Foundation Models Paper • 2402.03885 • Published Feb 6, 2024 • 8
TimeSeriesGym: A Scalable Benchmark for (Time Series) Machine Learning Engineering Agents Paper • 2505.13291 • Published May 19, 2025 • 3
STAMP: Spatial-Temporal Adapter with Multi-Head Pooling Paper • 2511.10848 • Published Nov 13, 2025 • 2
TimeSeriesGym: A Scalable Benchmark for (Time Series) Machine Learning Engineering Agents Paper • 2505.13291 • Published May 19, 2025 • 3
Investigating Compositional Reasoning in Time Series Foundation Models Paper • 2502.06037 • Published Feb 9, 2025 • 2