Think Before You Score: Thinking Reward Model for Visual Generation Paper • 2609.37372 • Published 4 days ago • 96
VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models Paper • 2609.04355 • Published 15 days ago • 9
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy Paper • 2609.28660 • Published 10 days ago • 14
Towards Full Pipeline FP8 Reinforcement Learning for LLMs Paper • 2609.22870 • Published 14 days ago • 17
One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents Paper • 2609.23377 • Published 13 days ago • 50
BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence Paper • 2609.20886 • Published 17 days ago • 30
From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention Paper • 2609.21788 • Published 15 days ago • 13
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 15 days ago • 138
MintAct: A Unified Visual Agent for Digital Environments Paper • 2609.22083 • Published 15 days ago • 35
EvoOntology: A Self-Evolving Ontology Layer for Data Agents Paper • 2609.15779 • Published 19 days ago • 159
AdityaaXD/Multi-Agent_Reinforcement_Learning_Trading_System_Data Viewer • Updated Feb 1 • 5.28k • 236 • 8
jarguello76/reinforcement_learning_lunar_landing Reinforcement Learning • Updated Aug 17, 2025 • 4 • 8
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning Paper • 2609.20784 • Published 16 days ago • 57