arxiv:2610.01509
๐ In a Training Loop
Changdae Oh
changdae
AI & ML interests
Generalization; Distribution Shift; Uncertainty Quantification; Reward Modeling; Post-training
Recent Activity
upvoted a paper about 5 hours ago
On-Policy Distillation with Negative-Policy Rollouts upvoted a paper 1 day ago
DiVeR: Decision-Critical Verifier Learning for VLA Test-Time Scaling