Ian Cole
codingiancole
AI & ML interests
Reinforcement learning, reward modeling, RLHF, policy optimization, offline RL
Recent Activity
upvoted a paper 17 minutes ago
Towards Full Pipeline FP8 Reinforcement Learning for LLMs upvoted a paper 17 minutes ago
One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents upvoted a paper 17 minutes ago
VideoGen-Agent: Reinforcing Video Generation AgentsOrganizations
None yet