AI & ML interests

Open RL Environments at Scale

Recent Activity

AdithyaSKย  updated a dataset about 2 hours ago
FineEnvs/openai-math
AdithyaSKย  published a dataset about 3 hours ago
FineEnvs/openai-math
sergiopaniegoย  updated a collection about 6 hours ago
FineEnvs Academy
View all activity

FineEnvs 's collections 12

Simulation RL Envs
Real-world work rebuilt as deterministic RL simulations from seed data. Part 1: PortSimEnv v1, berth planning at the Port of Barcelona.
SmolDataEnvs Multi-harness RL
Try the three environments, follow the tutorial, and explore the datasets, models and historical RL/SFT results.
MiMo-V2.6-RL in Harbor
All 7,780 of Xiaomi's MiMo-V2.6 RL environments as Harbor tasks, set up and graded like Xiaomi's harness.
Repo2RLEnv โ€” Verifiable RL Environments
Repo2RLEnv coding and terminal RL environments in Harbor format. Datasets include per-task quality labels, provenance and generation economics.
Data Agent
Deterministic data-analysis agent tasks from the jupyter-agent dataset โ€” verified answers, no LLM judge. Harbor env suites, plain dataset & SFT.
Multilingual Multimodal Envs
OpenEnv environments for multilingual OCR (22 languages) and speech recognition (102 languages), with GRPO recipes and Kannada runs.
SmolDataEnvs
5.5K+ RL tasks for hill-climbing small models in code and data science. Deterministic grading, no LLM judge.
LaTeX OCR
LaTeX OCR environment, model, dataset, and five-run training comparison across four models, including unstable and stabilized Gemma.
Paint with Code
An RL environment where the agent paints by writing p5.brush sketches, rewarded by an aesthetic preference model looking at the render.
FineEnvs Academy
A curated collection of articles, guides, tutorials, slides, and resources for learning how to build, train, and evaluate RL environments for Agents
Simulation RL Envs
Real-world work rebuilt as deterministic RL simulations from seed data. Part 1: PortSimEnv v1, berth planning at the Port of Barcelona.
Multilingual Multimodal Envs
OpenEnv environments for multilingual OCR (22 languages) and speech recognition (102 languages), with GRPO recipes and Kannada runs.
SmolDataEnvs Multi-harness RL
Try the three environments, follow the tutorial, and explore the datasets, models and historical RL/SFT results.
MiMo-V2.6-RL in Harbor
All 7,780 of Xiaomi's MiMo-V2.6 RL environments as Harbor tasks, set up and graded like Xiaomi's harness.
SmolDataEnvs
5.5K+ RL tasks for hill-climbing small models in code and data science. Deterministic grading, no LLM judge.
Repo2RLEnv โ€” Verifiable RL Environments
Repo2RLEnv coding and terminal RL environments in Harbor format. Datasets include per-task quality labels, provenance and generation economics.
LaTeX OCR
LaTeX OCR environment, model, dataset, and five-run training comparison across four models, including unstable and stabilized Gemma.
Paint with Code
An RL environment where the agent paints by writing p5.brush sketches, rewarded by an aesthetic preference model looking at the render.
Data Agent
Deterministic data-analysis agent tasks from the jupyter-agent dataset โ€” verified answers, no LLM judge. Harbor env suites, plain dataset & SFT.
FineEnvs Academy
A curated collection of articles, guides, tutorials, slides, and resources for learning how to build, train, and evaluate RL environments for Agents