PreciseCoder-27B-GRPO-U
Qwen3.6-27B post-trained with reinforcement learning for precise debugging: given buggy Python code, produce a minimal, correct fix (PreciseCoder project). Final checkpoint, global step 328 (2 epochs).
Recipe. GRPO from the base model (no SFT warm-start) with a unit-test-only reward (1 if all tests pass, else 0; no precision/recall shaping), on the PDB train split (buggy inputs only). This is the unit-only GRPO baseline of the PreciseCoder project. Trained with verl on 1 node of 8x H200 (n=6 rollouts, temperature 1.0, lr 1e-6, batch 32, max response 6144 tokens, KL loss 0.005, rollout TP=2). Minimal-edit prompt template.
Eval (PDB benchmark, BigCodeBench + LiveCodeBench debugging tasks):
| eval | result |
|---|---|
| pdb_test (buggy inputs, 4k think + 2k answer) | unit 0.763 · precision 0.432 · recall 0.710 · F1 0.504 · edit lines 7.9 |
Load with transformers / vLLM (>=0.17) as a standard Qwen3.6 model (thinking enabled by default).
Source: verl experiment qwen36-27b-GRPO_unit_h200 (Delta H200, Sep 2026).
- Downloads last month
- -
Model tree for PreciseCoder/PreciseCoder-27B-GRPO-U
Base model
Qwen/Qwen3.6-27B