PreciseCoder-27B-GRPO-0.5F1
Qwen3.6-27B post-trained with reinforcement learning for precise debugging: given buggy Python code, produce a minimal, correct fix (PreciseCoder project). Best-validation checkpoint, global step 328 (2 epochs).
Recipe. GRPO from the base model (no SFT warm-start) with reward = (unit + 0.5·F1) / 1.5, where F1 is the edit-block precision/recall F1 against the ground-truth diff (PDB evaluator), on the PDB train split (buggy inputs only). Trained with verl on 2 nodes (n=6 rollouts, temperature 1.0, lr 1e-6, batch 32, max response 6144 tokens, KL loss 0.005, rollout TP=2). Minimal-edit prompt template.
Eval (PDB benchmark, BigCodeBench + LiveCodeBench debugging tasks):
| eval | result |
|---|---|
| pdb_test (buggy inputs, 4k think + 2k answer) | unit 0.777 · precision 0.784 · recall 0.850 · edit lines 2.7 |
| gt_holdout (1512 clean inputs, 16k think + 4k answer) | no-op 47.2% · damage 10.2% · unit pass 89.7% |
| PDB-Wild (228 SWE-smith repo-level bugs) | unit 0.531 · precision 0.544 · recall 0.622 · F1 0.559 · edit lines 9.4 |
no-op = share of clean (bug-free) inputs returned unchanged; damage = share broken by an edit.
Computed over answers with an extractable code block.
Load with transformers / vLLM (>=0.17) as a standard Qwen3.6 model (thinking enabled by default).
Source: verl experiment qwen36-27b-GRPO_0.5F1 (Delta, Sep 2026).
- Downloads last month
- 19