PreciseCoder-27B-GRPO-U

Qwen3.6-27B post-trained with reinforcement learning for precise debugging: given buggy Python code, produce a minimal, correct fix (PreciseCoder project). Final checkpoint, global step 328 (2 epochs).

Recipe. GRPO from the base model (no SFT warm-start) with a unit-test-only reward (1 if all tests pass, else 0; no precision/recall shaping), on the PDB train split (buggy inputs only). This is the unit-only GRPO baseline of the PreciseCoder project. Trained with verl on 1 node of 8x H200 (n=6 rollouts, temperature 1.0, lr 1e-6, batch 32, max response 6144 tokens, KL loss 0.005, rollout TP=2). Minimal-edit prompt template.

Eval (PDB benchmark, BigCodeBench + LiveCodeBench debugging tasks):

eval result
pdb_test (buggy inputs, 4k think + 2k answer) unit 0.763 · precision 0.432 · recall 0.710 · F1 0.504 · edit lines 7.9

Load with transformers / vLLM (>=0.17) as a standard Qwen3.6 model (thinking enabled by default). Source: verl experiment qwen36-27b-GRPO_unit_h200 (Delta H200, Sep 2026).

Downloads last month
-
Safetensors
Model size
27B params
Tensor type
BF16
·
Video Preview
loading

Model tree for PreciseCoder/PreciseCoder-27B-GRPO-U

Base model

Qwen/Qwen3.6-27B
Finetuned
(409)
this model