PreciseCoder-27B-GRPO-0.5F1

Qwen3.6-27B post-trained with reinforcement learning for precise debugging: given buggy Python code, produce a minimal, correct fix (PreciseCoder project). Best-validation checkpoint, global step 328 (2 epochs).

Recipe. GRPO from the base model (no SFT warm-start) with reward = (unit + 0.5·F1) / 1.5, where F1 is the edit-block precision/recall F1 against the ground-truth diff (PDB evaluator), on the PDB train split (buggy inputs only). Trained with verl on 2 nodes (n=6 rollouts, temperature 1.0, lr 1e-6, batch 32, max response 6144 tokens, KL loss 0.005, rollout TP=2). Minimal-edit prompt template.

Eval (PDB benchmark, BigCodeBench + LiveCodeBench debugging tasks):

eval result
pdb_test (buggy inputs, 4k think + 2k answer) unit 0.777 · precision 0.784 · recall 0.850 · edit lines 2.7
gt_holdout (1512 clean inputs, 16k think + 4k answer) no-op 47.2% · damage 10.2% · unit pass 89.7%
PDB-Wild (228 SWE-smith repo-level bugs) unit 0.531 · precision 0.544 · recall 0.622 · F1 0.559 · edit lines 9.4

no-op = share of clean (bug-free) inputs returned unchanged; damage = share broken by an edit. Computed over answers with an extractable code block.

Load with transformers / vLLM (>=0.17) as a standard Qwen3.6 model (thinking enabled by default). Source: verl experiment qwen36-27b-GRPO_0.5F1 (Delta, Sep 2026).

Downloads last month
19
Safetensors
Model size
27B params
Tensor type
BF16
·
Video Preview
loading

Model tree for PreciseCoder/PreciseCoder-27B-GRPO-0.5F1

Base model

Qwen/Qwen3.6-27B
Finetuned
(410)
this model
Quantizations
3 models