compressor-reflex

Author: Eric Maddox

Model Description

compressor-reflex is a small (151M-parameter) extractive line-level compressor for tool output in agentic coding loops (IDEs such as Cursor and Antigravity). It scores every line of a tool result for keep/drop and runs post-tool, pre-context — cutting the tokens that reach the next model call while guaranteeing safety-critical lines are never dropped.

  • Architecture: ModernBERT encoder (heman10x/rlcd-modernbert-151m) with a per-line keep/drop head.
  • Format: INT8 ONNX (model_int8.onnx, 150MB). CPU-only via ONNX Runtime. Zero API tokens to run.
  • Decision rule: keep line iff P(keep) ≥ τ*, with τ*=0.50 calibrated on a held-out must-keep anchor set.
  • Fail-open policy: outputs of ≤5 physical lines or ≤64 tokens bypass the compressor and pass through verbatim; if scoring ever yields zero kept lines, the first and last lines are kept. The model never silently drops everything.

MCP server

The companion MCP server packages this model for IDE deployment (Cursor, Antigravity, Claude Desktop): install with pip install compressor-reflex-mcp (PyPI). It downloads these weights from this repo on first run and exposes compress_tool_output / compress_file tools plus a transparent proxy mode.

Intended Uses & Limitations

Intended use: compressing tool outputs (build logs, test results, grep/file reads, command output) before they enter an LLM's context window, in agentic coding workflows. Designed as a local MCP-side layer.

Not intended for: anything where a dropped line is catastrophic and unverifiable — the fail-open policy mitigates this, but the 100% retention figure is measured on the eval distribution, not a formal guarantee. Not a general text summarizer: it is extractive (keep/drop), never generative.

Training and Evaluation Data

  • Train: 2,200 items (400 must-keep anchors at 3× emphasis) — tool outputs across 6 projects (build logs, pytest output, grep results, file reads, error traces).
  • Validation: 250 items.
  • Eval (held-out): 340 items, 611 must-keep anchor lines across 244 items, including 172 deep anchor lines at line index ≥50. 0% verbatim item, intent, and anchor overlap with train/val (independently re-verified). Eval and training sets are never mixed; the eval set is not distributed with this release.

Training Procedure

Single-epoch fine-tune of the 151M base (per delivered training_history.json: loss 0.5681). Training logs are thin (one epoch of metrics; the log covers data loading only) — disclosed here because the weights are judged on the held-out eval, not on training claims. No training data, eval data, or weights were modified after the independent validation.

Evaluation

Held-out eval results below were independently reproduced: the shipped replay_harness.py was re-run on Linux against the shipped INT8 weights and eval set, reproducing eval_report.json bit-exactly.

Metric Result
Must-keep retention @ τ*=0.50 611/611 = 100%
Downstream next-action fidelity 100% (50/50 replayed steps)
Tool-output compression ratio 86.73% (163,716 → 21,726 tokens)
CPU latency, INT8, per chunk p50 ~0.3–0.8s, p99 ~0.9–5.7s (hardware- and length-dependent)

Real coding sessions (builder-measured, n=12 across 4 repositories — methodology reviewed, not independently re-run):

Metric Result
Tool-output compression ~90%
Total session token savings ~48% median (range 28–61%; 5/12 sessions ≥50%)

The 50% session-savings bar was narrowly missed on the median (48.30%) due to invariant per-turn system-prefix dilution (~3.5K tokens/turn the compressor cannot touch). Reported as measured, not rounded up.

Limitations

  • Closest call: the lowest-scoring must-keep anchor in eval scores 0.5449 — retained at τ*=0.50 but dropped at τ≥0.55. The calibrated threshold carries a thin margin on this anchor; disclosed.
  • Latency: 150ms/chunk-class latency is not achievable with this model size on CPU (p99 ~0.9–5.7s). Very long outputs should be chunk-capped in deployment. Latency scales with output length — which is also where the token savings are largest.
  • Distribution: retention figures hold for the tool-output distribution it was trained and evaluated on. Unusual output formats may behave differently; the fail-open policy bounds the downside.

How to Get Started

import numpy as np
import onnxruntime as ort
from transformers import AutoTokenizer
from serializer import CompressorSerializer  # shipped in this repo

tok = AutoTokenizer.from_pretrained("aialchemist-dev/compressor-reflex")
ser = CompressorSerializer(tok)
sess = ort.InferenceSession("model_int8.onnx", providers=["CPUExecutionProvider"])

TAU = 0.50  # calibrated keep threshold — do not raise without re-validating retention

def compress(intent: str, raw_text: str) -> list[str]:
    lines = raw_text.splitlines()
    if len(lines) <= 5 or len(tok.encode(raw_text, add_special_tokens=False)) <= 64:
        return lines  # fail-open bypass
    scores: list[float] = []
    for ch in ser.chunk_tool_output(intent, raw_text):
        ids = np.asarray(ch["input_ids"])
        mask = np.asarray(ch["attention_mask"])
        seq, nl = ids.shape[0], ch["num_lines"]
        lm = np.zeros((nl, seq), dtype=np.float32)
        for i, (s, e) in enumerate(ch["line_spans"]):
            s, e = min(s, seq - 1), min(max(e, s + 1), seq)
            if e > s:
                lm[i, s:e] = 1.0 / (e - s)
        probs = sess.run(None, {"input_ids": ids[None],
                                "attention_mask": mask[None],
                                "line_mask": lm})[0]
        scores.extend(np.asarray(probs).ravel().tolist())
    kept = [l for l, p in zip(lines, scores) if p >= TAU]
    return kept if kept else [lines[0], lines[-1]]  # fail-open: never return empty

See INSERTION.md for the full deployment spec and replay_harness.py for the reference evaluation.

Citation

@misc{maddox2026compressorreflex,
  author = {Eric Maddox},
  title = {compressor-reflex: a small local line-level compressor for tool output},
  year = {2026},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/aialchemist-dev/compressor-reflex}}
}

License

Apache-2.0

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Evaluation results

  • Must-keep retention @ τ*=0.50 on compressor-reflex held-out eval
    self-reported
    100.000
  • Downstream next-action fidelity on compressor-reflex held-out eval
    self-reported
    100.000
  • Tool-output compression ratio on compressor-reflex held-out eval
    self-reported
    86.730