narcolepticchicken/aco-specialists-verifier-gater-v2

Verifier gater for Agent Cost Optimizer — predicts whether an agent's code patch needs verification (testing/review). Goal: avoid running expensive verifiers on every output.

Model

  • Base: ModernBERT-base (answerdotai/ModernBERT-base, 149.6M params)
  • Architecture: ModernBertForSequenceClassification (2-class)
  • Training: Focal loss (γ=2.0, α=[0.949, 1.051]), 10 epochs, early stopping, dropout=0.1, threshold=0.52

Verified Metrics (independent evaluation, 2026-06-03)

Evaluated on 863 held-out R2E-Gym Verifier trajectories (class dist: fail=478, pass=385; 44.6% positive).

Metric v2 (ModernBERT) v1 (DistilBERT) Delta
Accuracy 66.7% 65.5% +1.3pp
F1 (macro) 0.667 0.654 +0.013

Per-class: neg P=0.778 R=0.559 F1=0.650 | pos P=0.594 R=0.803 F1=0.683

Assessment: The model predicts both classes (NOT collapsed). It shows modest bias toward predicting "pass" (high recall 0.803 on positive, lower recall 0.559 on negative). The 1.3pp accuracy gain is real but practically negligible. The feature extraction (600-char github issue + 600-char agent actions + 400-char patch) provides insufficient signal — the model is barely above a majority-class baseline.

NOT COLLAPSED: verified on 863 samples, both classes predicted, F1=0.667 ≠ base rate.

Training Data

R2E-Gym/R2EGym-Verifier-Trajectories (5,750 code-patching agent trajectories with reward scores). Preprocessed split available at narcolepticchicken/aco-traces (data/verifier_gater/).

Usage

from transformers import AutoModelForSequenceClassification, AutoTokenizer

model = AutoModelForSequenceClassification.from_pretrained(
    "narcolepticchicken/aco-specialists-verifier-gater-v2")
tokenizer = AutoTokenizer.from_pretrained(
    "narcolepticchicken/aco-specialists-verifier-gater-v2")

inputs = tokenizer("TASK: ...\nAGENT_ACTIONS: ...\nPATCH: ...",
                    return_tensors="pt", truncation=True, max_length=2048)
outputs = model(**inputs)
# Threshold calibrated at 0.52: prob[1] >= 0.52 → needs verification

Limitations

  • Features are too sparse (1,600 chars) to reliably discriminate patch quality
  • Needs enrichment with test execution results, error traces, and full agent trajectories
  • The R2E-Gym rewards-based labeling (≥1.0 = "verify") is crude; consider finer-grained labels
  • 66.7% accuracy means 1 in 3 decisions is wrong — too high for production
Downloads last month
4
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support