narcolepticchicken/aco-specialists-verifier-gater-v2
Verifier gater for Agent Cost Optimizer — predicts whether an agent's code patch needs verification (testing/review). Goal: avoid running expensive verifiers on every output.
Model
- Base: ModernBERT-base (
answerdotai/ModernBERT-base, 149.6M params) - Architecture: ModernBertForSequenceClassification (2-class)
- Training: Focal loss (γ=2.0, α=[0.949, 1.051]), 10 epochs, early stopping, dropout=0.1, threshold=0.52
Verified Metrics (independent evaluation, 2026-06-03)
Evaluated on 863 held-out R2E-Gym Verifier trajectories (class dist: fail=478, pass=385; 44.6% positive).
| Metric | v2 (ModernBERT) | v1 (DistilBERT) | Delta |
|---|---|---|---|
| Accuracy | 66.7% | 65.5% | +1.3pp |
| F1 (macro) | 0.667 | 0.654 | +0.013 |
Per-class: neg P=0.778 R=0.559 F1=0.650 | pos P=0.594 R=0.803 F1=0.683
Assessment: The model predicts both classes (NOT collapsed). It shows modest bias toward predicting "pass" (high recall 0.803 on positive, lower recall 0.559 on negative). The 1.3pp accuracy gain is real but practically negligible. The feature extraction (600-char github issue + 600-char agent actions + 400-char patch) provides insufficient signal — the model is barely above a majority-class baseline.
NOT COLLAPSED: verified on 863 samples, both classes predicted, F1=0.667 ≠base rate.
Training Data
R2E-Gym/R2EGym-Verifier-Trajectories (5,750 code-patching agent trajectories with reward scores).
Preprocessed split available at narcolepticchicken/aco-traces (data/verifier_gater/).
Usage
from transformers import AutoModelForSequenceClassification, AutoTokenizer
model = AutoModelForSequenceClassification.from_pretrained(
"narcolepticchicken/aco-specialists-verifier-gater-v2")
tokenizer = AutoTokenizer.from_pretrained(
"narcolepticchicken/aco-specialists-verifier-gater-v2")
inputs = tokenizer("TASK: ...\nAGENT_ACTIONS: ...\nPATCH: ...",
return_tensors="pt", truncation=True, max_length=2048)
outputs = model(**inputs)
# Threshold calibrated at 0.52: prob[1] >= 0.52 → needs verification
Limitations
- Features are too sparse (1,600 chars) to reliably discriminate patch quality
- Needs enrichment with test execution results, error traces, and full agent trajectories
- The R2E-Gym rewards-based labeling (≥1.0 = "verify") is crude; consider finer-grained labels
- 66.7% accuracy means 1 in 3 decisions is wrong — too high for production
- Downloads last month
- 4