Thanks — that's the direction we're building toward.
We agree on the premise: a single aggregate score is the wrong unit for a deployment decision. That's why every diagnosis already produces a per-item result across the 117-item catalog rather than one number — the aggregate is a summary of that, not a substitute for it.
Exposing those per-diagnostic traces publicly is on the roadmap. We'd rather ship it once coverage is uniform across the leaderboard than publish partial traces that invite the wrong comparisons.
On the harness: the probe sets are held out deliberately — publishing them would contaminate the behaviours they measure. What we can open is the layer above them: item schemas, category distributions, and the pass/fail criteria per axis. That should be enough to audit regressions without making the benchmark self-defeating.
Appreciate the push. It's the right one.