News

Field note 42 · Detection science

A Collusion Benchmark Needs Ground Truth

Without controlled channels and incentives, benchmarks cannot tell true coordination from coincidental correlation.

Editorial illustration for A Collusion Benchmark Needs Ground Truth

Benchmarks need known opportunities, controlled interventions, realistic false positives, and evidence-rich outcomes.

Why this question matters

A dataset of suspicious traces is not enough because intent is rarely observable and labels may reflect analyst assumptions. Synthetic environments provide ground truth but can become too simple. Production traces are realistic but difficult to label and share.

A strong benchmark combines controlled simulations, red-team scenarios, and privacy-preserving real patterns. It should score detection, calibration, explanation, time to intervention, and damage avoided rather than raw classification alone.

Signals worth observing

  • Benchmark labels depend only on behavioral similarity.
  • All negative cases lack shared models or infrastructure.
  • A detector succeeds by memorizing scenario-specific artifacts.

Practical control direction

  1. Include common-cause and legitimate-cooperation negatives.
  2. Hold out channels, incentives, and organizations, not only traces.
  3. Evaluate explanations against known causal structure.
AgentCollusion lensThe benchmark should reward evidence that supports governance decisions, not merely an alarm score.

Sources and further reading

Next field note: Collusion Monitoring Must Preserve Privacy