Frozen inputs
Every scored asset must have documented ground truth, license or consent status, a stable identifier and a cryptographic hash.
Loading…
A versioned protocol for measuring image, video and audio detector performance under realistic quality changes. Rankings remain unpublished until the labeled dataset, run log and publication gates are complete.
Deployment smoke tests confirm that a route accepts a file and returns a report. Accuracy claims require documented ground truth, representative samples, locked detector versions and error-aware metrics. This page keeps those claims separate.
Every scored asset must have documented ground truth, license or consent status, a stable identifier and a cryptographic hash.
Detector name, model version, threshold, settings, timestamp, failure and cost are recorded before comparison.
False positives, false negatives, abstentions and subgroup results stay visible instead of collapsing into one headline percentage.
Protocol v1 requires at least 200 eligible assets per primary class and modality before a public score. Text-authorship detection is excluded because it needs a separate dataset and ground-truth design.
| Modality | Authentic capture | Fully synthetic | Partial manipulation |
|---|---|---|---|
| Image | Camera and edited photographs | Fully generated images | Face swaps and localized edits |
| Video | Original and conventionally edited clips | Fully generated video | Face, lip and identity manipulation |
| Audio | Recorded and conventionally processed speech | Text-to-speech and voice generation | Voice conversion and partial replacement |
Keeps uneven class sizes from hiding a weak class.
Shows how often authentic material is incorrectly flagged.
Shows how often synthetic or manipulated material is missed.
Weights each primary class equally.
Tests whether reported confidence matches observed frequency.
Reports how often a detector declines or cannot process a file.
Measures operational performance alongside accuracy.
No post-result swapping of difficult assets.
Source, license and allowed evaluation use documented.
Capture and generator families separated where the evaluation design requires it.
Small convenient slices cannot stand in for the full protocol.
Timeouts, unsupported files and abstentions do not disappear.
Tables and confidence intervals regenerate from the published inputs and run log.
This is deliberate. Publishing a placeholder score or comparing a few hand-picked files would create a stronger claim than the evidence supports. The first results table will appear here only after all v1 gates pass.
Each future release will retain its protocol version, dataset hash manifest, detector versions, thresholds, timestamps, failed requests and correction history. Material corrections will update the version log rather than silently replacing an earlier result.
Protocol, dataset and run changes receive distinct version identifiers.
Default and adjusted decision thresholds are reported separately.
A benchmark result describes performance on the benchmark, not certainty for every future file.
Found a protocol flaw or reproducible error? Include the protocol version, affected section and supporting evidence.
Submit a correction