They may not be answering the same question

Imagine a genuine phone video with an AI-cloned voice laid over it. One service looks mainly for generated frames and reports a low score. Another analyses the soundtrack and calls it suspicious. A third searches for a face swap and finds nothing. The results appear to disagree, but each tool has examined a different part of the file.

The word detector hides a crowded category. Some systems look for fully generated images, some for facial manipulation, some for synthetic speech and some for statistical patterns in writing. Even two tools with an identical-looking percentage may define the positive class differently. Before comparing numbers, read the label beside them.

Models learn from different examples

A detector learns from examples. The people building it choose which real cameras, codecs, languages, editing tools and generators appear in the training data. They also choose what counts as fake. A model trained heavily on face swaps will notice different clues from one trained on diffusion-generated landscapes.

That history matters most when a new generator or an unusual real file arrives. Research on cross-dataset detection repeatedly finds that performance can fall when the test material comes from a different distribution. In plain English: a model can be excellent on the kinds of fakes it knows and hesitant - or confidently wrong - somewhere new.

The percentage is a product decision too

A raw model produces a signal. The service still has to turn that signal into a result. It may calibrate the score, combine several specialist models, set a threshold for 'suspicious' and decide when to return 'unable to evaluate.' Those choices are not decorative. Moving a threshold catches more fakes but can also flag more authentic files.

One provider may show confidence in its final classification; another may show estimated AI likelihood. Those are not interchangeable. An authentic classification with high confidence should not be read as a high probability of AI. Good reporting names the class, explains the scale and leaves room for an inconclusive result.

The file you upload is part of the experiment

A WhatsApp copy is not technically identical to the camera original. A social platform may resize it, reduce the frame rate, recompress the audio and strip metadata. A screenshot changes it again. Some forensic traces become weaker; new compression patterns appear. Two services may respond differently because their models tolerate those changes differently.

There is a quieter source of accidental disagreement: people do not always upload the same file. One test uses the downloaded clip, another a screen recording, and a third a trimmed export. If you want a fair comparison, begin with the exact same bytes, not three versions that merely look alike.

A fair way to compare results

Write down the question first. Are you checking whether the picture was generated, whether a face was swapped, whether the voice is synthetic, or whether the caption is true? Then use tools that actually claim to answer that question. Check supported formats and minimum lengths, and save the model or report version when it is available.

  • Use the same, highest-quality file in every service.
  • Compare status labels and definitions before comparing percentages.
  • Separate image, audio, face-manipulation and authorship results.
  • Record 'not applicable' and 'inconclusive' instead of converting them into zero.
  • Check the source and claim independently of every detector score.

When the tools still disagree

Do not average the percentages. They are not votes cast on a shared scale. Look for the reason one result may be better matched to the file: a supported language, a cleaner original, a specialist audio model or a clear provenance record. If no reason stands out, keep the conclusion open.

Disagreement is useful information. It tells you the case is sensitive to the method and probably should not support a public accusation, payment decision or legal conclusion on automation alone. The responsible sentence may be: 'The automated checks were mixed, and the source could not be independently confirmed.' That is less dramatic than 97%, but much more informative.

Automated results require source and context review.

Continue with independent verification.

Run a careful second opinion
Sources

Primary reading

We use original standards, regulators, public institutions and research papers wherever possible. Sources were last checked on 2 September 2026.