How to Choose a Deepfake Detection Service
Give each provider the same files, then see what a colleague can actually do with the results. This checklist covers photo and video checks, report quality, batches and API integration.
Define what a successful evaluation must show
An AI image checker, a video detector and a face-swap detector may appear in the same search results. They do not necessarily test the same thing. Ask what each service analyses: the whole image, a face, selected video frames or a sequence over time. A supported file extension tells you what you can upload, not which manipulations the service can detect.
Start with one decision: which submitted photos or short videos should a reviewer examine further? Write down the required coverage, acceptable errors, turnaround time and cost before testing. Separate fully generated images, face manipulation and local edits; success on one task does not establish coverage of another.
The UK DSIT market study published in March 2026 identifies inconsistent datasets and accuracy metrics as barriers to comparing services. The workflow below is our proposed buying method, not a vendor ranking or a claim that any service has passed it.
- Read the DSIT deepfake detection market study
- AI image detector: photo checks and their limits
- AI video detector: what a video check covers
- Deepfake checker: choose the relevant media check
- Coverage: record the formats, file sizes, clip lengths and manipulation types you need.
- Errors: set acceptable false-alert and missed-detection limits for your review process.
- Operations: set a completion-time target, reviewer workload limit and monthly budget.
- Evidence: specify required report fields, exports, retention and access controls.
- Decision: name the reviewer who can accept, restrict or reject the service.
Build a representative, labelled file set
Include camera originals with a documented source, ordinary edited photographs, and generated or manipulated examples whose creation history you know. Match the subjects, devices and upload channels in your actual queue. Keep disputed files in a separate exploratory group; a convincing appearance or another detector’s opinion is not a reliable label.
Keep each original linked to its resized, cropped, screenshotted or recompressed versions. Record the source, label basis, transformation, generator where known and file hash. Separate results for full generation, face swaps and local edits instead of combining them into one accuracy figure.
Split the material into a setup set and an untouched evaluation set. Keep related originals and derivatives in the same split. Where practical, also hold out a generator family from your setup set. That tests generalisation within your evaluation; it does not prove the generator was absent from the vendor’s training data. Ask the vendor what it can disclose.
Freeze the rules before scoring the holdout
Use the setup set to verify requests, understand labels and choose your operating rule. A provider may expose a threshold, a fixed classification or an inconclusive state. Record which you use, then freeze that rule and the service version where available before opening the evaluation labels.
Submit identical files to every candidate under the agreed test conditions. Keep the labels hidden during scoring and record the test date, mode, status and raw output. If a vendor tunes its system on those files, treat them as development material and obtain a fresh holdout. Sandbox fixtures check integration behaviour; they cannot measure detection performance.
Count errors and missing decisions separately
For files receiving a decided result, calculate the false-positive rate as authentic files incorrectly flagged divided by authentic files with a decided result. Calculate the false-negative rate as known generated or manipulated files incorrectly given a negative result divided by those target files with a decided result. Publish the counts and denominators.
Alongside those rates, report inconclusive, unsupported and failed results separately for each class, plus the share of all submitted files receiving a decision. A service that abstains frequently can appear accurate on its remaining results while leaving most work to people. Never convert an unavailable score into zero.
- NIST AI 100-4: detection evaluation and performance measures
- Deepfake detector accuracy: false positives and real-world evaluation
- Break results down by task, image or video, quality and transformation.
- Report uncertainty: no false positives observed in a small pilot does not establish a zero false-positive rate.
- Estimate alerts and review time using your real mix of incoming files; a balanced test set does not predict that workload.
Have a second reviewer use the exported report
Give another reviewer the exported record and ask them to identify the file, finding, limitations and next step. Look for a report ID, analysis timestamp, source identity or recorded hash, and clearly separated detector and contextual findings. A hash identifies bytes; it does not prove that the depicted event happened.
Compare the PDF with the structured response. Scores, classification, uncertainty and material disagreements should agree. Check what happens when metadata or a preview is unavailable, whether past results remain stable and whether exporting again triggers another analysis. Keep your decision notes linked to the report.
A convincing AI report can still describe something that is not there
Read the explanation with the original file open beside it. If a report describes a detail you cannot see, stop and ask where that claim came from. A neat layout, a confident sentence and a valid JSON response do not answer that question.
In a preprint submitted on 18 September 2026, Laurent Colbois and Sébastien Marcel tested explanations from vision-language models comparing faces. Some explanations described facial regions hidden by masks or sunglasses. Models with similar face-verification performance could differ in explanation quality. This was a face-comparison study, not a deepfake-detection benchmark or a test of DeepfakePolicy. Its findings cannot be used as error rates for our service.
Our recommendation for buyers is to test the explanation as well as the detector result. This is an additional review check, not a new way to calculate detection accuracy. NIST’s 2021 principles for explainable AI provide useful background: an understandable explanation and one that accurately reflects the system’s process are different requirements.
Use a small set of authorised, non-sensitive files with deliberately missing information. For a media check that offers visual commentary, try a crop that removes the detail being discussed. For photo-and-document comparison, try a label with an unreadable serial number or a receipt with the date cut off. These are proposed test cases, not results from the face-comparison study. A useful response identifies the missing evidence and leaves the question open; it does not borrow a value from another file and present it as visible in both.
Keep three things separate: what the detector returned, what any visual or contextual review observed, and what your colleague decided to do. A visual description is not automatically an explanation of a detector’s internal score. Ask the provider which claims come from which process. Before declining a claim, restricting an account or accusing someone of fraud, review the source material and obtain the evidence needed for that decision.
- Read the face-comparison explanation study (preprint, 18 September 2026)
- NIST: Four Principles of Explainable AI (2021)
- Photo and document cross-checks: source-linked findings and limits
- JSON and PDF reports: what should stay the same
- Take each material factual claim and find its supporting file, page or frame. Record unsupported claims separately from wrong detector classifications.
- Check missing, cropped and unreadable information. Does the report say what could not be assessed, or quietly fill the gap?
- If two files disagree, check that both observations remain visible. A summary should not settle the conflict without a stated basis.
- Compare the web report, JSON and PDF. Formatting or a client logo must not add facts, remove uncertainty or turn an AI signal into a fraud finding.
- Set an acceptance rule before testing. One invented detail that changes the decision deserves investigation even if the overall detection score looks good.
Test the workflow your team will actually use
For an API integration, test supported uploads, asynchronous completion, retries, signed webhooks and explicit errors. Confirm that repeated requests cannot silently create duplicate charges and that failed uploads remain distinguishable from completed low-signal results. Measure upload-to-report time under a representative workload.
For browser batches, test file selection, filename-to-report matching, return to a completed batch and CSV reconciliation. Before uploading real cases, read the processing terms: originals, report data, thumbnails, logs and backups may have different retention periods. Check company access, deletion and training-use terms separately from personal-account terms.
Record acceptance, restrictions and cost
Complete one worksheet for each provider. Enter your acceptance rules before testing, then attach the evidence and observed result for every criterion. The linked CSV is a blank buyer worksheet, not benchmark results. Mark each requirement as met, not met or unresolved; do not average away a failed essential requirement.
Compare the cost of completing your workflow: analysis, chargeable inconclusive outcomes, retries, exports and reviewer time. Record the quoted currency, unit, date and any minimum commitment. Approve only the tested scope, document who owns unresolved cases and repeat the relevant checks after material model or workflow changes.
Evaluating DeepfakePolicy with this checklist
This guide is written by DeepfakePolicy, which sells photo and video checks and has a commercial interest in your choice. Apply the same criteria to our service. Our API provides structured reports and PDF exports; company browser batches support selected photos and short videos. The linked samples are illustrative records, not measured accuracy evidence.
Create a company workspace to test the integration with fixed sandbox responses, then connect company billing before evaluating real files. Keep the final decision with your reviewers: media signals do not establish identity, claim validity, full document authenticity or the truth of a caption.
Frequently asked questions
How do I compare AI image detectors?
Submit the same labelled photos to each service, including camera originals, ordinary edits and known AI-generated examples. Record false positives, missed detections and files that receive no usable result. Then check whether the report gives a reviewer enough information to act. A headline accuracy percentage is not a comparison on your own files.
Can I use an image detector to evaluate AI-generated video?
A frame check assesses that frame. It does not establish coverage of the whole clip, motion, face swaps or audio. Ask how the video service selects frames, handles unsupported content and reports partial results. Test image and video performance separately.
Does a detailed AI detection report prove that a file is fake?
No. Keep the detector finding separate from visual commentary and source research. Check factual statements against the submitted material and preserve uncertainty. A PDF is a record of the analysis, not an authenticity certificate or a finding of fraud.
When do I need an API rather than a browser checker?
A browser checker suits occasional files; a batch workflow helps reviewers handle a queue. An API makes sense when your own system needs to submit files, track processing and collect results. Test retries, access controls, spending limits and report exports before connecting it to live work.
Continue with independent verification.
Evaluate the photo and video APIPrimary reading
We use original standards, regulators, public institutions and research papers wherever possible. Sources were last checked on 23 September 2026.
- UK DSIT: Deepfake detection technology market study (2026)
- NIST: Guardians of Forensic Evidence
- NIST AI 100-4: Reducing Risks Posed by Synthetic Content
- Colbois and Marcel: Explanatory quality of vision-language models in face recognition (preprint, 18 September 2026)
- NISTIR 8312: Four Principles of Explainable Artificial Intelligence (2021)