What launched on 29 September 2026

OpenAI released GPT-6.1 Sol on 29 September 2026. Standard API pricing is $2 per million input tokens, $0.10 for cached input and $10 for output. Access starts in the API, ChatGPT Work and Codex; the launch announcement says ordinary Chat access is still pending.

OpenAI positions it close to GPT-6 Astra on several professional and agent tasks, with standard input and output prices at one-fifth of Astra’s. Those are vendor-reported comparisons. We have not independently benchmarked either model for this article.

Our assessment: a promising candidate for routine review

We would trial Sol first for bounded extraction, evidence summaries and preparing a reviewer’s queue. The interesting opportunity is to examine more cases with the same budget, while preserving the option to stop, ask for missing evidence or escalate.

We would not declare Sol universally better than Astra, or choose Astra merely because it costs more. The comparison that matters is whether a model produces fewer unsupported claims on your documents at a tolerable total cost. Different case types can produce different answers.

A reasonable pilot starts with Sol on routine cases and tests Astra on the difficult subset. Promote the larger model only where measured gains justify the extra expense. This is our proposed evaluation strategy, not a description of DeepfakePolicy’s production model routing.

Pricing: a worked example, not a quote for a fraud check

An illustrative request with 10,000 uncached input tokens and 2,000 total billed output tokens would cost $0.02 + $0.02 = $0.04 at Sol’s standard launch rates. This is arithmetic from stated token assumptions, not a measured document-processing cost.

A real case may need multiple requests, billable reasoning tokens, OCR, tools, retries and a human reviewer. Cached rates apply only to qualifying cached input. Track cost per correctly reviewed case, not just cost per API request; savings disappear quickly when someone has to repair an unsupported conclusion.

What GDP.pdf and AutomationBench actually measure

GDP.pdf tests professional reasoning using complex PDF documents. AutomationBench tests complete business workflows across simulated tools and scores the resulting business state. These are useful measures of document understanding and execution.

Neither scope establishes a model’s fraud-detection sensitivity, deepfake-detection rate or identity-proofing reliability. A PDF question can be answered correctly while the underlying PDF is forged. A workflow can update the intended record while accepting a false premise.

For model selection, keep the benchmark version, reasoning setting and fallback policy beside every score. Do not combine figures from different setups into a single claim of overall superiority.

Safety findings: read the conditions beside the numbers

OpenAI’s system card treats Sol as Critical in cybersecurity and High in biological/chemical capability. These describe preparedness capability thresholds, not a certification for fraud review. The card also says no reviewer-bypass attempts were observed in the specified auto-review evaluation.

A separate warning-respect test found unwanted persistence in 23.5% of Sol rollouts versus 17.4% for Astra. It ran without system-level controls and does not measure how often circumvention would succeed with those controls. OpenAI also cautions that research/API evaluation conditions may differ from production.

Our interpretation is operational: useful improvement in one test does not remove the need for access limits and action approval. A model allowed to read an invoice should not automatically gain permission to change a supplier account or release a payment.

A fictional invoice example: correct reading, unresolved authority

Suppose an invoice, a delivery photo and a message all mention the same order. The invoice also gives a replacement bank account. A model extracts the amounts correctly and produces a persuasive explanation that the documents agree.

That still leaves the account change unresolved. Agreement between supplied files is not independent confirmation that the supplier authorised the change. Both models could read every field correctly and still support the wrong business decision.

The useful review packet names the changed field, cites its source, shows the earlier trusted record and states what has not been independently confirmed. An authorised reviewer checks the change through a contact established before the suspicious message. This example is hypothetical; it is not a reported Sol incident.

A pilot benchmark that can answer your buying question

Build a held-out set that reflects the files you actually receive. Include ordinary records as well as ambiguous and adversarial cases. Establish reference answers through qualified review before comparing model outputs. Keep the same source files, task, tool access and output schema for both models.

Measure useful work and costly errors separately. An elegant summary is a failure if its main assertion has no supporting passage. A refusal or insufficient-evidence result can be correct when the source really is unreadable or incomplete.

  • Include scans, cropped pages, inconsistent dates and amounts, duplicate records and lookalike names.
  • Test changed payment details and instructions embedded inside source documents.
  • Count unsupported claims, missing contradictions and incorrect citations.
  • Record correct abstentions, latency, reviewer corrections and total billed cost.
  • Keep high-impact actions behind approval and validate a small controlled rollout before scaling.

Document review and deepfake detection need different evidence

A general assistant can explain what a photograph appears to show. That description does not establish the photograph’s origin, whether pixels were edited, or whether the caption is true. Preserve the original file and assess the media claim separately from the document’s text.

DeepfakePolicy Cross-check compares supplied records and highlights consistencies, discrepancies and insufficient evidence. Its output supports review; it does not determine identity, supplier authority or whether a person committed fraud. Media-detector results are separate probabilistic signals.

Our recommendation is to keep source citations, file identity, unresolved conflicts and the reviewer’s decision visible. Changing the language model should not silently turn a consistency finding into an authenticity verdict.

Our verdict

Sol deserves a controlled trial for document-heavy review. Our preference is to buy measured reliability: start with a bounded task, preserve the evidence and pay for escalation only when it solves an observed failure. The model name should never become the reason a document is accepted as genuine.

FAQ

Frequently asked questions

Is GPT-6.1 Sol better than GPT-6 Astra?

There is no universal winner for document and fraud workflows. Our recommendation is to compare both on a held-out set of your own cases and choose by unsupported claims, missed contradictions, reviewer corrections and total cost.

Can GPT-6.1 Sol detect forged documents or deepfakes?

The cited PDF and business-workflow benchmarks do not validate that capability. Document understanding, consistency comparison, media detection and independent verification address different questions.

Should a fraud-review system switch from GPT-6 Sol immediately?

Run a controlled comparison before switching. Check the new model against a fixed reference set, keep existing evidence records intact and establish rollback criteria from the errors you observe.

Does a stronger model remove human review?

No. A reviewer still needs the original evidence, cited fields, unresolved questions and authority to stop consequential actions. The required controls should follow the business decision and its impact.

Automated results require source and context review.

Continue with independent verification.

Explore document and photo Cross-check
Sources

Primary reading

We use original standards, regulators, public institutions and research papers wherever possible. Sources were last checked on 29 September 2026.