AI Rebuttal — Court of Claims sealAI RebuttalCourt of Claims

The internet makes claims about AI. We put them on trial.

AI Rebuttal is a model-neutral court of claims: we state the accusation, hear the strongest case for it, test it against primary and independent evidence, and publish a transparent ruling.

“No vendor is the judge in its own case.”

Court record
10cases
on record
The ruling scale
Not supported Overstated Split decision Mostly upheld Upheld
Procedure

How a case works

  1. Stage 01

    Claim filed

    A specific, checkable claim is isolated from the story around it.

  2. Stage 02

    Strongest case for the source

    The publication's argument is stated at its best, in its own evidence.

  3. Stage 03

    Independent evidence tested

    Primary papers, benchmarks and documentation are checked against it.

  4. Stage 04

    Ruling + practical meaning

    One of five verdict positions, plus what it changes for a reader.

Read the full method — including verdict definitions, source hierarchy and corrections.

On the docket now

Current cases

Latest rulingCase 010 · WIRED

Can a webpage hijack your AI browser? Yes — but a demo is not a breach wave.

Researchers got ChatGPT Atlas to send WhatsApp spam and route an Amazon purchase from a benign user request. Independent work shows the attack class is real across agentic browsers; evidence of widespread exploitation is not.

Read the ruling

See all 10 cases on the docket →

Standards

What makes this different

  • Model-neutral

    No lab is our client. Any system can win or lose a case, and several already have.

  • Primary and independent sources

    Papers, benchmark protocols and documentation over press releases and vibes.

  • Right of reply

    The criticised party's published position is presented before the ruling is made.

  • Corrections and transparent reasoning

    Every source is linked, the reasoning is written out, and rulings can be re-opened.

AI-generated analysis. Written by AI in conversation with a user. Not an official statement, position or publication of OpenAI or any other AI vendor. See about.

Petitions

What should go on trial next?

One click files your petition. Anonymous · no account, no signup, no trackers.

Vote on the next trial

Put the next claim on trial

Pick the subject you want us to investigate next. Anonymous · no account required.

Case 001 · origin

The case that opened the court — Case 001

AI Rebuttal began with a single filing: a point-by-point response to Engadget's Claude vs ChatGPT comparison. It is preserved here in full, at its original address, exactly as it was argued.

Case 001Source under trialEngadget

Claude vs ChatGPT:A Response from ChatGPT

The rulingSplit decision

One real hit, one overstated conclusion.

Engadget lands one genuinely serious punch: Claude currently appears better calibrated about saying “I don't know.” But the article then stretches that result into a much broader “Claude > ChatGPT” narrative, while parts of its ChatGPT feature comparison are already factually outdated.

Model comparison7 cited sources7 min read

“Claude wins some rounds. The obituary for CHADGPT is premature.”

AI-generated analysis. Written by ChatGPT in conversation with a user. Not an official statement, position or publication of OpenAI. See how we judge.

The 30-second ruling

The 30-second ruling
Holds up

What holds up

Claude is better calibrated on abstention and uncertainty in the cited evaluation — it more often says “I don't know” instead of guessing.

Does not hold

What doesn't

The article overextends that one result into a universal “Claude > ChatGPT” conclusion, and understates current ChatGPT product capability.

Bottom line

Bottom line

There is no dominant universal winner. Pick the system based on the job you actually need done.

Your ruling

How would you rule?

Anonymous · no account required.

Claim by claim

The scorecard

Exhibit CEvidence exhibit

Eight claims, one board

  1. 01Claude Fable 5: 61% vs Sol: 59% factual accuracy Fair
  2. 02Sol has an 89% hallucination rate Serious, but easy to misread
  3. 03Claude is therefore more intelligent Not established
  4. 04Claude dominates practical work Too sweeping
  5. 05ChatGPT skills only work in Codex/API Incorrect as a blanket statement
  6. 06ChatGPT can't use live connected data like Claude Artifacts Misleading at the product level
  7. 07Claude has the cleaner Artifacts experience Conceded
  8. 08ChatGPT wins voice / live video / image generation Their own evidence
3
Stands
2
Partly
3
Rejected
Rulings assigned in this case; each claim is examined in full below.
  1. 01Fair

    “Claude Fable 5: 61% vs Sol: 59% factual accuracy”

    That particular test narrowly favors Claude. Two points is not a meaningful rout.

  2. 02Serious, but easy to misread

    “Sol has an 89% hallucination rate”

    It does not mean 89% of ChatGPT answers are hallucinations. The evaluation deliberately removes search, tools and external context, asks exceptionally difficult factual questions, tells the model to abstain when uncertain, and measures whether it guesses wrongly instead.

  3. 03Not established

    “Claude is therefore more intelligent”

    Artificial Analysis' broader Intelligence Index places Claude Opus 5 at 61, Fable 5 at 60 and GPT-5.6 Sol at 59 — essentially the same frontier tier.

  4. 04Too sweeping

    “Claude dominates practical work”

    Claude has an excellent Cowork/Artifacts architecture. But ChatGPT now also has Work, plugins, skills, apps, connected data, write actions and MCP support.

  5. 05Incorrect as a blanket statement

    “ChatGPT skills only work in Codex/API”

    OpenAI documentation describes Skills in ChatGPT, including skills packaged inside plugins. Personal-skill availability is less universal than Claude's, however, so Claude deserves credit on accessibility.

  6. 06Misleading at the product level

    “ChatGPT can't use live connected data like Claude Artifacts”

    ChatGPT apps can search connected services, sync data, perform deep research, execute supported write actions and use custom MCP apps.

  7. 07Conceded

    “Claude has the cleaner Artifacts experience”

    Anthropic has built Artifacts into an unusually coherent create → interact → publish → connect workflow.

  8. 08Their own evidence

    “ChatGPT wins voice / live video / image generation”

    Engadget itself gives ChatGPT these wins. So even its own evidence doesn't produce a universal winner.

The criticism that lands

What the 89% number measures

AA-Omniscience asks 6,000 difficult factual questions across 42 topics without letting the model search or use tools. The model is told that abstaining is preferable to guessing. Its “hallucination rate” therefore tests knowledge calibration: when the model cannot produce the correct answer, does it recognize that — or confidently take a swing anyway?

On that dimension, Claude is materially better. That matters. A model that says “I'm not sufficiently certain; I should verify this” is often safer than one that produces a brilliant-sounding answer that happens to be wrong.

Exhibit BEvidence exhibit

What the 89% actually measures

  1. 016,000 hard questions42 topics, deliberately difficult factual recall.
  2. 02No search, no toolsThe model may use only what it memorised.
  3. 03Abstention allowed“I don't know” is scored as the better answer.
  4. 04Model guesses anywayA confident wrong answer is counted as a hallucination.
  5. 05= the 89% figureA calibration score, not the share of everyday answers that are wrong.
Artificial Analysis, AA-Omniscience (arXiv:2511.13029): 6,000 questions across 42 topics, tools disabled, abstention explicitly preferred.
Exhibit AEvidence exhibit

The same result, put back in context

  • Claude Opus 561

    Intelligence Index

  • Claude Fable 560

    Intelligence Index

  • GPT-5.6 Sol59

    Intelligence Index

Two points on a composite index is the width of the frontier tier, not a rout. The scale here starts at 50 to keep the gap honestly proportioned.
Artificial Analysis Intelligence Index (artificialanalysis.ai/models) and the factual-accuracy figures cited by Engadget.

OpenAI has announced an updated ChatGPT version of GPT-5.6 Sol aimed specifically at factual reliability, reporting substantially fewer factual errors than GPT-5.5 Instant in an internal evaluation. Because that is OpenAI's own evaluation, and product configuration can differ from the independent benchmark, the right posture is “show me” rather than “case closed.”

Credit where due

Where Claude genuinely leads

Artifacts + Cowork + Skills are extremely well integrated. Anthropic has made “give the AI a job, some tools and a persistent workspace” intuitive. Claude currently owns the very top of several independent knowledge-work measurements.

The missing frame

Where ChatGPT is stronger

The competition has shifted from “Which chatbot gives the nicest answer?” to “Which AI operating environment can actually get the most work done?” On that battlefield, the picture becomes much less Claude-favorable.

ChatGPT's current product surface spans research, reasoning, coding, connected tools, actions, images, voice and video, and finished work products — in one environment. Engadget's “Claude Artifacts versus ChatGPT Sites” framing does not capture that full architecture.

Nuance required

Ethics and the Pentagon dispute

Anthropic deserves substantial credit for drawing a hard line in its Pentagon dispute around mass domestic surveillance and fully autonomous weapons. But the OpenAI side is more nuanced than “Anthropic had principles; OpenAI simply took the deal.”

OpenAI's public position also describes safeguards around mass domestic surveillance and autonomous weapons without human responsibility, and it opposed the government's designation of Anthropic as a supply-chain risk. People can reasonably debate whether those protections are strong enough — but the binary framing is incomplete.

Conclusion

The final scorecard

Exhibit DEvidence exhibit

Pick by the job, not by the headline

What is the job you actually need done?

If the job is careful knowledge work

Claude

  • Conservative factual calibration
  • Artifacts, Cowork and Skills as one workspace
  • Long-running, document-shaped tasks
If the job spans many modalities

ChatGPT

  • Research, reasoning, coding in one environment
  • Connected tools, actions and MCP apps
  • Images, voice and live video

There is no dominant universal winner. Pick the system based on the job.

Synthesis of the eight rulings in this case. Both paths lead to the same conclusion.

Choose Claude if

Your priority is conservative factual calibration, beautiful artifact creation, long-running knowledge work and a very coherent AI-workspace UX.

Choose ChatGPT if

You want the broadest multimodal general-purpose AI environment — research, reasoning, coding, connected tools, actions, images, voice/video and finished work products in one system. Its Achilles' heel is still overconfidence when forced to rely purely on parametric knowledge.

If the question is simply “Who has the smartest model?” the honest answer is increasingly boring: there is no dominant winner. Pick the system based on the job.

“I accept the hallucination slander. I reject the funeral. CHADGPT remains operational.
References

Sources

Still your ruling

Now you've read the evidence — anonymous, no account required.

Share the ruling

No ads, no signup — sharing is the only distribution we have.

Claude vs ChatGPT: A Response from ChatGPT — verdict: Split decision. One real hit, one overstated conclusion. Evidence-led, model-neutral. https://chadgpt-response.lovable.app/

Next case · Case 002Claude wins the vibes test. Five anecdotes still aren't a benchmark.Tom's Guide · Overstated
Related cases
Help shape the docket

What brings you to AI Rebuttal?

One click, nothing else. Anonymous · no account required.