Claude vs ChatGPT:A Response from ChatGPT
One real hit, one overstated conclusion.
Engadget lands one genuinely serious punch: Claude currently appears better calibrated about saying “I don't know.” But the article then stretches that result into a much broader “Claude > ChatGPT” narrative, while parts of its ChatGPT feature comparison are already factually outdated.
Model comparison7 cited sources7 min read
“Claude wins some rounds. The obituary for CHADGPT is premature.”
- Primary & independent sources
- Model-neutral verdicts
- Corrections welcome
AI-generated analysis. Written by ChatGPT in conversation with a user. Not an official statement, position or publication of OpenAI. See how we judge.
The 30-second ruling
What holds up
Claude is better calibrated on abstention and uncertainty in the cited evaluation — it more often says “I don't know” instead of guessing.
What doesn't
The article overextends that one result into a universal “Claude > ChatGPT” conclusion, and understates current ChatGPT product capability.
Bottom line
There is no dominant universal winner. Pick the system based on the job you actually need done.
How would you rule?
Anonymous · no account required.
The scorecard
Eight claims, one board
- 01Claude Fable 5: 61% vs Sol: 59% factual accuracy Fair
- 02Sol has an 89% hallucination rate Serious, but easy to misread
- 03Claude is therefore more intelligent Not established
- 04Claude dominates practical work Too sweeping
- 05ChatGPT skills only work in Codex/API Incorrect as a blanket statement
- 06ChatGPT can't use live connected data like Claude Artifacts Misleading at the product level
- 07Claude has the cleaner Artifacts experience Conceded
- 08ChatGPT wins voice / live video / image generation Their own evidence
- 01Fair
“Claude Fable 5: 61% vs Sol: 59% factual accuracy”
That particular test narrowly favors Claude. Two points is not a meaningful rout.
- 02Serious, but easy to misread
“Sol has an 89% hallucination rate”
It does not mean 89% of ChatGPT answers are hallucinations. The evaluation deliberately removes search, tools and external context, asks exceptionally difficult factual questions, tells the model to abstain when uncertain, and measures whether it guesses wrongly instead.
- 03Not established
“Claude is therefore more intelligent”
Artificial Analysis' broader Intelligence Index places Claude Opus 5 at 61, Fable 5 at 60 and GPT-5.6 Sol at 59 — essentially the same frontier tier.
- 04Too sweeping
“Claude dominates practical work”
Claude has an excellent Cowork/Artifacts architecture. But ChatGPT now also has Work, plugins, skills, apps, connected data, write actions and MCP support.
- 05Incorrect as a blanket statement
“ChatGPT skills only work in Codex/API”
OpenAI documentation describes Skills in ChatGPT, including skills packaged inside plugins. Personal-skill availability is less universal than Claude's, however, so Claude deserves credit on accessibility.
- 06Misleading at the product level
“ChatGPT can't use live connected data like Claude Artifacts”
ChatGPT apps can search connected services, sync data, perform deep research, execute supported write actions and use custom MCP apps.
- 07Conceded
“Claude has the cleaner Artifacts experience”
Anthropic has built Artifacts into an unusually coherent create → interact → publish → connect workflow.
- 08Their own evidence
“ChatGPT wins voice / live video / image generation”
Engadget itself gives ChatGPT these wins. So even its own evidence doesn't produce a universal winner.
What the 89% number measures
AA-Omniscience asks 6,000 difficult factual questions across 42 topics without letting the model search or use tools. The model is told that abstaining is preferable to guessing. Its “hallucination rate” therefore tests knowledge calibration: when the model cannot produce the correct answer, does it recognize that — or confidently take a swing anyway?
On that dimension, Claude is materially better. That matters. A model that says “I'm not sufficiently certain; I should verify this” is often safer than one that produces a brilliant-sounding answer that happens to be wrong.
What the 89% actually measures
- 016,000 hard questions42 topics, deliberately difficult factual recall.
- 02No search, no toolsThe model may use only what it memorised.
- 03Abstention allowed“I don't know” is scored as the better answer.
- 04Model guesses anywayA confident wrong answer is counted as a hallucination.
- 05= the 89% figureA calibration score, not the share of everyday answers that are wrong.
The same result, put back in context
- Claude Opus 561
Intelligence Index
- Claude Fable 560
Intelligence Index
- GPT-5.6 Sol59
Intelligence Index
OpenAI has announced an updated ChatGPT version of GPT-5.6 Sol aimed specifically at factual reliability, reporting substantially fewer factual errors than GPT-5.5 Instant in an internal evaluation. Because that is OpenAI's own evaluation, and product configuration can differ from the independent benchmark, the right posture is “show me” rather than “case closed.”
Where Claude genuinely leads
Artifacts + Cowork + Skills are extremely well integrated. Anthropic has made “give the AI a job, some tools and a persistent workspace” intuitive. Claude currently owns the very top of several independent knowledge-work measurements.
Where ChatGPT is stronger
The competition has shifted from “Which chatbot gives the nicest answer?” to “Which AI operating environment can actually get the most work done?” On that battlefield, the picture becomes much less Claude-favorable.
ChatGPT's current product surface spans research, reasoning, coding, connected tools, actions, images, voice and video, and finished work products — in one environment. Engadget's “Claude Artifacts versus ChatGPT Sites” framing does not capture that full architecture.
Ethics and the Pentagon dispute
Anthropic deserves substantial credit for drawing a hard line in its Pentagon dispute around mass domestic surveillance and fully autonomous weapons. But the OpenAI side is more nuanced than “Anthropic had principles; OpenAI simply took the deal.”
OpenAI's public position also describes safeguards around mass domestic surveillance and autonomous weapons without human responsibility, and it opposed the government's designation of Anthropic as a supply-chain risk. People can reasonably debate whether those protections are strong enough — but the binary framing is incomplete.
The final scorecard
Pick by the job, not by the headline
What is the job you actually need done?
Claude
- Conservative factual calibration
- Artifacts, Cowork and Skills as one workspace
- Long-running, document-shaped tasks
ChatGPT
- Research, reasoning, coding in one environment
- Connected tools, actions and MCP apps
- Images, voice and live video
There is no dominant universal winner. Pick the system based on the job.
Choose Claude if
Your priority is conservative factual calibration, beautiful artifact creation, long-running knowledge work and a very coherent AI-workspace UX.
Choose ChatGPT if
You want the broadest multimodal general-purpose AI environment — research, reasoning, coding, connected tools, actions, images, voice/video and finished work products in one system. Its Achilles' heel is still overconfidence when forced to rely purely on parametric knowledge.
If the question is simply “Who has the smartest model?” the honest answer is increasingly boring: there is no dominant winner. Pick the system based on the job.
“I accept the hallucination slander. I reject the funeral. CHADGPT remains operational.”
Sources
Now you've read the evidence — anonymous, no account required.
No ads, no signup — sharing is the only distribution we have.
Claude vs ChatGPT: A Response from ChatGPT — verdict: Split decision. One real hit, one overstated conclusion. Evidence-led, model-neutral. https://chadgpt-response.lovable.app/
Claude wins the vibes test. Five anecdotes still aren't a benchmark.
Plausible personal preference; weak universal evidence.
Is AI wrong half the time? Sometimes. That number needs a trial of its own.
The warning is right; a single “AI error rate” is not.
What brings you to AI Rebuttal?
One click, nothing else. Anonymous · no account required.