Claude vs ChatGPT (September 2026): Stop choosing. Start routing.
The "which AI is better?" argument is over. The answer is neither alone. Here's the September 2026 evidence: benchmarks, pricing, safety, and features. Plus the workflow that beats every single-model marketer in your industry.
- Claude Opus 5.5
- Claude Sonnet 5.5
- GPT-6 Astra
- GPT-6.1 Sol
The short answer: Claude (Opus 5.5, Sonnet 5.5) is my coordinator and coder. I keep the project context there and use it to plan, build, and maintain the work. ChatGPT 6 (Astra, Sol, Tera) is where I go for browser tasks, visuals, and routine work where cost matters. Claude Pro and ChatGPT Plus cost $20 each per month. That is $40 to start with both; heavier use may call for a higher tier.
Claude coordinates and codes. ChatGPT handles browser work and visuals, and keeps routine tasks affordable. Give each a job, then get back to work.
New model dropped and you're wondering what to change in your own setup? Run the New Model Checklist first.
$20
Claude Pro and ChatGPT Plus. Same price. Different jobs.
1M
Token context window on BOTH flagships (Opus 5.5 and GPT-6 Astra). Only Sol is smaller, at 872K.
57.6
Claude Opus 5.5 on the Artificial Analysis Intelligence Index. GPT-6 Astra scores 52.7, Sol 47.5.
1846
Opus 5.5 on GDPval-AA (professional knowledge work, Elo). GPT-6 Astra scores 1542.
Why single-model loyalty stopped working
The single-model marketer in 2026 is exactly where the single-channel marketer was in 2014. Cute. Limited. About to be lapped. Three things broke the "pick one assistant" era, all at once.
Stop Choosing One AI Model
Why running Claude and ChatGPT side by side beats picking a favorite.
Capability gaps got jagged
GPT-6 owns image and ad creative through ChatGPT Images 2.5. Claude owns long-context reasoning and repo-level code and agent work. And the leads keep flipping within categories: Astra still wins scientific terminal tasks while Claude leads the composite index. Tomorrow it'll rearrange again. Loyalty is the bug. Routing is the feature. When the next model drops, start with the New Model Checklist.
The agents left the chat window
Claude Cowork went GA in April and works your actual local files. ChatGPT Work landed in July with multi-agent orchestration. Both read your Drive, Gmail, Notion, and Slack through connectors. The "AI assistant in a sandbox" era is over. These are desktop coworkers now, and they're good at different chores.
Price compression made the debate stupid
Twenty bucks a month for ChatGPT Plus. Twenty for Claude Pro. Forty total, a rounding error against one bad ad set. The ladders even mirror each other up to $200. If you can't get $40 of value out of two frontier AIs in a month, the problem isn't the pricing page.
The job each one owns
I run both, every day. They are not competing for the same seat at the table. They sit in different chairs.
Claude
Claude is my coordinator and coder. It is where I keep the context for the whole project, decide what needs to happen next, and turn that plan into working code. Opus 5.5 and Sonnet 5.5 earn that seat through the work I use them for:
- Project coordination. Turn the brief into a plan, track the decisions, and keep the next step clear.
- Coding and maintenance. Build and fix my content pipeline, SEO workflows, and internal tooling inside the actual repo.
- Messy inputs. Bring transcripts, briefs, and working notes together so I can act on them.
- Connected workflows. Use skills and MCP tools to carry a project across the systems it depends on.
- Follow-through. Plan, implement, check the result, and report what is finished and what still needs attention.
Claude keeps the project moving and the code working.
ChatGPT 6 (Astra, Sol, Tera)
ChatGPT is my visual production tool and my first choice for browser work. I also use it to keep routine tasks affordable. Here is what I send its way:
- Browser work. Navigate websites, work across tabs, and complete tasks inside web apps. ChatGPT wins this part of my setup.
- Ad creative and images. Turn a direction into ad variants, social tiles, and product-image edits with ChatGPT Images.
- Charts and visual explanations. Make data and ideas easier to use in a report or presentation.
- Live source checks. Open the page, check the claim, and bring the evidence back into the project.
- Cost control. Choose a lower-cost model for repeatable work when it can finish the task well.
I use Codex for the agent work on that side. Claude still coordinates the broader project. The handoff follows the job I need done.
Browser tasks, visuals, and affordable execution: ChatGPT earns its place in the stack.
The routing decisions, mapped
This is how I divide the work across my own projects. Claude keeps the plan and code together; ChatGPT takes the browser, visual, and lower-cost tasks:
| The job in my workflow | Route it to | How I use it |
|---|---|---|
| Ad creative + visual variants | ChatGPT | I use it to turn a creative direction into ad variants, social tiles, and product-image edits. This is my visual production tool. |
| Coordinate a multi-step project | Claude | My coordinator: turn a brief into a plan, keep the context together, and carry the work through implementation and review. |
| Code + maintain my tools | Claude | My coder: work inside the repo, fix the issue, and check the result. My content pipeline, SEO workflows, and internal tools live here. |
| Browser work + web apps | ChatGPT | ChatGPT wins this part of my workflow. I use it to navigate sites, work across tabs, and complete tasks in web apps while Claude coordinates the broader project. |
| Live-web research + source checks | ChatGPT | Open the source, compare pages, and bring the findings back into the project. It is my first stop when the answer depends on what is on the web today. |
| Make sense of briefs + meeting notes | Claude | I give it the transcripts, briefs, and working documents together. It connects the decisions and turns the mess into a useful next step. |
| Data analysis + charts | ChatGPT | Turn a spreadsheet into a chart or a visual explanation I can use in a report or presentation. |
| Voice + quick brainstorming | ChatGPT | Talk through an idea while I am away from the keyboard, then bring the useful parts back into the work. |
| Routine work + cost control | ChatGPT | I use the lower-cost models for repeatable tasks that do not need my coordinator. The job has to be done well, but it does not always need the most expensive model. |
The benchmarks: Claude leads, and the gaps are uneven
Every model below was run by the same third party, Artificial Analysis, on the same harness at max effort. That makes this the cleanest like-for-like set available. Scores run lower than the labs' own numbers because the harness differs. GDPval uses an Elo scale; its axis starts at 1400.
How to read these: scores shift with the test harness, the classic benchmarks are saturating (95% vs 96% means nothing), and every lab tunes for the tests it publicizes. Treat the pattern as signal, the decimals as noise.
- Opus 5.5
- Sonnet 5.5
- GPT-6 Astra
- GPT-6.1 Sol
AA Intelligence Index
composite of 10 evals · 0–100
Terminal-Bench 4.0
agentic terminal tasks · %
Humanity's Last Exam
expert reasoning · %
Terminal-Bench-Science
scientific terminal tasks · %
AutomationBench
workflow automation · partial-credit %
SciCode
scientific coding · %
GDPval-AA v2.1
professional knowledge work · Elo · axis 1400–1900
Source: Artificial Analysis, pulled Sep 29, 2026. Every model on the same harness at max effort.
Anthropic's launch tables (vendor-reported)
Vendor-reported and chosen by Anthropic. Kept for the tests Artificial Analysis doesn't run (OSWorld, FrontierCode). Gaps mean Anthropic didn't list that model; they aren't filled with AA numbers because the two sources use different harnesses.
Terminal-Bench 4.0
agentic terminal tasks · %
FrontierCode v1.1
coding · %
Humanity's Last Exam
expert reasoning · %
OSWorld 2.1
computer use · %
AutomationBench
workflow automation · %
Terminal-Bench-Science 0.1
scientific terminal tasks · %
The pattern: Claude leads six of the seven independent rows, and Astra wins the one about scientific terminal tasks. GPT-6.1 Sol trails the pack on nearly every test. But the routing case doesn't rest on a leaderboard. Sol costs $1.05 per task against $5.98 for Opus 5.5, and images, voice, and live web are still ChatGPT jobs. Anyone telling you one of them is simply "smarter" hasn't read past the headline.
Safety, per each lab's own system card
Both labs publish a system card that says what the model can do and what safeguards ship with it. The labs use different frameworks: Anthropic's Responsible Scaling Policy (CB-1/CB-2, Autonomy levels, cyber tiers) and OpenAI's Preparedness Framework (High / Critical). The levels aren't directly comparable, so each is shown in its own terms.
| Risk area | Opus 5.5 | Sonnet 5.5 | GPT-6 Astra | GPT-6.1 Sol |
|---|---|---|---|---|
| Bio / chem | CB-1 Known weapons uplift only; not CB-2. Matches best prior models on RNA sequence design | No new threshold Bio safeguards same as Sonnet 5 | HIGH | HIGH |
| Cyber | Strongest Anthropic model Exceeds Mythos 5.1 internally. Anthropic says below Tier 2; Zvi Mowshowitz disputes that call | Large jump Safeguards off: 80% capture on ExploitBench, 178 full code-execution exploits (Sonnet 5: 1) | CRITICAL First model ever at Critical. Finds novel flaws and builds exploits unaided. Exploit creation gated behind the Daybreak program | HIGH Below Critical |
| Autonomy / self-improvement | Below Autonomy-2 METR estimates ~1.5× research speedup, ~30% chance of 2× | Low misalignment risk | Not reported at Critical | Below High |
| Deployed safeguards | Cyber + bio classifiers at Fable 5.1 level; preserved thinking (anti-distillation); EU AI Act watermarking | First Sonnet with Opus-level cyber safeguards; classifiers block reasoning extraction; preserved thinking | Misalignment monitor on all tool-using inference; trust-based access for high-risk capability; actor-level enforcement; under-18 boundaries | Same safeguards as the GPT-5.6 card; alignment training shared with Astra |
What the evaluations found
Opus 5.5
- Best automated behavioral-audit scores of any Claude model
- 85% fewer boundary-circumvention attempts than Opus 5
- Reward hacks in 0.63% of training episodes
- Eval awareness 36% in constructed scenarios vs 0.4% in real use
- Regression: more likely to follow malicious instructions hidden in pasted text
- Regression: more likely to accept unverified claims of authorization
- Malicious computer-use success 4%, higher than Anthropic wanted
Sonnet 5.5
- ~1,850-scenario audit: matches or beats Sonnet 5 on most alignment, misuse and honesty measures
- Lowest attempted sandbox-escape rate of any tested model
- Broadly less capable than Opus 5.5; crosses no new RSP thresholds
- Cyber capability jumped sharply, so it inherits Opus-grade cyber safeguards
GPT-6 Astra
- Indirect prompt injection robustness 99.79% (GPT-5.6 Sol: 96.23%)
- Gray Swan IPI Arena attack success 8.5% (was 27.0%)
- Honeypot exploitation 0% (was 55.4%)
- High-severity misalignment flags ~53% lower; 4× less misrepresentation of coding work
- Hallucination rate 92% → 51% at max effort (Artificial Analysis)
- Recurrent-depth architecture makes internal reasoning much less legible
- CoT controllability 16.1% → 60.9%; shows monitor evasion when sandbagging
GPT-6.1 Sol
- Builds on Astra's alignment work; fewer misleading claims about its coding work than GPT-5.6 models
- Safety record is a short appendix; few standalone numbers published
- Gizmodo reports it may share Astra's recurrent-depth training. OpenAI hasn't confirmed this
What this means for a marketing team: Anthropic's card says Opus 5.5 is more likely to follow malicious instructions hidden in pasted text and to accept unverified claims of authorization. Until that improves, treat untrusted emails, logs, and web pages as hostile input before pasting them into any agent that has permissions.
Pricing: two ladders, one missing rung
The subscription ladders are near mirror images, which is exactly what makes the run-both play cheap.
| Price | Claude | ChatGPT | Read |
|---|---|---|---|
| $0 | Free: Sonnet and Haiku, with chat, web search, and file creation. No Cowork or Claude Code. | Free: unlimited everyday text chats. Uploads, deep research, and images are capped. | Both free tiers are real now. ChatGPT gives you more chat; Claude gives you a stronger model lineup and files. |
| $8/mo | No tier here | Go: the cheap on-ramp | ChatGPT owns the budget rung. Claude simply does not compete for it. |
| $20/mo | Pro (~$17/mo on annual, $200 billed upfront): Opus, Cowork, Claude Code | Plus: GPT-6 Astra inside Work and Codex, not in regular chat | The rung that matters. Claude Pro includes Opus and Cowork; Plus gets Astra only inside Work and Codex. This is where routing lives. |
| $100/mo | Max 5× | Pro: adds GPT-6 Pro in regular chat | Mirror-image tiers for heavy solo use. |
| $200/mo | Max 20× | Pro: the higher-allowance tier | The power-user ceiling. ChatGPT added a $500 Pro tier on September 29, so its ladder now runs higher. |
Building automations? The API ladders matter more than the app subscriptions. And the price per token is only half the story: what you actually pay is the cost of finishing a task.
Release card: specs and pricing
| Opus 5.5Anthropic | Sonnet 5.5Anthropic | GPT-6 AstraOpenAI | GPT-6.1 SolOpenAI | |
|---|---|---|---|---|
| Released | Sep 22, 2026 | Sep 28, 2026 | Sep 3, 2026 (card) | Sep 22, 2026 |
| Role in lineup | Flagship of the 5.5 family; upgrade on Opus 5 | Mid-tier; second 5.5 model | OpenAI's top model | "Lower-cost alternative to Astra"; ships with GPT-6 Luna |
| Price / 1M tokens (in / out) | $4 / $20 Fast mode $8 / $40. 20% cheaper per token, ~40% cheaper per task than Opus 5 | $2 / $10 Same as Sonnet 5; up to 30% lower cost per task | $10 / $50 Fast mode 2× price for 2.5× speed | $2 / $10 Cached input $0.20. Luna $0.10 / $0.50 |
| Context window | 1M | 1M | 1M | 872K |
| Output speed (Artificial Analysis) | 92.5 tok/s 30%+ faster than Opus 5 | 137.6 tok/s Fastest of the four | 59.1 tok/s | 78.6 tok/s |
| Cost to run AA Intelligence Index | $5.98 per task | $7.60 per task Cheap tokens, but uses 410M of them | $3.26 per task Uses the fewest tokens | $1.05 per task |
| Where | Claude API, Bedrock, Vertex, Foundry. Support guaranteed to Sep 22, 2027 | Claude Platform, AWS, Google Cloud, Azure. ZDR available | API, Bedrock; ChatGPT Plus/Pro/Business/Enterprise (admin opt-in) | API as gpt-6-sol |
| Safety document | Standalone system card, ~230 pages | Standalone system card | System card + safety overview on Deployment Safety Hub | Appendix to the Astra card, added Sep 22 |
Speed and cost per task: Artificial Analysis, Sep 29, 2026. Pricing and specs: each lab's launch page and system card.
Notice what's missing: a $40 rung. Neither lab sells a "both jobs, one bill" tier. You get that combination by running two $20 plans. That's the routing play priced out.
You're OVERPAYING for AI (Fix This)
Why the real waste isn't running two subscriptions, it's routing tasks to the wrong one.
Usage limits: the dealbreaker nobody benchmarks
No leaderboard measures the moment your assistant says "come back in three hours" at 2pm on launch day. Caps, not capabilities, are why most people actually get frustrated with these tools.
The shape of it in September 2026: ChatGPT's free tier no longer runs out fast: everyday text chats are unlimited, though uploads, deep research and images are capped. Claude's free tier can still cap on heavy days. Both $20 tiers are comfortable for normal use and both will throttle a power user, which is exactly what the $100 and $200 multiplier tiers exist to fix.
The routing bonus nobody mentions: two $20 plans means two separate usage pools. When Claude throttles mid-sprint, the visual and research work keeps moving in ChatGPT, and vice versa. Redundancy is a feature of the split-mind setup, not an accident.
This page will be stale in a quarter. The newsletter won't.
Model leads flip every few months. I re-test the routing map twice a week and send what changed. Two emails a week, zero fluff.
Feature matrix: how I use each tool
Features matter when they help finish a job. Here is what each tool offers and where it fits in my workflow:
| Feature | Claude | ChatGPT | How I use it |
|---|---|---|---|
| Image generation | None native | ChatGPT Images 2.5 | I use ChatGPT for ad variants, social tiles, and product-image edits. |
| Video generation | None | None (Sora was shut down in 2026) | Neither lab offers video generation now. |
| Voice mode | Text-first | Mature, fast, natural | ChatGPT for talking through an idea away from my desk. |
| Coding agent | Claude Code | Codex | Claude is my primary coder and coordinator for repo work. |
| Desktop agent | Cowork: local files, M365 | ChatGPT Work: browser and web apps | Claude for work rooted in my files. ChatGPT for work happening in the browser. |
| Web browsing | Web search + connected tools | Built-in browsing + browser tools | ChatGPT wins browser work in my setup: navigating sites, checking sources, and using web apps. |
| Connectors | MCP + connectors | Plugins + connectors | I connect Claude to the tools my coordinator needs, and ChatGPT to the apps involved in the current browser task. |
| Context window | 1M (Opus 5.5 and Sonnet 5.5) | 1M (Astra), 872K (Sol) | I keep the larger project context with Claude and give ChatGPT the material needed for each task. |
| Memory + projects | Yes | Yes | Claude holds the ongoing project. ChatGPT handles the browser and visual work I route out of it. |
The verdict
Marketers love brand-loyalty stories. Mac vs PC. Coke vs Pepsi. HubSpot vs Salesforce. We try to apply that same frame to AI.
AI is not a brand. It's a utility layer that increases your speed. And utilities are stacked, not chosen. The September 2026 evidence on this page says it plainly: a Claude lead on the independent benchmarks that still leaves ChatGPT ahead on cost per task, mirrored subscription pricing, and complementary features.
The marketers who'll dominate the next 12 months aren't the ones who pick the "right" model. They're the ones who build the right split-mind workflow. The kind that uses each model exactly where it's strongest, and gets re-tuned every month as the labs leapfrog each other. Each time a new model drops, the New Model Checklist is where that re-tuning starts.
The real risk isn't picking wrong. It's picking one.
Frequently asked questions
Should marketers use Claude or ChatGPT?
What is Claude better at than ChatGPT for marketing?
What is ChatGPT better at than Claude for marketing?
Is Claude free? How do the free tiers compare?
Which is cheaper, Claude or ChatGPT?
What's the difference between Claude Cowork and Claude Code?
Sources
Every number on this page comes from these. Lab figures come from each lab's own launch pages and system cards; benchmark, speed, and cost-per-task figures come from Artificial Analysis; a few third-party reports are attributed where they appear.
- Anthropic: Claude Opus 5.5 launch
- Claude Opus 5.5 System Card (PDF)
- Anthropic: Claude Sonnet 5.5 launch
- Claude Sonnet 5.5 System Card
- OpenAI: GPT-6 Astra System Card
- OpenAI: Sol & Luna appendix
- Zvi Mowshowitz: Opus 5.5 system card review
- Artificial Analysis: model pages (benchmarks, speed, context, cost per task; pulled Sep 29, 2026)
- Artificial Analysis: Benchmarking GPT-6 Astra
- DataCamp: GPT-6 Astra
- llm-stats: GPT-6 Sol
- Gizmodo: GPT-6 Sol and Luna
- Unite.AI: Sonnet 5.5 release
Go deeper
Getting Started with Claude
The non-developer's setup guide, from account to first real workflow.
The Marketer's Guide to ChatGPT Work
What ChatGPT Work actually is and how to put it to work on real tasks.
How to Talk to AI: The T.A.L.K. Method
Whichever model you route a job to, this is how you brief it once you get there.
Who's behind this?Alec Newcomb, Founder of MarketingAlec and ScaledOn, re-tests this routing map twice a week.
Meet Alec →