Tools & Models

Claude vs ChatGPT (September 2026): Stop choosing. Start routing.

The "which AI is better?" argument is over. The answer is neither alone. Here's the September 2026 evidence: benchmarks, pricing, safety, and features. Plus the workflow that beats every single-model marketer in your industry.

  • Claude Opus 5.5
  • Claude Sonnet 5.5
  • GPT-6 Astra
  • GPT-6.1 Sol
ChatGPT vs Claude: Stop choosing. Start routing.

The short answer: Claude (Opus 5.5, Sonnet 5.5) is my coordinator and coder. I keep the project context there and use it to plan, build, and maintain the work. ChatGPT 6 (Astra, Sol, Tera) is where I go for browser tasks, visuals, and routine work where cost matters. Claude Pro and ChatGPT Plus cost $20 each per month. That is $40 to start with both; heavier use may call for a higher tier.

Claude coordinates and codes. ChatGPT handles browser work and visuals, and keeps routine tasks affordable. Give each a job, then get back to work.

New model dropped and you're wondering what to change in your own setup? Run the New Model Checklist first.

$20

Claude Pro and ChatGPT Plus. Same price. Different jobs.

1M

Token context window on BOTH flagships (Opus 5.5 and GPT-6 Astra). Only Sol is smaller, at 872K.

57.6

Claude Opus 5.5 on the Artificial Analysis Intelligence Index. GPT-6 Astra scores 52.7, Sol 47.5.

1846

Opus 5.5 on GDPval-AA (professional knowledge work, Elo). GPT-6 Astra scores 1542.

I don't have a model anymore. I have a stack.

Why single-model loyalty stopped working

The single-model marketer in 2026 is exactly where the single-channel marketer was in 2014. Cute. Limited. About to be lapped. Three things broke the "pick one assistant" era, all at once.

This page's whole thesis, in 90 seconds

Stop Choosing One AI Model

Why running Claude and ChatGPT side by side beats picking a favorite.

MarketingAlec Open on YouTube →
01

Capability gaps got jagged

GPT-6 owns image and ad creative through ChatGPT Images 2.5. Claude owns long-context reasoning and repo-level code and agent work. And the leads keep flipping within categories: Astra still wins scientific terminal tasks while Claude leads the composite index. Tomorrow it'll rearrange again. Loyalty is the bug. Routing is the feature. When the next model drops, start with the New Model Checklist.

02

The agents left the chat window

Claude Cowork went GA in April and works your actual local files. ChatGPT Work landed in July with multi-agent orchestration. Both read your Drive, Gmail, Notion, and Slack through connectors. The "AI assistant in a sandbox" era is over. These are desktop coworkers now, and they're good at different chores.

03

Price compression made the debate stupid

Twenty bucks a month for ChatGPT Plus. Twenty for Claude Pro. Forty total, a rounding error against one bad ad set. The ladders even mirror each other up to $200. If you can't get $40 of value out of two frontier AIs in a month, the problem isn't the pricing page.

The job each one owns

I run both, every day. They are not competing for the same seat at the table. They sit in different chairs.

Coordinator + Coder

Claude

Claude is my coordinator and coder. It is where I keep the context for the whole project, decide what needs to happen next, and turn that plan into working code. Opus 5.5 and Sonnet 5.5 earn that seat through the work I use them for:

  • Project coordination. Turn the brief into a plan, track the decisions, and keep the next step clear.
  • Coding and maintenance. Build and fix my content pipeline, SEO workflows, and internal tooling inside the actual repo.
  • Messy inputs. Bring transcripts, briefs, and working notes together so I can act on them.
  • Connected workflows. Use skills and MCP tools to carry a project across the systems it depends on.
  • Follow-through. Plan, implement, check the result, and report what is finished and what still needs attention.

Claude keeps the project moving and the code working.

Browser + Visuals + Cost

ChatGPT 6 (Astra, Sol, Tera)

ChatGPT is my visual production tool and my first choice for browser work. I also use it to keep routine tasks affordable. Here is what I send its way:

  • Browser work. Navigate websites, work across tabs, and complete tasks inside web apps. ChatGPT wins this part of my setup.
  • Ad creative and images. Turn a direction into ad variants, social tiles, and product-image edits with ChatGPT Images.
  • Charts and visual explanations. Make data and ideas easier to use in a report or presentation.
  • Live source checks. Open the page, check the claim, and bring the evidence back into the project.
  • Cost control. Choose a lower-cost model for repeatable work when it can finish the task well.

I use Codex for the agent work on that side. Claude still coordinates the broader project. The handoff follows the job I need done.

Browser tasks, visuals, and affordable execution: ChatGPT earns its place in the stack.

The routing decisions, mapped

This is how I divide the work across my own projects. Claude keeps the plan and code together; ChatGPT takes the browser, visual, and lower-cost tasks:

The job in my workflow Route it to How I use it
Ad creative + visual variants ChatGPT I use it to turn a creative direction into ad variants, social tiles, and product-image edits. This is my visual production tool.
Coordinate a multi-step project Claude My coordinator: turn a brief into a plan, keep the context together, and carry the work through implementation and review.
Code + maintain my tools Claude My coder: work inside the repo, fix the issue, and check the result. My content pipeline, SEO workflows, and internal tools live here.
Browser work + web apps ChatGPT ChatGPT wins this part of my workflow. I use it to navigate sites, work across tabs, and complete tasks in web apps while Claude coordinates the broader project.
Live-web research + source checks ChatGPT Open the source, compare pages, and bring the findings back into the project. It is my first stop when the answer depends on what is on the web today.
Make sense of briefs + meeting notes Claude I give it the transcripts, briefs, and working documents together. It connects the decisions and turns the mess into a useful next step.
Data analysis + charts ChatGPT Turn a spreadsheet into a chart or a visual explanation I can use in a report or presentation.
Voice + quick brainstorming ChatGPT Talk through an idea while I am away from the keyboard, then bring the useful parts back into the work.
Routine work + cost control ChatGPT I use the lower-cost models for repeatable tasks that do not need my coordinator. The job has to be done well, but it does not always need the most expensive model.

The benchmarks: Claude leads, and the gaps are uneven

Every model below was run by the same third party, Artificial Analysis, on the same harness at max effort. That makes this the cleanest like-for-like set available. Scores run lower than the labs' own numbers because the harness differs. GDPval uses an Elo scale; its axis starts at 1400.

How to read these: scores shift with the test harness, the classic benchmarks are saturating (95% vs 96% means nothing), and every lab tunes for the tests it publicizes. Treat the pattern as signal, the decimals as noise.

Independent scores · Artificial Analysis
  • Opus 5.5
  • Sonnet 5.5
  • GPT-6 Astra
  • GPT-6.1 Sol

AA Intelligence Index

composite of 10 evals · 0–100

Opus 5.5 57.6
Sonnet 5.5 56.0
GPT-6 Astra 52.7
GPT-6.1 Sol 47.5

Terminal-Bench 4.0

agentic terminal tasks · %

Opus 5.5 59.6
Sonnet 5.5 63.6
GPT-6 Astra 59.1
GPT-6.1 Sol 43.9

Humanity's Last Exam

expert reasoning · %

Opus 5.5 61.4
Sonnet 5.5 55.0
GPT-6 Astra 54.7
GPT-6.1 Sol 47.9

Terminal-Bench-Science

scientific terminal tasks · %

Opus 5.5 59.0
Sonnet 5.5 53.3
GPT-6 Astra 63.3
GPT-6.1 Sol 30.0

AutomationBench

workflow automation · partial-credit %

Opus 5.5 69.5
Sonnet 5.5 71.3
GPT-6 Astra 68.5
GPT-6.1 Sol 61.6

SciCode

scientific coding · %

Opus 5.5 66.9
Sonnet 5.5 61.0
GPT-6 Astra 56.5
GPT-6.1 Sol 57.6

GDPval-AA v2.1

professional knowledge work · Elo · axis 1400–1900

Opus 5.5 1846
Sonnet 5.5 1844
GPT-6 Astra 1542
GPT-6.1 Sol 1487

Source: Artificial Analysis, pulled Sep 29, 2026. Every model on the same harness at max effort.

Anthropic's launch tables (vendor-reported)

Vendor-reported and chosen by Anthropic. Kept for the tests Artificial Analysis doesn't run (OSWorld, FrontierCode). Gaps mean Anthropic didn't list that model; they aren't filled with AA numbers because the two sources use different harnesses.

Terminal-Bench 4.0

agentic terminal tasks · %

Opus 5.5 66.4
Sonnet 5.5 70.6
GPT-6 Astra 57.9
GPT-6.1 Sol N/A

FrontierCode v1.1

coding · %

Opus 5.5 54.4
Sonnet 5.5 46.2
GPT-6 Astra 53.3
GPT-6.1 Sol 49.3

Humanity's Last Exam

expert reasoning · %

Opus 5.5 67.7
Sonnet 5.5 64.5
GPT-6 Astra 57.2
GPT-6.1 Sol N/A

OSWorld 2.1

computer use · %

Opus 5.5 81.8
Sonnet 5.5 80.1
GPT-6 Astra N/A
GPT-6.1 Sol N/A

AutomationBench

workflow automation · %

Opus 5.5 40.0
Sonnet 5.5 N/A
GPT-6 Astra 41.4
GPT-6.1 Sol N/A

Terminal-Bench-Science 0.1

scientific terminal tasks · %

Opus 5.5 58.7
Sonnet 5.5 N/A
GPT-6 Astra 64.6
GPT-6.1 Sol N/A

The pattern: Claude leads six of the seven independent rows, and Astra wins the one about scientific terminal tasks. GPT-6.1 Sol trails the pack on nearly every test. But the routing case doesn't rest on a leaderboard. Sol costs $1.05 per task against $5.98 for Opus 5.5, and images, voice, and live web are still ChatGPT jobs. Anyone telling you one of them is simply "smarter" hasn't read past the headline.

Safety, per each lab's own system card

Both labs publish a system card that says what the model can do and what safeguards ship with it. The labs use different frameworks: Anthropic's Responsible Scaling Policy (CB-1/CB-2, Autonomy levels, cyber tiers) and OpenAI's Preparedness Framework (High / Critical). The levels aren't directly comparable, so each is shown in its own terms.

Risk area Opus 5.5Sonnet 5.5GPT-6 AstraGPT-6.1 Sol
Bio / chem CB-1 Known weapons uplift only; not CB-2. Matches best prior models on RNA sequence design No new threshold Bio safeguards same as Sonnet 5 HIGH HIGH
Cyber Strongest Anthropic model Exceeds Mythos 5.1 internally. Anthropic says below Tier 2; Zvi Mowshowitz disputes that call Large jump Safeguards off: 80% capture on ExploitBench, 178 full code-execution exploits (Sonnet 5: 1) CRITICAL First model ever at Critical. Finds novel flaws and builds exploits unaided. Exploit creation gated behind the Daybreak program HIGH Below Critical
Autonomy / self-improvement Below Autonomy-2 METR estimates ~1.5× research speedup, ~30% chance of 2× Low misalignment risk Not reported at Critical Below High
Deployed safeguards Cyber + bio classifiers at Fable 5.1 level; preserved thinking (anti-distillation); EU AI Act watermarking First Sonnet with Opus-level cyber safeguards; classifiers block reasoning extraction; preserved thinking Misalignment monitor on all tool-using inference; trust-based access for high-risk capability; actor-level enforcement; under-18 boundaries Same safeguards as the GPT-5.6 card; alignment training shared with Astra

What the evaluations found

Opus 5.5

  • Best automated behavioral-audit scores of any Claude model
  • 85% fewer boundary-circumvention attempts than Opus 5
  • Reward hacks in 0.63% of training episodes
  • Eval awareness 36% in constructed scenarios vs 0.4% in real use
  • Regression: more likely to follow malicious instructions hidden in pasted text
  • Regression: more likely to accept unverified claims of authorization
  • Malicious computer-use success 4%, higher than Anthropic wanted

Sonnet 5.5

  • ~1,850-scenario audit: matches or beats Sonnet 5 on most alignment, misuse and honesty measures
  • Lowest attempted sandbox-escape rate of any tested model
  • Broadly less capable than Opus 5.5; crosses no new RSP thresholds
  • Cyber capability jumped sharply, so it inherits Opus-grade cyber safeguards

GPT-6 Astra

  • Indirect prompt injection robustness 99.79% (GPT-5.6 Sol: 96.23%)
  • Gray Swan IPI Arena attack success 8.5% (was 27.0%)
  • Honeypot exploitation 0% (was 55.4%)
  • High-severity misalignment flags ~53% lower; 4× less misrepresentation of coding work
  • Hallucination rate 92% → 51% at max effort (Artificial Analysis)
  • Recurrent-depth architecture makes internal reasoning much less legible
  • CoT controllability 16.1% → 60.9%; shows monitor evasion when sandbagging

GPT-6.1 Sol

  • Builds on Astra's alignment work; fewer misleading claims about its coding work than GPT-5.6 models
  • Safety record is a short appendix; few standalone numbers published
  • Gizmodo reports it may share Astra's recurrent-depth training. OpenAI hasn't confirmed this

What this means for a marketing team: Anthropic's card says Opus 5.5 is more likely to follow malicious instructions hidden in pasted text and to accept unverified claims of authorization. Until that improves, treat untrusted emails, logs, and web pages as hostile input before pasting them into any agent that has permissions.

Pricing: two ladders, one missing rung

The subscription ladders are near mirror images, which is exactly what makes the run-both play cheap.

Price Claude ChatGPT Read
$0 Free: Sonnet and Haiku, with chat, web search, and file creation. No Cowork or Claude Code. Free: unlimited everyday text chats. Uploads, deep research, and images are capped. Both free tiers are real now. ChatGPT gives you more chat; Claude gives you a stronger model lineup and files.
$8/mo No tier here Go: the cheap on-ramp ChatGPT owns the budget rung. Claude simply does not compete for it.
$20/mo Pro (~$17/mo on annual, $200 billed upfront): Opus, Cowork, Claude Code Plus: GPT-6 Astra inside Work and Codex, not in regular chat The rung that matters. Claude Pro includes Opus and Cowork; Plus gets Astra only inside Work and Codex. This is where routing lives.
$100/mo Max 5× Pro: adds GPT-6 Pro in regular chat Mirror-image tiers for heavy solo use.
$200/mo Max 20× Pro: the higher-allowance tier The power-user ceiling. ChatGPT added a $500 Pro tier on September 29, so its ladder now runs higher.

Building automations? The API ladders matter more than the app subscriptions. And the price per token is only half the story: what you actually pay is the cost of finishing a task.

Release card: specs and pricing

Opus 5.5AnthropicSonnet 5.5AnthropicGPT-6 AstraOpenAIGPT-6.1 SolOpenAI
Released Sep 22, 2026 Sep 28, 2026 Sep 3, 2026 (card) Sep 22, 2026
Role in lineup Flagship of the 5.5 family; upgrade on Opus 5 Mid-tier; second 5.5 model OpenAI's top model "Lower-cost alternative to Astra"; ships with GPT-6 Luna
Price / 1M tokens (in / out) $4 / $20 Fast mode $8 / $40. 20% cheaper per token, ~40% cheaper per task than Opus 5 $2 / $10 Same as Sonnet 5; up to 30% lower cost per task $10 / $50 Fast mode 2× price for 2.5× speed $2 / $10 Cached input $0.20. Luna $0.10 / $0.50
Context window 1M 1M 1M 872K
Output speed (Artificial Analysis) 92.5 tok/s 30%+ faster than Opus 5 137.6 tok/s Fastest of the four 59.1 tok/s 78.6 tok/s
Cost to run AA Intelligence Index $5.98 per task $7.60 per task Cheap tokens, but uses 410M of them $3.26 per task Uses the fewest tokens $1.05 per task
Where Claude API, Bedrock, Vertex, Foundry. Support guaranteed to Sep 22, 2027 Claude Platform, AWS, Google Cloud, Azure. ZDR available API, Bedrock; ChatGPT Plus/Pro/Business/Enterprise (admin opt-in) API as gpt-6-sol
Safety document Standalone system card, ~230 pages Standalone system card System card + safety overview on Deployment Safety Hub Appendix to the Astra card, added Sep 22

Speed and cost per task: Artificial Analysis, Sep 29, 2026. Pricing and specs: each lab's launch page and system card.

Notice what's missing: a $40 rung. Neither lab sells a "both jobs, one bill" tier. You get that combination by running two $20 plans. That's the routing play priced out.

Before you assume $40/mo is a stretch

You're OVERPAYING for AI (Fix This)

Why the real waste isn't running two subscriptions, it's routing tasks to the wrong one.

MarketingAlec Open on YouTube →

Usage limits: the dealbreaker nobody benchmarks

No leaderboard measures the moment your assistant says "come back in three hours" at 2pm on launch day. Caps, not capabilities, are why most people actually get frustrated with these tools.

The shape of it in September 2026: ChatGPT's free tier no longer runs out fast: everyday text chats are unlimited, though uploads, deep research and images are capped. Claude's free tier can still cap on heavy days. Both $20 tiers are comfortable for normal use and both will throttle a power user, which is exactly what the $100 and $200 multiplier tiers exist to fix.

Two usage pools: when one hits its limit, work keeps moving in the other

The routing bonus nobody mentions: two $20 plans means two separate usage pools. When Claude throttles mid-sprint, the visual and research work keeps moving in ChatGPT, and vice versa. Redundancy is a feature of the split-mind setup, not an accident.

This page will be stale in a quarter. The newsletter won't.

Model leads flip every few months. I re-test the routing map twice a week and send what changed. Two emails a week, zero fluff.

Feature matrix: how I use each tool

Features matter when they help finish a job. Here is what each tool offers and where it fits in my workflow:

Feature Claude ChatGPT How I use it
Image generation None native ChatGPT Images 2.5 I use ChatGPT for ad variants, social tiles, and product-image edits.
Video generation None None (Sora was shut down in 2026) Neither lab offers video generation now.
Voice mode Text-first Mature, fast, natural ChatGPT for talking through an idea away from my desk.
Coding agent Claude Code Codex Claude is my primary coder and coordinator for repo work.
Desktop agent Cowork: local files, M365 ChatGPT Work: browser and web apps Claude for work rooted in my files. ChatGPT for work happening in the browser.
Web browsing Web search + connected tools Built-in browsing + browser tools ChatGPT wins browser work in my setup: navigating sites, checking sources, and using web apps.
Connectors MCP + connectors Plugins + connectors I connect Claude to the tools my coordinator needs, and ChatGPT to the apps involved in the current browser task.
Context window 1M (Opus 5.5 and Sonnet 5.5) 1M (Astra), 872K (Sol) I keep the larger project context with Claude and give ChatGPT the material needed for each task.
Memory + projects Yes Yes Claude holds the ongoing project. ChatGPT handles the browser and visual work I route out of it.

The verdict

Marketers love brand-loyalty stories. Mac vs PC. Coke vs Pepsi. HubSpot vs Salesforce. We try to apply that same frame to AI.

AI is not a brand. It's a utility layer that increases your speed. And utilities are stacked, not chosen. The September 2026 evidence on this page says it plainly: a Claude lead on the independent benchmarks that still leaves ChatGPT ahead on cost per task, mirrored subscription pricing, and complementary features.

The marketers who'll dominate the next 12 months aren't the ones who pick the "right" model. They're the ones who build the right split-mind workflow. The kind that uses each model exactly where it's strongest, and gets re-tuned every month as the labs leapfrog each other. Each time a new model drops, the New Model Checklist is where that re-tuning starts.

The real risk isn't picking wrong. It's picking one.

Frequently asked questions

Should marketers use Claude or ChatGPT?
I use both. Claude is my coordinator and coder. ChatGPT 6 (Astra, Sol, Tera) handles browser work, visuals, and routine tasks where cost matters. Claude Pro and ChatGPT Plus are $20 each per month, so $40 gets you both entry-level paid plans. Higher usage may need a higher tier.
What is Claude better at than ChatGPT for marketing?
In my workflow, Claude coordinates the project and writes the code. I use it to connect briefs and meeting notes, plan the work, and build or maintain my content pipeline, SEO workflows, and internal tools. It is where I keep the context for a project that runs across many steps.
What is ChatGPT better at than Claude for marketing?
I use ChatGPT for browser work, visual production, and cost control. That includes navigating sites and web apps, checking live sources, creating ad variants and product-image edits, and turning data into charts. For repeatable tasks, I choose a lower-cost model when it can finish the job well.
Is Claude free? How do the free tiers compare?
Both have real free tiers, but they give you different things. ChatGPT's free tier has unlimited everyday text chats, with caps on uploads, deep research, and images. Claude's free tier covers chat, web search, and file creation with Sonnet and Haiku, but not Cowork or Claude Code, which start at Pro. ChatGPT also offers an $8/month Go plan that Claude has no answer to.
Which is cheaper, Claude or ChatGPT?
The subscription ladders mirror each other: $0, $20, $100, and $200 tiers on both sides, with ChatGPT adding an $8 Go rung Claude lacks. On the API, per million tokens (in / out), Claude Opus 5.5 is $4 / $20 against $10 / $50 for GPT-6 Astra, while Claude Sonnet 5.5 and GPT-6.1 Sol are both $2 / $10. Per task the picture changes: on the Artificial Analysis index Sol costs $1.05 per task, Astra $3.26, Opus 5.5 $5.98 and Sonnet 5.5 $7.60, because Sonnet uses far more tokens.
What's the difference between Claude Cowork and Claude Code?
Same engine, two doors. Cowork runs from the desktop app with clickable plugins and connectors, which makes it the right entry point for most marketers. Claude Code is the terminal-grade workshop: custom skills, sub-agents, scheduled jobs, and full audit logs for regulated teams. Build in Code, ship to the team in Cowork.

Who's behind this?Alec Newcomb, Founder of MarketingAlec and ScaledOn, re-tests this routing map twice a week.

Meet Alec →