Cited

Paying the API directly versus paying for the app

For several common AI tasks we priced the subscription against the raw API bill. The crossover point is lower than most people assume.

Close-up of vintage kilowatt, volt, and ampere gauges in Essen's industrial setting.
Photo: Wolfgang Weiser / Pexels

Part of What AI tools actually cost, in time as well as money

Most AI subscriptions are a thin interface wrapped around a model call, sold at a markup. That's not a complaint — building the interface is real work, and the markup pays for it — but it means the honest question is never "is $20 a month worth it," it's "what is the $20 buying, on top of what the model itself costs." We ran five tasks we actually do each month through both routes — the subscription, and the same volume priced at the vendor's own published API rate — to find where that line sits. In most cases it sits lower than the marketing suggests. In a couple of cases the subscription wins outright and it isn't close.

Ordinary chat and writing: the subscription loses more often than people expect

Start with the least controversial case: a working professional who uses an AI chat tool daily for drafting, editing and thinking out loud — thirty-odd exchanges a day, most with the growing context of a conversation attached rather than a single clean prompt. Over twenty-two working days, call it 2 million input tokens and 500,000 output tokens once every reply re-sends the conversation so far.

ChatGPT Plus is a flat $20 a month. Priced at GPT-5.2's published API rate — $1.75 per million input tokens, $14 per million output — that same volume of conversation comes to roughly $3.50 for the input and $7 for the output, about $10.50 total. That is genuinely less than half the subscription price, for a use pattern most people would call heavy. Prices checked August 2026.

The gap only closes at the low end. A light user having five short exchanges a day is spending closer to a dollar a month in raw tokens either way, and at that volume the question of which is cheaper stops mattering — nobody manages an API key and writes their own chat interface to save a dollar.

What the subscription is actually selling at this volume isn't the tokens. It's voice mode, image generation, a place your conversation history lives without you building storage for it, and file uploads that don't require your own document-parsing code. Whether that bundle is worth the roughly $10 premium over raw compute depends on whether you'd use any of it — and most people who chat with an AI daily do use at least one, which is the honest reason the subscription stays the sane default for almost everyone reading this.

Coding assistants: this is where the wrapper earns its keep

GitHub Copilot Pro is $10 a month and includes unlimited inline completions plus $15 of monthly credit toward chat and agent-mode requests on premium models. Compare that to running an agentic coding tool wired straight to a model API — the heaviest realistic AI task most people run: constant tool calls, files re-read into context on every turn, edits proposed and revised in a loop. A developer using an agent for a meaningful part of their month can plausibly push 5 million input tokens and 1 million output tokens, because agentic coding resends far more context per exchange than a chat window does.

At Claude Sonnet 5's current introductory rate — $2 per million input tokens, $10 per million output, in effect through the end of August 2026 — that volume prices out to $10 for the input and $10 for the output, $20 total. Compare that to Copilot Pro's flat $10 plus $15 of included credit: for a developer whose usage stays inside that allowance, the subscription is cheaper and includes something the raw API doesn't hand you for free — multi-file awareness, the editor integration, a permission model scoped through your GitHub account rather than a bare API key in an environment variable. Once usage runs past the included credit, GitHub's own next tier is Pro+ at $39 a month for four times the allowance, which is roughly the point where paying the API directly and building your own tool integration starts to look competitive again, provided you're willing to maintain it. Prices checked August 2026.

This is the cleanest case in the whole comparison for what a wrapper genuinely buys: not the tokens, but the file-aware context and the editor plumbing, which is expensive to rebuild and cheap to rent.

Support tickets: the API is startlingly cheap, and that's not the whole cost

Anthropic publishes its own worked example for this one, which is unusually candid for a vendor: roughly 3,700 tokens per conversation on Claude Haiku 4.5, at $1 per million input tokens and $5 per million output, works out to about $37 to answer 10,000 support tickets. That's close to free per ticket, and it's real. Prices checked August 2026.

It's also not the comparison most small teams are actually making. The realistic alternative isn't a $37-per-10,000-tickets pipeline versus nothing — it's a person manually pasting each incoming ticket into ChatGPT Plus or Claude Pro at $20 a month, because building a pipeline is a project and copy-pasting is not. The subscription route is bounded by how many tickets one person can paste in a shift, which doesn't scale, but it also doesn't need building. The API route scales indefinitely and needs someone to write the code that pulls a ticket in, calls the model, handles the reply, and decides what happens when the call fails or the answer is wrong. None of that plumbing shows up in the per-token price. It's the same lesson we ran into building our own automations, documented properly in self-hosting n8n for a fortnight: the software cost drops close to zero and the maintenance hours are where the real bill moves to.

The honest rule for this one: the API is the right choice once ticket volume is high enough that pasting them by hand has become someone's whole job anyway, because at that point you were going to spend engineering time on the problem regardless.

Narration: this is where a subscription buys something you cannot get any other way

ElevenLabs Creator is $22 a month (discounted to $11 for the first month) and includes 121,000 credits, roughly 121,000 characters of standard speech. A raw alternative exists: OpenAI's text-to-speech API, at $15 per million characters on the standard tts-1 model, no subscription required.

For a small monthly narration job — a weekly video script, four scripts a month at around 7,000 characters each, 28,000 total — the raw API comes to about 42 cents. ElevenLabs' own Free tier, at 10,000 credits, doesn't quite cover that volume; the cheapest paid step up is Starter at $6 a month. Priced by the character, the gap looks enormous, and if the only thing you need is a voice reading your script, it is. Prices checked August 2026.

But that comparison only holds if you don't care which voice you get. ElevenLabs' actual product, the thing the subscription pays for, is a specific voice — a clone of your own voice or a chosen character voice, with commercial usage rights attached — not "some voice that reads text aloud." Raw TTS APIs give you a library of stock voices and nothing resembling cloning. If the whole point of the narration is that it sounds like a particular person, the wrapper isn't a markup on the same product, it's a different product, and $6 to $22 a month is a reasonable price for owning a voice.

Bulk summarisation: this is where per-token pricing wins by an order of magnitude

The last task is deliberately the most lopsided, because it's the clearest illustration of what "the wrapper is a margin on the model" actually means at scale. A small team needs to summarise 500 internal documents a month — reports, call notes, long email threads — each roughly 4,000 tokens of input, producing a 400-token summary.

Run through a per-seat AI subscription like ChatGPT Business at $25 per user per month, five people doing this work costs $125 a month regardless of how much summarising anyone does, because seat pricing doesn't care about volume. Run through the API directly on a lightweight model built for exactly this kind of task — GPT-5-mini, at $0.25 per million input tokens and $2 per million output — the same 500 documents cost about 50 cents in input tokens and 40 cents in output, call it 90 cents for the whole team's month. Prices checked August 2026.

Ninety cents against $125 isn't a rounding difference, it's two pricing models meeting a task that suits one of them perfectly: high-volume, low-complexity, no chat interface needed because nobody is having a conversation, just documents in and summaries out. A script calling a cheap model directly costs less in a month than one seat of a general-purpose subscription costs in four days.

The variable that decides it, and how to estimate it before you commit

Every example above turns on the same number: how many tokens does your actual monthly volume represent. The estimate is rougher than it sounds necessary, and rough is fine, because you're choosing between two pricing models that differ by a wide margin, not optimising the last ten percent.

Take your monthly word count and divide by roughly 0.75 to get tokens, or your character count and divide by 4 — either gets you close enough. Split it into input tokens (what you send) and output tokens (what comes back), because output is priced at four to six times the input rate on almost every current model, and a task that mostly generates text is more expensive than the input-heavy tasks above suggest. Multiply by the per-million-token rate of the model you'd actually use, and compare the total to the subscription's flat fee. If they're within a few dollars of each other, stop — you're not deciding on price, you're deciding on which bundle of extra features you'd actually use, the same conclusion the chat-and-writing example reached. How vendors turn this into their own credit systems, and where that makes budgeting harder than it needs to be, is covered in how credit-based pricing actually works; the wider worksheet this single line sits inside — subscription, usage, learning, verification, exit — is laid out in what AI tools actually cost.

A decision rule that fits in one paragraph

Go direct to the API when the task is high-volume and low-drama: bulk summarisation, batch classification, anything running on a schedule rather than in a conversation. Stay on the subscription when volume is moderate, when a non-technical colleague needs the tool without an API key, or when what you actually want — a file-aware coding agent, a cloned voice, a memory that persists across sessions — doesn't exist as a raw model call at any price, and you'd be rebuilding the wrapper yourself for free. Treat technical appetite as its own line item, separate from cost: the API route is usually cheaper in dollars and never cheaper in time, because someone now owns the calls that fail and the outage that lands during your busiest hour.

Questions people ask

Is it cheaper to use the OpenAI or Anthropic API directly instead of paying for ChatGPT Plus or Claude Pro?
For a single person doing ordinary text chat, often yes — a heavy daily user can land near ten dollars a month in raw token cost against a twenty-dollar subscription. The subscription still wins the moment you want voice mode, image generation or a place your history lives without you building one.
How do I estimate how many tokens my AI usage actually is?
Divide your monthly word count by roughly 0.75 to get tokens, or your character count by 4. Multiply by the per-million-token rate of the model you'd use directly, remembering that output tokens are billed at four to six times the input rate on most models.
When does a subscription clearly beat paying per token?
When the volume is low enough that the subscription's flat fee is close to what a handful of API calls would cost anyway, or when the product bundles something you cannot buy separately — a coding assistant's file-aware context, a cloned voice, a place non-technical colleagues can use without an API key.
What's the biggest hidden cost of switching from a subscription to a raw API?
You inherit the parts the subscription was quietly doing for you — retries when a call fails, storing conversation history, handling the day the model provider has an outage. None of that shows up in the per-token price, and all of it is now your job.

Cited — We use a tool for a fortnight before we write a word about it, and we say where every number came from.