Stefan Maritz··14 min read

The ultimate guide to what AI models to use for what

No model wins everything any more. Here is what to use for code, writing, images, video, audio, and high-volume work, with real pricing and the expiry dates attached to it.

TL;DR

  • No model wins everything any more, and picking one for all your work is the expensive mistake.

  • Code: Claude Fable 5.1 sits at the frontier, Opus 5 gives the best quality per pound, Sonnet 5 is the sensible default.

  • Writing: Fable 5.1 for prose, Sonnet 5 for volume, Gemini 3.1 Pro when the draft needs live research inside it.

  • Images: GPT Image for prompt accuracy, Nano Banana for free and fast, Midjourney for art direction, FLUX.2 to self-host, Recraft for real vectors.

  • Video: Veo 3.1 for native audio, Kling 3.0 Turbo for value, Runway Gen-4.5 for control. Sora is being switched off, so do not start there.

  • High volume: DeepSeek V4 Flash at $0.14 per million input tokens makes bulk work close to free.

  • Prices in here are correct for September 2026 and several expire on fixed dates, so check before you budget.

  • Model choice decides quality of output. It decides nothing about whether that output sounds like your company.

How to read this

Three things have changed about choosing an AI model, and they all point the same way.

The frontier has split into lanes. There is no longer one best model, there are specialists, and the price difference between the right one and the wrong one for a given job runs to 50 times. The pricing has also fragmented: the same model can cost half or double its list price depending on whether you use batch processing, fast mode, or a long context window. And the release cadence has gone weekly, so any guide including this one starts decaying the day it publishes.

Everything below was checked against provider pricing pages and independent leaderboards in the week of 15 September 2026. Where boards disagree, and they do, that is noted rather than smoothed over. Where a price has an expiry date attached, that date is given.

The text model table

Text generation is where most spend happens, so start here. Prices are US dollars per million tokens on the standard tier, input first, output second.

ModelInputOutputContextBest at
Claude Fable 5.1$10$501MFrontier coding, prose quality
GPT-6 Astra$10$501MLong-horizon agentic work
Claude Opus 5$5$251MQuality per pound at the frontier
GPT-5.6 Sol$4$201.05MAll-round workflow, tool use
Claude Sonnet 5$2$101MEveryday default, natural long-form
Gemini 3.1 Pro$2$122MManuscript-length context, research
Grok 4.6$2$6-Cheap frontier-adjacent reasoning
Qwen 3.8 Max$2$61MCoding value, native video input
GPT-5.6 Terra$2$121.05MBalanced mid-tier
Kimi K3$3$151MBest open-weight all-rounder
Claude Haiku 4.5$1$5-Classification, routing, volume
Gemini 3.8 Flash$0.75$3.751MVolume with a real context window
DeepSeek V4 Pro$0.435$0.871MCheap capable work
GPT-5.6 Luna$0.20$1.201.05MBulk work in the OpenAI stack
DeepSeek V4 Flash$0.14$0.281MThe price floor

Four footnotes that will cost you money if you skip them. Gemini 3.8 Flash is on introductory pricing until 31 December 2026, then doubles to $1.50 and $7.50. GPT-5.6 Sol's rate is promotional and guaranteed only through 21 November 2026. GPT-6 Astra reprices the whole request to $20 and $75 once a prompt passes 272,000 input tokens, and Gemini 3.1 Pro steps up to $4 and $18 past 200,000. Claude Sonnet 5 was scheduled to rise to $3 and $15 on 1 September, and as of mid-September it is still listed at $2 and $10, which is the sort of detail worth confirming on Anthropic's own pricing page rather than from a blog.

There is two numbers that decide your bill, and the second one is the one people forget. Output tokens cost five times input on most frontier models, so a model that reasons at length before answering is more expensive than its headline rate suggests. Batch processing halves both on every major provider, and cached input runs at roughly 10% of standard, dropping to $0.25 per million on Fable 5.1.

Five for code

Claude Fable 5.1 is the raw peak. It holds the highest WebDev Arena score of any model in any tool at 1762 Elo, 73 points ahead of the Fable 5 it replaced, and one September ranking puts it near 95% on SWE-bench Verified. The 5.1 release cut cache reads by 75% to $0.25 per million, which lowers real agentic cost by roughly a quarter on typical work. Use it when a failed agent loop costs more than the tokens.

Claude Opus 5 is the value answer at the frontier. At 1688 Elo and $5 and $25 it does most of what Fable does for half the price, which is why it holds second place on the dev-tool rankings rather than slipping out of them.

GPT-6 Astra arrived on 3 September 2026 built for long-horizon agentic work: computer use, multi-step tasks, research runs that previous models abandoned halfway. It matches Fable 5.1 exactly at $10 and $50. Treat the launch benchmarks with care, since the headline 99.9% on ARC-AGI-3 came from a provider-specific harness that preserved reasoning state between actions, while the neutral harness produced 62.7% at much higher compute cost.

Qwen 3.8 Max is the surprise. It entered the top three at 1686 Elo, matching Grok 4.6 at $2 and $6 while offering a 1M context window and native video input. For teams where cost per task decides the stack, it changes the maths.

Claude Sonnet 5 is what most people should run most days. At $2 and $10 it handles the bulk of production coding, and you escalate to Opus 5 or Fable 5.1 only for the hard parts. The routing pattern beats picking one model, every time.

Five for writing

Claude Fable 5.1 is the best raw writer available. Writers rank it ahead of GPT-5.6 Sol and the Opus line for pure prose, and it tops the creative-writing boards for tone, character voice, and subtext. The honest caveat is that its output runs dense, so budget an editing pass.

Claude Sonnet 5 is the value default just behind it. It needs little cleanup, sounds human, and holds tone across a full article at a fifth of Fable's cost. For a content programme publishing weekly, this is the one doing the work.

Claude Opus 5 takes the top spot on the EQ-Bench creative writing leaderboard in the most recent aggregate ranking, ahead of Kimi K3 and GPT-5.6 Sol. That it disagrees with the boards putting Fable first tells you something useful: prose quality has no clean benchmark, and the difference between the top three is smaller than your brief.

GPT-5.6 Sol is the stronger all-round writing workflow rather than the stronger writer. Precise formatting, reliable structure, good tool use, which is what you want for structured marketing copy with a strict template.

Gemini 3.1 Pro earns its place on context. A 2M token window means an entire manuscript, style guide, and research folder in one pass, and Google Search grounding puts live facts into the draft. Slower and pricier than Sonnet for plain drafting, worth it when the draft has to be current.

One thing no leaderboard measures: voice consistency across 2,000 words. Plenty of models write one excellent paragraph and then drift into their house register. Test on your own material before you commit, using the same brief across three models, and read the second half of each output rather than the first.

Five for images

GPT Image holds the top of the blind-vote arenas. The line grew a reasoning step in April 2026, planning layout and self-checking output before it draws, which is why prompt adherence improved so sharply. The successor powering ChatGPT image generation landed on 8 September 2026.

Nano Banana 2 and Nano Banana Pro are Google's answer and the best free option going. Nano Banana 2 is free in the Gemini app, fast, edits by conversation, and now outscores its Pro sibling on the blind-vote board. Pro is the one for 4K output, identity lock across images, and multi-image blends.

Midjourney V8 keeps the most distinctive default aesthetic in the market, and it remains useless for automation because there is no official public API, only third-party resellers. Pick it for hero art, not for pipelines.

FLUX.2 is where self-hosting went. Open weights, multi-reference conditioning, a distilled variant that runs on consumer hardware in well under a second, and hosted tiers when you would rather not own the GPUs.

Recraft V4 does the one thing frontier raster models still cannot: true editable vector output. If you need an SVG rather than a picture of one, this is the whole list.

ModelTypical API costBest at
Nano Banana Pro~$0.15 per image4K, identity consistency, editing
Nano Banana 2~$0.06 per image, free in GeminiVolume, speed, chat editing
FLUX.2 [pro]~$0.03 per megapixelPhotorealism, self-hosting
Seedream~$0.04 per imageReference-based generation
Recraft V4~$0.04, $0.08 for vectorSVG, brand assets, text
Midjourney V8$10 to $120 per monthArt direction, no API

Per-image prices vary by host, so treat these as the shape of the market rather than a quote. Text rendering, which broke every model two years ago, is largely solved at the frontier now, and Ideogram's old monopoly on legible typography has gone.

Five for video

Start with a warning. OpenAI closed the Sora consumer app on 26 April 2026 and the API sunsets on 24 September 2026, with no direct replacement named in the deprecation table. Anything you build on it this month stops working next week.

Veo 3.1 is the safest Western pick and the only model in its tier shipping native audio in the output. That removes an entire text-to-speech step from an explainer or ad pipeline, which matters more to a production budget than a few cents per second.

Kling 3.0 Turbo is the value leader at roughly $0.11 to $0.14 per second, with unusually strong multi-angle subject consistency. For generating 20 variations of a social clip, nothing else is close on cost per usable output.

Seedance 2.5 is what creators call the overall king, with native 30-second 4K and support for up to 50 reference images. Access is still patchy, which is the only reason it sits behind the more available models.

Runway Gen-4.5 remains the professional workflow choice. Video-to-video, image-to-video, motion controls, and the editing environment around them. You pay for control rather than for leaderboard position.

Wan 3.0 and LTX-2.3 are the open options. Apache-licensed, self-hostable, with LTX-2.3's fast tier among the cheapest listed rates anywhere. If you have GPU budget, the per-second cost goes to zero.

ModelRoughly per secondNative audioBest at
Veo 3.1 (fast)~$0.15YesAudio-video sync, 4K
Veo 3.1 Lite$0.03 at 720p, $0.05 at 1080pNoDrafting cheaply
Kling 3.0 Turbo$0.11 to $0.14Yes on base modelValue, character consistency
Runway Gen-4.5~$0.15 to $0.20NoCamera control, editing
LTX-2.3 (fast tier)~$0.06NoOpen source, native 4K
Grok Imagine Video 1.5~$0.05Voice cloningAnything living inside X

Video pricing is the least stable category here. The same model can cost double through one gateway and half through another, and resolution, duration, and audio each move the meter. Price a real clip before you plan a campaign around a number.

Five for voice, music, and transcription

ElevenLabs is still the reference for voice cloning and text-to-speech, across roughly 30 languages, and is now facing credible competition for the first time from OpenAI's voice models, Cartesia, and PlayHT. Credits are shared across their whole platform, so heavy music use eats your voice quota.

Suno leads consumer music generation, with vocals that pass casual listening and a studio tier offering stem separation and MIDI export into a DAW. It struggles on rap and spoken word, where rhythmic precision gives it away.

Udio earns its place on inpainting, which lets you regenerate one section of a track rather than rerolling the whole thing. That single feature saves more time than a quality difference would.

ElevenLabs Music is the commercially cautious option, built on licensed training data, which is a real consideration for anyone putting generated music behind a client's advert. Vocal realism is its strength, editing tools are its weakness.

Whisper remains the practical default for transcription: open, accurate across a hundred-odd languages, and free to run yourself. For a content team turning calls and interviews into source material, this is the model quietly doing most of the work.

Audio is the category where my figures are thinnest, since the public comparisons lag the other categories by a few months. Check current plan pricing directly.

Five for volume and self-hosting

DeepSeek V4 Flash is the price floor at $0.14 and $0.28 with a 1M context window. At that rate, classifying 100,000 support tickets costs less than lunch.

DeepSeek V4 Pro is the cheapest genuinely high-quality option at $0.435 and $0.87, and its cache-hit input rate of $0.003625 per million makes repeated-context agent loops almost free.

Kimi K3 leads the open-weight boards on overall, reasoning, and coding, with 1M context and native multimodality. At $3 and $15 hosted it is the premium end of open, and its 2.8 trillion parameters make self-hosting a rack-scale exercise rather than a weekend one.

GLM-5.3 replaced GLM-5.2 in August 2026 as the coding-focused open checkpoint. List price sits at $1.40 and $4.40, though the provider median runs closer to $0.55 and $1.85 because third parties host the weights.

Qwen 3.8 27B is the practical local default: Apache 2.0, multimodal, and small enough to run on one workstation, which is the constraint that decides most self-hosting projects.

One caution on open weights. The same model name buys a different product depending on the host. Providers serve different quantizations and often less context than the model advertises, with some capping output at a fraction of the documented maximum. Ask what precision and context you are actually getting.

How to route in practice

The winning pattern is three tiers, not one model.

Put a cheap model at the bottom for classification, routing, extraction, and anything high-volume: Haiku 4.5, Gemini 3.8 Flash, or DeepSeek V4 Flash. Run a mid-tier model as the default for real work: Sonnet 5, GPT-5.6 Terra, or Gemini 3.8 Flash again if the context window is doing the heavy lifting. Reserve the frontier for tasks where being wrong is expensive: Fable 5.1, Opus 5, GPT-6 Astra.

Then turn on the two discounts everyone forgets. Batch processing halves the bill for anything that does not need an answer this second, which covers most content work. Prompt caching cuts repeated input to a tenth or less, which covers any workflow sending the same brand guidelines or codebase on every call.

For a content team specifically, the split usually lands like this: drafting on Sonnet 5, final prose passes on Fable 5.1 when quality is worth the tokens, research and long-document work on Gemini 3.1 Pro, images on Nano Banana with Recraft for anything vector, and video on Kling for volume with Veo when the clip needs audio.

What will be wrong about this by Christmas

Most of the prices. Gemini's introductory Flash rates double on 1 January 2027, GPT-5.6 Sol's promotional rate expires on 21 November 2026, and Sora's API disappears this month. Frontier releases have been arriving within days of each other, with Anthropic shipping Fable 5.1 and OpenAI answering with GPT-6 Astra inside a week.

What will not change is the shape. Cheap models will keep getting good enough for more of the work, frontier models will keep being worth it for less of it, and the teams that win will be the ones who route rather than the ones who pick.

The other thing that stays constant is the part no model solves. Every model here can write a competent paragraph about your product. None of them know your positioning, your customers' objections, or the phrase your founder refuses to use. That context has to live in a knowledge base the model reads every time, with a review step between the output and the internet, which is the difference between a content programme and a slop machine with a good model behind it. Where that output lands matters too, which is a CMS question rather than a model one.

Frequently asked questions

Which AI model is best overall in 2026?

There is no longer a single answer, which is the defining feature of this year. Claude Fable 5.1 leads coding and prose, GPT-6 Astra leads long-horizon agentic work, Gemini 3.1 Pro leads on context length, and DeepSeek leads on price. Pick per task, not per company.

What is the cheapest capable AI model?

DeepSeek V4 Flash at $0.14 input and $0.28 output per million tokens, with a 1M context window. GPT-5.6 Luna at $0.20 and $1.20 is the cheapest current-generation option inside the OpenAI stack, and Gemini 3.8 Flash at $0.75 and $3.75 sits just above while that introductory rate lasts.

Which AI model is best for writing blog posts?

Claude Sonnet 5 for most of it, at $2 and $10 per million tokens with prose that needs little cleanup. Move to Claude Fable 5.1 for pieces where voice carries the work, and to Gemini 3.1 Pro when the draft needs current research pulled in during writing.

Is Sora still available?

No. OpenAI closed the Sora consumer app on 26 April 2026 and the API sunsets on 24 September 2026, with no direct replacement listed. Migrate to Veo, Kling, Seedance, or Runway.

What does AI video generation cost?

Between roughly $0.03 and $0.75 per second of output depending on model, resolution, and whether audio is generated. A ten-second clip runs from about 50 cents on the cheap end to several pounds on the premium end, before the failed generations you do not use.

Are open-weight models good enough to replace frontier models?

For a growing share of work, yes. Kimi K3 leads the open boards and DeepSeek V4 undercuts everything on price. The frontier still pulls ahead on the hardest agentic and reasoning tasks, and the practical constraint on self-hosting is hardware rather than quality.

How do I cut my AI bill without dropping quality?

Three levers, in order of size. Route cheap work to cheap models instead of defaulting everything to the frontier. Turn on batch processing for anything not needed instantly, which halves the rate. Cache repeated input such as system prompts and reference documents, which drops that portion to roughly a tenth.

How often does this change?

Monthly, and sometimes weekly. Several prices in this guide have published expiry dates in November and December 2026. Treat any model comparison older than a quarter as history rather than guidance, and verify rates on the provider's own pricing page before you commit a budget.

Content like this, in your own voice

The workflows behind this blog run inside Contengi. A strategist sets up your knowledge base, and the agents write to your rules.

Request beta access