The right large language model in August 2026 depends on data sensitivity, ecosystem, volume, and task. Claude and GPT-5.6 lead for coding and general work, Gemini offers the best Workspace fit, and DeepSeek's MIT-licensed V4 replaced Llama as the standard open-weight, self-hosted option after Meta stopped updating Llama.
Summary: Five months after this guide first ran, almost every price in it changed. OpenAI split its Pro plan into two tiers, Claude Sonnet 5's discount expires September 1, and Meta walked away from open-weight models entirely. DeepSeek's MIT-licensed V4 is now the practical self-hosted option, not Llama.
Meta walked away from open-weight frontier models on April 8, 2026, and most guides to choosing an LLM still point readers toward Llama as the default self-hosted pick.
Five months ago this guide named three frontier leaders and one safe open-weight fallback. That fallback is gone. OpenAI has shipped two new model families since, Anthropic replaced its entire consumer lineup, and a Chinese lab's freshly open-sourced model has become the more credible self-hosted pick.
For a four-person startup team revisiting the subscription decision it made back then, almost none of the specifics below match what this guide said in March, even though the underlying logic for picking a model has barely moved.
Three changes matter more than the rest. First, OpenAI split its Pro tier in two: a $100-a-month plan launched April 9, 2026, sitting below the existing $200 plan, both running the same underlying reasoning model at different usage ceilings, according to reporting from TechCrunch.
Second, Claude Sonnet 5 is on an introductory rate that expires September 1, 2026, when its price rises from $2/$10 to $3/$15 per million tokens, per Anthropic's own pricing documentation. Third, and the one no competing guide has caught up to yet: Meta's new Superintelligence Labs division launched a closed-weight model called Muse Spark on April 8, rather than a new Llama release, according to Meta's own announcement.
That third change breaks a load-bearing assumption. Anyone who read the March version of this guide and picked Llama 4 for a regulated workload is now running a model line Meta itself has stopped extending.
Does the work have to stay on infrastructure you control? If yes, your list shrinks to open-weight models you can self-host, which today means DeepSeek rather than Llama.
Are you already paying for a Google or Microsoft seat? Workspace shops get Gemini embedded for free at the Business tier; Microsoft shops get Microsoft Copilot the same way. Neither integration is worth switching platforms to chase on its own.
What is your monthly token volume? Under a few million tokens a month, the sticker price barely matters. Past that, DeepSeek's $0.14 input rate versus GPT-5.6 Sol's $5 rate is a real line item, not a rounding error.
Is the task coding, writing, or research? Coding and structured output still favor Claude. Research that needs citations still favors Perplexity. Broad drafting and agent orchestration still favor ChatGPT.
Answer those four honestly before opening a pricing page. Most of the mistakes teams make with this decision happen because they compare benchmark scores first and their own constraints last, which is backward. A model that scores two points higher on a leaderboard is worthless if it cannot legally touch your data.
OpenAI. Best for teams that want the broadest agent tooling and do not want to think hard about which model to call. The lineup now runs three deep: a $5/$30 flagship tier, a $2/$12 mid tier, and a $0.20/$1.20 budget tier for high-volume work, per the company's own developer pricing page.
On the consumer side, ChatGPT Plus is still $20 a month, but a new $100 mid tier now sits below the existing $200 plan. The caveat: three flagship pricing tiers in one family is one more decision than most four-person teams want to make before lunch.
Anthropic. Best for developers and knowledge workers who need output that follows instructions without drifting. The whole consumer lineup turned over. The flagship model holds the same $5/$25 list price its predecessor charged, and the mid-tier model is priced at an introductory $2/$10 through August, stepping up after.
Claude Pro costs $17 a month on the annual plan or $20 billed monthly, and the Max plan starts at $100 for five times the usage. Anthropic's own product announcement leads with two proprietary benchmark suites rather than a single external coding leaderboard, a shift in how the company pitches itself. The caveat: the price you budget for this month is not the price you pay in October.
Google. Best for any team already paying for Workspace seats. Consumer plans were restructured this spring: the renamed mid tier costs $19.99 a month, and the top tier split into $99.99 and $199.99 options, down from a single $250 plan, per the company's own announcement.
On the API side, the flagship model still lists at $2/$12 per million tokens under 200,000 tokens, and a newer, cheaper option runs $1.50/$7.50. The caveat: developer sentiment still rates this lab's prompting as needing more explicit instructions than its two closest rivals to hit equivalent output quality.
DeepSeek. Best for high-volume workloads and anyone who needs to self-host without paying last year's infrastructure premiums. Its previous generation was retired on April 22, 2026, replaced by a new release as two checkpoints, both MIT-licensed and published directly to Hugging Face.
The budget checkpoint prices at $0.14 per million input tokens on a cache miss and $0.28 output; the larger one runs $0.435/$0.87, per the vendor's official documentation. That is roughly a third of what the prior generation charged for the same tier five months ago. The caveat: the English-language support ecosystem is thinner than the two big American labs, and the sourcing question some procurement teams ask has not gone away.
Meta. Best for almost nobody right now, and that is the point of this entry. This is the section that no longer holds. No new open-weight flagship shipped in 2026; the newest research effort launched closed instead, reachable only through a paid API.
The prior generation, dated April 2025, is still downloadable and still powers Meta AI, the free assistant inside Facebook, Instagram, and WhatsApp, but the practical open-weight successor for a self-hosted build today comes from the lab above, not a new release here. The caveat: if this changes before year end, this entry flips back.
Perplexity. Best for research, fact-checking, and any task where a citation matters more than a clever answer. Pricing has not moved: Pro is still $20 a month and Max is still $200, with a $325-per-seat enterprise tier, according to a July 2026 review of the company's plans by eesel AI.
What changed is context. The premium tier's headline extra five months ago was bundled video generation from a partner model whose consumer rollout has since been pulled back. The caveat: Perplexity is a research layer, not a model, and cannot be fine-tuned or self-hosted.
Most of the traffic to a guide like this is really one question: Claude or GPT, for a small team, right now.
For a two-person team shipping an AI feature into a product and needing the widest agent tooling and third-party library support, GPT-5.6 Terra is the more defensible default. It undercuts Sol on price, its function calling is mature, and OpenAI's ecosystem is still the largest of any provider.
For a four-person engineering team that lives inside an AI coding tool such as Cursor running Claude models, Claude Sonnet 5 at its expiring $2/$10 rate is difficult to beat before September 1. Anthropic's own benchmark framing leans on tasks like Zapier's AutomationBench and OSWorld 2.0 rather than a single leaderboard number, which makes a like-for-like comparison harder than it was in March, when Opus 4.6's SWE-bench score was the headline figure everyone quoted.
Neither is a clean sweep. If your team already routes agent tool calls through Groq for latency-sensitive inference, that infrastructure choice will often matter more than which frontier model sits behind it.
If your team has three chat subscriptions active at once and cannot name a specific task each one covers that the others do not, that is not redundancy for safety, it is a subscription problem. Cancel two, keep the one your team actually opens daily, and revisit in three months.
If your primary need is sourced, current research rather than generation, Perplexity alone, paired with NotebookLM for document-grounded work, covers most of that job without a frontier chat subscription at all.
And if your team is under three people and nobody is building an AI-native product, the honest answer is that a single $20-a-month Plus or Pro subscription, used consistently, beats a carefully optimized four-model stack that gets opened twice a week. Optimization only pays off once usage is high enough for the per-token gap to show up on an invoice.
The solo builder or two-person startup usually needs one primary subscription and one gap-filler, not four. A workable pair right now: Claude Pro at $20 a month for coding and written output, plus Perplexity Pro at $20 a month for anything that needs a current, sourced answer. That is $40 a month total, roughly what one seat of ChatGPT Team costs alone, and it covers the two jobs a small team actually does every day.
A four-person engineering team with real API spend looks different. Route routine, high-volume calls through DeepSeek-V4-Flash at $0.14/$0.28 per million tokens, and reserve Claude Sonnet 5 or GPT-5.6 Terra for the requests where instruction-following or agent tool use actually matters. At 50 million tokens a month, the difference between routing everything through a $5-input flagship and splitting the load is thousands of dollars, not a rounding error on an invoice.
Neither stack should include four subscriptions running in parallel. If you cannot name the specific job each tool does that the others cannot, you are paying for redundancy, not coverage.
The strongest objection to reshopping every few months is that model quality has converged enough that switching costs, not capability, now decide the question. A team that picked GPT-5.4 in March has no urgent reason to migrate to GPT-5.6 Terra simply because a newer number exists on a pricing page. That argument holds for most of this guide.
It does not hold for Claude Sonnet 5. Its price rises 50 percent on September 1, 2026, whether or not your workflow has changed at all. That is not a switching-cost question. It is a calendar, and it is the one line item in this guide with a hard deadline attached.
Watch two dates. September 1 is when Sonnet 5's discount expires, and whether Anthropic extends it again says more about how competitive the frontier tier actually is than any benchmark released this year.
Also watch whether Meta ships an open-weight Llama update at all before year end. If it does not, DeepSeek's MIT-licensed V4 will have become the default self-hosted answer for reasons that have nothing to do with which model scores higher, and everything to do with which vendor bothered to keep shipping the version you can actually run yourself.
Claude Sonnet 5 and Claude Opus 5 remain the models most developers reach for first, and Anthropic's own Opus 5 announcement leans on Frontier-Bench and OSWorld 2.0 rather than a single coding leaderboard number. GPT-5.6 Terra is the closest competitor on price and has the larger third-party tool ecosystem. For high-volume, cost-sensitive coding work, DeepSeek-V4-Flash is now MIT-licensed and self-hostable.
No, not by default. Meta has not shipped a new Llama model in 2026 and moved its newest research effort, Muse Spark, to a closed-weight release in April 2026. DeepSeek's V4, released the same month under an MIT license and published on Hugging Face, is now the more actively maintained open-weight option for teams that need to self-host.
Claude Sonnet 5 is priced at an introductory $2 per million input tokens and $10 per million output tokens through August 31, 2026. Standard pricing of $3/$15 per million tokens takes effect on September 1, 2026, according to Anthropic's own pricing documentation. That is a 50 percent increase on a fixed date, not a gradual drift.
DeepSeek-V4-Flash prices at $0.14 per million input tokens on a cache miss and $0.28 per million output tokens, roughly a third of what its V3.2 predecessor charged five months earlier. Gemini 3.6 Flash at $1.50/$7.50 is the cheapest option among the frontier-lab consumer brands.
Most teams under five people need one primary subscription and one secondary tool for a specific gap, such as Claude Pro plus Perplexity Pro at $40 a month combined. Paying for ChatGPT Plus, Claude Pro, and Gemini AI Pro simultaneously without a distinct daily use case for each is a subscription problem, not a research strategy.
All AI guides · Browse the AI directory
Still deciding? Get matched.
Smart Match checks every listing in the directory against how you work and what you can spend, then hands you the shortlist and the reason behind each pick.
Start Smart Match