AI Model Comparison Table
Last updated: 2026-09-24
Model lineups now change every few weeks, so this page describes the families and what each is known for. For the current versions, per-token prices, context windows and benchmark scores, use the live model leaderboard and the model directory — every figure there carries its source and the date it was checked.
The major families
OpenAI (GPT) — General-purpose models with broad tool use, strong coding, and a separate reasoning tier for multi-step problems. The widest third-party ecosystem.
Anthropic (Claude) — Strong on long documents, analysis, writing and agentic coding. A lineup from a large flagship down to a fast, low-cost tier.
Google (Gemini) — Very long context windows and native multi-modal input. The Flash tier is often among the best value for high-volume work, and it integrates deeply with Google Workspace and Cloud.
Meta (Llama) — Open-weight models you can self-host or run through hosted providers. The usual choice when data control or fine-tuning matters more than having the newest frontier model.
Mistral — Open and proprietary models from a European provider, a common pick when EU hosting is a requirement.
DeepSeek — Strong reasoning and coding at aggressive prices, with open weights for many releases. Check its compliance posture before regulated use.
Qwen (Alibaba) — Open and API models with strong multilingual and Asian-language performance.
Cohere — Enterprise-focused, known for embeddings and retrieval, and common in RAG pipelines.
How to choose
Complexity of the task — Hard, multi-step work benefits from a family's flagship or reasoning tier. Simpler, high-volume tasks can usually drop to a fast tier with little loss.
Cost at scale — Token pricing compounds quickly. Compare input and output prices against your real usage pattern on the leaderboard, not the headline input price.
Context requirements — Long documents need long context. Check the context window on each model's page, and test recall on your own documents before relying on the maximum.
Multi-modal needs — If you need vision, audio or video, check each model's modality support on its page before committing.
Open vs. proprietary — Open weights give you self-hosting, privacy control and fine-tuning. Proprietary models give you managed infrastructure and usually the newest frontier capability.
Integration fit — Your existing stack matters. Google Workspace users may get more from Gemini; many developers prefer Claude or GPT for tool use and agent workflows.