AI Model Comparison Table

Last updated: 2026-09-24

Model lineups now change every few weeks, so this page describes the families and what each is known for. For the current versions, per-token prices, context windows and benchmark scores, use the live model leaderboard and the model directory — every figure there carries its source and the date it was checked.

The major families

OpenAI (GPT) — General-purpose models with broad tool use, strong coding, and a separate reasoning tier for multi-step problems. The widest third-party ecosystem.

Anthropic (Claude) — Strong on long documents, analysis, writing and agentic coding. A lineup from a large flagship down to a fast, low-cost tier.

Google (Gemini) — Very long context windows and native multi-modal input. The Flash tier is often among the best value for high-volume work, and it integrates deeply with Google Workspace and Cloud.

Meta (Llama) — Open-weight models you can self-host or run through hosted providers. The usual choice when data control or fine-tuning matters more than having the newest frontier model.

Mistral — Open and proprietary models from a European provider, a common pick when EU hosting is a requirement.

DeepSeek — Strong reasoning and coding at aggressive prices, with open weights for many releases. Check its compliance posture before regulated use.

Qwen (Alibaba) — Open and API models with strong multilingual and Asian-language performance.

Cohere — Enterprise-focused, known for embeddings and retrieval, and common in RAG pipelines.

How to choose

Complexity of the task — Hard, multi-step work benefits from a family's flagship or reasoning tier. Simpler, high-volume tasks can usually drop to a fast tier with little loss.

Cost at scale — Token pricing compounds quickly. Compare input and output prices against your real usage pattern on the leaderboard, not the headline input price.

Context requirements — Long documents need long context. Check the context window on each model's page, and test recall on your own documents before relying on the maximum.

Multi-modal needs — If you need vision, audio or video, check each model's modality support on its page before committing.

Open vs. proprietary — Open weights give you self-hosting, privacy control and fine-tuning. Proprietary models give you managed infrastructure and usually the newest frontier capability.

Integration fit — Your existing stack matters. Google Workspace users may get more from Gemini; many developers prefer Claude or GPT for tool use and agent workflows.