| Zhipu GLM |
GLM-4.5 |
2025-07-28 |
128K |
A native-agent hybrid-reasoning flagship, 355B / 32B active, MIT-licensed. Coding and tool use are the core selling points. |
| Qwen |
Qwen3-Coder |
2025-07-22 |
256K |
A repo-scale agent coding specialist. Long context and tool calling are optimized for software engineering. |
| Kimi |
Kimi K2 |
2025-07-11 |
128K |
An open-source MoE flagship of about 1T parameters, with about 32B active. Focused on agents, tools, and coding, and seen as a major open-source milestone. |
| xAI |
Grok-4 |
2025-07-09 |
256K |
A multimodal flagship with native tool use and realtime search. The API offers 256K context, and Grok-4 Heavy launched with it. |
| xAI |
Grok-4 Heavy |
2025-07-09 |
256K |
The highest-compute Grok-4 tier for SuperGrok Heavy. It spends more compute for higher reasoning quality. |
| Gemini |
Gemini 2.5 Flash |
2025-06-17 |
1M |
The official 2.5 high-speed tier. Thinking budget is adjustable, for a flexible tradeoff between cost and intelligence. |
| DeepSeek |
DeepSeek-R1-0528 |
2025-05-28 |
128K |
A mid-cycle R1 update with stronger reasoning and fewer hallucinations. It also added JSON output and function calling. |
| Anthropic |
Claude Opus 4 |
2025-05-22 |
200K |
The Claude 4 flagship, focused on long-horizon agents and coding, with extended thinking. It was positioned as the strongest coding model at the time. |
| Anthropic |
Claude Sonnet 4 |
2025-05-22 |
200K |
The Claude 4 balanced tier. A broad upgrade over 3.7 in coding, instruction following, and thinking stability, also available to free users. |
| Qwen |
Qwen3 |
2025-04-29 |
131072 |
A hybrid thinking / non-thinking generation, including MoE sizes such as 235B-A22B, fully Apache-2.0 open source. |
| OpenAI |
o3 |
2025-04-16 |
200K |
The full o-series reasoning flagship. It can agentically use all tools in ChatGPT and reached the then-highest level on math, science, and complex tasks. |
| OpenAI |
o4-mini |
2025-04-16 |
200K |
A cost-efficient reasoning model released with o3. Full tool-calling support, for cases that need reasoning without o3's cost. |
| OpenAI |
GPT-4.1 |
2025-04-14 |
1M |
The API flagship non-reasoning model, strong at coding. It officially brought context to 1 million tokens, with mini and nano released at the same time. |
| OpenAI |
GPT-4.1 mini |
2025-04-14 |
1M |
A mid-small GPT-4.1 size that keeps million-token context and strong coding ability. A good fit for cost-efficient production use. |
| OpenAI |
GPT-4.1 nano |
2025-04-14 |
1M |
The smallest GPT-4.1 size, with the lowest latency and price. Suited to classification, extraction, simple completion, and other high-traffic work. |
| Gemini |
Gemini 2.5 Pro |
2025-03-25 |
1M |
A 2.5 flagship preview with a thinking process. It led then-current Gemini on math, code, and long-context retrieval. |
| DeepSeek |
DeepSeek-V3-0324 |
2025-03-24 |
128K |
A major hosted V3 update. Reasoning, coding, writing, search, and function calling all improved, MIT-licensed. |
| Qwen |
QwQ-32B |
2025-03-06 |
131072 |
A Qwen open-source reasoning model. A 32B dense model that approached larger models on math and code reasoning. |
| OpenAI |
GPT-4.5 |
2025-02-27 |
128K |
A research-preview extra-large non-reasoning chat model that emphasized world knowledge and writing quality. Its lifespan was short, and it was later absorbed into the GPT-4.1 / GPT-5 line. |
| Anthropic |
Claude 3.7 Sonnet |
2025-02-24 |
200K |
The first hybrid-reasoning Sonnet. It can switch between instant answers and extended thinking, and Claude Code launched with it. |