| DeepSeek |
DeepSeek-V3.2 |
2025-12-01 |
128K |
The closing V3 generation. Both thinking and non-thinking can use tools, long context is more efficient, and it became the default before V4. |
| Anthropic |
Claude Opus 4.5 |
2025-11-24 |
200K |
The mid-4.x flagship, emphasizing top-tier reasoning. It shipped with product capabilities such as context compression and Claude for Excel. |
| xAI |
Grok-4.1 Fast |
2025-11-19 |
2M |
The high-speed long-context version of 4.1, for retrieval, browsing, and low-latency chat. |
| Gemini |
Gemini 3 |
2025-11-18 |
1048576 |
The start of Gemini 3. Multimodal and agent capability stepped up again, with about 1 million context. |
| xAI |
Grok-4.1 |
2025-11-17 |
256K |
Hallucinations are clearly lower than Grok-4, with better reasoning and creative writing. It once ranked near the top of LMSYS Arena. |
| OpenAI |
GPT-5.1 |
2025-11-12 |
400K |
The first major GPT-5 upgrade. Instant is warmer and follows instructions better. Thinking is faster on easy problems and more persistent on hard ones. |
| Kimi |
Kimi K2 Thinking |
2025-11-06 |
256K |
A K2 reasoning-only tier with deep thinking on by default. Fits math, complex planning, and long-horizon agents. |
| Anthropic |
Claude Haiku 4.5 |
2025-10-15 |
200K |
The fastest 4.5-generation model. Coding and computer use approach Sonnet 4 at a lower price, for browser agents and high concurrency. |
| Zhipu GLM |
GLM-4.6 |
2025-09-30 |
200K |
Context rose from 128K to 200K, with continued gains in coding, reasoning, tools, and agents. |
| Anthropic |
Claude Sonnet 4.5 |
2025-09-29 |
200K |
Then the strongest real-world agent / coding / computer-use Sonnet. It became the default model for products such as Claude in Chrome. |
| Qwen |
Qwen3-Max |
2025-09-24 |
256K |
The hosted Qwen3 flagship, strong at coding and knowledge work, with 256K context. |
| DeepSeek |
DeepSeek-V3.1-Terminus |
2025-09-22 |
128K |
Fixes mixed Chinese/English and stray characters, stabilizing V3.1 language consistency and agent behavior. |
| xAI |
Grok-4 Fast |
2025-09-19 |
2M |
A high-speed version that uses far fewer thinking tokens. Context can expand to 2 million, with much lower cost per unit of intelligence than early frontier models. |
| OpenAI |
GPT-5-Codex |
2025-09-15 |
400K |
A GPT-5 variant optimized for agentic coding. It became the default Codex model and is strong at repo-scale changes and long-horizon coding. |
| Kimi |
Kimi K2-0905 |
2025-09-05 |
256K |
A K2 instruction-tuning update. Context rose to 256K, with more stable coding and tool calling. |
| DeepSeek |
DeepSeek-V3.1 |
2025-08-21 |
128K |
One model that supports both thinking and non-thinking modes. Tool use and agent tasks are clearly stronger than V3/R1. |
| OpenAI |
GPT-5 |
2025-08-07 |
400K |
A generation that unified the fast model and deep-thinking model. A realtime router sets thinking depth. It became ChatGPT's default, and the API offers 400K context. |
| OpenAI |
GPT-5 mini |
2025-08-07 |
400K |
A mid-size GPT-5 that keeps unified routing and reasoning. Price and latency are a better fit for large-scale apps. |
| OpenAI |
GPT-5 nano |
2025-08-07 |
400K |
The smallest GPT-5 size, aimed at high throughput and low cost, while still offering basic reasoning. |
| Anthropic |
Claude Opus 4.1 |
2025-08-05 |
200K |
An incremental Opus 4 upgrade. More stable on complex reasoning, analysis, and creative work, and the transitional flagship before 4.5. |