| Zhipu GLM |
GLM-4-9B |
2024-06-05 |
128K |
A compact open-source GLM-4, plus an experimental 1M version, for local and private deployment. |
| Gemini |
Gemini 1.5 Flash |
2024-05-14 |
1M |
A high-speed model distilled from 1.5 Pro. Low price and low latency, for high concurrency and realtime apps. |
| OpenAI |
GPT-4o |
2024-05-13 |
128K |
The first native omni multimodal model, handling text, image, and audio in one system with lower latency. It was ChatGPT's default model for a time. |
| DeepSeek |
DeepSeek-V2 |
2024-05-06 |
128K |
An MLA-architecture MoE. Performance approached then-current mid-tier closed models at a very low price, and it directly drove a domestic API price war. |
| xAI |
Grok-1.5 |
2024-03-28 |
128K |
Clearly stronger than Grok-1 on reasoning and instruction following. Context grew to 128K, and vision (Grok-1.5V) followed soon after. |
| Kimi |
moonshot-v1-128k |
2024-03-20 |
128K |
An early flagship on the Kimi open platform. 128K context, suited to Chinese long-document Q&A and summarization. |
| Anthropic |
Claude 3 Haiku |
2024-03-13 |
200K |
The fastest and cheapest Claude 3 tier, nearly instant. Fits classification, extraction, customer support, and other high-concurrency work. |
| Anthropic |
Claude 3 Opus |
2024-03-04 |
200K |
The Claude 3 flagship. At the time it matched GPT-4-class reasoning, math, and vision, with 200K context for complex analysis. |
| Anthropic |
Claude 3 Sonnet |
2024-03-04 |
200K |
The Claude 3 balanced tier, trading off intelligence and speed. It was the most commonly deployed Sonnet for enterprise use. |
| Gemini |
Gemini 1.5 Pro |
2024-02-15 |
1M |
The first to bring production context to 1 million (later experimented to 2 million). Long video, codebases, and extra-long documents are strengths. |
| Gemini |
Gemini 1.0 Ultra |
2024-02-08 |
32768 |
The strongest 1.0 tier, launched with Gemini Advanced. It targeted then-current GPT-4-class reasoning and multimodality. |
| Qwen |
Qwen1.5 |
2024-02-04 |
32768 |
A full generation aligned to ChatML. Sizes from 0.5B to 110B MoE, and the ecosystem started to take shape. |
| Zhipu GLM |
GLM-4 |
2024-01-16 |
128K |
Zhipu's API flagship with 128K context. All Tools can automatically choose browsing, Python, images, and custom functions. |
| Gemini |
Gemini 1.0 Pro |
2023-12-06 |
32768 |
Google's first public Gemini generation, natively multimodal. Pro targeted developers and the Bard upgrade. |
| Qwen |
Qwen-72B |
2023-11-30 |
32768 |
An early large open-source flagship. Chinese, English, and long-text ability were clearly stronger than 7B. |
| DeepSeek |
DeepSeek LLM |
2023-11-29 |
4096 |
7B/67B general dense models. DeepSeek's first base as it moved from code into general chat. |
| Anthropic |
Claude 2.1 |
2023-11-21 |
200K |
Doubled context to 200K tokens, reduced long-document hallucinations, and improved tool-use reliability. |
| OpenAI |
GPT-4 Turbo |
2023-11-06 |
128K |
Announced at the first DevDay with 128K context, JSON Mode, and lower pricing. The Assistants API and GPTs launched alongside it. |
| xAI |
Grok-1 |
2023-11-03 |
8192 |
xAI's first chat model opened to X Premium users. In March 2024 it open-sourced 314B MoE weights. |
| DeepSeek |
DeepSeek Coder |
2023-11-02 |
16384 |
DeepSeek's first public model, focused on code completion and repo understanding. It started the later Coder line. |