中文

AI Model Comparison

Compare mainstream AI models side by side: capability scores, pricing, context windows and benchmark rankings

55

Models

30

Open source

21

Multimodal

10

Vendors

Claude 3 Haiku

Anthropic

Multimodal

Anthropic's fastest and cheapest model, ideal for real-time conversations and high-frequency API calls with extremely low cost.

Context window

200K

Input price

$0.25 / 1M

Output price

$1.25 / 1M

View details

Claude 3.5 Sonnet

Anthropic

Multimodal

Anthropic's strongest coding model, ranked first on SWE-bench, with top-tier code quality and instruction-following capabilities, excelling in agent tasks.

Context window

200K

Input price

$3 / 1M

Output price

$15 / 1M

View details

Claude Haiku 4.5

Anthropic

Multimodal

Claude Haiku 4.5, the fastest Claude, with near-frontier intelligence and 200K context, ideal for high-concurrency and low-latency scenarios.

Context window

200K

Input price

$1 / 1M

Output price

$5 / 1M

View details

Claude Opus 4.5

Anthropic

Multimodal

Claude Opus 4.5, 200K context, high-quality reasoning and coding, relatively better cost-effectiveness.

Context window

200K

Input price

$5 / 1M

Output price

$25 / 1M

View details

Claude Opus 4.6

Anthropic

Multimodal

Claude Opus 4.6, 1M context, supports extended thinking, stable performance on complex tasks.

Context window

1M

Input price

$5 / 1M

Output price

$25 / 1M

View details

Claude Opus 4.7

Anthropic

Multimodal

Claude Opus, the previous generation flagship, offers 1M context, strong complex reasoning and agentic coding capabilities, and remains a top-tier choice.

Context window

1M

Input price

$5 / 1M

Output price

$25 / 1M

View details

Claude Opus 4.8

Anthropic

Multimodal

Anthropic's current strongest model, top-tier in complex reasoning, long-cycle agentic coding, and highly autonomous tasks, ranked first in the Intelligence Index.

Context window

1M

Input price

$5 / 1M

Output price

$25 / 1M

View details

Claude Sonnet 4.5

Anthropic

Multimodal

Claude Sonnet 4.5, with 200K context, balances speed and intelligence, excelling in coding and agent tasks.

Context window

200K

Input price

$3 / 1M

Output price

$15 / 1M

View details

Claude Sonnet 4.6

Anthropic

Multimodal

The best balance of speed and intelligence, with 1M context, offering great value for daily development and agent tasks.

Context window

1M

Input price

$3 / 1M

Output price

$15 / 1M

View details

GPT / OpenAI

Series comparison

Qwen2.5-72B

Alibaba

Open sourceMultimodal

Alibaba's Tongyi Qianwen latest flagship, with the strongest Chinese language capabilities domestically, fully open-source, and supports multimodal.

Context window

128K

Input price

开源免费

Output price

开源免费

View details

Qwen2.5-Coder

Alibaba

Open source

Alibaba's specialized code model surpasses Claude 3.5 Sonnet in coding ability, achieving 98.5% on HumanEval, fully open-source.

Context window

128K

Input price

开源免费

Output price

开源免费

View details

Qwen2.5-Max

Alibaba

Tongyi Qianwen 2.5 Max: A large-scale MoE flagship model with comprehensive capabilities comparable to mainstream closed-source models.

Context window

128K

Input price

Output price

View details

Qwen3-Coder

Alibaba

Open source

Tongyi Qianwen 3 Code Special: 480B-A35B MoE open source (Apache 2.0), strong code capabilities.

Context window

Input price

Output price

View details

Qwen3-Max

Alibaba

Tongyi Qianwen 3 Max: A closed-source flagship with over 1T parameters, representing the ceiling of the Tongyi series' capabilities.

Context window

Input price

Output price

View details

Qwen3.5

Alibaba

Open source

Alibaba Tongyi Qianwen 3.5, open-source, fast, and extremely low-cost (starting from approximately $0.01/1M), available in multiple sizes.

Context window

Input price

$0.01 / 1M

Output price

View details

Qwen3.6

Alibaba

Open source

Tongyi Qianwen 3.6: 35B-A3B MoE open-source model (Apache 2.0), the latest generation in 2026.

Context window

Input price

Output price

View details

Qwen3.6-Plus

Alibaba

Tongyi Qianwen 3.6 Plus: Closed-source flagship version, released in 2026, with comprehensive capabilities comparable to mainstream closed-source models.

Context window

Input price

Output price

View details

Step (StepFun)

Series comparison

Other models

Benchmark rankings

GAIAAgent

Measures the ability of AI agents to complete real-world tasks, including multi-step reasoning, tool use, and information retrieval.

Claude 3.5 Sonnet

53.6%

SWE-bench VerifiedCode

Tests AI's ability to fix code bugs based on real GitHub issues, considered the closest evaluation to real-world development scenarios.

Claude 3.5 Sonnet

49%

HumanEvalCode

Code generation capability benchmark, containing 164 programming problems, testing the ability to generate functions directly from descriptions.

DeepSeek-V3

90.2%

MMLUKnowledge

A comprehensive knowledge understanding test covering 57 subjects, including mathematics, science, law, medicine, etc., to evaluate the model's broad knowledge base.

GPT-4o

88.7%

Chatbot ArenaUser preference

A preference ranking based on blind voting by real users, which is the most accurate reflection of actual user satisfaction.

Claude 3.5 Sonnet

ELO 1268

HumanEval+Code

Code generation benchmark released by OpenAI, evaluating the model's ability to write Python functions to solve algorithmic problems.

Qwen2.5-Coder 32B

98.5%

MATHReasoning

A test set of math problems ranging from high school to competition level, designed to evaluate the model's mathematical reasoning and problem-solving abilities.

o1-preview

94.8%

MMLU ProKnowledge

A multi-task language understanding benchmark covering 57 subjects, the most widely used knowledge evaluation set.

GPT-4o

88.7%

Pricing note

Prices are indicative; check each vendor's official site for the latest. Some models offer free tiers or API trials. Open-source models can be self-hosted, paying only compute cost.