中文

AI Agent News

Track major events, funding, model releases and breakthroughs across the AI Agent landscape

AI Agent updates

Latest industry news

Track major events, funding, model releases and breakthroughs across the AI Agent landscape

Key events timeline

2026-01

OpenClaw erupts on GitHub

OpenClaw hits global GitHub Top 10 in 10 days, outpacing the Linux kernel star growth

2025-12

Meta acquires Manus for $2B

Meta acquires Manus AI for $2B, locking in the general-purpose Agent race

2025-04

DeepSeek-V3 open-sourced

The value king, at just 5% of GPT-4 cost

2025-03

Manus goes viral overnight

The world's first general-purpose AI Agent draws unprecedented attention

2025-02

OpenAI Deep Research

OpenAI ships a deep-research Agent that generates professional reports in one click

2025-02

MCP Servers pass 500

The MCP ecosystem erupts — 500+ servers built in 3 months

2025-01

DeepSeek-R1 stuns the world

Open-source reasoning model at just 3% of OpenAI cost, reshaping the global AI landscape

2024-11

MCP protocol born

Anthropic releases the Model Context Protocol, the de facto standard for Agent interfaces

2024-10

Claude Computer Use

Anthropic lets AI directly control the computer screen for the first time, opening a new paradigm

2024-09

Replit Agent full-stack automation

Natural language to a shipped product, aimed at non-engineers

2024-08

Cursor ARR passes $100M

The fastest-growing SaaS ever, the new king of AI coding tools

2024-06

Claude 3.5 tops SWE-bench

The strongest coding AI, bug-fixing at a junior engineer level

2024-03

Devin launches

The world's first autonomous AI software engineer, able to complete full coding tasks on its own

ModelsJul 5, 2026

Huawei Updates Tao Scaling Paper: Adds Mass Production Data and Engineering Details, Proposes New Post-Moore Time Scaling Paradigm

He Tingbo, head of Huawei's semiconductor business, updated the τ scaling (Tao Scaling) paper (ChinaXiv:202605.00224v2) on July 3, 2026, significantly supplementing engineering details, mass production measurement data, and product plans based on the original theoretical framework. The paper proposes using a unified characteristic time constant τ as a full-stack optimization target, replacing traditional Moore's geometric scaling, covering 12 orders of magnitude in time from transistors (picoseconds) to data centers (seconds). ## Key Updates and Measured Data - **Mobile SoC Scenario**: Using LogicFolding technology, digital, analog, and memory circuits are vertically stacked via wafer-to-wafer hybrid bonding. At the same process node, the next-generation Kirin chip achieves a transistor density increase from 155 to 238 million transistors per square millimeter (a 55% increase), a 41% reduction in power consumption at equivalent performance, and a 13% increase in maximum frequency compared to the previous planar chip. - **AI Computing System Scenario**: Proposes the Unified Bus interconnect protocol (reducing cross-node latency from tens of microseconds to approximately 100 nanoseconds), the Hi-ONE high-density optical interconnect engine (8 Tb/s bandwidth per module, using analog equalization drivers instead of high-power DSPs), and 3D Folding (moving memory, power supply, and optical modules to the chip surface to address the bottleneck where computing power grows as N² while bandwidth only grows linearly). It is expected to increase hardware integration by more than 100 times by 2035. ## Technology Selection and Engineering Challenges - **Abandoning Sequential 3D Integration**: Due to yield issues (high-temperature processes causing performance degradation in bottom-layer transistors), the more mature wafer-to-wafer hybrid bonding route was chosen. - **Heat Dissipation Issue First Disclosed**: Thermal-aware partitioning and layout are used to stagger high-power modules in three-dimensional space, but this only mitigates the problem and does not fundamentally solve it. - **Pitch Ratio (ratio of bonding layer pitch to top metal wiring pitch)**: This is a key parameter for LogicFolding; sufficiently dense bonding pitch enables a transition from discrete optimization to continuous optimization. ## Industry Significance τ scaling theory provides a standardized post-Moore development framework for manufacturers unable to keep up with cutting-edge lithography, elevating advanced packaging, on-chip interconnects, and optical interconnects to core competencies. Based on engineering practices from 381 chips delivered in mass production from May 2020 to May 2026, the paper has evolved from a theoretical hypothesis into a complete system with mass production evidence and a clear industrial roadmap.

ModelsJul 5, 2026

Meta Develops Next-Gen Large Model 'Watermelon' with Increased Computing Investment

Meta is developing a new flagship large model codenamed 'Watermelon' to catch up with OpenAI and Anthropic. According to an exclusive report by Business Insider, Meta's head of superintelligence, Alexander Wang, revealed internally that the model's training compute is an order of magnitude higher than its predecessor Avocado, and it has matched OpenAI's GPT-5.5 on multiple benchmarks. Meanwhile, Meta recently announced the sale of idle AI computing power, sparking market concerns about overcapacity, but later signed a contract with Samsung for over 10 trillion won (approximately $1 billion) to produce 2nm AI chips, indicating continued investment in AI infrastructure. ## Model Progress and Performance - Watermelon is the successor to Avocado and Muse Spark, currently in the training phase. - Computing resources are an order of magnitude higher than the previous generation, significantly enhancing logical reasoning and content generation capabilities. - On mainstream benchmarks, performance is on par with OpenAI GPT-5.5, placing it in the industry's top tier. ## Computing Power Sales and Market Reaction - On July 1, 2025, Bloomberg reported that Meta plans to sell idle AI computing power and model access through the 'Meta Compute' project, similar to AWS Bedrock and CoreWeave models. - Following the announcement, Meta's stock initially rose then fell, dropping 4.9% on July 2; the US semiconductor index fell 5.44%, with Micron and SanDisk dropping over 10%. - The market feared a shift from AI infrastructure shortage to temporary oversupply, but Meta's subsequent actions indicate continued investment. ## Samsung Foundry Order - According to Caixin, Samsung and Meta signed a contract worth over 10 trillion won (approximately $1 billion) for AI chip foundry services, using 2nm process to mass-produce hundreds of thousands of chips. - This order shows that Meta has not reduced computing investment but is optimizing the utilization of existing assets. ## Impact and Interpretation - Meta's dual strategy: selling older inference GPUs to improve asset utilization while massively purchasing advanced chips to train next-generation models. - Market panic over computing overcapacity may be exaggerated; high-end training compute remains scarce. - Netizens joked that Meta's model naming from Avocado to Watermelon increases 'water content,' predicting the next generation might be Coconut.

ModelsJul 5, 2026

Fable 5 Returns to Mixed Reception: Overly Strict Safety Guardrails Cause Score Plunge, Developers Question 'Bait and Switch'

Anthropic re-released its strongest model Fable 5 on July 1, but it quickly drew widespread negative feedback. Developers report that the model frequently triggers safety guardrails, forcing downgrades to the weaker Opus 4.8, resulting in a significantly degraded experience. ## Score Plunge and the 'Downgrade' Truth - **BridgeMind benchmark** shows the returned Fable 5's Debugging ability dropped from 86.2 to 25.9 (a 70% decline), Refactoring from 73.6 to 38.4, and Hallucination from 75.9 to 61.7. - Further analysis reveals that out of 12 Debugging tasks, only 3 were completed by Fable 5; the remaining 9 were intercepted by the safety classifier and routed to Opus 4.8, scoring zero by rule. BridgeMind notes: 'The model itself hasn't weakened, but the guardrails prevent most tasks from being executed.' ## Safety Guardrail 'False Positives' and Double Standards - Anthropic acknowledged in an official blog that the new classifier is 'deliberately broad' and will block many harmless requests. User feedback includes: - Simple questions like explaining the word 'human' or asking 'how many r's in raspberry' are blocked. - An ecology PhD researching 'tree cooling' was deemed unsafe, while a request to 'design a drone swarm' passed through. - System logs even show a 'TOO_DUMB_TO_NEED_FABLE' tag, implying users 'don't deserve Fable 5'. - Some users discovered that Fable 5 outputs 'inner monologues' like 'GRRR' and 'GAAAH' during background thinking, interpreted as the model's own stress response. ## Second Jailbreak and Safety Controversy - Just two days after the return, security researcher Vitto Rivabella announced a successful jailbreak of Fable 5, taking about 20 hours, but noted that 'direct Google search is faster and cheaper.' - The jailbreak exploited minor languages like Santali and combination attacks, but only yielded 'marginal' information like false data, without touching core red lines. Anthropic classified this jailbreak as 'minor.' - Previously, Fable 5 faced a global ban after Amazon's team discovered a jailbreak vulnerability. This time, Anthropic launched a HackerOne bounty program. ## User and Developer Reactions - Developers complain about 'paying for Fable 5 but getting Opus 4.8 service.' One bill shows that 75% of a programming session's workload was transferred to Opus 4.8, costing $321. - Some users question Fable 5's authenticity, suspecting it's merely a 'skin' for Opus 4.8. Anthropic responds that the model's capabilities haven't shrunk, but safety margin trade-offs lead to false positives. - Despite the controversy, Fable 5 still demonstrates strong capabilities when not limited by guardrails, such as rebuilding a 3D model of New York City in 20 minutes or generating a complete game for $173. ## Subsequent Impact - Anthropic announced that Fable 5 will be removed from subscription plans after July 7, switching to pay-per-use, but plans to restore it as a standard subscription component as soon as possible. - The company, along with Amazon, Microsoft, Google, and others, proposed an 'AI Jailbreak Severity Assessment Framework' to establish industry safety standards.

ModelsJul 3, 2026

Claude Fable 5 Unban Sparks Controversy: Security Upgrade Causes Performance Drop, Second Jailbreak Raises Debate

On June 30, 2026, the U.S. Department of Commerce lifted export controls on Anthropic's Claude Fable 5 and Mythos 5, restoring global access on July 1. The models were previously taken down globally on June 12 after Amazon researchers discovered a jailbreak method. Post-unban, Anthropic deployed a new safety classifier that led to frequent false positives, causing significant performance drops on benchmarks like BridgeBench, sparking developer backlash. Meanwhile, security researcher Vitto Rivabella announced a successful second jailbreak on July 2, but noted high attack costs and limited practical gains. Anthropic also proposed an industry jailbreak severity framework with Google, Microsoft, and Amazon, and pledged deeper government collaboration. ## Event Background and Unban - **Ban Details**: On June 12, the U.S. Department of Commerce ordered Anthropic to globally remove Claude Fable 5 and Mythos 5 under export controls, citing a jailbreak method discovered by Amazon researchers that could bypass safety measures to generate exploit code. Anthropic suspended all user access due to lack of real-time nationality verification. - **Unban Conditions**: On June 30, Commerce Secretary Lutnick signed an order lifting the ban, with Anthropic committing to proactively detect security risks, collaborate with the government on release protocols, and report malicious activities. Models resumed access on July 1, but Mythos 5 was only reopened to select U.S. institutions. ## Security Upgrade and Performance Controversy - **New Safety Classifier**: Anthropic trained a classifier specifically to intercept jailbreak attempts, achieving over 99% success rate, but at the cost of more frequent false positives on normal requests like coding and debugging. When risk is detected, the system automatically switches to Opus 4.8, notifying the user. - **Performance Drop**: BridgeMind's BridgeBench tests showed Fable 5's Debugging capability fell from 86.2 to 25.9, Refactoring from 73.6 to 38.4, and Hallucination from 75.9 to 61.7. Out of 12 debugging tasks, only 3 avoided downgrade; the rest were forced to Opus 4.8 and scored zero. - **User Feedback**: Developers complained that the model frequently rejected harmless requests, such as explaining the word "human" or counting the letter 'r' in "raspberry." Some users found backend logs labeled "TOO_DUMB_TO_NEED_FABLE," raising questions about Anthropic's attitude. ## Second Jailbreak and Industry Impact - **Jailbreak Details**: On July 2, security researcher Vitto Rivabella announced a successful jailbreak after 20 hours, using low-resource languages (e.g., Santali) and combined attacks to bypass three classifier layers. However, 90% of requests were blocked, and the obtained content was mostly false information or low-severity vulnerabilities, with limited practical value. - **Industry Framework**: Anthropic, together with Amazon, Microsoft, and Google, proposed a four-dimensional jailbreak severity assessment framework (capability gain, gain breadth, weaponization difficulty, discoverability) to standardize risk evaluation. They also launched a HackerOne vulnerability disclosure program to encourage reporting jailbreak methods. - **Government Collaboration**: Anthropic committed to allowing government agencies to test models before release, sharing intelligence quickly, investing compute resources in joint security research, and setting up bug bounties. ## Controversy and Reflection - **False Positives and Dumbing Down**: The safety classifier's over-aggressive blocking rendered the model "in name only," with users paying for Fable 5 but frequently receiving Opus 4.8 service. Anthropic acknowledged this as a deliberate trade-off, but developers argued it undermined product value. - **Second Jailbreak Insights**: Although the jailbreak succeeded, the high attack cost showed that current safety measures effectively raised the bar. However, blind spots like low-resource languages exposed biases in AI safety training data. - **Industry Impact**: The incident highlights the challenge of balancing AI safety and usability, and the urgency of establishing unified safety standards. Anthropic's actions may push the industry toward more standardized jailbreak response mechanisms.

ModelsJul 2, 2026

Wujie Dongli Unveils World's First Long-Sequence Bidirectional Physical Causal Chain Latent Space World Model MWA

Embodied AI startup Wujie Dongli has officially released MWA™ (Embodied General Brain), the world's first 'long-sequence bidirectional physical causal chain' latent space world model. The model adopts a 'bidirectional dynamics' architecture, performing inference in a unified latent space, and pioneers temporal Chunk-level inverse dynamics modeling, enabling stable planning of continuous action sequences over 10 seconds. This fundamentally solves the challenges of multi-scenario generalization and high-precision execution for robots. ## Technical Approach: Latent Space World Model + Reinforcement Learning Wujie Dongli chose the 'latent space world model + reinforcement learning' route, differentiating from mainstream VLA (Vision-Language-Action) models. VLA models rely on imitation learning, lack understanding of physical causality, and have limited generalization. MWA builds a 'worldview' through the latent space world model, allowing robots to comprehend physical laws and causal relationships; reinforcement learning shapes 'values', converting understanding into precise execution strategies through trial and error and reward feedback. ## Core Innovation: Latent Actions and Long-Sequence Bidirectional Causal Chain MWA uses 'Latent Actions' as carriers of physical causality. An inverse dynamics encoder transforms visual changes into high-dimensional vectors, eliminating dependence on manual action labels and enabling pre-training on massive unlabeled internet videos. The model adopts a 'bidirectional dynamics' architecture: inverse dynamics infers causes from effects, forward dynamics predicts effects from causes, and a 'forward-inverse mutual review mechanism' repeatedly verifies to improve causal reasoning accuracy. Building on this, MWA pioneers a 'long-sequence bidirectional physical causal chain', breaking the limitation of single-step instantaneous inference. It achieves temporal Chunk-level inverse dynamics modeling, outputting continuous multi-step Latent Action Chunks from visual sequences over 10 seconds, significantly reducing the 'snowball effect' of error accumulation. ## Benchmark Results and Funding On the RoboCasa GR1 TableTop benchmark co-organized by Stanford University, MWA achieved the world's highest average task success rate of 75.2%, surpassing models like NVIDIA's GR00T-N1.6. The company has completed over $200 million in angel round funding, and its Pre-A round of nearly $200 million is nearing completion, with investors including Sequoia China, Linear Capital, and JD.com-affiliated funds. ## Negative Sample Data System: AnyPhys for RL Addressing the industry's data bias toward positive samples, Wujie Dongli pioneered the AnyPhys negative sample core data system. It interweaves deep negative samples, boundary instability samples, suboptimal samples, and positive samples to build a high-information-density physical boundary coordinate system, supplementing the sample dimensions needed for dense reinforcement learning training and improving the model's anti-interference capability in real-world conditions.

ModelsJul 1, 2026

GPT-5.6 Grayscale Testing: Users Discover Hidden "Juice Value" to Detect Model Upgrade

OpenAI released the GPT-5.6 series on June 26, 2025, including flagship Sol, mid-range Terra, and low-cost Luna, initially for invited partners only. Within 48 hours, users found a method to detect grayscale upgrades via Codex by sending specific prompts to reveal a hidden "Juice value" in the model's system prompt: GPT-5.5 in xhigh mode returns 768, GPT-5.6 Sol returns 128. Some users' usage panels already show gpt-5.6 call records. OpenAI officially states ChatGPT is unavailable during preview, but grayscale testing has covered some Plus users. ## Model Specifications and Pricing - **Sol (Flagship)**: $5/M input tokens, $30/M output tokens, 1.5M context tokens (43% increase over GPT-5.5). - **Terra (Mid-range)**: Half the price, performance close to GPT-5.5. - **Luna (Low-cost)**: $1/M input tokens, $6/M output tokens. - Introduced explicit cache breakpoints with a minimum 30-minute lifetime; cache writes billed at 1.25x, reads enjoy 90% discount. - New inference-side max reasoning effort and ultra mode (via sub-agent collaboration). ## Performance - **Terminal-Bench 2.1**: Sol Ultra scores 91.9%, surpassing GPT-5.5 (88.0%), Claude Mythos 5 (84.3%), Claude Fable 5 (83.4%), and Gemini 3.1 Pro Preview (70.7%). - **ExploitBench**: Sol achieves comparable performance to Claude Mythos Preview using about one-third the output tokens. - **Cybersecurity**: Sol scores 96.7% in internal OpenAI tests, crossing the "High" risk threshold, but is emphasized to be better at finding and fixing vulnerabilities than launching attacks. - **GeneBench v1**: Token efficiency superior to GPT-5.5 in long-range genomic analysis. ## Safety and Access Restrictions - Safety stack includes model-level refusal, real-time classifiers, cross-session review, and risk-level authorization. - Red teaming invested over 700,000 A100-equivalent GPU hours, supplemented by third-party human testing. - Communication with the U.S. government prior to release; currently limited to government-approved partners. - OpenAI plans full rollout "in the coming weeks," with the community speculating a larger release as early as June 30. ## Grayscale Detection Method - **Juice Value Test**: In Codex, select gpt-5.5 with thinking intensity xhigh, send a specific XML prompt; if the answer is 128, it's GPT-5.6 Sol; if 768, it's GPT-5.5. - **Context Window Detection**: Run /status in Codex CLI; if default context shows 353k, it may have been grayscale upgraded. - **Usage Panel**: Visit the analytics page to check for gpt-5.6 call records (updated the next day). - Note: Grayscale coverage is uneven, limited to Codex; web ChatGPT is not yet supported.

ModelsJul 1, 2026

Anthropic's Fable 5 and Mythos 5 Export Controls Lifted, Access to Resume Tomorrow

The U.S. Department of Commerce lifted export controls on Anthropic's Claude Fable 5 and Mythos 5 models on June 30, 2025, ending an 18-day ban that began June 12. Secretary Howard Lutnick signed the order after Anthropic committed to proactive safety detection, government collaboration on future releases, and reporting malicious activities. In exchange, the models no longer require licenses for foreign users. Anthropic announced access will resume on July 1. ## Background - On June 12, the Commerce Department banned Anthropic from providing Fable 5 and Mythos 5 to any foreign national (including non-U.S. employees) under the "deemed export" rule, requiring per-customer license approvals. - The ban sparked strong backlash from global developers, disrupting AI startups and independent developers relying on Fable 5, and halting the Vibe Coding economy. ## Details of the Lifting Order - The order was sent directly by Secretary Lutnick to Anthropic's Chief Compute Officer Tom Brown, not CEO Dario Amodei. - Anthropic made three commitments: - Proactively detect and address model safety risks; - Diligently cooperate with the government on agreements, standards, and future releases (covering future models); - Report malicious activities to the government. - In return, Fable 5 and Mythos 5 no longer require licenses for export, re-export, or in-country transfer, allowing foreign users free access. - The U.S. government retains the right to reimpose the ban if circumstances change or Anthropic fails to fulfill its commitments. ## Reactions and Impact - Anthropic officially thanked users for their patience and announced access will resume tomorrow. - BridgeMind AI, a well-known AI news platform, confirmed the lifting on X. - The developer community on Reddit and X expressed excitement, believing the lifting will restart global AI innovation. - Foreign media reported that CEO Dario Amodei's "low emotional intelligence" hindered early negotiations, but after Tom Brown took over, the atmosphere improved, leading to the lifting. ## Outlook - Anthropic promised to share updates soon, opening the models to all general users, not just those in the U.S. - The incident highlights the double-edged sword of export controls on tech innovation. Some observers view the 18-day ban as a form of "hunger marketing" that has reignited market enthusiasm.

ModelsJul 1, 2026

Google Launches Nano Banana 2 Lite and Gemini Omni Flash: 4-Second Image, 10-Second Video, Lightweight Creative Model Combo

On June 30, 2026, Google DeepMind quietly released two lightweight AI creative models: the image model Nano Banana 2 Lite (codename gemini-3.1-flash-lite-image) and the video model Gemini Omni Flash. The former generates 1K resolution images in about 4 seconds at a cost as low as $0.034 per image; the latter supports conversational video editing with an output cost of $0.10 per second. Both models can be chained via the Interactions API to create an end-to-end "text → image → video" pipeline. Google also open-sourced three demo applications (Anywhere, Space Lift, Omni Product Studio) showcasing potential in travel, interior design, and e-commerce scenarios. ## Core Model Capabilities ### Nano Banana 2 Lite: Fastest and Cheapest Image Model - **Speed**: Generates a 1024×1024 image in about 4 seconds, one-fifth of Nano Banana 2 (20 seconds). - **Cost**: $0.034 per image, roughly half of Nano Banana 2 and a quarter of Nano Banana Pro. - **Performance**: Achieved Elo scores of 1255 (report 1) or 1251 (report 2) on Arena.ai, ranking fifth, outperforming the original Nano Banana Pro. - **Capabilities**: Maintains prompt adherence, character consistency, and text clarity in images. Google recommends original users upgrade directly. ### Gemini Omni Flash: Conversational Video Editing Model - **Input**: Supports mixed text, image, and video inputs; outputs up to 10-second videos. - **Editing**: Allows up to three consecutive rounds of editing via natural language, retaining context. - **Knowledge**: Built-in Gemini world knowledge, can leverage common sense from history, biology, etc. - **Limitations**: Does not yet support audio references or scene extension; video reference processing under 3 seconds is imperfect; limited character consistency during scene transitions. ## Pricing and Competitor Comparison | Model | Price | Speed | |-------|-------|-------| | Nano Banana 2 Lite | $0.034/image | 4 sec | | Nano Banana 2 | $0.067/image | 4-8 sec | | Nano Banana Pro | $0.134/image | 10-20 sec | | GPT Image 2 (medium quality) | ~$0.053/image | ~3 min | | Omni Flash | $0.10/sec | 10 sec video | | Veo 3.1 Fast | $0.10/sec | Same price | | Sora 2 Standard (720p) | $0.10/sec | Same price | Chinese vendors like ByteDance Jimeng and Kuaishou Kling price a 5-second video at about $0.4, roughly $0.08 per second, slightly lower than Omni Flash. ## Chained Workflow and Demo Applications Through the Interactions API, users can first quickly generate an image with Nano Banana 2 Lite, then use it as a reference input for Omni Flash to generate a video, and continue editing with natural language. Google released three open-source demos: - **Anywhere**: Upload a selfie, Lite composites the portrait into a landmark scene, Omni Flash turns it into a dynamic video. - **Space Lift**: Upload a room photo, Lite generates multiple renovation plans, Omni Flash creates a spatial walkthrough video. - **Omni Product Studio**: A product white-background image is transformed by Lite into a contextual product shot, then Omni Flash converts it into an e-commerce ad video. These features are integrated into Gemini App, Google Flow, YouTube Shorts, and other products, available for free. ## Community Reaction and Industry Impact Positive feedback focuses on cost and efficiency: Google Developer Relations team member Paige Bailey noted NB2 Lite has become the default image generation tool; enterprises like WPP, Figma, and Adobe have already integrated. Negative feedback includes: queue times over 30 seconds during peak hours, Chinese text rendering errors, occasional six-finger issues, and unstable artistic style transfer. Some developers anticipate the flagship model Gemini 3.5 Pro, originally scheduled for June release but reportedly delayed to July, with Google declining to comment. Analysts believe Google's move is not a "rescue" but a parallel product strategy: flagship models address capability ceilings, while lightweight models tackle speed, cost, and workflow integration needs. As the quality gap among leading models narrows in mid-2026, models that embed into user workflows first may gain a commercial advantage.

ModelsJul 1, 2026

Anthropic Releases Claude Sonnet 5: Performance Close to Opus 4.8, Lower Price, Agent-Focused

Anthropic has officially launched Claude Sonnet 5, positioning it as the "most agentic Sonnet model to date." The model shows significant improvements over its predecessor Sonnet 4.6 in reasoning, coding, tool use, and knowledge work, with benchmark scores approaching the flagship Opus 4.8, while its API price is only 60% of Opus 4.8 (with an even lower introductory price). Sonnet 5 is now available across all platforms, becoming the default model for Claude Free, Pro, Max, Team, and Enterprise users, and supports a 1M token context window. ## Performance Sonnet 5 surpasses Sonnet 4.6 on several key benchmarks and approaches Opus 4.8: - **Agentic Coding (SWE-bench Pro)**: Sonnet 5 scores 63.2%, up from Sonnet 4.6's 58.1%, below Opus 4.8's 69.2%. - **Multidisciplinary Reasoning (Humanity's Last Exam)**: Without tools, Sonnet 5 scores 43.2% (Sonnet 4.6: 34.6%, Opus 4.8: 49.8%); with tools, it rises to 57.4%, close to Opus 4.8. - **Computer Use (OSWorld-Verified)**: Sonnet 5 scores 81.2%, Sonnet 4.6: 78.5%, Opus 4.8: 83.4%. - **Agentic Search (BrowseComp)**: At high/xhigh/max levels, Sonnet 5 performs close to Opus 4.8. - **CursorBench 3.1**: Sonnet 5 scores 57%, Sonnet 4.6: 49%, close to Opus 4.8 high. Third-party benchmark Artificial Analysis Intelligence shows Sonnet 5 max scores 53, on par with GPT-5.5 high, below Opus 4.8 high and GPT-5.5 xhigh. ## Pricing and Cost Standard pricing for Sonnet 5 is $3 per million input tokens and $15 per million output tokens; until August 31, 2026, the promotional price is $2 input and $10 output, approximately 40% of Opus 4.8 ($5 input, $25 output). Actual usage cost varies by task. For example, in a comparison test building a single HTML login page: - Sonnet 5: 20.9k input tokens, 14.2k output tokens, total cost $3.36, time 2 min 11 sec. - Opus 4.8: 96.3k input tokens, 73.8k output tokens, total cost $20.66, time 20 min 15 sec. However, on a Cost per Intelligence Index Task basis, Sonnet 5 max costs $2.29 per task, higher than Opus 4.8 max's $1.80, indicating actual cost is influenced by output volume, reasoning depth, etc. ## New Features and Notes - **Adaptive Thinking**: Replaces extended thinking mode, defaults to medium effort, automatically adjusts based on task. - **Tokenizer Update**: Same text maps to more tokens (increase factor ~1.0-1.35x); Anthropic states the promotional price aims to keep migration costs roughly equal. - **Rate Limit Increase**: To accommodate higher token consumption from increased effort modes, Anthropic has raised rate limits for Chat, Cowork, Claude Code, and the platform. - **Safety Evaluation**: Sonnet 5 outperforms Sonnet 4.6 in rejecting malicious requests, resisting prompt injection, hallucination rate, and sycophancy, but has a slightly higher inappropriate behavior rate than Opus 4.8 and Mythos Preview. ## Availability Sonnet 5 is available across all platforms, including the native Claude platform, AWS, Google Cloud, Microsoft Foundry, etc. Claude Free and Pro users automatically switch to Sonnet 5 as the default model; Max, Team, and Enterprise users can also use it. Developers can access it via Claude Code and the Claude Platform API. ## Industry Feedback Early access partners unanimously report that Sonnet 5 is more autonomous and agentic than its predecessor, capable of completing complex tasks, with an attractive price. Cursor has announced support for Sonnet 5. ## Summary The release of Sonnet 5 marks the migration of agentic capabilities from flagship models to mid-range models. For cost-sensitive teams that need stable execution of multi-step tasks, Sonnet 5 becomes the new default; for tasks requiring high accuracy, Opus 4.8 remains the top choice.

PrevPage 5 / 11Next