Bullish

DeepSeek-V4-Pro Agent Scores Surge Nearly 50 Points in Internal Tests

11:03

DeepSeek-V4-Pro shows massive agent capability gains, with DeepSWE up 49.9 points. It now outperforms Claude Opus 4.8 on multiple benchmarks while maintaining unchanged API pricing.

Woofun AI data shows that DeepSeek-V4-Pro-0813 demonstrates substantial improvements in agent capabilities compared to its preview version. Internal metrics indicate DeepSWE scores jumped from 12.8 to 62.7, a gain of 49.9 points. CyberGym results increased from 52.7 to 83.3, and AutomationBench rose from 12.8 to 31.8.

The updated model exceeds Claude Opus 4.8 in several evaluations, including Terminal Bench 2.1 (87.9 vs 85.0) and CyberGym (83.3 vs 78.3). AutomationBench also surpasses Fable 5 with a score of 31.8 against 29.1. API pricing remains at 3 yuan per million input tokens and 6 yuan for output. These figures derive solely from internal testing, with independent verification pending.

WOOFUN AI

Impact Assessment · Quick Read

The significant leap in agent benchmark scores suggests DeepSeek is closing the gap with leading Western models in autonomous task execution. Maintaining low pricing despite performance gains could disrupt the competitive landscape for enterprise AI services. However, reliance on internal metrics introduces verification risk until third-party audits confirm the results.
Generated by WOOFUN AI · For reference only, not investment advice

Comments

Me
Replying to @User
0/800

No comments yet.

Notifications

Sign in to view messages
View all messagesManage subscriptions