#Claude
ClawBench Benchmark Exposes AI Agent Weaknesses: Top Real-Web Success Rate Reaches Just 33%
The new ClawBench benchmark evaluated autonomous web agents on live websites. The best-performing model successfully completed only one-third of the practical tasks.
Anthropic Releases Claude Sonnet 5: Focus on Autonomous Agents, Coding, and Hidden Costs
Anthropic has introduced Claude Sonnet 5, a new language model optimized for multi-step autonomous workflows and coding. Despite attractive base pricing, reasoning patterns and a redesigned tokenizer may make real-world usage more expensive than expected.