#Benchmarks
Remote Labor Index Benchmark Shows Modern AI Agents Can Handle Only 2.5% of Real Freelance Tasks
Researchers from the Center for AI Safety and Scale AI have developed the RLI benchmark based on hundreds of real Upwork projects. Even leading language models and agents were able to successfully complete no more than 2.5% of the jobs.
Too Good to Be Bad: Study Shows Why AI Struggles to Convincingly Roleplay Villains
Researchers from Tencent and Sun Yat-sen University have found that built-in safety mechanisms prevent large language models from effectively roleplaying negative characters, replacing subtle manipulation with generic aggression.
GLM-5 Released: Zhipu AI's Open Model Reaches Top-Tier Performance in Agentic Coding
Zhipu AI and Tsinghua University have introduced GLM-5, a powerful open MoE model with 744 billion parameters capable of autonomously writing code, conducting deep web search, and tackling complex multi-step tasks on par with proprietary flagships.
ClawBench Benchmark Exposes AI Agent Weaknesses: Top Real-Web Success Rate Reaches Just 33%
The new ClawBench benchmark evaluated autonomous web agents on live websites. The best-performing model successfully completed only one-third of the practical tasks.
OpenAI Unveils First Custom AI Chip 'Jalapeño' to Challenge Nvidia on Inference
OpenAI has published the first performance benchmarks for Jalapeño, its in-house inference chip developed alongside Broadcom, reporting significant gains in speed and power efficiency over Nvidia hardware.