Research
News in category “Research”
Remote Labor Index Benchmark Shows Modern AI Agents Can Handle Only 2.5% of Real Freelance Tasks
Researchers from the Center for AI Safety and Scale AI have developed the RLI benchmark based on hundreds of real Upwork projects. Even leading language models and agents were able to successfully complete no more than 2.5% of the jobs.
Too Good to Be Bad: Study Shows Why AI Struggles to Convincingly Roleplay Villains
Researchers from Tencent and Sun Yat-sen University have found that built-in safety mechanisms prevent large language models from effectively roleplaying negative characters, replacing subtle manipulation with generic aggression.
ClawBench Benchmark Exposes AI Agent Weaknesses: Top Real-Web Success Rate Reaches Just 33%
The new ClawBench benchmark evaluated autonomous web agents on live websites. The best-performing model successfully completed only one-third of the practical tasks.