HAVE AI NEWS HAVE AI NEWS
ru

Research

News in category “Research”

Remote Labor Index Benchmark Shows Modern AI Agents Can Handle Only 2.5% of Real Freelance Tasks

Researchers from the Center for AI Safety and Scale AI have developed the RLI benchmark based on hundreds of real Upwork projects. Even leading language models and agents were able to successfully complete no more than 2.5% of the jobs.

Too Good to Be Bad: Study Shows Why AI Struggles to Convincingly Roleplay Villains

Researchers from Tencent and Sun Yat-sen University have found that built-in safety mechanisms prevent large language models from effectively roleplaying negative characters, replacing subtle manipulation with generic aggression.

ClawBench Benchmark Exposes AI Agent Weaknesses: Top Real-Web Success Rate Reaches Just 33%

The new ClawBench benchmark evaluated autonomous web agents on live websites. The best-performing model successfully completed only one-third of the practical tasks.