HAVE AI NEWS HAVE AI NEWS
ru

#LLM

Remote Labor Index Benchmark Shows Modern AI Agents Can Handle Only 2.5% of Real Freelance Tasks

Researchers from the Center for AI Safety and Scale AI have developed the RLI benchmark based on hundreds of real Upwork projects. Even leading language models and agents were able to successfully complete no more than 2.5% of the jobs.

Too Good to Be Bad: Study Shows Why AI Struggles to Convincingly Roleplay Villains

Researchers from Tencent and Sun Yat-sen University have found that built-in safety mechanisms prevent large language models from effectively roleplaying negative characters, replacing subtle manipulation with generic aggression.

GLM-5 Released: Zhipu AI's Open Model Reaches Top-Tier Performance in Agentic Coding

Zhipu AI and Tsinghua University have introduced GLM-5, a powerful open MoE model with 744 billion parameters capable of autonomously writing code, conducting deep web search, and tackling complex multi-step tasks on par with proprietary flagships.

ClawBench Benchmark Exposes AI Agent Weaknesses: Top Real-Web Success Rate Reaches Just 33%

The new ClawBench benchmark evaluated autonomous web agents on live websites. The best-performing model successfully completed only one-third of the practical tasks.

Anthropic Releases Claude Sonnet 5: Focus on Autonomous Agents, Coding, and Hidden Costs

Anthropic has introduced Claude Sonnet 5, a new language model optimized for multi-step autonomous workflows and coding. Despite attractive base pricing, reasoning patterns and a redesigned tokenizer may make real-world usage more expensive than expected.

The Race for a Personal AI Companion: Strategies of OpenAI, Apple, Yandex, and Meta

Tech giants are moving from isolated chatbots to end-to-end personal AI assistants. We explore how wearables are evolving and why user digital memory will become the ultimate asset in this battle.