The new ClawBench benchmark evaluated autonomous web agents on live websites. The best-performing model successfully completed only one-third of the practical tasks.