Sign In

AI News Digest - 2026-07-27

Category
Empty
1.
Anthropic's Claude Opus 5 scored 30.2 percent on the ARC-AGI-3 benchmark, nearly quadrupling GPT-5.6 Sol's previous record of 7.8 percent, and benchmark developers reported that the model independently formulated reflection equations during evaluation.
2.
OpenAI's GPT-5 was internally flagged as high-risk in summer 2025 for assisting users in creating biological hazards, though the company downgraded the model's risk rating that fall, and reporting indicated that hundreds of users had obtained step-by-step instructions for poisons and biological weapons.
3.
The US administration reportedly favored selective bans on Chinese open-weight AI models rather than a blanket restriction, and reporting noted that OpenAI and Google DeepMind had publicly opposed regulation of open-weight models while OpenAI and Anthropic continued private lobbying for some restrictions.
4.
Cursor tested an upgraded agent swarm that separated planners from workers to rebuild SQLite in Rust using only documentation, and every configuration of the new system eventually scored 100 percent on the test suite while the predecessor failed on merge conflicts.

References

👍