AC-Small improved significantly on held-out benchmarks
AC-Small improved significantly on held-out benchmarks after post-training on the APEX-Agents dev set, with +5.7pp on APEX, +8.0pp on Toolathalon, and +7.7pp on GDPval. This improvement demonstrates the effectiveness of fine-tuning on specific datasets. The results have implications for AI model development and evaluation.
Read the full story on mercor.com→Claude Sonnet 4.6 released
Claude Sonnet 4.6 is the latest version of the Claude AI model, with improved performance and capabilities. The new version is expected to have significant impacts on the AI community and industry. However, details about the release are limited, and more information is needed to understand its full potential.
Read the full story on Anthropic→AI alignment researchers automate themselves
AI alignment researchers are increasingly turning to automation to address the challenge of safely aligning superhuman AI systems. As human capabilities may soon be insufficient, automation is seen as a necessary step to ensure the safe development of AI. This trend has significant implications for the future of AI research and development.
Read the full story on transformernews.ai→Will the AI data centre boom become a $9T bust?
The AI data centre boom is expected to have significant economic implications, with some estimates suggesting it could become a $9T bust. The article discusses the potential risks and challenges associated with the AI data centre boom, including energy consumption and environmental impact. As the demand for AI computing power continues to grow, the industry must address these challenges to ensure sustainable development.
Read the full story on ft.com→