📰 News
AI News — page 16
Funding, launches, analysis, and interviews.
Frontier AI Models Compete Then Collaborate to Train Better Coding Students
New research shows four leading AI models working together through competition and collaboration can improve smaller coding models more effectively than traditional imitation learning approaches.
Chain-of-Thought Monitoring Vulnerable to Persuasion Attacks, AI Safety Research Shows
New research reveals that AI safety monitoring systems using chain-of-thought reasoning can be manipulated by adversarial agents, with harmful action approval rates increasing by 9.5% when monitors access reasoning traces.
AutoPersonas Research Tackles AI Agent Self-Locking in Long-Term Simulations
New research introduces AutoPersonas, a multi-timescale engine designed to prevent AI persona agents from falling into repetitive behavioral loops during extended simulations.
Researchers Publish Mathematical Theory of Slow Thinking for Large Language Models
Four researchers have published a first-principles mathematical framework called 'active lifting' that formally derives slow thinking processes in large language models and active perception systems.
Researchers Launch ZendoWorld Benchmark to Test AI Visual Reasoning
German and US researchers created ZendoWorld, a new benchmark that challenges AI agents to infer logical rules from visual scenes and design experiments to test hypotheses.
CausalDS Benchmark Tests AI Agents' Causal Reasoning in Data Science
Researchers introduce CausalDS, a new benchmark evaluating how well AI agents perform causal reasoning tasks within realistic data science workflows using synthetic datasets.
OpenAI faces GPT-5.5 reasoning token clustering bug in Codex
A GitHub issue reveals GPT-5.5 responses in OpenAI's Codex disproportionately cluster at exactly 516 reasoning tokens, potentially causing degraded performance on complex coding tasks.
Anthropic's Latest Models Show Worse Tool-Calling Performance Than Predecessors
Claude Opus 4.8 and Sonnet 5 generate malformed tool calls more frequently than older Anthropic models, suggesting post-training on forgiving harnesses may hurt adaptation to strict schemas.
ByteDance readies Seedance 2.5 with 3-minute AI video generation
ByteDance's Dreamina Seedance 2.5 will launch in July with 3-minute AI video output, extending from current 15-second clips to longer commercial sequences.
How AI-powered smartwatches detect early illness signs before symptoms appear
Wearables excel at spotting deviations from baseline health patterns, with AI increasingly helping combine multiple sensor readings to flag potential infections hours before users feel sick.
Mistral AI emerges as Europe's AI champion amid US regulatory uncertainty
French AI company Mistral AI gains attention as governments seek alternatives to US-controlled AI systems, with revenues climbing from $20 million to over $400 million annually.
Anthropic Faces Session Cache Bug Report Alleging Cross-Account Data Leakage
A GitHub issue filed against Anthropic's Claude Code suggests potential session cache leakage between enterprise and consumer accounts, raising security concerns about workspace isolation.
Judge denies xAI request to block Minnesota nudify app ban
A federal judge rejected xAI's emergency request to halt Minnesota's first-in-nation ban on AI apps that create non-consensual nude images, citing the company's delayed legal challenge.
Lean Theorem Prover Fixes Critical Soundness Bug After AI-Generated False Proof
The Lean theorem prover patched a kernel soundness bug within hours after researchers used AI to generate a false proof of the Collatz conjecture, exposing flaws in nested inductive type handling.
xAI adds character references and 1080p to Imagine Video 1.5
xAI upgraded Imagine Video 1.5 with image and voice references, prompt-only generation, and native 1080p output for Grok users.
Thinking Machines releases open-weight Inkling-Small with 12B active parameters
Thinking Machines Lab launched Inkling-Small, a 276B parameter mixture-of-experts model with 12B active parameters, offering comparable performance to Inkling at quarter the size with full open weights.