AI agents can't yet do open-ended AI research
Early evidence from two case studies
- AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
- Early evidence from two case studies
Source profile
AI News Hub tracks AI Snake Oil AI updates with visible source status, reuse boundaries, collection method, and published articles.
AI analysis newsletter; summary-only unless authorization is obtained.
Early evidence from two case studies
This is the keynote from ICML 2025, arguing that AI should be viewed as a 'normal technology' whose impacts unfold gradually through invention, innovation, diffusion, and adaptation. While recursive self-improvement is a serious possibility, it won't suddenly render everyone jobless. The future of work will require radical adaptation and human-AI 'co-superintelligence'.
This article argues that despite fears, AI has not led to mass layoffs in software engineering. It presents evidence that layoffs attributed to AI are often financial in nature, and that AI compresses execution but not decision-making and delivery. The 'decide-execute-deliver sandwich' model explains why coding agents haven't displaced workers: the bottlenecks are deciding, verifying, and deep understanding.
Google claimed its AI agents built an entire operating system with a single prompt and about $900 in API costs, but this analysis highlights multiple issues: the prompt was actually thousands of lines long, the scaffold may be overfitted, and critical details like code, logs, and methodology are missing. The article underscores the need for independent evaluation and proposes norms for 'open-world evaluations'.
Let’s not skip the hard work of AI governance
Introducing CRUX, a collaborative project that conducts open-world evaluations—long, real-world tasks—to measure frontier AI capabilities. The first experiment shows an AI agent autonomously publishing an iOS app, highlighting both progress and risks like app store spam.
Researchers propose a framework to measure AI agent reliability, decomposing it into 12 dimensions across four categories. Testing 14 models over 18 months reveals rapid capability improvements but only modest reliability gains, calling for reliability-specific optimization.
Applying the AI as Normal Technology framework to legal services, this article argues that advanced AI will not by default help consumers achieve desired legal outcomes at lower costs due to three bottlenecks: regulatory barriers, adversarial dynamics, and human involvement. It also discusses potential institutional reforms.
This famous aphorism is neither true nor useful
This article explores the 'AI as Normal Technology' framework, contrasts it with AI 2027, addresses common confusions, and discusses the slow diffusion and adoption challenges of AI.
AI might slow scientific progress by exacerbating the production-progress paradox, introducing software errors, entrenching flawed theories, and undermining human understanding. The article calls for reforms in incentives, meta-science investment, and AI tool design.
Artificial General Intelligence (AGI) is not a milestone because it does not represent a discontinuity in AI properties or impacts. AGI definitions are vague and unobservable, economic impacts take decades through diffusion, and risks stem from design choices rather than capabilities. Businesses and policymakers should focus on gradual diffusion rather than chasing AGI.
A new paper argues that AI should be viewed as a normal technology, not as a superintelligent entity. It emphasizes slow adoption, gradual economic impact, and the importance of human control, contrasting with utopian/dystopian narratives.
The article examines the debate over whether AI capability progress is slowing. Authors argue model scaling isn't dead, insider predictions are unreliable, inference scaling has promise but limits, and capability gains weakly translate to real-world impact due to product and adoption lags.
An analysis of AI use in 2024 global elections reveals that over half of deepfakes lack deceptive intent, and most deceptive content can be cheaply replicated without AI. Misinformation spreads due to demand, not supply.