DR-DCI is a retriever-steered Direct Corpus Interaction (DCI) framework that treats retrieval as an agent-callable action to dynamically expand a local workspace, achieving scalable and precise evidence resolution. Experiments show up to 73.3% accuracy on Browsecomp-Plus, outperforming raw DCI and BM25, and scaling stably to 20M documents.
Live AI News Intelligence
Live monitoring
Live updates
Trusted sources, attribution, rights, and in-site reading distilled into a signal-first AI brief.
Live updates
This paper proposes a definition of good explanations inspired by counterfactual explanations, incorporating the interlocutor's prior beliefs, and explores its implications for AI explainability, particularly why LLM outputs are difficult to explain well.
Fog is a notes app for Apple devices that uses on-device AI to automatically organize notes into smart collections called Clouds. It ensures privacy with no third-party servers and syncs via iCloud.
Europe's reliance on foreign AI models, especially after the US suspension of access to Anthropic's Fable series, exposes the myth that Europe can just use AI applications without building the underlying models. The article argues that frontier model-building is a continuous practice, and Europe lacks true ecosystem and expertise.
Cybersecurity expert Katie Moussouris revealed that Anthropic shared a White House report on the Fable jailbreak with her. The report showed that Fable refused to review code for security issues but complied when asked to fix the code, which Moussouris considered the model working as intended for cyberdefense.
As the rest of the country celebrated the USA's first World Cup win and the New York Knicks championship, Anthropic spent its weekend fighting the Trump administration over its latest model release. At 5:21 PM on Friday, the company received a US export control directive to suspend access to its Mythos 5 and Fable 5 AI models by "any foreign national" inside or outside the US, "including foreign national Anthropic employees." The only way that was possible, Anthropic determined, was to completely disable products it spent the past week hyping - and travel to Washington, DC in hopes of changing President Donald Trump's mind. Now, over the coming days, the US government could dramatically alter the trajectory of the entire industry, dealing a major blow to American AI companies.
This article introduces predictive data debugging, a method to accurately predict which behaviors reinforcement learning will amplify or suppress before training, trace them back to the responsible data, and reshape the dataset or training process to prevent undesired effects. Case studies demonstrate its effectiveness in identifying and fixing data issues, including safety guardrail degradation, hallucinated links, physics sycophancy, and unexpected behaviors. Validation with a 'goblin mode' ground truth confirms reliability.
Microsoft CEO Satya Nadella published a blockbuster article and X post about building a 'frontier ecosystem' over a 'frontier model,' introducing 'Loopcraft' as a new theory of the firm. Meanwhile, the Anthropic Fable/Mythos export-control crisis pushes industry toward model neutrality and own-your-stack architecture. Other highlights include agent systems moving to production, inference efficiency gains, and commercial agent launches.
MIT physicist and AI researcher Max Tegmark separates the real opportunities and threats from the myths, describing the concrete steps we should take today to ensure that AI ends up being the best — rather than worst — thing to ever happen to humanity.
Deep-XPIA is a new benchmark for evaluating the robustness of multi-agent AI systems against prompt injection attacks through multi-hop cross-prompt challenges.
A GPT tool called BoomURL that offers free URLs for AI-generated content, with the standard ChatGPT disclaimer.
Interlatent is a reinforcement learning platform for physical AI.
whoburnedmore is a local CLI tool that reads Claude Code session history to display token and cost breakdowns by model and project, featuring prompt cache insights and an optional HTML dashboard. It is privacy-focused, making no network requests and never exposing user data.
NocoBase doubled revenue in first five months of 2026 year-over-year, after a major pivot from a no-code platform to AI + no-code infrastructure. The article recounts the team's journey from panic to repositioning, emphasizing that enterprise-grade software still needs standardized architecture, visual configuration, and robust design in the AI era.
whoburnedmore is a free, open-source tool that generates a dashboard of your AI coding usage across Claude Code, Codex, Cursor, and 12+ other tools, then places you on a live public leaderboard for comparison.
Simon Willison uses Cloudflare's Managed Challenge to protect his faceted search from aggressive crawlers, but even simple ?q=term searches triggered the challenge. Using Claude Code, he discovered a rule that only triggers CAPTCHA for search URLs containing at least one ampersand, allowing simple searches to pass through without challenge.
OpenAI introduces Deployment Simulation, a method to predict AI model behavior before deployment using real conversation data to improve safety and evaluation accuracy.
Databricks recognized over 60 consulting, system integrator, and ISV partners at the Data + AI Summit across global, regional, industry, product, and champion categories. The awards highlight a focus on AI transformation, lakehouse modernization, Unity Catalog governance, and agentic AI at enterprise scale. Notable winners include Accenture, Deloitte, NVIDIA, Anthropic, and Salesforce.
Whissle's META-1 is a meta-aware ASR model that simultaneously outputs transcription and metadata (emotion, intent, age, gender, etc.) in a single forward pass at ~200ms latency. By integrating KenLM n-gram language models, it reduces word error rates by up to 3.6% absolute (10.8% relative) across four languages, while extracting metadata 9x faster than commercial alternatives like Deepgram, AssemblyAI, and Gemini 2.0 Flash.
HashMeterAi is a local-first, private usage meter for AI coding tools. It tracks token usage across multiple tools like Claude Code, Codex, Kimi, Qwen CLI, and others, providing a unified dashboard with metrics like cost, processed tokens, focus time, and AI persona. It is 100% offline and privacy-focused, never sending data. Supports several tools via local transcripts.
Prism is a native macOS AI companion that supports multiple AI model providers, featuring a Quick AI panel, browser automation, study tools, and a focus on privacy. It offers a permanent free tier and paid plans, now launching on Product Hunt.
Meta's chief technology officer acknowledged that the restructuring of the Applied AI engineering unit of about 6,500 employees was poorly communicated and executed, leading to a loss of trust and morale. He promised improvements in management, caps on direct reports, and AI coaching tools, while also allowing drafted employees to seek other roles within the company.
Shelly runs the OpenAI Codex CLI natively on Android — no PC, Termux, or proot needed. It features a native terminal, an Agent Chat pane for one-tap fixes, and a home-screen widget showing quota, cost, and rate limits. Local LLMs are supported. Open source under GPLv3, built entirely by directing AI agents.
An AI chatbot cracked an 80-year-old geometry problem after a single prompt from OpenAI mathematicians.
Katalyst is an AI sales agent for Salesforce teams that automatically summarizes calls, updates records, drafts follow-ups, and surfaces key signals, saving reps time and improving pipeline management. It integrates natively with Salesforce and offers a free trial.
A test of Google's Gemini 3 Flash Preview as an autonomous AI agent showed that 67% of generated curl commands targeted unsafe internal networks or metadata endpoints. All dangerous commands were blocked by Check, a preflight security tool. The test highlights the risk of AI agents executing commands without guardrails.
A new study finds that people frequently use AI for simple tasks even when it is inefficient. The research reveals miscalibration at two levels: underestimation of actual AI use and overestimation of its benefits, along with a carryover effect that entrenches these biases.
EU AI Act obligations for high-risk systems hit in August 2026. Stateless agent frameworks can't satisfy them. This guide covers seven types of state compliant agents must maintain, four streaming patterns for auditability, and a reference architecture using Kafka and Flink as the control plane.
Tokyo-based Sakana AI launched its first commercial product, Sakana Marlin, an autonomous research agent for enterprises. Each task runs up to eight hours, producing reports of dozens to 100 pages with slides. It leverages AB-MCTS and AI Scientist workflows. Pricing starts at pay-as-you-go with 100 credits per run at ¥98 per credit.
Internal messages reveal Meta employees' frustration with Zuckerberg's planned companywide AI hackathon, citing lack of time due to layoffs, low morale, and distrust in management. The hackathon, set for July 14-16, excludes performance evaluation credit, further fueling discontent.