This study examines how mathematicians use AI for formalizing mathematical proofs. Through surveys and a controlled user study, it finds that people desire AI assistance while retaining high-level control, using AI improves formalization accuracy, and users tend to employ multiple AI tools flexibly.
Live AI News Intelligence
Live monitoring
Live updates
Trusted sources, attribution, rights, and in-site reading distilled into a signal-first AI brief.
Live updates
A new benchmark, Curation-Bench, tests whether generalist coding agents can autonomously curate training data. While out-of-the-box agents match published baselines within ten iterations, they tend to tune local variants rather than explore new policy families. Scaffolded agents that require citing and adapting prior methods autonomously discover a policy that outperforms baselines at one-tenth the data budget, suggesting that structured method adaptation is key for reliable data curation automation.
StepPRM-RTL is a novel framework that combines stepwise trajectory modeling, process-reward modeling (PRM), and retrieval-augmented fine-tuning (RAFT) to enhance LLM-based RTL code generation. It achieves over 10% improvement in functional correctness and reasoning fidelity over prior methods on Verilog and VHDL benchmarks.
Multimodal large language models are increasingly capable of complex reasoning, yet their performance often degrades when they must externalize a problem through a tool and then reason over the tool's output, specifically when relying on visual aids. To study this discrepancy, we introduce VAMPS (Visual-Assisted Mathematical Problem Solving), a benchmark for graph-assisted mathematics. VAMPS contains 1,168 multimodal, bilingual multiple-choice question-answer pairs drawn from Iranian University Entrance Exam algebra and calculus problems, expanded with human-reviewed LLM-generated synthetic variants. Overall, we found that direct analytical solving surprisingly outperforms tool-enabled visual solving, even on problems where plotting is a natural strategy.
SMAC-Talk extends the StarCraft Multi-Agent Challenge with a natural language communication channel to evaluate LLM-based agents in cooperative multi-agent settings. It features decentralized control, partial observability, long-horizon decision making, and scenarios with deceptive communicators. Benchmarking using Qwen3.5 models reveals how reasoning, memory, and scale affect coordination.
LLMs are reshaping research while eroding epistemic accountability. The PEEL framework combines deterministic distant reading with LLM interpretation, grounded in Peircean semiotics, to reveal systematic distortions in AI-generated text. Key implications: deterministic tools must accompany AI, fluency does not equal fidelity, and epistemic authority must be designed in.
This paper challenges the common assumption that AI emotional support is a deliberate act. It argues that AI emotional support often emerges incidentally during task-oriented interactions on general-purpose platforms, and these positive experiences can path-dependently shift preferences away from humans toward AI. A 28-day longitudinal study with OpenAI found that daily 5-minute AI conversations led to a 10.3% decrease in preference for human support and an 11.6% increase for AI. The authors conclude that current policies focused on companion apps are inadequate and must be extended to general-purpose AI systems to address cumulative behavioral changes.
This paper proposes an ontology-grounded verification framework for pre-deployment assurance of enterprise AI agents. The framework combines an Agent Operational Envelope, an ontology-to-scenario generation pipeline, and a Trust Certificate. A pilot in four regulated industries across the US and Vietnam generated 1,800 scenarios and showed ontology-grounded generation achieved 48.3% regulatory coverage versus 33.1% for a persona-based baseline, with the highest domain specificity.
The European Commission's new EU Cloud and AI Development Act (CADA) introduces an ambitious 'open source first' principle, a EUR 2 billion investment, and a sovereignty certification framework. However, it falls short of a binding requirement for open source over proprietary software, lacking an enforceable procedural mandate.
OpenAI CEO Sam Altman is offering $2 million in OpenAI tokens to every startup in Y Combinator's current batch in exchange for equity, marking a pilot program for spring and summer 2026. The deal uses an uncapped SAFE without an MFN provision, reflecting the rising importance of compute tokens in AI startup economics.
Today's AI news highlights include Microsoft's MAI-Thinking-1 technical report with unprecedented transparency, the release of Gemma 4 12B open multimodal model, Ideogram 4.0 going open weights, and various developments in AI agents, model routing, and cost controls.
Explore why model neutrality is critical for AI agents. Learn how labs lock you in at the harness layer—and why a neutral, open-source framework is the answer.
A systematic review of 97 studies categorizes large AI models in dentistry into language-generative, discriminative vision foundation, and dental-specific foundation models, highlighting their complementary roles and three key barriers: hallucination, data scarcity, and lack of standardized benchmarks.
According to The Information, Meta is developing an AI pendant and plans to release up to four new smart glasses models before the end of the year, in an aggressive push to offset massive losses from its Reality Labs division. The company is also launching a business subscription service called 'Wearables for Work.'
AI Gauge is an open-source desktop widget that monitors usage limits for Claude, ChatGPT Codex, GitHub Copilot, and OpenRouter. It displays session and weekly usage, reset times, account balances, and spend in a compact always-visible view. Supports Windows, macOS, and Linux with platform-native UIs.
Boxes.dev is a cloud-only agentic dev environment (ADE) that gives each Claude Code and Codex agent its own cloud computer. It allows coding from any device with full-featured desktop, CLI, and mobile apps, using your existing subscription. Features include easy setup, parallel agent workflows, scheduled automations, Slack integration, and 10 free box-hours for new users.
Piece is a mobile app that helps couples settle everyday arguments. Each person records their side, AI weighs both perspectives fairly, and one of you apologises. It's free, takes 5-10 minutes, and prioritises privacy.
JackHamr is a cloud platform that provides hosted environments, specialist AI agents, and full pipeline orchestration to help teams ship software faster. Agents have personality and autonomy, handling end-to-end tasks from spec to release. Developers interact via chat or voice, and the platform supports custom LLMs, skills, and flexible resource configurations.
At UC Berkeley, CS 10 and CS 61A had abnormally high failure rates in spring 2026, attributed by professors to increased AI usage and declining math skills. Students are over-relying on AI and underprepared, leading to calls for reinstating standardized tests and changes in teaching.
An action plan for AI-powered biological resilience
NVIDIA Nemotron 3 Ultra is a 550 billion parameter (55B active) open model designed for long-running agentic workflows, with 1M token context and NVFP4 optimization, leading in agentic benchmarks and cost efficiency.
Hugging Face has redesigned its hf CLI to be optimized for both human users and AI coding agents. The CLI automatically detects agent environments and adapts its output format, provides next-command hints, and ensures non-blocking, idempotent operations. Benchmarks show agents using the hf CLI consume up to 6× fewer tokens on complex tasks and achieve higher success rates compared to using curl or the Python SDK directly.
Microsoft showcased its own AI models at Build, Florida sued OpenAI and went after Sam Altman personally, new research and Workday products showed no one trusts AI agents yet, and Alphabet raised a record $85 billion as the Fed flagged AI as a systemic risk. Money moves faster than trust.
Lookspan is a local-first observability dashboard for AI agents, supporting MCP, LangGraph, CrewAI, and OpenTelemetry. All data stays in local SQLite, no cloud required. Features include real-time tracing, cost tracking, alerts, replay evaluation, and dataset experiments. Launch with one command.
The article explores responsible ways to use AI in writing, including using LLMs for editing drafts, as a self-study tool, and generating lists. It warns about models' tendency to pander and advises users to actively counteract it.
This NBER working paper by Mert Demirer, Leon Musolff, and Liyuan Yang explores the productivity effects of AI coding tools across different generations. It distinguishes between the writing code and shipping code phases, providing insights into how AI tools impact software development efficiency.
ChatGPT has evolved from a simple chatbot into a versatile tool for writing, research, image generation, file analysis, and more. This guide covers getting started, free vs. paid plans, and essential features like web search, deep research, file uploads, app integrations, image creation, GPTs, projects, voice mode, and memory.
DNS-AID leverages existing DNS infrastructure for AI agent discovery, publishing, and verification without new overlay networks, using DNSSEC for trust. It supports MCP, A2A, and HTTPS, and integrates into existing DNS zones.
Google has released the Gemma 4 12B model, a 12-billion-parameter AI model that can run on consumer laptops with 16GB of RAM, filling a gap in the Gemma 4 lineup between mobile and high-performance models.
MIT researchers use the classic game as a test bed for AI agents, finding a small AI model can outperform the biggest ones at 1 percent of the cost.