MacArena is a new benchmark of 421 manually verified tasks across 50 applications for evaluating computer-use agents (CUAs) on macOS. It combines curated OSWorld tasks, content from macOSWorld, and 49 new macOS-native tasks, running on Apple's native Virtualization framework on Apple Silicon. Evaluation reveals that macOS presents distinct GUI challenges, and model rankings invert between ported and macOS-native tasks, with a leading model trailing by over 26% on the MacArena subset, suggesting that macOS is a genuinely harder environment for current GUI agents.
Live AI News Intelligence
Live monitoring
Live updates
Trusted sources, attribution, rights, and in-site reading distilled into a signal-first AI brief.
Live updates
New research proposes MSFAN deep learning architecture, combined with terahertz dual-comb spectroscopy, to classify 12 types of polymers with 85.2% accuracy, promising to improve plastic recycling sorting efficiency.
Diffusion Large Language Models (dLLMs) refine tokens iteratively but commit them irreversibly, causing a "stability lag" where early decisions remain fragile. Post-Training Quantization (PTQ) error easily flips these borderline decisions at the write frontier, locking them in. FAIR-Calib, a two-stage PTQ framework, probes a full-precision teacher for a position prior and performs off-policy layer-wise calibration with a reweighted hidden-state MSE, protecting fragile frontier states without expensive end-to-end rollouts. Theoretically justified as a surrogate for output KL divergence, FAIR-Calib outperforms baselines on LLaDA and Dream (W4A4), reducing frontier flips and post-commit mismatches.
This paper introduces Elmes*, an end-to-end framework for automatically constructing, refining, and applying fine-grained evaluation rubrics for LLMs in educational contexts. It combines a multi-agent engine with a self-evolving module SceneGen to co-optimize evaluation criteria and test data. The resulting Edu-330 benchmark covers 330 scenarios across 11 subjects, 3 grade levels, and 10 task types with over 1,000 indicators. Experiments reveal that educational capability is multidimensional: top LLMs differ mainly in creativity and values integration, knowledge-strong models may fail at Socratic scaffolding, and the education-specialized InnoSpark achieves the best human-evaluated score. LLM judges preserve human-comparable rankings but exhibit biases like self-preference. The framework provides scalable diagnostic infrastructure for pedagogy-grounded LLM evaluation.
This paper studies parallel Continuous Local Search (CLS) on Boolean satisfiability problems with symmetric pseudo-Boolean constraints. By relaxing the problem to continuous optimization on a hypercube, experiments reveal that redundant constraints inhibit convergence; CLS can quickly complete partial assignments in hybrid settings; and local search rapidly converges to a stable solution quality distribution due to saddle-dense objectives. These findings inform practical uses of CLS for SAT on modern accelerator hardware.
We present Accelerated Fourier SAT (AFSAT), a GPU-accelerated solver for pseudo-Boolean satisfiability based on continuous local search (CLS). AFSAT realises the proof-of-concept approach, FastFourierSAT, into a fully-engineered solver supporting any heterogeneous mixture of symmetric constraint types and lengths within a single problem instance. Using the JAX compiler, AFSAT leverages pure function composition, automatic vectorisation, automatic differentiation, and just-in-time (JIT) compilation to perform massively parallel CLS across batches of candidate assignments. We demonstrate substantially improved numerical stability, runtime performance, and memory efficiency over the proof-of-concept. We achieve this by way of identifying and addressing various limitations that arise from memory latency and floating-point representation, as well as leveraging automatic parallelisation and compact representations. The inherent representational and stability limitations of floating point are partially addressed by a tailored discrete Fourier transform implementation. We achieve near-linear throughput when scaling to multiple accelerators via JAX array sharding.
This position paper argues that AI research should focus on studying the training dynamics that produce model behavior, rather than only analyzing final models post-hoc. It calls for a science of AI that can predict, intervene, and design training procedures to reliably achieve desired properties.
CARVE-Q is a quantum-AI hybrid architecture for certified repair of vetoed driving maneuvers. It uses quantum minimum finding to accelerate repair enumeration while keeping safety certification classical, achieving provable speedup and 100% compliance.
A new study shows that if attackers strategically choose when to attack, the measured safety of AI control evaluations drops significantly. The researchers decompose attack decisions into start and stop policies, and demonstrate on BashArena and LinuxArena that both policies reduce safety by 20 to 28 percentage points at a 1% audit budget, without changing the underlying attack capability. Existing evaluations may overestimate safety, and the study recommends incorporating attack selection for more realistic estimates.
Large language models have made substantial progress on mathematical reasoning, but existing benchmarks typically evaluate well-specified problems with final answers or complete proofs, missing collaborative open-problem solving. CrowdMath is a dataset of 164 expert-annotated progress chains from the MIT PRIMES-AoPS CrowdMath program (2016-2025). Each chain tracks multi-participant forum discussions from problem statement to completed proof, with posts labeled by functional roles. Six frontier models achieve 83-88% accuracy on next-post prediction but only 0.42 macro-F1 on post-role classification, highlighting a gap in understanding collaborative mathematical progress.
This paper proposes Lean4Agent, the first framework that uses Lean4, a dependent-type formal language, to model and verify LLM agent behavior. It includes FormalAgentLib for verification and LeanEvolve for workflow revision. Experiments show verification-passing workflows outperform failing ones by 11.94%, and LeanEvolve further improves SWE performance by 7.47%.
Existing Sudoku solvers either lack correctness guarantees (learning-based) or suffer from long-tail search (symbolic). DiBS combines a complete symbolic solver with a diffusion model for branch ordering, reducing search cost significantly. Theoretical proofs and experiments on the Royle 17-clue benchmark show effectiveness in nodes, backtracks, and long-tail percentiles. Code is open-source.
This paper formalizes bias as a symmetry breaking operation and uses loss-based regularization to restore symmetry, achieving over 90% violation reduction with only about 5% accuracy cost on synthetic datasets. The framework requires no causal graph knowledge, is computationally light, and generalizes to any bit-flip definable sensitive attribute.
Cerebro's whitepaper proposes that small operators can run AI-native companies by implementing the company itself as infrastructure. All aspects—memory, decisions, workflows, agents, models, and governance—are versioned, inspectable, and operator-owned. It is a companion to Libro, an open-source Ops scaffold for Claude Code.
This post shares a custom Claude Code status line that displays context usage, token count, code churn, and rate limit budgets at a glance. Includes a shell script and setup instructions.
NVIDIA and LG Group are building an AI factory to accelerate LG's AI-driven businesses in robotics, autonomous driving, data center technologies, and GPU cloud services. The collaboration integrates NVIDIA's full-stack AI factory platform with LG's leadership in consumer electronics and robotics, aiming to create a unified workflow for physical AI systems.
TeardownHQ offers verified revenue data and deep-research playbooks for indie SaaS founders, detailing how startups grew through channels, pricing, and go-to-market strategies. The directory is free; full teardowns are paid. A founding member lifetime deal is available for the first 500.
Solarch creates interactive diagrams with AI, keeping your code always in sync.
Preseason.ai is an open-source benchmark that tracks which tools AI models pick across a frozen panel of vibe-coding prompts at every level, from beginners to expert engineers. The platform ranks tools for various advanced scenarios and provides direct comparisons between popular options.
PandaProbe Cloud is a fully managed agent engineering platform.
A vision for the future of AI, focusing on access, safety, and shared prosperity as OpenAI works to ensure AGI benefits everyone.
A vision for the future of AI, focusing on access, safety, and shared prosperity as OpenAI works to ensure AGI benefits everyone.
Every reader deserves to be informed about whether what they are reading is human or AI. A few weeks ago, Dr Kylie Moore-Gilbert, an academic in political science at Macquarie University, wrote an opinion piece in the Sydney Morning Herald in which she reported on excessive use of AI chatbots by students to write their essays. In it, she raised her concern that universities are qualifying lawyers, nurses, financial advisers, engineers and teachers who do not have the essential skills required to perform their roles. If that is the case, the societal consequences are obvious.
EMILIAProtocol defines a portable, vendor-neutral protocol for trust evaluation and pre-action trust enforcement in high-risk workflows. It answers the specific question of whether a particular action should be allowed under given policy, actor, and context, using core objects like Trust Receipt, Trust Profile, and Trust Decision, along with Handshake and Accountable Signoff for enforcement.
AI Code Stitcher is an open-source tool that intelligently applies AI-generated code to existing codebases. The latest v1.74 introduces auto-update and import hoisting, while reinforcing a philosophy that prioritizes user control over autonomous AI agents.
Apple announces its third generation of Foundation Models, a family of five models built with Google, including on-device and server-based models, with a focus on privacy and new architectures like sparsely activated models and instruction-following pruning. The models power new Siri and intelligent tools, and show significant quality improvements in evaluations.
OpenAI launches the Economic Research Exchange to study AI’s impact on jobs, productivity, and the economy. Applications are now open for selected research projects.
OpenEnv is a tool for creating an agentic execution environment like terminals, browsers, or anything an agent can interact with. Today, we’re excited to announce that OpenEnv is becoming even more open, to make the future of training agents open source. Starting today, OpenEnv will be coordinated by a committee that so far includes Meta-PyTorch, Reflection, Unsloth, Modal, Prime Intellect, Nvidia, Mercor, Fleet AI, and Hugging Face. OpenEnv now lives at huggingface/OpenEnv. The project focuses on being an interoperability layer for RL environments, not a reward framework or trainer.
matchmind is an AI-driven job matching tool that scans 5 job boards every morning at 6am, scores listings against your profile using GPT-4, and emails your top 10 matches with explanations. It learns from your feedback over time and only shows fresh jobs from the last 24 hours.
Simon Willison releases datasette-agent-edit 0.1a0, a base plugin for Datasette Agent that provides storage-agnostic file-editing tools (view, str_replace, insert) inspired by Claude's text editor.