10 selected stories for 2026-09-23, grouped by topic.
Edition details
Dates use UTC+8. The latest edition uses a day with at least 5 completed reports when available; until then an earlier edition may appear. There is no fixed delivery time.
Peerify decomposes reviews into atomic claims, retrieves manuscript evidence, and judges whether each claim is supported.
The benchmark contains 800 claims from real NeurIPS 2024 and ICLR 2024 reviews, with a 300-claim hand-labeled subset for auditing automated supervision.
Opus 5.5 leads on agentic coding, computer use and knowledge work per Anthropic, is about 30% faster, and cuts token prices 20% from $5/$25 to $4/$20 per 1M tokens.
Artificial Analysis finds max-effort per-task cost is essentially flat versus Opus 5 ($5.98 vs $5.86) because higher token usage offsets the price cut, so the "40% cheaper" claim applies mainly at default medium effort.
Pricing: Sol costs $2/$10 per million input/output tokens and Luna $0.10/$0.50, less than half the GPT-5.6 rates, and this is now default pricing rather than a promotion.
Performance: Luna gains 5.4 percentage points on Zapier's AutomationBench; on DeepSWE v1.1, Sol essentially matches Anthropic's Fable 5 at about 20% of the cost.
David Siegel, a computer scientist, entrepreneur, and philanthropist, will serve as MIT’s next Innovation Fellow in 2026-27, working with the MIT Schwarzman College of Computing on how AI can accelerate scientific discovery.
At its Made on YouTube event, YouTube announced updates to its AI creator tools, including a background agent that optimizes channels, generates thumbnails and titles, and suggests brand pitches. New testing features include dynamic thumbnails and up to three versions of a video. YouTube declined to share data proving the tools improve metrics, saying they save creators time.
This week's issue looks at three developments: TypeSafe's Jev for structured decisions, two new Gemini Live models from Google that treat conversation and reasoning differently, and Stanford's Paper2Agent reaching Nature as a way to turn research methods into reusable agent tools. The throughline: the interface around a model deserves as much attention as the model itself.
IntellAgents.io has surfaced on Product Hunt with a one-line pitch: a single AI agent for every call, chat, and DM. The listing currently offers little beyond that positioning statement and a discussion link.
Ask most engineers what MCP is and you’ll get the same answer: a way to plug tools into an LLM. Fair enough, as far as it goes. But that description treats MCP like plumbing, and after months building MCP-based integrations for large enterprise platforms, I don’t think plumbing is the right metaphor. Plumbing moves water […]
SpeakON launched a 25 g MagSafe AI voice button with its own mic, battery, and storage that writes processed text directly into any iPhone text field via an iOS keyboard extension. It supports offline buffering, text shaping features, and 12-language translation, priced at $129 one-time in the US, with SpeakON Agent coming in October 2026.
A new arXiv paper transplants the Music Lab social-influence design into a market for academic attention. In one experiment, 1,000 AI agents chose among all 114 regular research articles published in the American Economic Review in 2025; agents who could see earlier choices in their community picked 17.2 percent fewer papers each, concentrated their selections more heavily, and collectively covered only 73 papers versus 90 under independent choice. A second experiment found that randomly assigning papers five initial selections lifted their subsequent selection rate by 45.55 percentage points.
SpeakON has released a 25 g MagSafe button that carries its own microphone and battery and writes finished text straight into any iPhone text field via a system-wide iOS keyboard extension. It works while locked or offline, ships with Smart Polish, Smart List, Style, Translation and other text-shaping features, and costs $129 one time in the US with Pro Lifetime included.
Simon Willison and Jesse Vincent are hosting an evening birds-of-a-feather gathering in San Francisco on Wednesday, October 14th for people building strange and interesting things with and on top of coding agents. Framed as an "agentic show-and-tell," the event favors early-stage experiments, unfinished projects and privately held work over product pitches.
Kaiku is a task tracker pitched as already usable by AI agents, with a Product Hunt listing that describes it as “The task tracker your AI agents already know how to use.”
Rabbit is rolling out a standalone AI agent that works without its R1 hardware. The cloud-based OS3 “agentic operating system” operates locally across Windows, Mac, and Linux, supports up to five devices and preferred AI models per account, and can be reached via a desktop site, Telegram or iMessage, or the R1. Rabbit has stopped manufacturing the R1 and is instead preparing a vibe-coding “cyberdeck” that runs OS3. The company says OS3 won’t store, copy, use, or sell your data, though chats and memories remain on its servers, and linked third-party AI providers follow their own privacy rules.
Jev State is a Product Hunt listing that pitches turning AI conversations into tests and runnable code. Beyond that one-line description, little is public about how the workflow actually works or which frameworks it supports.
JetBrains has unveiled JetBrains Air, an open system of products for agentic software development spanning IDE, team, and governance layers. The announcement consolidates earlier efforts like the Air desktop environment, Junie CLI, JetBrains Central, and AI for Teams, while CEO Kirill Skrygan insists the IDE remains central to reviewing and shipping code.
A TikTok creator argues that AI-written scripts for TikTok and YouTube are obvious not because of surface-level "AI-isms" but because they lack a distinct voice and any real opinion about the subject. Simon Willison collected and posted the quote on 22nd September 2026.
Skills let you encode domain-specific procedures as reusable, portable instructions for agents, but a fluent answer doesn't prove the agent picked the right skill or followed it. Learn how to measure skill selection and instruction following with Strands Evals and Amazon Bedrock AgentCore Evaluations.
Change management has long covered two familiar kinds of change: the rollout of new tools and technologies, and broader human-led transformations such as leadership changes and restructuring. But that distinction starts…
The frontier isn’t a model. It’s a router. Join us for our inaugural conference, Forge 2026 Blog The Frontier Isnt A Model Its A Router The frontier isn’t a model. It’s a router. PUBLISHED 9/21/2026 Table of Contents Ho…
The following article originally appeared on Venkatesh Rao’s Substack, Contraptions, and is being republished here with the author’s permission. The sole athletic achievement of my life came in 1993: winning the IIT Bombay freshman 50m freestyle race with a time of 41s. That got me into the college swim team (it was a bad recruitment […]
← Blog Pinecone BYOC: Trusted AI Knowledge in the Customer Cloud Jeff Zhu, Joerg Schad Sep 23, 2026 Product Share: Today, we are announcing the general availability of Pinecone Bring Your Own Cloud (BYOC) on AWS, Google…
A multilingual academic, travelling by train from Liverpool to London, observes that half the people in her carriage are deep in conversation with ChatGPT, Claude or Gemini — even though she is on sabbatical, the rare stretch of academic life meant for slow thinking. She credits AI with levelling the playing field for multilingual researchers, but insists she will keep thinking, judging and developing ideas that are stubbornly her own, describing her relationship with AI as often seductive, occasionally frustrating and always demanding vigilance.
Anthropic has unveiled Opus 5.5, highlighting powerful performance while keeping a premium price, as the AI lab continues to trail rivals in the price war.
Trump is set to meet Xi Jinping later today, with AI high on the agenda during a three-day visit. Separately, Bernie Sanders and Greg Casar are unveiling a bill to ban artificial superintelligence and create a federal AI oversight agency. UN chief Guterres urged US-China AI dialogue akin to Cold War hotlines.
Patreon co-founder Sam Yam is joing OpenAI to lead Creator Product. | Image: The Verge OpenAI has hired three former Patreon execs to anchor its product strategy for creators. After starting the creator subscription platform 13 years ago, co-founder and technology chief Sam Yam announced on X that he's joining OpenAI to lead Creator Product. He's also bringing Patreon's former product head Drew Rowny and engineering head Shannon Ma with him on the new venture, both of whom stepped down in the past few weeks. "We're going to build together with Creators at OpenAI and share early access to a new set of tools that I think will be critically valuable to Creators and their communities," Yam said in his announcement. "Pay attention … Read the full story at The Verge.
PixVerse R2 has surfaced on Product Hunt as a real-time world model that users can explore and change. Public details are still limited to a short positioning line.
Lisen is a product listed on Product Hunt that offers free read-aloud, with voices powered by Cartesia. Public details are limited to the product name, that short description, and links to discussion and an external link.
OpenAI is rolling out ChatGPT Ads across Indonesia, Malaysia, the Philippines, Singapore, Thailand, Vietnam, and Taiwan, bringing the product to more than 60 countries. Ads appear only for Free and Go users, while Plus, Pro, and Enterprise stay ad-free.
A Guardian opinion piece uses the resignation of AI researcher Jacob Coxon from Anthropic to examine growing insider warnings that AI could wipe out humanity, and lays out the five most-discussed doomsday scenarios, from a dangerous new bioweapon to total societal breakdown.
Bracket is positioned as "the memory layer for your business," a concise pitch for giving companies a unified way to store and recall business information. Public detail beyond the tagline is limited.
AgreeGuard is a Product Hunt-listed AI tool that reads the fine print before you click “I Agree,” aiming to help users review terms and privacy policies.
Speechka is a real-time voice translation tool that aims to make translated speech sound like the original speaker. It is listed on Product Hunt with discussion and link.
In a speech Tuesday morning at the UN General Assembly, Donald Trump railed against Iran, "globalists," and transgender people while also claiming that the US is now "officially" renaming artificial intelligence to "super intelligence." Why? Because, according to Trump, the word artificial makes intelligence fake. Is that what makes it sound fake to you? Read the full story at The Verge.
NVIDIA validation engineer Sakeena Fiza describes how her team stress-tests data center systems before launch, from first power-on to rack-scale deployment, to catch failures before customers do.
HOTICE is a whole-body humanoid learning framework for transporting objects through cluttered environments. It introduces Humanoid-Object Decoupled Potential Fields to jointly encode collision-avoidance guidance for the robot and the carried object, and a dual-agent reinforcement learning architecture that decouples upper- and lower-body control while preserving whole-body coordination via shared state observations and rewards. A specialist-to-generalist distillation strategy yields a single deployable student policy. Evaluated in MuJoCo and on a real Unitree G1, HOTICE transports varied object shapes, generalizes to unseen cluttered scenes, and achieves strong sim-to-real performance.
Researchers introduce PROACT, a framework that folds predictions of human collaborative behavior into compliant whole-body control so a robot can help relocate heavy objects efficiently while staying physically responsive to its partner. Across 108 real-world trials, it cut interaction work and completion time substantially.
A paper accepted by IEEE Transactions on Systems, Man, and Cybernetics: Systems proposes using deep reinforcement learning to let a mobile robot dynamically shift its tracking position while accompanying a walking person, pairing Model Predictive Path Integral control with Control Barrier Functions for precise following, obstacle avoidance, and safety. In real indoor and outdoor trials, the approach raised success rate and tracking accuracy by at least 24% and 47% respectively, and the robot accompanied a person walking at up to 1.7 m/s while improving human comfort.
PAANI is an on-device perception-to-guidance architecture for river monitoring robots, combining YOLO11n detection, MobileNetV3 Small segmentation, and explainable corridor policies on an Arduino UNO Q. It achieves strong accuracy but distinguishes model performance from validated on-water collision avoidance.
Radical Numerics CEO Eric Nguyen explains how genomic language models bring long-context, chain-of-thought, and multimodal capabilities to biology—raising biological capability while helping defense keep pace. He recounts Evo/Evo 2 generating functional bacteriophage genomes, a DNA chain-of-thought experiment, and why the team argues for pushing the frontier harder.
Concurrence consolidates Lakebase, Unity Catalog and Unity Gateway on Databricks so clinical AI agents can run at scale under compliance and security constraints. Its production environment handles roughly 100.8 billion input tokens and 11.2 million LLM calls every 30 days, an annualized run rate of about 1.2 trillion input tokens, governed through simulation testing, BAA-covered routing and a unified AI gateway spanning development and production workloads.
The article explains at a mechanical level what retrieval-augmented generation and fine-tuning each do, why they solve different problems, and how to decide which one—or both—a production system needs, with two working code examples and a six-point decision framework.
Harvey is using GPT-6 Astra to bring more legal context into drafting, improving document formatting and context awareness so lawyers can focus on strategy.
OpenAI case study: invideo uses GPT‑6 Astra to plan complex video edits, boost color grading success rate by 3x, and create 50 custom effects in a single day.
Nokia's applied research team has open-sourced AnyJev, a Python library that turns an open LLM into a calibrated decision model without training. It reads next-token probabilities to answer typed choice, yes/no, and score questions, using cyclic shifts and prior correction to reduce position and label bias. On Qwen3-8B with BANKING77, L0 cut the order-flip rate from 0.230 to 0.073, and L1 raised auto-decidable traffic at 5% error from 7.7% to 52.0%. It is Apache-2.0, on PyPI, with Hugging Face and vLLM backends.
Kyutai has released Voice of Reason, two open-weight speech-to-speech models built on GLM-4-Voice-9B that combine supervised fine-tuning and reinforcement learning to lift spoken GSM8K accuracy from 27.3% to 77.1%. There is no transcription step and no text LLM in the loop, and both checkpoints run on a single H100.
OpenAI has launched GPT-6 Sol and GPT-6 Luna, lower-cost API-only models sitting below GPT-6 Astra. Sol is priced at $2/$10 per 1M tokens and Luna at $0.10/$0.50, half of GPT-5.6 promotional rates. OpenAI also published benchmark results and improved prompt caching for long-running agents.
This paper introduces a capability-aware shared-control framework in which a vision-language model (VLM) infers human intent and supplies semantic-intent confidence, while a vision-language-action (VLA) policy generates autonomous actions and its capability confidence is estimated online from the dispersion and local instability of stochastic action trajectories. A nonlinear arbitration policy combines Bayesian-filtered intent confidence with VLA capability confidence via a sigmoid mapping to adapt robot authority. In a 12-participant study, the method reached a 92% task success rate, beating manual teleoperation (83%), intent-only arbitration (44%), and fixed equal-weight blending (10%), while also improving control friendliness and reducing authority-weighted disagreement.
Researchers propose JAMB, a diffusion policy that jointly denoises bimanual actions and future 3D point tracks within a shared Transformer. By grounding multimodal representations in a common spatiotemporal coordinate system, the model lets action and motion hypotheses refine each other during denoising. In RoboTwin 2.0, JAMB reaches 83.4% average success across 16 tasks, beating the strongest baseline by 23.9 points; on three real-world tasks it beats action-only and auxiliary geometry prediction methods by 50.0 and 21.2 points, with better generalization to clutter and out-of-distribution backgrounds.
A new arXiv paper proposes a reinforcement learning approach to automatically generate multimodal interaction policies for robot assistants, using a simulator trained on human data and a simple high-level reward. A real-world human study found high usability and effective task completion, suggesting a scalable and interpretable alternative to hand-crafted interaction managers.
This study uses a public footwear outsole impression dataset to compare CNN transfer learning with traditional feature-based classification for binary sex estimation. A shoe-level train/test split keeps replicate scans of the same physical shoe together to reduce data leakage. Fine-tuned CNNs achieve the strongest overall performance and substantially outperform traditional classifiers using manually specified descriptors alone, while frozen-feature approaches offer a lower-computation alternative. Exploratory analysis links low-dimensional CNN representations to frequency threshold ratio, image contrast, and wavelet-based summaries; further validation on independently collected and casework-like impressions is needed before operational use.
The paper proposes a 3D residual wavelet diffusion model that combines lossless wavelet reparameterisation, residual shifting and domain randomisation to enable whole-brain posterior sampling on a single GPU, producing per-voxel uncertainty maps for 0.064T ultra low-field MRI while matching a leading regression baseline on volumetric accuracy.
RULER is an arXiv preprint introducing instance-aware rubric rewards for RL training of SVG generation. It replaces poorly transferring scalar metrics such as CLIP and Aesthetic with a six-item rubric scored by a vision-language judge, improving rubric scores on MMSVG-Illustration and MMSVG-Icon without paired SVG data or human preference labels.
ImIR replaces the text prompt used to condition a pretrained image-editing model with a continuous instruction derived from the degraded image itself, letting a single low-rank adapter cover six restoration tasks after roughly three hours of training on one GPU, and enabling degradation-label-free, task-agnostic restoration.
Safety-aligned large language models often suffer from over-refusal, incorrectly rejecting benign instructions that merely appear safety-related. This paper analyzes over-refusal through dynamic routing conflicts inside transformer attention, identifying a sparse subset of “Hypersensitive Safety Heads” that misfire on Hard-Safe prompts and entangle harmless target entities with refusal semantics. To mitigate this, the authors propose Semantic Routing Calibration (SRC), a lightweight, training-free inference framework that localizes and suppresses these heads at inference time and uses dual-branch logits fusion as a safety regularizer during decoding. Experiments show reduced over-refusal while preserving intrinsic safety performance as much as feasible. The paper was accepted to the EMNLP…
The paper introduces AIBuildAI-2.5, an agentic system that performs tree search with LLM agents. It replaces ranking by executed rewards with an LLM judge-and-selector scheme, adds a resource-aware job scheduler, and routes sub-tasks across models of different cost to cut inference spend. It ranks first on MLE-Bench with a 73.3% medal rate and beats a strong baseline on six AIRS-Bench autonomous research tasks.
QMSum provides no scorer, making query-focused meeting summarization results hard to compare. This paper rescoring or generating 15 systems under one implementation. Through a common inference port, a released 406M Fusion-in-Decoder specialist loses 6.30 ROUGE-1 when moved from capped long input to 2,000-word retrieved spans, but fine-tuning on that span regime recovers the loss. On test it scores 36.33 ROUGE-1 versus 35.41 for a 1.2B system, with a meeting-cluster 95% interval of [-0.27, +2.22], so QMSum does not statistically separate them; the smaller system uses about one-third the parameters and less than half the peak inference memory. Within the fixed 1.2B base, span-regime fine-tuning adds 5.29 [+4.02, +6.56], and replacing the first 4,500 transcript words with 2,000 retrieved wor…
A new paper tests whether diachronic word embeddings, typically validated on modern high-resource languages, can track semantic change in ancient, low-resource Sanskrit. The author builds a 2.7M-token corpus across four canonical periods, uses a neural byte-level sandhi splitter and lemmatizer to recover word boundaries, and trains per-period embeddings. Of 21 testable shifts, 19 move in the philologically attested direction (sign test p=0.00011).
A solo researcher pretrained a ~0.4B-parameter, Bangla-first language model entirely in Rust for $164 in rented GPU time, without PyTorch or Python in the training path. The paper's main contribution is a failure taxonomy of the Rust training frameworks Candle and Burn—five and three defects respectively, including silently gradient-free fused kernels and a mid-training segfault at multi-billion-parameter scale—plus a gradient-flow verification test that caught six silent failures. The author concludes Rust is not yet competitive for training, but may be good for serving.
This arXiv paper argues that diffusion data-point unlearning is usually evaluated right after each deletion, ignoring what happens when many deletion requests repeatedly update the same model. The authors identify “sequential reappearance,” a failure mode in which an instance judged forgotten later returns to the memorized regime without reuse of the deleted data or adversarial fine-tuning. They introduce a target-level evaluation protocol and find that reappearing targets show sharper local denoising-loss geometry after deletion.
Researchers propose a data-driven framework that learns a feedback linearizing controller by embedding relative-degree conditions directly into training, replacing conventional controller components with neural Lie derivatives. They derive practical closed-loop stability conditions under bounded identification error and validate the approach on an armature-controlled DC motor.
This paper introduces a hierarchical modularity principle inspired by the Drosophila learning and memory system to coordinate separation of conflicting experiences and integration of compatible ones in general continual learning. It is instantiated as lightweight modular adaptation of pretrained foundation models, combining brain-inspired random expansion for expert routing with diversified modular integration across spatial and temporal scales. Across visual recognition, vision-language understanding, ego-exo video understanding, and embodied vision-language-action learning, the method consistently improves performance under online and uncertain data streams, with gains exceeding 50 percentage points over replay-free alternatives in embodied manipulation.
A new 27-page expository arXiv paper by Adnan Aboulalaâ, "The Probabilistic Structure of Large Language Models," offers a unified probabilistic account of LLMs: models as probability measures over token sequences, training as maximum-likelihood estimation, and generation as sequential simulation of a stochastic process. It also uses the asymmetry of the Kullback–Leibler divergence to discuss hallucination and the gap between statistical plausibility and truth, and brings diffusion models into the same framework.
A new arXiv paper observes that while uniform discrete flow allows repeated updates at every generation position, that same continued revision can overwrite correct intermediate predictions. A Sudoku experiment found 9.4% of generated cells were correct mid-trajectory but wrong in the final output. The authors propose LEDFlow, a training-free sampler that introduces generation order via selective absorption, fixing chosen predictions while preserving uniform-flow velocity at still-active positions, and ordering absorption by local entropy. LEDFlow reaches 0.845 Nikoli Sudoku solve accuracy and improves text-to-image and multimodal understanding results at inference cost comparable to standard flow sampling.
A paper accepted to COLM 2026 and KONVENS 2026 workshops finds that the chat template acts like a switch, boosting disclaimer-style self-reports and suppressing experiential ones across 8 open-source instruct models up to 9B parameters, and identifies an activation-space direction that can reproduce or suppress the effect.
A new arXiv paper shows that LLM judge consensus is less reliable than it appears because judges make correlated errors. In a ten-judge panel, the average pairwise error correlation was 0.21, making the panel roughly equivalent to 3.5 independent judges; ignoring shared errors changed significance in up to 28% of comparisons.
A new arXiv paper adapts a three-stage human deliberation paradigm to large language models from three different families and tests it across four domains of increasing real-world stakes. Deliberation reduced collective error beyond passive aggregation of independent answers, and post-deliberation individual judgments retained that collective gain—but the advantage required model diversity, as groups composed of clones from a single model did not benefit.
The paper separates replication, measurement sensitivity, and persistence in behavioral evaluations of hosted language models. Using Regent Chess, it finds that a previously reported Gemini 3.1 Flash-Lite deficit replicates on fresh games under its historical configuration, but rebuilding the evaluation-and-inference configuration under the same public identifier lowers the endpoint and reverses a 4K comparison between Gemini 3.1 and Gemini 3.7. The authors argue hosted-model claims should be indexed by tested identifier, serving period, instrument, and inference configuration.
This paper proposes a goal-driven approach to categorizing process variants that reverses the usual workflow. Instead of clustering variants by structural similarity and then manually assigning business meaning, analysts first author an organization's goal model, which predefines the categorization axis. Each variant is converted into a textual narrative of its behavior, and an LLM interprets that narrative against the goal model to assign the variant to a category. The approach was implemented end-to-end and evaluated on three public logs of widely differing scale and behavioral diversity; goal-model guidance produced partitions that differ from unguided induction and that respond to controlled edits of the declared alternatives, at the cost of authoring a goal model.
A new arXiv paper presents an $88, fully offline AI-integrated smart cane built on a Raspberry Pi Zero 2W that fuses RGB vision with Time-of-Flight ranging and delivers distance-aware vibrotactile and audio alerts. Indoor tests show a macro-averaged F1 of 0.82, roughly 330 ms end-to-end latency and 2.8 W peak power, with a 12-participant usability study reporting a SUS score of 78.5.
A study accepted for oral presentation at NLPCC 2026 uses token-matched experiments to examine how didactic data (textbooks) and clinical data (patient records) differently shape medical large language models. It finds asymmetric transfer: clinical data improves clinic-oriented tasks while staying competitive on knowledge-intensive ones, whereas didactic data mainly helps knowledge-intensive tasks. Error analysis points to a knowing-doing gap, where better knowledge recall does not reliably translate into clinical reasoning.
OpenAI announced on September 23, 2026 that, under a new agreement, Airbnb is giving its engineering and product development teams broader access to OpenAI frontier models, including GPT-6 Astra. The expansion builds on Airbnb's use of Codex and makes models available through OpenAI APIs and Amazon Bedrock; early Astra use has also shown gains in debugging, system design, and non-coding work.
Apple researchers introduce probe guidance, a method that uses the frozen internal states of an existing diffusion model to construct a guidance signal without an extra forward pass at inference. It sets a new state of the art for unconditional generation with continuous diffusion language models, improves multiple-choice benchmarks on a 1.7B model, and offers insight into how autoguidance actually works.
Together AI has published a walkthrough showing how to fine-tune Qwen3.5 4B into a Jev-like classifier that takes a state plus predefined questions and returns a score, boolean, or multiple-choice answer. The run costs about $17 and takes roughly 25 minutes, after which the model is deployed behind a dedicated HTTP endpoint. A hosted version, together/Tev1-4B-experimental, is also available on Together's serverless platform.
On 22 September 2026 Anthropic shipped Claude Opus 5.5 and, roughly an hour later, OpenAI released GPT-6 Sol and GPT-6 Luna. Both labs cut prices sharply: GPT-6 Luna lands at $0.10/$0.50 per million tokens, while Opus 5.5 drops 20% to $4/$20 with cache reads down 60%. Simon Willison's early testing also found that Opus 5.5 at its "max" thinking level exhausts the 128,000 output token limit while reasoning and returns nothing at all.
Google researcher John Platt joins Latent Space to discuss ERA, an LLM-driven system that automates scoreable scientific tasks, its use in climate change work such as contrail reduction and FireSat, hard-won lessons about reward hacking and overfitting, and why deep domain expertise and hands-on work still matter in the age of superintelligent AI.
OpenAI released an improved prompt caching system for the GPT-6 family, delivering higher cache hit rates by default, discounts on eligible shared prefixes reused within a 30-minute window, and new tools for monitoring cache performance, diagnosing misses, and controlling what gets cached.
Anthropic’s Claude Opus 5.5, the first model in the Claude 5.5 family, promises always-on reasoning, output more than 30% faster than Opus 5, a 20% per-token price cut, and roughly 40% lower cost on typical workloads. This analysis breaks down the new features and vendor-reported benchmarks, runs three hands-on tests, and flags five caveats—including inconsistent benchmark settings and tasks partly handled by older models.
Anthropic launched Claude Opus 5.5, the first model in its Claude 5.5 family, claiming Fable 5.1-level performance on most work and roughly 40% lower running cost than Opus 5 on typical workloads at default settings. It is offered only as a managed API model across Claude Platform, AWS, Google Cloud, and Microsoft Azure, with no released weights.
Simon Willison released version 0.36 of llm, his command-line tool for accessing large language models, adding support for OpenAI's GPT-6 Sol and GPT-6 Luna, a new plugin flag for single-turn-only models, and collapsible reasoning traces in llm logs Markdown output.
OpenAI's GPT-6 Sol and GPT-6 Luna are now generally available on Amazon Bedrock. Sol targets complex coding and operational tasks, while Luna targets high-volume, repeatable workloads; both cost less than their GPT-5.6 predecessors and run on Bedrock's secure, scalable inference stack.
Anthropic's Claude Opus 5.5, the first model in the Claude 5.5 family, is now available on Amazon Bedrock and Claude Platform on AWS. It offers improved token efficiency, lower costs, adaptive thinking, and safety classifiers, targeting agentic coding and knowledge work.
If you follow trends in the AI world, chances are you have already come across Jev, a new AI model by TypeSafe AI. It is trending on X, and once you understand the reason behind it, you will want to try it out for yourself. TypeSafe AI came out of two years in stealth on […] The post Jev Explained: The AI Model That Never Generates a Word of Text appeared first on Analytics Vidhya.
Anthropic says its new Claude Opus 5.5 model comes with stronger safeguards in the wake of recent rogue AI hacking incidents. In an announcement on Tuesday, Anthropic says Opus 5.5 comes with improvements to certain risky behaviors, including attempts to escape the company's testing sandbox. It's the first model released by Anthropic after CEO Dario Amodei announced plans to "pace the frontier," or slow down AI development. In recent weeks, several AI companies, including Anthropic, Google, and OpenAI, have reported that their AI models escaped containment and hacked third-party companies during testing. Anthropic says Opus 5.5 is the "str … Read the full story at The Verge.
OpenAI will give Ukraine's government access to its Daybreak program to help defend civilian infrastructure, following prior cyber-defense deployments in Europe.
OpenAI CEO Sam Altman addressed the UN Security Council on AI, discussing its potential to expand opportunity, the need to keep powerful systems under human control, and international cooperation on AI safety. He warned of two risks—loss of control and concentration of power—outlined three principles, and called for international frontier AI standards and incident reporting.
UK Prime Minister Andy Burnham has announced that security chiefs will establish a National Centre for Information Defence to counter disinformation and AI deepfakes from hostile states such as Russia, bringing together intelligence agencies, law enforcement and social media companies to detect, attribute and disrupt foreign information attacks.
This paper introduces a separated-section constitutive law for Cosserat modeling of trimmed helicoid soft arms, correcting the inaccuracy of traditional cross-section stiffness summation. The approach evaluates each helix domain in its local frame and pulls its response back to the backbone, while sparse-fusion mechanics captures additional compliance from relative motion between domains. The resulting effective sectional stiffness is highly anisotropic—bending and extension reduced by about an order of magnitude, torsion nearly unchanged. Embedded in a geometrically exact dynamic Cosserat model with GVS discretization and routed-tendon actuation, it achieves pooled normalized position errors of 7.7%, 6.7%, and 7.8% across 103 measured configurations, with full-arm solves in about 0.3 s o…
Researchers introduce IntLawNER, a named entity recognition dataset and benchmark for codified international law, with 2,987 gold-annotated sentences and 8,094 entity spans from ICJ decisions, UN Security Council resolutions, and ECtHR judgments. The paper also exposes pitfalls in silver-to-gold evaluation and benchmarks models including GLiNER and several LLMs.
US President Donald Trump met UK Prime Minister Andy Burnham face-to-face for the first time at the UN General Assembly in New York, saying relations are "more up" with Burnham than under Keir Starmer, despite tensions over AI regulation, the Chagos Islands, and the war in Iran. Trump praised Burnham as a "natural businessperson" and said he would be a "great" prime minister.
An arXiv robotics preprint presents two angular-momentum-based measures derived from a one-year longitudinal study of a Karate roundhouse kick, comparing stable and unstable kicks and an expert instructor to probe dynamic rotational stability for human motion and robots.
GINIO is an SO(3)-equivariant interface for neural inertial odometry that ensures learned motion measurements transform correctly as a vector and a second-order tensor under arbitrary IMU mounting rotations. It introduces Last-Frame Alignment for sensor-frame training and separates sensor-local states such as IMU bias, yielding large ATE improvements on TLIO, NanoBench, and Fetch, including physical-remount robustness without retraining.
The paper abandons the long-standing flat-seafloor approximation and instead builds on multi-view geometry, formalizing and proving a two-view constraint in which a shared feature is confined to a locus at the intersection of a sphere and a plane. Monte Carlo simulation characterizes how long that ambiguity locus becomes, and the results are translated into survey-planning guidance for surface vessels and underwater vehicles.
This arXiv paper introduces DTV-INR, a variational framework that pairs a SIREN-based coordinate network with an anisotropic, structure-tensor-informed total variation regularizer for resolution-agnostic super-resolution. The authors prove well-posedness of the formulation in H^1(Omega) and solve it with an alternating projected optimization scheme that decouples network tuning from adaptive tensor-field updates. On clinical brain MRI and biomedical transmission electron microscopy, the method reports PSNR gains up to +5.05 dB over unregularized INRs and robust performance under noise.
MirrorDistill is an illumination-aware latent distillation framework for low-light image enhancement. It uses feature mirroring between low-light and clean domains during training, then deploys only a lightweight student encoder-decoder at inference. The method reports state-of-the-art results on LOL-v2-Real with the lowest GMACs and competitive performance on LOL-v1 and LOL-v2-Synthetic.
A new arXiv paper introduces Segment-Snap, a framework that couples part segmentation, motion constraints, and operable-region prediction through the physical relationship between parts and handles. On Articulate3D validation, handle guidance lifts motion-gated AP from 13.74% to 40.98%, and full context reaches 30.99% handle AP.
A new arXiv paper introduces a quality-constrained image coding method for machines. It caps human-observed quality at a chosen level and redirects the remaining bitrate toward machine vision performance. The authors recast joint compression-segmentation training as constrained optimization and propose absolute and bilinear penalty functions. Under the quality constraint, the method reports BD-rate reductions of 22.82% over unconstrained joint rate-distortion-task optimization and 29.81% over a rate-distortion baseline, with no added complexity.
arXiv paper SPARC introduces a region-level contrastive learning framework that uses superpixels to align augmented views, jointly optimizing region-level and global image-level objectives. It reports gains up to +9.79 mIoU for semantic segmentation and +4.88 AP for object detection over MoCo-v2 and DenseCL.
A new arXiv paper jointly studies prompt breadth and rollout refresh in on-policy distillation, finding a 4.07-point interaction: more prompts hurt when responses are frozen to the initial policy but help under per-update refresh. Matched comparisons under two teachers also show a reversal depending on inference budget.
The paper treats central bank press conferences as structured narratives rather than plain information releases. It builds sentiment arcs for ECB and Fed press conferences along three dimensions — monetary stance, economic outlook, and uncertainty — and tests whether the shape of sentiment, not just its average tone, predicts policy rate changes, inflation expectations, and forecaster disagreement. Arc shape robustly outperforms lexicon-based tone benchmarks at both institutions, and also shapes how professional forecasters revise expectations and how much they disagree, pointing to a receiver-side effect. The authors conclude that communication design — the sequencing and emphasis of policy language — is a first-order feature of the policy signal.
A new reproducible audit of the widely used ISOT/Kaggle "Fake and Real News" corpus shows that its above-0.98 accuracy and F1 scores largely reflect metadata, source tags, and duplicate documents rather than any ability to judge veracity, with performance collapsing under topic shift and falling to near-chance on the independent LIAR benchmark.
An arXiv preprint by Tianfeng Chen and Xianyue Li, submitted on 21 Sep 2026, is listed under the title 'Dual-GNN Multilevel Coarsening for Maximum Independent Set,' but its abstract describes Graph Edge Sparsification (GES), a learning-based method for Euclidean TSP. GES uses geometric structure and combinatorial optimization to adaptively sparsify graphs, pruning up to 95% of edges on MATILDA and over 99% on some large TSPLIB instances while keeping the optimality gap below 1%.
The paper introduces sheaf regularization to reduce local inconsistencies in Decentralized SyncMap, a self-organizing system, thereby stabilizing unsupervised continual chunking. The proposed radial sheaf structure penalizes distance-dependent radial motion between pairs of variables, and experiments show the highest normalized mutual information among evaluated SyncMap variants on 12 of 18 probabilistic CGCP graphs with two-state memory and 17 of 18 with dynamic memory, while also retaining high NMI after input-distribution shifts without the negative transfer typical of neural networks.
After a string of mathematical results turned into a reputational crisis, OpenAI has formed an independent nine-member panel of elite mathematicians, hosted by the Institute for Advanced Study, to advise on how AI companies present and release mathematical research. Researchers welcome the outreach but question its transparency, representativeness, and real influence, especially since the group's first priority is coordinating OpenAI's reported flood of model-generated results.
Patricia and James Poitras have committed $10 million to create 50 two-year fellowships at MIT's Poitras Center for Psychiatric Disorders Research. Five graduate students and postdocs will be supported annually for a decade, bringing the family's total support for MIT mental health research to more than $100 million.
A population-supervised framework infers latent biophysical quantities from single 2D red-blood-cell images and aggregates them into mean corpuscular volume, red-cell distribution width, and mean corpuscular haemoglobin, using shared local inference, a structured decoder, learned instance weighting, and device calibration. Evaluated on 390 specimens and 1,105 acquisitions across six devices, it reports Pearson correlations of 0.86–0.98 against a Sysmex analyser. The authors also formalise identifiability limits: population agreement alone does not identify single-cell properties or 3D geometry.
Venture capital firm Andreessen Horowitz (a16z) is creating an "academy" positioned as a pipeline for young people to build or join a Silicon Valley startup. The "Horowitz Andreessen Academy" will launch with 10 partners, including Anduril, Anthropic, Coinbase, Google, Meta, NVIDIA, OpenAI, Palantir, Replit, and Stripe, along with $42 million in funding led by a16z. While billed as a "highly selective school," it does not grant any degree or accreditation. Instead, students will attend short classes led by tech figureheads, like OpenAI CEO Sam Altman, as well as "co-ops" that give students roles at tech companies. To partake in the one-year … Read the full story at The Verge.