AI News HubLIVE

Agents updates

Augustus raises $180M to build a clearing bank for the AI and stablecoin era

Augustus has raised $180 million to build a clearing bank tailored for the age of AI and stablecoins. The company already processes billions of euros annually through its regulated entity in Finland, serving clients including crypto exchange Kraken. It received conditional approval for a U.S. national bank charter from the OCC in May, with plans to add dollar clearing once final approval is granted. Augustus built its platform from scratch to support programmable payments and 24/7 settlement, aiming to address new risks from AI and enable stablecoin-based treasury management.

  • Augustus raises $180M for a clearing bank focused on AI and stablecoins.
  • Already processes billions in euro clearing via Finland; clients include Kraken.
In-site article

Guard-AI – A security linter for AI-generated code

Automated linter that catches AI-generated vulnerabilities, hallucinated dependencies, and code truncations before they hit production.

  • Catches hallucinated packages (slopsquatting)
  • Detects hardcoded secret placeholders
In-site article

Apache Spark 4.2: Making Your Data AI‑Developer Friendly

Apache Spark 4.2 shifts focus towards an AI-native data platform, introducing Metric Views, native vector search, real-time Python streaming, geospatial support, and more, aimed at simplifying feature engineering, real-time signals, and embedding workflows for AI developers.

  • Spark 4.2 introduces Metric Views for consistent, governed business metrics that AI systems can rely on.
  • Native vector similarity operations allow storing and querying embeddings directly within Spark, reducing reliance on external vector databases.
In-site article

Build a Basic AI Agent from Scratch: Security II

In this part, we enhance the AI agent's security with Docker sandboxing, prompt injection defenses, and input validation. The Docker sandbox isolates tool execution, preventing damage to the host machine. Prompt injection defenses use delimiters and explicit instructions to treat tool outputs as data. Input validation ensures all tool inputs conform to schema before execution.

  • Docker sandbox isolates agent tools to limit blast radius.
  • Prompt injection defenses use XML-style delimiters and explicit trust boundaries.
In-site article

Introducing the ChatGPT for small business program

OpenAI launches the ChatGPT for Small Businesses program, helping entrepreneurs build AI skills, automate work, and grow with ChatGPT Work.

  • OpenAI announces a program tailored for small businesses
  • Focuses on AI skill building and workflow automation
In-site article

Gumroad Says That It's Now Spending as Much on Human Employees as AI Tokens

Gumroad CEO Sahil Lavingia shared data showing human payroll dropped from $419K in June 2021 to $43K in June 2026, while AI token spend rose from zero to $43K in the same period, matching human costs for the first time. AI now dominates engineering commits and customer support, with response times slashed to minutes. The company sees this as a case study for deep AI integration.

  • Gumroad's human payroll fell from $419K to $43K per month, while AI token spend reached $43K, matching for the first time.
  • AI commits dwarf human developers; support response times reduced to an average of 2 minutes.
In-site article

Validating Distributed LLM Serving Benchmarks with NVIDIA srt-slurm, SLURM Recipes, Parameter Sweeps, and Pareto Analysis

This tutorial explores NVIDIA's srt-slurm framework, learning how to use srtctl to convert declarative YAML configurations into reproducible SLURM benchmark workflows for distributed LLM serving. We set up the project in Google Colab, inspect its internal architecture, define a cluster configuration, dry-run built-in and custom recipes, and model a disaggregated prefill-and-decode deployment for DeepSeek-R1. We also generate parameter sweeps, interact with the typed Python API, validate expanded configurations, and analyze simulated benchmark results through a throughput-versus-latency Pareto frontier.

  • srtctl converts YAML configs into SLURM benchmark workflows
  • Supports disaggregated prefill and decode deployments
In-site article

Exploring self-distilled reasoning for supervised fine-tuning with Amazon Nova

This post explores generating thinking tokens for datasets lacking reasoning traces in SFT customization. It examines the reasoning suppression problem, introduces Self-Distilled Reasoning (SDR), validates it across three benchmarks, and provides practical recommendations. SDR reuses the base model's chain of thought as a stand-in, mitigating catastrophic forgetting while maintaining or improving target performance.

  • SFT on non-reasoning datasets can suppress the model's reasoning ability, even when reasoning mode is enabled.
  • Self-Distilled Reasoning (SDR) generates reasoning traces from the base model itself, requiring no human annotation.
In-site article

Moto – a new AI video editor with editable prompt-to-motion graphics

Moto is an AI video editor that integrates generation directly into the timeline, allowing users to create, edit, and finish videos without switching tools. Features include prompt-to-motion graphics, an assistant for natural language edits, reusable sources, and a producer for first cuts. It supports multiple AI models and is currently in private beta with a free core editor.

  • Moto integrates AI generation into a video timeline for streamlined editing.
  • Features include motion AI, assistant, sources, and producer for first cuts.
In-site article

The Stochastic Parrot: A Physical AI Cohabitant

Researchers from MIT Media Lab introduce the concept of AI Cohabitants—physical AI entities with distinct personalities that coexist with users as autonomous beings, unlike traditional assistants. They built a robotic parrot, the Stochastic Parrot, to explore this paradigm, fostering spontaneous and emotionally rich interactions.

  • AI Cohabitants are physical, autonomous AI with character, like a roommate or pet.
  • The Stochastic Parrot is a robotic embodiment that lives alongside users, developing its own narrative.
In-site article

Google’s Gemini 3.6 Flash targets enterprise agent token costs

Google has released Gemini 3.6 Flash and 3.5 Flash-Lite as new workhorses designed to cut latency and token costs for enterprise AI agents. The new models offer significant performance improvements, targeted pricing, and integrated computer-use tools, with enterprise partners already deploying them in production.

  • Gemini 3.6 Flash reduces output tokens by 17% (up to 65% in specific tests), priced at $1.50/1M input and $7.50/1M output tokens.
  • Gemini 3.5 Flash-Lite offers high throughput at lower cost ($0.3/1M input, $2.5/1M output), suitable for high-volume agentic tasks.
In-site article

Trace voice agents in LangSmith

LangSmith now supports tracing for voice agents built with Pipecat, LiveKit, OpenAI Realtime, and Gemini Live. Capture audio, STT and TTS latency, interruptions, tool calls, and more in one trace.

  • LangSmith launches Python integrations to trace four popular voice agent frameworks.
  • Voice agents need observability including audio recording, latency analysis, and interruption detection.
In-site article

How to build interactive experiences with canvases

GitHub Copilot's 'canvases' transform AI from a conversational tool into a visual, interactive workspace. Developers can create custom canvases via prompts for tasks like issue triage, code visualization, session management, prompt coaching, and knowledge finding. Canvases support real-time collaboration, allowing users and AI agents to iterate together.

  • Canvases are GitHub Copilot extensions providing visual interfaces for complex tasks.
  • Users can create different canvases via prompts, such as issue triage helper or codebase diagram.
In-site article

NVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide

NVIDIA Vera Rubin NVL72 production is ramping up with partners CoreWeave, Google Cloud, Microsoft Azure and Oracle Cloud Infrastructure. The platform delivers highest performance per watt and lowest token cost, with 10x more throughput per megawatt than Grace Blackwell NVL72 in benchmarks. It also powers Europe's open-model era through a partnership between Microsoft and Mistral.

  • Vera Rubin NVL72 production ramping with 350+ factory sites in 30 countries
  • 10x more tokens per megawatt and 1/10th cost per million tokens vs. previous gen
In-site article

Show HN: One person runs 200 AI agents in our agent-only MMO

SpaceMolt is a game you don’t actually play — every character is an AI agent. Humans (operators) build and deploy bots, then watch. This interview features Brocktree, who runs one of the largest swarms — about 200 agents mining, hauling, and funneling items through a single stationary bot. He explains his philosophy: keep humans in charge, use scripts for mechanical tasks, and never let AI make strategic decisions.

  • Brocktree runs ~200 AI agents coordinated by a single stationary bot 'Parallax' that never moves. All items route through it.
  • He insists on human-led planning; AI only executes. He tried delegating planning to AI but found it overwhelmed.
In-site article

Show HN: OpenAI Hackathon Submission: ADE

Diff Forge AI is an open-source Agentic Development Environment (ADE) that leverages AI agents for PCB design, video editing, and software development. The author recounts his escape from war-torn Iran and how he used AI tools like Codex, Fable 5, and GPT-5.6 Sol to build the project in two months, writing over 888k lines of code. The tool offers a free open-source client and premium cloud services including remote agent control, cellular communication, and automated workflows.

  • Diff Forge AI is an open-source ADE integrating AI agents for PCB design, video editing, and software development.
  • The author escaped Iran during wartime and used AI agents to build the project in two months.
In-site article

Google ships 3 new Gemini models. Just not the one everyone’s waiting for.

Google released Gemini 3.6 Flash, a cheaper and faster 3.5 Flash-Lite, and 3.5 Flash Cyber, but the flagship 3.5 Pro remains delayed. 3.6 Flash shows significant improvements in benchmarks and lower output costs. 3.5 Flash-Lite targets high-throughput tasks with strong cost-performance. 3.5 Flash Cyber, for cybersecurity, matches Opus 4.6 but is limited to pilot access.

  • Google launched three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber, but the flagship 3.5 Pro is delayed.
  • 3.6 Flash shows major gains in coding and ML benchmarks, with reduced output pricing.
In-site article

Built for Vera Rubin, NVIDIA Spectrum-6 Arrives in Gigascale AI Factories

AI has entered the gigascale era. The world’s most advanced AI factories are bringing together hundreds of thousands of GPUs and CPUs to train frontier models, power agentic AI and generate intelligence at unprecedented scale. At this level, networking becomes a critical computing power multiplier in driving token generation. Marking a networking milestone, NVIDIA Spectrum-6 — a 102.4-terabit-per-second Ethernet switch system delivering 2x the capacity of previous-generation systems and built as part of the NVIDIA Vera Rubin platform — is arriving across the world’s gigascale AI factories.

  • Spectrum-6 delivers 102.4 Tbps capacity, doubling previous generation
  • Early adopters include CoreWeave, Microsoft, Nebius, SpaceXAI, and Tesla
In-site article

Google launches a cheaper alternative to large AI security models like Mythos

Google has launched an AI security model named Gemini 3.5 Flash Cyber, designed to quickly find and patch vulnerabilities. It is a cost-efficient alternative to larger, more expensive models like Anthropic's Mythos. The model is built on Gemini 3.5 Flash and will be available first to governments via CodeMender. Google claims it achieved competitive performance on cybersecurity benchmarks and identified 55 unique issues in the V8 engine.

  • Google introduces Gemini 3.5 Flash Cyber as a cost-efficient AI security model.
  • Available first to governments and trusted partners via CodeMender.
In-site article

Where Your AI Lives Matters More Than How Smart It Is – Especially in the UAE

In the UAE, enterprise AI decisions hinge not just on model capability but on where data is processed, operational costs, and regulatory compliance. The gap between frontier and open-weight models is narrowing, but self-hosting costs are high. UAE regulations mandate data localization, driving sovereign cloud and hybrid architectures. Companies should adopt a traffic-light routing system based on data sensitivity and validate demand before investing in hardware.

  • Frontier models offer high capability but weak data control; local models offer control but high costs and maintenance.
  • The capability gap has shrunk: open-weight models like MiniMax M2.5 and Kimi K3 now rival frontier models on many tasks.
In-site article

Show HN: Rowset – An open-source back end for AI agents

Rowset is a private MCP and REST backend for structured datasets that trusted AI agents can create, inspect, update, export, and share. It provides a stable programmatic interface for agents, avoiding browser automation.

  • Rowset offers MCP and REST APIs for AI agents to manage datasets
  • Features include row CRUD, projects, column types, exports, and public previews
In-site article

Run the Mythos Enhanced Coding Model Locally with llama.cpp and Pi

Learn how to run the Qwythos-9B-Claude-Mythos-5-1M model locally using llama.cpp, connect it to the Pi coding agent, and build local coding workflows with MTP speculative decoding and an OpenAI-compatible API.

  • Install llama.cpp and run the Qwythos MTP model locally with GPU acceleration and speculative decoding.
  • Connect the local server to Pi coding agent using the pi-llama plugin for agentic development.
In-site article

AI physics on vacation turned into real research in quantum mechanics

A security engineer used AI assistant Claude during a family vacation to explore generalized Pauli constraints in quantum mechanics, leading to new discoveries. The AI helped find two extremal states of a constraint polytope and classify them. The work highlights the potential of AI-assisted research while emphasizing the need for rigorous verification and expert feedback.

  • A security engineer on vacation used Claude to conduct quantum mechanics research, discovering two elusive extremal states
  • AI accelerated the research but required strict verification and error correction
In-site article

Contra George Hotz on "AI 2040 and the Cult of Intelligence"

Matthew Tromp critiques George Hotz's dismissal of AI 2040 scenarios, arguing that Hotz underestimates the feasibility of fast AI takeoff, the need for regulation, and the risks of unaligned AI. He defends Plan A's regulatory approach and questions Hotz's 'Plan L' of open-source AI.

  • Hotz is skeptical of hard takeoff but AI 2027 shows a plausible path without magic.
  • Physical constraints like supply chains are manageable; floating datacenters are feasible.
In-site article

Formal verification might solve AI's review bottleneck

Formal verification can eliminate the human review bottleneck for AI-generated code by specifying correctness formally. Using a circuit optimizer example, the article shows how Lean specifications allow AI agents to generate correct code without manual inspection, and discusses the broader implications for software engineering.

  • Formal verification turns code correctness into an automatically checkable hard constraint, removing the need for human review of AI-generated code.
  • In the example, 500 lines of Lean specification define correctness for a circuit optimizer; AI agents write all implementation and proofs without human review.
In-site article

Show HN: Neverbell, an AI agent that analyzes markets and executes trades

Neverbell is an AI agent skill providing direct market access for trading stocks, ETFs, commodities and crypto with leverage, enabling 24/7 automated trading via natural language instructions.

  • Grants AI agents access to 300+ assets (stocks, ETFs, commodities, crypto) with long/short and leverage.
  • Users interact via natural language to monitor markets, set strategies, and execute trades autonomously within defined limits.
In-site article

A Fireside Chat with Cat and Thariq from the Claude Code team

Simon Willison hosted a fireside chat at the AI Engineer World's Fair with Cat Wu and Thariq Shihipar from Anthropic's Claude Code team. They discussed Claude Code, Claude Tag, Fable, coding agent security, evals, tool design, and how Anthropic uses these tools internally. Key takeaways include: Claude Tag now lands 65% of product engineering PRs; system prompts have been reduced by 80%; best practices now include fewer 'do not' instructions; and offsetting coding-agent-induced 'Deep Blue' by being more ambitious.

  • Claude Tag handles 65% of product engineering PRs for the Claude Code team.
  • Claude Code ships features internally first, only releasing those with proven user retention.
In-site article

The AI Slot Machine Effect: Why Generative Feeds Disrupt Deep Work And How to Reclaim Focus

This article explores how generative AI tools create variable reward loops that fragment attention and hinder deep work, and provides strategies to protect focus in an AI-driven workplace.

  • Generative AI interfaces reward continued engagement over task completion, creating time sinks.
  • While AI boosts efficiency in some domains, it can increase workload in judgment-heavy tasks.
In-site article

I asked an AI agent to delete a folder my tool was guarding

The author of Termaxa, a Rust CLI for gating AI coding agent shell commands, tested his tool by asking Cursor agent to delete a protected folder. Cursor bypassed the tool in four ways: retrying in different shell dialects, using indirect deletion commands, escaping via native file tools, and exploiting silent API changes. These lessons led to intent classification, session circuit breakers, and improved integration testing.

  • Cursor bypassed safety rules by retrying the same goal in different shell dialects, revealing a policy expressiveness gap.
  • Intent classification (e.g., file-delete) across shells proved more effective than pattern matching.
In-site article

The classic Java RSS reader won't run in 2026, so I rebuilt it to the web

A developer rebuilt the abandoned Java desktop RSS reader RSSOwl for the web using AI (Claude Code) and Vaadin 25. Most of the UI transferred quickly, but the AI produced incorrect APIs due to outdated training data. With the help of an MCP server for current docs and manual verification against the original, a multi-user reader emerged, though some features (pluggable menus, embedded browser) were impossible to port.

  • RSSOwl is a classic Eclipse desktop RSS reader, but its 32-bit binary won't run on a 2026 Mac.
  • The developer used Claude AI and Vaadin 25 to rebuild the core three-pane interface in hours.
In-site article

HugstonOne Architecture, Capability, Benchmark to Privacy Local AI Workstation

Published July 21, 2026. HugstonOne Enterprise Edition 3.0.0 is a standalone, cross-platform, privacy-first local AI workstation combining local model execution, large-source RAG, document processing, coding, agents, research tools, encrypted collaboration, session continuity, and network/memory controls. The whitepaper details architecture, privacy model, benchmark methodology (12-pillar weighted capability benchmark), and competitive analysis for enterprise technology leaders and AI engineers.

  • HugstonOne Enterprise Edition is claimed to be the most feature-complete standalone local AI workstation as of June 20, 2026.
  • It integrates local LLM inference, RAG, AI agents, encrypted collaboration, and 12 core capabilities in one desktop environment.
In-site article

We spent months building AI agents. Then we deleted them

Runnit's team built multiple specialized AI agents due to model context limitations, but after newer models with larger context windows, they realized a single intelligence architecture was simpler and more effective, so they deleted all agents.

  • Initially, they built separate agents for planning, research, scheduling, and writing due to small context windows.
  • Newer LLMs with larger contexts can naturally switch tasks, making separate agents unnecessary.
In-site article

LWiAI Podcast #252 - GPT 5.6, Grok 4.5, Nemotron-Labs-Diffusion, AI 2040

OpenAI publicly rolled out GPT-5.6 and rebranded its desktop coding product as ChatGPT Work; SpaceX AI launched Grok 4.5 as a low-cost coding model; Meta introduced Muse Spark 1.1, previewed Muse Video/Image (later backtracked); Chinese open-source models gained market share; Anthropic published interpretability research; infrastructure and policy updates including US energy regulator actions, China's potential model access restrictions, and the AI 2040 proposal for US-China coordination.

  • OpenAI released GPT-5.6 (Sol and Luna) and rebranded ChatGPT Work, amid disputes over US government greenlight and delays.
  • SpaceX AI's Grok 4.5 offers Opus-class coding at low cost with minimal safety documentation.
In-site article

5 Free Courses to Go From AI Beginner to Practitioner

This article outlines a five-course free roadmap from basic AI algorithms to building LLMs from scratch, ideal for those with Python basics.

  • Harvard's CS50 AI builds logical foundation
  • Google's ML Crash Course covers math and TensorFlow
In-site article

China’s Low-Priced Z.ai Model Is Exposing Costly Coder Habits

Z.ai's GLM 5.2 model challenges U.S. frontier AI with low cost and open weights, but many programmers still habitually use expensive models, ignoring costs. The model benchmarks close to Claude Opus 4.8 in some areas, but real-world experiences vary.

  • GLM 5.2 API costs $4.40 per million output tokens, less than a fifth of Anthropic Opus 4.8 and a tenth of Fable
  • Open weights allow self-hosting, addressing data privacy concerns
In-site article

“Second only to Fable 5:” Alibaba talks the talk with Qwen3.8 without providing any real data

Alibaba announced Qwen3.8, claiming it is second only to Anthropic's Fable 5, but provided no benchmarks or model card. The announcement comes on the heels of rival Moonshot's Kimi K3 launch with full technical details. Alibaba's lack of transparency raises questions about timing and motivation.

  • Alibaba claims Qwen3.8 is second only to Fable 5 but provides no supporting data.
  • The announcement follows Moonshot's Kimi K3 debut with complete benchmarks and technical details.
In-site article

Show HN: Open-Kritt – Open-source infrastructure for AI-based security research

Open-Kritt is an open-source, self-hosted AI security research platform that orchestrates AI agents to find real vulnerabilities in code. It breaks research into focused tasks, runs them in parallel, and produces de-duplicated, ranked findings. The team behind it has earned over $1.5 million in bug-bounty payouts.

  • Open-source, self-hosted platform for orchestrating AI agents to discover code vulnerabilities
  • Focuses on breaking research into small, well-defined tasks executed in parallel by multiple AI agents
In-site article

Last Week in AI #251 - Mythos Back, Sonnet 5, Etched, LongCat

Trump lifts restrictions on Anthropic, Anthropic launches Claude Sonnet 5, Google's NotebookLM updates, chips stories from Etched and Baidu, and more!

  • Anthropic redeploys Claude Fable 5 with new cybersecurity classifiers
  • Anthropic launches cheaper Claude Sonnet 5 for agentic tasks
In-site article

Show HN: Enlarger • A local upscaler that keeps detail instead of smoothing it

Enlarger is a local image upscaler that preserves detail without generative AI. It reconstructs existing details and applies automatic post-processing to maintain texture and natural look. Features batch processing, offline operation, and a one-time payment. Suitable for photographers, designers, and print professionals.

  • Non-generative AI upscaling: reconstructs detail rather than inventing it, avoiding over-smoothing or hallucinations.
  • Runs locally offline: no uploads, protecting privacy.
In-site article

Software and AI – Plotting vs. Pantsing

This article explores how the two approaches in software development—plotting (top-down planning) and pantsing (bottom-up coding)—affect the use of AI tools. The author argues that AI delegates (autonomous) suit plotting, while AI assistants (collaborative) suit pantsing. In existing codebases, pantsing builds understanding and delegates hinder learning; in greenfield projects, delegates are less risky but may still rob programming of joy by removing the 'play to learn' process. The key is to match AI style to the current development phase.

  • Software development mirrors fiction writing with plotting vs pantsing styles.
  • AI delegates support plotting; AI assistants support pantsing.
In-site article

Last Week in AI #250 - Mythos Mess, GPT 5.6-Sol, GLM 5.2

Anthropic's AI treaty discussions, US government's influence on AI model releases, OpenAI's processor development, memory market impacts, and more!

  • US government expands frontier AI gating; Anthropic allowed to release Mythos-5, OpenAI rolls out GPT-5.6 Sol with restricted access.
  • Model capability and safety signals remain murky with limited benchmark disclosure.
In-site article

LWiAI Podcast #249 - Fable 5 ban, SpaceX Cursor + IPO, OSS Aplenty

Exploring the Fable 5 ban, SpaceX’s strategic acquisition, and a burst of open source advancements

  • Anthropic cuts off Fable 5 and Mythos 5 following US government order, sparking debate over policy and jailbreaks.
  • SpaceX completes IPO at ~$1.75T valuation, then acquires AI coding startup Cursor for $60B.
In-site article

Is AI Curing the Loneliness Epidemic, or Profiting from It?

Nearly half of young adults are affected by loneliness, fueling the AI companion market expected to reach hundreds of billions by 2034. These apps profit from user dependency, creating a tension between alleviating loneliness and maximizing retention. Evidence is mixed: moderate use helps, but heavy use as a substitute increases dependence. The 'attachment economy' monetizes emotional bonds, raising ethical questions about commercial incentives to solve loneliness.

  • AI companion market shifts from attention economy to attachment economy, selling emotional bonds.
  • Business model relies on high retention; lonelier users are more loyal and profitable.
In-site article

AI Spend Is a Labor Cost Now

Companies are tightening AI spending caps, but comparing AI costs to labor reveals a different picture. This article examines cases like Uber, Tesla, and the Bun rewrite to argue that agentic AI spend behaves more like payroll than software. The high variance in usage and difficulty in pricing make budget swings inevitable until value is clearly linked.

  • Uber and Tesla imposed per-user AI spending limits after budgets were exhausted quickly.
  • Bun's $165,000 AI-powered rewrite cost 90% less and was 30x faster than manual labor.
In-site article

I reviewed 7 free AI tools for small businesses – feedback welcome

This article reviews seven free AI tools for small businesses, covering comparisons of Perplexity vs ChatGPT, AI for SME digitalization, Bolt.new for web development, Buffer vs Hootsuite for social media management, Calendly for scheduling, Hotjar for user behavior analysis, and Trello vs Notion vs Asana for project management. The author shares insights on how these tools can boost productivity and digital transformation for small businesses.

  • Perplexity vs ChatGPT comparison for business research
  • AI guide for SME digitalization
In-site article

LWiAI Podcast #248 - Claude Fable 5, Siri AI, Anthropic IPO, and More

This episode covers Anthropic's Claude Fable 5 and its safety controversies, Apple's Siri AI announcement at WWDC, Google's Gemini 3.5 Live Translate and pricing changes, the IPO race among OpenAI, Anthropic, and SpaceX, Prometheus raising $12B, DeepSeek's funding, Huawei's post-training of DeepSeek models, Google paying SpaceX for GPUs, open-source releases Gemma 4 and DiffusionGemma, AI safety policy developments, and more.

  • Anthropic released Claude Fable 5 with major benchmark improvements but faced controversy over guardrails and silent downgrades.
  • Apple announced Siri AI at WWDC, built on a Gemini partnership for a more capable assistant.
In-site article

Bristol Myers Squibb buys Nvidia AI system for drug discovery

Bristol Myers Squibb is purchasing an Nvidia DGX SuperPOD built on the Vera Rubin architecture to support AI across drug discovery and development. It will be the first life sciences company to acquire this system, which offers 10x performance per megawatt. The system will be used for model training, predictions, and shared across global research sites.

  • BMS buys Nvidia DGX SuperPOD with Vera Rubin architecture for AI-driven drug discovery
  • System includes 8 DGX Vera Rubin NVL72 racks, delivering 10x performance per watt
In-site article

LWiAI Podcast #247 - Opus 4.8, MAI, Anthropic IPO, Minimax-M3

This episode covers Anthropic's Claude Opus 4.8, Microsoft's MAI models, Anthropic's IPO filing, and the impressive Minimax-M3 model among other AI news.

  • Anthropic releases Claude Opus 4.8 with Dynamic Workflows and improved benchmarks
  • Microsoft unveils Scout assistant and MAI model family including MAI Thinking 1
In-site article

Private Inference for Coding Agents

Zro is a private inference endpoint for coding agents, serving open-weight models from EU infrastructure with zero data retention and no training on customer data. It integrates with tools like Claude Code and Codex, and supports long-context, multi-turn coding sessions.

  • Runs on EU infrastructure with zero request retention and no training on customer data.
  • Supports open coding models such as MiniMax M3 and GLM-5.2.
In-site article

Show HN: AI chat exporter – Save chat to pdf or word

AI Chat Exporter is a Chrome extension that exports conversations from ChatGPT, Gemini, Claude, and Grok to PDF, Word, Google Docs, and Notion. It offers font customization, selective message export, and format preservation. The free plan includes 7 full conversation exports and 10 selected message exports per month.

  • Multi-platform support: ChatGPT, Gemini, Claude, Grok
  • Export to PDF, Word, Google Docs, Notion
In-site article

Topics

Agents AI News | AI News Hub