Shippy is a maritime AI agent built for high-stakes decisions, where the wrong answer has real impacts. The article covers its architecture—soul, skills, config—and key design decisions like using a deterministic CLI for API access, sandboxed hosting for user isolation, and a custom evaluation system that scores the whole agent against live data. Lessons learned and future plans are also discussed.
Model routing in AI agents is more complex than it seems. It is not a classification problem but a systems optimization problem involving cost, complexity, and latency. The article shares three key challenges and explains IBM Research's optimization-based approach.
Existing benchmarks suggest voice AI is nearing human-level performance but real-world conversations tell a different story. Hume AI introduces Real World VoiceEQ, a benchmark evaluating over 40 voice models across 15+ dimensions and 60+ metrics, based on over 1 million human ratings. Key findings include: progress is becoming specialized, models are better at speaking than listening, traditional benchmarks overestimate real-world performance, and human evaluation remains essential.
NVIDIA emphasizes the importance of open data and synthetic data for building agentic AI, highlighting data inspectability, quality, and trust. The article details Nemotron datasets, the Prompt Atlas visualization tool, and the use of synthetic personas for local diversity.
The transformers vLLM backend is now as fast (or faster) than custom vLLM implementations for many LLM architectures. Model authors can automatically leverage their transformers implementations to get ultra fast vLLM inference, for free.
Hugging Face and Amazon SageMaker AI announce a deep-link integration enabling one-click transition from model discovery to SageMaker Studio. The integration pre-configures permissions, surfaces GPU quotas, and supports model customization and deployment, streamlining the path from inspiration to enterprise deployment.
SkyPilot and Hugging Face collaborate to allow users to store models and datasets on the Hub while running compute on any cloud without egress fees for reads.
LeRobot v0.6.0 introduces world model policies (VLA-JEPA, FastWAM, LingBot-VA), new VLAs (GR00T N1.7, MolmoAct2, etc.), reward model API (Robometer, TOPReward), six new simulation benchmarks, and a deployment CLI with DAgger corrections, depth sensing, automatic language annotation, up to 2x faster data loading, cloud training, and a leaner install—all aimed at closing the robot learning loop.
This article details the data pipeline behind PRX, a 7B text-to-image model. Key aspects include assembling a diverse pre-training dataset from public and internal sources, using long accurate captions generated by a VLM, and employing Lance for dataset building and MDS for streaming. The team explains their choice of JPEG encoding at quality 92, on-the-fly text latent computation, and lessons learned about data fragmentation.
Hugging Face's Kernels project, aimed at standardizing custom kernel packaging, distribution, and consumption, has undergone a major redesign. This post summarizes key updates: a new 'kernel' repository type for better discoverability; enhanced security through trusted publishers and code signing; revamped CLIs with clearer separation of concerns; expanded framework support including Torch Stable ABI and Apache TVM FFI; a foundation for agentic kernel development; and miscellaneous improvements like simplified environment setup and compatibility checking.
Hugging Face and Cerebras have collaborated to create a real-time voice AI system powered by Gemma 4, achieving dramatically lower latency through an open modular architecture. The pipeline integrates Nvidia's speech recognition, Cerebras's fast inference, and Alibaba's text-to-speech, and is already deployed in over 9,000 Reachy Mini robots.
IBM Research introduces ScarfBench, an open benchmark for evaluating AI agents on cross-framework migration tasks in Enterprise Java. The benchmark includes 34 applications, 102 framework implementations, and 204 migration tasks. Current top agents achieve less than 10% behavioral success, highlighting the difficulty of preserving behavior during migration.
This article argues that specialization is an inevitable consequence of finite resources and selection pressure, drawing from optimization theory (No Free Lunch theorems), evolutionary biology, competitive markets, and machine learning. It distinguishes specialization from domain knowledge and addresses the Bitter Lesson, concluding that scaling does not eliminate the need for focused systems.
Every Eval Ever (EEE) and Hugging Face Community Evals are now intercompatible, allowing cross-posting and interpretation of evaluation results with links to open models, leaderboards, and a unified standardized metadata store.
DiScoFormer is a transformer that estimates both density and score of a distribution from a set of data points in a single forward pass without retraining. It uses cross-attention, a shared backbone with two heads, and a consistency loss to adapt to new distributions. It significantly outperforms KDE in high dimensions.
Spin up a private, OpenAI-compatible LLM endpoint on Hugging Face infrastructure with a single command — no servers to provision, no Kubernetes, pay-per-second. Covers the full process from launch, querying, cleanup, scaling to larger models, creating a chat UI, SSH debugging, and using as a coding agent backend, with a comparison to Inference Endpoints.
Ai2 compares its 7B transformer Olmo 3 and hybrid Olmo Hybrid, finding the hybrid excels on content words (nouns, verbs, adjectives) and tokens requiring context, but loses advantage on repeated tokens and closing brackets. Token-level loss filtering reveals architectural differences.
NVIDIA NeMo AutoModel builds on HuggingFace Transformers v5, adding Expert Parallelism, DeepEP fused all-to-all dispatch, and TransformerEngine kernels to achieve 3.4-3.7x higher training throughput and 29-32% less GPU memory for fine-tuning MoE models, with no API changes.
CUGA is IBM's open-source agent harness that handles the plumbing of building agentic apps, leaving developers to write only a tool list and a prompt. This article walks through one example — an IBM Cloud advisor app — and explains how CUGA's planning, reflection, and policy system enable robust, production-ready agents.
This article explores the Cross-Origin Storage (COS) API proposal, which enables web apps to share large files (like AI models and Wasm runtimes) across origins using cryptographic hashes instead of URLs. Using Transformers.js as an example, it highlights the redundancy caused by current cache partitioning and how COS addresses it with hash-based identification, flexible access control, and integrity verification.
Hugging Face revamped the release process for huggingface_hub, using AI and open tools to ship weekly releases instead of monthly, while keeping a human in the loop for final review. The new pipeline costs about $0.25 per release and has improved release note quality and discovery of integration issues.
PP-OCRv6 is PaddleOCR's latest universal OCR model family, scaling from 1.5M to 34.5M parameters across three tiers, supporting 50 languages. It delivers a +4.6 percentage point improvement in text detection Hmean and +5.1 in recognition accuracy over PP-OCRv5_server. New architecture includes PPLCNetV4 backbone, RepLKFPN for detection, and EncoderWithLightSVTR for recognition. Supports multiple inference backends: Paddle Inference, Transformers, and ONNX Runtime.
A maintainer of OpenClaw built a system using local open-weight models (Gemma, Qwen) in an agent harness to triage issues and pull requests in real-time, achieving competitive performance with closed models while running on local hardware for minimal cost.
Deep research agents that combine private documents with web search can inadvertently leak sensitive information through their query logs. The MosaicLeaks benchmark quantifies this privacy risk and proposes a training method called Privacy-Aware Deep Research (PA-DR) that reduces information leakage by over 3x while maintaining task performance.
LoRA is the most popular parameter-efficient fine-tuning (PEFT) technique, but research shows other methods can outperform it on certain tasks. This article introduces Hugging Face's PEFT library and its benchmarks, discussing how to choose the right PEFT technique based on specific needs, and points out that LoRA is not always the best choice.
A new benchmark harness evaluates the entire process of AI agents using software libraries, using Hugging Face Transformers as a case study. By measuring token usage, time, and error rates across different models and tooling tiers, the authors uncover tradeoffs between ease of use and resource consumption, providing insights for library maintainers and agent users.
MolmoMotion is a new 3D motion forecasting model that predicts future 3D point trajectories of objects given a video frame, 3D points on an object, and a language instruction. It outperforms existing methods in robotics planning and controllable video generation. The model is accompanied by the MolmoMotion-1M dataset and PointMotionBench benchmark.
AWS's open-source SDK Strands Robots integrates LeRobot, enabling developers to train from Hub datasets and deploy policies on simulated or real robots through a single Agent workflow. This post walks through five steps with a runnable example on a laptop.
Z.AI introduces GLM-5.2, a flagship model for long-horizon tasks with a solid 1M-token context, advanced coding capabilities with flexible effort levels, and an open-source MIT license. It achieves top-tier performance on long-horizon coding benchmarks, rivaling closed-source models.