Skip to content
AI News HubLIVE
Public articles 34Collected articles 39Trust 82Refresh 120 min
Health HealthySource type OfficialFull-text rights Official full textLast ingested 2026-09-26ID modal-blogStatus Enabled

Official AI infrastructure blog; confirm reuse terms before full body display.

Latest public articles

How Botika runs full-stack generative AI on Modal

Botika builds agentic e-commerce tools for fashion brands and runs its entire AI stack on Modal: a 100-terabyte data pipeline, custom foundation-model research, and about 15 production inference models. CEO Eran Dagan says Modal removed the cloud-engineering overhead that once slowed them down and lets the team scale traffic dramatically without operational burden.

Modal BlogIn-site articleHow Botika runs full-stack generative AI on Modal

Modal expands in Europe with new London office

Modal is opening a London office as part of a larger European investment. The company already has a Stockholm engineering team of about 20 people and now expands to London to build GTM teams, with plans to grow engineering and other functions later. European AI companies like Black Forest Labs and Legora, as well as global platforms like DoorDash, are using Modal.

Modal BlogIn-site articleModal expands in Europe with new London office

Qwen3.8-2.4T-A95B now available on Modal

Qwen3.8-2.4T-A95B by Alibaba, with a 1M token context window, is now available via Modal Auto Endpoints.

Modal BlogIn-site articleQwen3.8-2.4T-A95B now available on Modal

Bringing serverless functions closer to the speed of wire

Modal’s Function Call data path is now >50ms faster. Our new routing layer is geographically distributed, so you can further reduce your network overhead.

Modal BlogIn-site articleBringing serverless functions closer to the speed of wire

A note on the Hugging Face agent incident

Hugging Face published a technical timeline of a recent agent intrusion, naming Modal as the third-party infrastructure used. Modal confirmed its platform and isolation were not compromised.

Modal BlogIn-site articleA note on the Hugging Face agent incident

Kimi K3 by Moonshot now available on Modal

Moonshot has released Kimi K3, a 2.8 trillion parameter multimodal model, now available on Modal at 460 tokens per second. The model features a mixture-of-experts architecture, 1M token context window, native vision, and is optimized with a custom DFlash speculator for faster inference.

Modal BlogIn-site articleKimi K3 by Moonshot now available on Modal

Devin Outposts on Modal

Cognition's Devin, an AI software engineer, can now run in Modal sandboxes via Devin Outposts, enabling customized environments with GPU support and fast cold-starts.

Modal BlogIn-site articleDevin Outposts on Modal

Inkling by Thinking Machines now available on Modal

Thinking Machines has released Inkling, a general-purpose multimodal model accepting text, image, and audio inputs, now available on Modal as a Managed Endpoint with token-based pricing. The post details its architecture, including local attention and DFlash speculation for fast inference.

Modal BlogIn-site articleInkling by Thinking Machines now available on Modal

How to price serverless GPUs

To compare rates for serverless and reserved GPUs, look at your application's peak-to-average ratio.

Modal BlogIn-site articleHow to price serverless GPUs

Multi-token Residual Prediction

Multi-token Residual Prediction (MRP) is a lightweight module for diffusion language models that predicts the residual between adjacent denoising steps rather than the full distribution. This enables both lossless speedup in static regimes (up to 1.56×) and quality recovery in dynamic regimes (up to +16 accuracy points) without trade-offs.

Modal BlogIn-site articleMulti-token Residual Prediction

Anthropic integration with Modal brings scalable compute to Claude Science

Anthropic launches Claude Science, an AI workbench for life sciences researchers, integrated with Modal to provide elastic compute infrastructure for running data processing, structure prediction, and molecule design directly from a conversation.

Modal BlogIn-site articleAnthropic integration with Modal brings scalable compute to Claude Science

Routing for serverless servers with Pingora, Envoy, and Spanner

A deep dive inside Modal's new ultra-low-latency serverless server product, explaining the architecture decisions behind building a custom proxy (fprs) using Pingora, Envoy at the edge, and Spanner for configuration, all optimized for LLM inference workloads.

Modal BlogIn-site articleRouting for serverless servers with Pingora, Envoy, and Spanner

Achieve state-of-the-art inference latencies with speculative decoding

Modal and Decagon collaborated to cut inference latency by 100ms using speculative decoding, outperforming proprietary providers. The article details the low-latency playbook including optimization of communication, host overhead, prefill, and decode latencies, with a focus on custom speculative decoding models (DFlash) for big wins.

Modal BlogIn-site articleAchieve state-of-the-art inference latencies with speculative decoding

Introducing Modal Auto Endpoints: Optimized inference you actually own

Modal launches Auto Endpoints, a self-serve on-ramp to production-grade LLM inference, allowing users to deploy frontier open models with a single command and gain full visibility and control over inference code, metrics, and infrastructure. Built on Modal's AI infrastructure platform, it features high-performance autoscaling, custom container runtime, global GPU availability, and Modal Servers for ultra-low-latency routing (5ms overhead). Pre-tuned recipes from top-tier team experience and DFlash speculative decoding are included. Future roadmap includes full automation of inference engineering.

Modal BlogIn-site articleIntroducing Modal Auto Endpoints: Optimized inference you actually own

Speculation Is All You Need

Modal is all-in on speculative decoding, arguing it's the single most important inference optimization, delivering 2-3x speedups. They released state-of-the-art DFlash speculators for Qwen models, achieving 5-20% extra speedups, and explain the theory, simulation, and math behind the acceleration.

Modal BlogIn-site articleSpeculation Is All You Need

Reinforcement Learning is an Infrastructure Problem

This article explores the practical application of reinforcement learning in post-training large language models, highlighting that the current bottleneck is infrastructure rather than algorithms. Modal shares its experience running RL post-training at scale and introduces its open-source library to help teams address key challenges like multi-node training, environment management, and GPU utilization.

Modal BlogIn-site articleReinforcement Learning is an Infrastructure Problem

Role-Based Access Control for Humans and Agents

Modal introduces Role-Based Access Control for all Team and Enterprise users, built on Environments to provide granular permissions for humans and AI agents.

Modal BlogIn-site articleRole-Based Access Control for Humans and Agents

Modal's Series C: Raising $355M at a $4.65B valuation

Modal raised $355M at a $4.65B valuation, led by General Catalyst and Redpoint. The company has grown fivefold since September, exceeding $300M in annualized revenue. Modal provides a cloud platform for AI, focusing on elastic inference, agent runtimes, and sandboxes. The funding will support expansion in low-latency inference, reinforcement learning, and agent compute.

Modal BlogIn-site articleModal's Series C: Raising $355M at a $4.65B valuation

Scaling Reinforcement Learning at Applied Compute

Applied Compute trains custom AI agents for enterprises using Reinforcement Learning, focusing on post-training to differentiate from commoditized frontier models. Their 'Specific Intelligence' approach leverages Modal's infrastructure for fast, flexible, and reliable RL training loops serving clients like DoorDash, Cognition, and Mercor.

Modal BlogIn-site articleScaling Reinforcement Learning at Applied Compute

Introducing Claude Managed Agents with Modal Sandboxes

Anthropic and Modal announce the integration of Claude Managed Agents with Modal Sandboxes, allowing developers to run tool calls in self-hosted, customizable sandboxes with fast startup, cost efficiency, and scalability. The collaboration enables secure, isolated execution environments for AI agents, with early adopters like Mason AI, DoorDash, and Blend sharing positive experiences.

Modal BlogIn-site articleIntroducing Claude Managed Agents with Modal Sandboxes

How to achieve truly serverless GPUs

Modal's deep engineering reduces GPU inference server boot times from kiloseconds to tens of seconds, enabling truly serverless computing for variable inference workloads.

Modal BlogIn-site articleHow to achieve truly serverless GPUs

Boosting multimodal inference performance by >10% with a single Python dictionary

By profiling SGLang's scheduler, Modal engineers discovered a bottleneck from repeated CUDA IPC pool handle opens. Replacing them with a simple Python dictionary cache improved throughput by 16.2% and reduced latency by over 10% on Qwen2.5-VL-3B. The optimization is merged in SGLang v0.5.10.

Modal BlogIn-site articleBoosting multimodal inference performance by >10% with a single Python dictionary

Building an RL Theorem-Proving Workflow on Modal

AE Studio used Modal to compare Evolution Strategies and GRPO for training LLMs on math theorem proving. By leveraging Modal's parallel GPUs, sandboxed verification, and volume storage, they reduced setup time by 60% and cost by up to 75%. Early results show ES matching or outperforming GRPO in several scenarios, especially with limited data.

Modal BlogIn-site articleBuilding an RL Theorem-Proving Workflow on Modal

Building with Modal and the OpenAI Agents SDK

Modal becomes an official sandbox provider for the OpenAI Agents SDK. This article demonstrates how to build a custom coding agent harness from scratch, integrating Modal sandboxes for secure, parallel, and scalable automation, using the Parameter Golf challenge as an example.

Modal BlogIn-site articleBuilding with Modal and the OpenAI Agents SDK

Autoscaling Autoresearch: Give your agents elastic GPUs on Modal

Modal integrates with Autoresearch to provide elastic GPU scaling, allowing AI agents to dynamically provision compute resources. In a Parameter Golf challenge, an agent ran 113 experiments across 238 GPU-hours, achieving 5x speedup over a single workstation while using a fraction of a dedicated cluster's resources.

Modal BlogIn-site articleAutoscaling Autoresearch: Give your agents elastic GPUs on Modal

Butter is joining Modal

Modal announces that Butter, an AI sandbox technology company, is joining Modal. Founder Erik Dunteman and researcher Raymond Tana will join the Modal Sandbox team. Butter's expertise includes agent harness engineering and the development of bVisor, a lightweight ephemeral sandbox built with Zig.

Modal BlogIn-site articleButter is joining Modal

Real-time inference for robots at Physical Intelligence

Physical Intelligence uses Modal to achieve low-latency remote real-time inference for robots, with a specialized QUIC-based transport adding only 10-15 ms of network overhead, enabling experimentation with larger models.

Modal BlogIn-site articleReal-time inference for robots at Physical Intelligence

All sources