Announcing Baseten for Model Labs
Baseten launches a new platform for closed-weight model labs to distribute and monetize models, featuring production-grade inference infrastructure, model library distribution, IP protection, and go-to-market support. The platform already has 15 lab partners including Cartesia, Gradium, Inception, NVIDIA, and more.
Product
Announcing Baseten for Model Labs
We're excited to launch a new platform designed specifically for model labs.
Authors
Marylise Tauzia
Bola Malek
Last updated
July 29, 2026
Share
TL;DR
We're excited to launch Baseten for Model Labs: a set of products and services designed to help closed-weight model labs distribute and monetize their models quickly and easily. Baseten for Model Labs is the central hub for model consumers to access models and for model builders to distribute them.
On May 6, we introduced the Frontier Gateway, a managed inference gateway built to give closed-weight model labs a way to serve their model in production using a white-labeled API. Since then, many labs like Poolside, Subconscious, Trajectory, and Writer have successfully leveraged the gateway to monetize their models.
As these labs gained traction with developers, we quickly saw a different need emerge. Developers wanted access to these specialized and powerful models, but they didn’t want to sign up for a new API or onboard another sub-processor to use a new model in production. They wanted to access it from their preferred inference platform easily.
At the same time, our lab partners saw the opportunity to list their models on the Baseten Model Library as a compelling new distribution channel beyond the white-labeled API we provide them with the Frontier Gateway.
Seeing the strong demand from both sides is what led us to collaborate with many labs to build a platform that addresses the needs of both model labs and their downstream customers, which we are excited to launch today.
A new distribution platform for closed models
With Baseten for Model Labs, we are expanding the services we offer labs with a new set of capabilities built to help them easily monetize and scale.
We built a distribution platform for labs that want to bring their model to production but don’t want to stand up the technical and operational infrastructure to do so:
Inference infrastructure built for production out of the box: Model labs don’t need to spend months building billing, remittance, API keys, compliance authentication, authorization, compute procurement, regional support and more, to bring their model to production. Baseten handles the infrastructure and operational complexity.
Increased visibility with a highly qualified developer audience: Baseten is well-known by AI developers around the world as the fastest, most reliable, and scalable inference provider. Labs benefit from increased visibility with Baseten’s growing developer community and enterprise customers, expanding awareness among highly qualified AI practitioners.
Strong IP protection for labs’ closed-weights: Lab model weights available through the Model Library are protected by our strict distribution agreement and secure infrastructure, so end-customers can only consume the model and can’t modify it or extract its weights.
GTM Support: When a model is published on the Model Library, model labs join our growing ecosystem and benefit from an awareness boost through joint marketing and co-selling engagements with Baseten’s sales team when their model is a strong fit for a customer’s need.
In short, model labs can get a direct line to monetize their model without having to build the infrastructure, compliance, and GTM motion themselves.
Labs that are already using Baseten for Labs
We are proud to launch alongside 15 lab partners and many users who are already leveraging these models for a wide range of use cases:
Cartesia builds state-space models (SSMs) for real-time voice intelligence, featuring Sonic (TTS) and Ink (STT). Replacing traditional transformers for hyper-efficiency, it delivers sub-90ms latency across cloud, edge, or on-premise deployments. Its top use cases include real-time conversational AI agents for customer support, sales outreach, marketing, training simulations, and recruiting.
Gradium builds real-time speech models: TTS, STT, voice cloning, and voice design. Streaming-native, built for voice agents, with latency low enough for natural turn-taking and reliable pronunciation on alphanumerics. The voice library is tuned for agent use, with cloning and voice design for teams that need their own accents, pitches, and tones. Spun out of the Kyutai research lab in 2025, from the team behind Moshi.
Inception is the creator of Mercury 2, a new class of diffusion LLMs that match the quality of frontier speed-optimized models, with sub-250 ms time to first token, 1,000+ tokens per second throughput, and 70% lower cost per task. Served on Baseten, top use cases include real-time voice applications, search and retrieval pipelines, and AI coding sub-agents (e.g., context compaction).
NVIDIA provides open models such as Nemotron ASR and BioNeMo under permissive licenses for commercial use. For customers who want production-ready deployments, NVIDIA NIM packages optimized inference runtimes, standard APIs, and enterprise-grade operational capabilities. Through its distribution agreement with NVIDIA, Baseten makes it easy to deploy NIM on its inference platform while also respecting NVIDIA’s commercial license.
PyannoteAI is the creator of Pyannote, a speaker intelligence foundation model for audio diarization. It can identify who spoke, when, and how, even in complex audio, and can pair with any STT model to turn transcripts into clean, speaker-attributed metadata. With centisecond boundary precision, Pyannote delivers 300ms real-time diarization latency, whether in cloud, on-premise, or at the edge, and up to 52% lower transcription errors. Top use cases include meeting transcription, healthcare AI scribing, financial compliance, and media dubbing.
SID.ai is the creator of SID-1, a specialized search LLM that excels at multi-step agentic retrieval across unstructured datasets. SID-1 surfaces ~2x more relevant documents than embedding search without requiring reindexing. Providing agentic search at ~100x lower cost and ~20x lower latency than frontier models, it powers high-recall retrieval for legal, healthcare, customer support, general knowledge, and more.
Subconscious builds specialized inference systems for long-horizon AI agents. Using near-lossless message compression and efficient prefix/suffix caching, it expands stateful context windows by 10x, increases throughput by 3.5x, and reduces token consumption by up to 50% without sacrificing accuracy. Serving its TIM models alongside open-weight options, top use cases include coding, workflow, and browser automation, research, and legal document generation.
Synthefy’s Nori is a tabular foundation model that outperforms XGBoost in accuracy without requiring retraining, feature engineering, or hyperparameter tuning. Operating in seconds on a single GPU via Baseten through zero-shot in-context learning, its instant, in-context predictions streamline high-volume workflows like demand forecasting, fraud detection, predictive maintenance, risk scoring, churn protection, and high-frequency trading.
The list is long already, but there are many more model labs that we’re proud to partner with: Bria, Canopy Labs (Orpheus TTS), Cosine (Lumen), Krea (Turbo), MongoDB (Voyage 4), Musubi (PolicyLM-1b), Scaled Cognition (APT-1). We look forward to working with many more and seeing how our customers leverage these models to build awesome applications.
✕
The Baseten distribution platform is bringing together our lab partners and our customers. We are excited to be at the center of this thriving ecosystem.
The future is a model ecosystem
We believe the future of AI won’t run on just a few large frontier models. It will be powered by a diverse ecosystem of models, open and closed.
We are investing in a future where open-weight models coexist with dozens of specialized closed-weight models, each optimized for a specific domain, modality, or use case. For that future to thrive, the labs building those models need an efficient way to bring their innovation to production, reach the right customers, and deploy them securely at scale.
Baseten for Model Labs makes this future possible. By combining production-ready inference infrastructure with distribution through the Baseten Model Library and joint go-to-market support, we help model labs focus on building great models while we power their inference growth for the long term.
Visit Baseten for Model Labs to learn more and join our growing ecosystem of lab partners.
Talk to us
Connect with our product experts to see how we can help.
Talk to an engineer