跳到主要内容
AI News HubLIVE
公开文章 58采集文章 71可信度 82刷新频率 120 分钟
健康状态 健康来源类型 官方原文权限 官方原文最近入库 2026-09-28ID baseten-blog运行状态 已启用

Official AI inference and deployment platform blog; confirm reuse terms before full body display.

最新公开文章

待翻译:Securing the open frontier with NVIDIA OpenShell and Blaxel sandboxes

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:News Securing the open frontier with NVIDIA OpenShell and Blaxel sandboxes Baseten is a launch partner for the NVIDIA Agent Safety Platform and supports OpenShell in the latest generation of Blaxel Sandboxes. Authors Ni…

Baseten Blog站内正文待翻译:Securing the open frontier with NVIDIA OpenShell and Blaxel sandboxes

待翻译:Sheila Vashee joins Baseten as CMO

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:News Sheila Vashee joins Baseten as CMO Welcome Sheila Vashee! Authors Tuhin Srivastava Last updated September 25, 2026 Share The next era of AI will be defined by what companies can build with intelligence that is open…

Baseten Blog站内正文待翻译:Sheila Vashee joins Baseten as CMO

待翻译:Fine-tune on your LangSmith traces with Baseten Loops

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:News Fine-tune on your LangSmith traces with Baseten Loops LangSmith Fine-Tuning is in public beta today, and its open-source CLI, smithtune, trains on Baseten Loops. Authors Mudith Jayasekara Aaron Ellis-Bloor Last upd…

Baseten Blog站内正文待翻译:Fine-tune on your LangSmith traces with Baseten Loops

待翻译:NVIDIA Nemotron 3 Diarization: real-time speaker labels at a cent per audio hour

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:AI models NVIDIA Nemotron 3 Diarization: real-time speaker labels at a cent per audio hour NVIDIA Nemotron 3 Diarization answers who spoke when, 300ms–1s behind live, alongside any ASR. Authors Ansel Erol Last updated S…

Baseten Blog站内正文待翻译:NVIDIA Nemotron 3 Diarization: real-time speaker labels at a cent per audio hour

待翻译:Introducing Baseten Hosted Tools

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Product Introducing Baseten Hosted Tools Hosted tool execution on the inference backend augments model capabilities and lowers latency. Authors Sai Maddali Marius Killinger Marylise Tauzia Last updated September 16, 202…

Baseten Blog站内正文待翻译:Introducing Baseten Hosted Tools

待翻译:LangChain trains custom models for LangSmith Engine with Baseten Loops

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:News LangChain trains custom models for LangSmith Engine with Baseten Loops Loops provides managed infrastructure for fine-tuning models through an API, with support for SFT, RL, and long-context workloads. Authors Aaro…

Baseten Blog站内正文待翻译:LangChain trains custom models for LangSmith Engine with Baseten Loops

待翻译:DeepSeek-V4.1-Flash: more efficient prefill for coding agents

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:AI models DeepSeek-V4.1-Flash: more efficient prefill for coding agents Discover DeepSeek-V4.1-Flash, a 552B multimodal MoE model featuring a Causal Encoder-Decoder architecture and efficient KV caching. Authors Albert…

Baseten Blog站内正文待翻译:DeepSeek-V4.1-Flash: more efficient prefill for coding agents

待翻译:Blaxel is joining Baseten to build the future of agentic infrastructure

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:News Blaxel is joining Baseten to build the future of agentic infrastructure Baseten has acquired Blaxel Authors Amir Haghighat Tuhin Srivastava Paul Sinai Last updated September 10, 2026 Share Today, Blaxel is joining…

Baseten Blog站内正文待翻译:Blaxel is joining Baseten to build the future of agentic infrastructure

待翻译:How Baseten makes pyannote’s diarization models 9.6x faster

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Model performance How Baseten makes pyannote’s diarization models 9.6x faster Using quality-aware quantization, index-based clustering, and scheduling optimizations Authors Matte Lim Ansel Erol Last updated September 9,…

Baseten Blog站内正文待翻译:How Baseten makes pyannote’s diarization models 9.6x faster

面向编码代理的新 MCP 与技能,助力使用 Baseten

Baseten 推出远程 MCP 服务器和技能,使 Claude Code、Codex、Cursor 等编码代理能更高效地操作其平台。基准测试显示,MCP 使任务平均耗时与成本降低 7.5%,在偏重后端的任务中降幅可达约 50%。

Baseten Blog站内正文面向编码代理的新 MCP 与技能,助力使用 Baseten

待翻译:Best open-source models for post-training

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:AI models Best open-source models for post-training Compare top open-source models for post-training. Learn what drives cost and which model fits use case. Authors Chloe Florit Last updated September 2, 2026 Share TL;DR…

Baseten Blog站内正文待翻译:Best open-source models for post-training

待翻译:The efficient frontier of LLM inference

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Model performance The efficient frontier of LLM inference Inference techniques either move a deployment along the latency–throughput frontier or push the entire frontier out, creating more efficiency to allocate. Author…

Baseten Blog站内正文待翻译:The efficient frontier of LLM inference

待翻译:Agentic kernels in production

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Model performance Agentic kernels in production Baseten's agentic kernel optimization framework cuts latency by 42.3% on Qwen-Image and by 15.2% on FLUX.2. Authors Brian Li Faraz Shahsavan Pankaj Gupta Last updated Augu…

Baseten Blog站内正文待翻译:Agentic kernels in production

待翻译:GLM 5.3: Scaling with post-training, intuitively explained

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:AI models GLM 5.3: Scaling with post-training, intuitively explained GLM 5.3's gains came entirely from post-training. An intuitive look at the environment design, RL algorithms, and infrastructure behind the leap. Auth…

Baseten Blog站内正文待翻译:GLM 5.3: Scaling with post-training, intuitively explained

待翻译:The two AI gateway patterns in production inference

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:AI engineering The two AI gateway patterns in production inference Access gateways and serving gateways explained, and how they handle identity, tenancy, limits, and metering for teams serving AI models. Authors Amit Ga…

Baseten Blog站内正文待翻译:The two AI gateway patterns in production inference

待翻译:How to run any open model inside DeepSeek Harness

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:AI engineering How to run any open model inside DeepSeek Harness Learn how to power DeepSeek Harness with Baseten Model APIs and run open models like Kimi K3, GLM 5.2, and DeepSeek V4 Pro in under 5 minutes. Authors Ale…

Baseten Blog站内正文待翻译:How to run any open model inside DeepSeek Harness

待翻译:How leading platforms ensure observability for LLM inference

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Infrastructure How leading platforms ensure observability for LLM inference Learn how metrics, logs, and traces work together to catch slow responses, errors, and failed deployments before users do. Authors Chloe Florit…

Baseten Blog站内正文待翻译:How leading platforms ensure observability for LLM inference

待翻译:Inference engineering for DeepSeek V4 Pro 0813

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Model performance Inference engineering for DeepSeek V4 Pro 0813 DeepSeek V4 Pro 0813 is a 1.7T-parameter open frontier model under the MIT license and is available for inference today on Baseten model APIs. Authors Mod…

Baseten Blog站内正文待翻译:Inference engineering for DeepSeek V4 Pro 0813

待翻译:Baseten delivers open-source inference for You.com

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Community Baseten delivers open-source inference for You.com Baseten is proud to power the inference behind You.com's search and answer stack. Authors Marylise Tauzia Last updated August 13, 2026 Share TL;DR You.com bui…

Baseten Blog站内正文待翻译:Baseten delivers open-source inference for You.com

待翻译:Qwen3.8-Max: Alibaba's new frontier reasoning model

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:AI models Qwen3.8-Max: Alibaba's new frontier reasoning model Qwen3.8-Max is Alibaba's new open-weight frontier model, featuring a 1M token context window and multimodal input. Authors Albert Lee Last updated August 12,…

Baseten Blog站内正文待翻译:Qwen3.8-Max: Alibaba's new frontier reasoning model

待翻译:Introducing NVIDIA Nemotron 3.5 Lightning

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:AI models Introducing NVIDIA Nemotron 3.5 Lightning NVIDIA’s Nemotron 3.5 Lightning, now on Baseten, delivers high-throughput, efficient reasoning for faster and more accurate agentic workflows. Authors Marylise Tauzia…

Baseten Blog站内正文待翻译:Introducing NVIDIA Nemotron 3.5 Lightning

待翻译:Introducing NVIDIA Nemotron 3.5 ASR Streaming

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:AI models Introducing NVIDIA Nemotron 3.5 ASR Streaming Deploy NVIDIA Nemotron 3.5 ASR for low-latency, production-ready speech recognition with 6x higher throughput and multilingual support. Authors Ansel Erol Ian Carr…

Baseten Blog站内正文待翻译:Introducing NVIDIA Nemotron 3.5 ASR Streaming

待翻译:Laguna S 2.1 goes Greek: a repository-scale game transformation

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:AI models Laguna S 2.1 goes Greek: a repository-scale game transformation We put Poolside’s new Laguna S 2.1 model to the test, tasking it with a repository-scale transformation of the open-source game Hypersomnia. Auth…

Baseten Blog站内正文待翻译:Laguna S 2.1 goes Greek: a repository-scale game transformation

微调Qwen3-TTS实现高质量语音克隆

本文介绍如何微调Qwen3-TTS实现高质量语音克隆,对比ICL、仅说话人嵌入和微调三种方式,并给出完整训练流程。基于约1.5小时LJ Speech数据在单张H100上训练约1小时,微调模型TTFA约130ms,比ICL快约16%,并省去推理时的参考音频处理。

Baseten Blog站内正文微调Qwen3-TTS实现高质量语音克隆

22,580:从GPT-2到Kimi K3,解读其发展历程

本文追溯了从GPT-2(1.24亿参数)到Kimi K3(2.8万亿参数)的架构演进,七年间规模增长22580倍。文章解释了关键创新,包括KV缓存、线性注意力和DeltaNet,揭示了基础机制如何演化以支撑巨大规模。

Baseten Blog站内正文22,580:从GPT-2到Kimi K3,解读其发展历程

如何在任何框架中运行 Kimi K3:使用 Baseten Switch 进行路由

Baseten Switch 是一款本地 Mac 应用,可让您在流行的人工智能框架(如 Claude Code 和 Codex)中无缝混合运行开源和闭源模型,只需切换开关即可在 Kimi K3、GLM 5.2 等模型之间切换,并跟踪成本、性能和使用情况。

Baseten Blog站内正文如何在任何框架中运行 Kimi K3:使用 Baseten Switch 进行路由

Baseten宣布推出Model Labs平台

Baseten发布了专为闭权模型实验室设计的新平台Baseten for Model Labs,提供推理基础设施、模型分发、知识产权保护和市场推广支持,帮助实验室快速实现模型货币化。已有15个合作伙伴如Cartesia、Gradium、Inception、NVIDIA等参与。平台旨在促进多样化的模型生态系统发展。

Baseten Blog站内正文Baseten宣布推出Model Labs平台

将Kimi K3分词速度提升18倍,用于百万令牌的代理工作负载

Baseten推出了新的分词器(Basetenkenizer),将Kimi K3的分词速度提升多达18倍,适用于百万令牌的代理工作负载。它结合了基于Rust的优化技术,包括专用预分词、BPE合并、多核分块和零拷贝NumPy传输,同时保持与tiktoken完全一致的令牌ID。这一改进显著降低了长输入序列的首令牌时间,尤其是在前缀缓存命中率高的情况下。

Baseten Blog站内正文将Kimi K3分词速度提升18倍,用于百万令牌的代理工作负载

如何为Kimi K3构建Day-0 API

Baseten 为 Kimi K3 提供了 Day-0 支持,这是一个 2.8T 参数的新前沿开放模型。本文介绍了在发布日前运行该模型所需的技术工作,包括硬件配置、权重加载、推理引擎集成、质量验证、性能优化和规模化部署。

Baseten Blog站内正文如何为Kimi K3构建Day-0 API

推出GLM 5.2 Fast

我们推出了一个新的模型API服务层级——GLM-5.2 Fast,专为实时代理工作负载优化,提供与标准GLM-5.2相同的权重,但基础设施针对每用户吞吐量进行了调整,以实现低延迟和稳定的性能。

Baseten Blog站内正文推出GLM 5.2 Fast

全部来源