跳到主要內容
AI News HubLIVE
公開文章 58採集文章 71可信度 82刷新頻率 120 分鐘
健康狀態 健康來源類型 官方原文權限 官方原文最近入庫 2026-09-28ID baseten-blog運行狀態 已啟用

Official AI inference and deployment platform blog; confirm reuse terms before full body display.

最新公開文章

待翻譯:Securing the open frontier with NVIDIA OpenShell and Blaxel sandboxes

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:News Securing the open frontier with NVIDIA OpenShell and Blaxel sandboxes Baseten is a launch partner for the NVIDIA Agent Safety Platform and supports OpenShell in the latest generation of Blaxel Sandboxes. Authors Ni…

Baseten Blog站內正文待翻譯:Securing the open frontier with NVIDIA OpenShell and Blaxel sandboxes

待翻譯:Sheila Vashee joins Baseten as CMO

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:News Sheila Vashee joins Baseten as CMO Welcome Sheila Vashee! Authors Tuhin Srivastava Last updated September 25, 2026 Share The next era of AI will be defined by what companies can build with intelligence that is open…

Baseten Blog站內正文待翻譯:Sheila Vashee joins Baseten as CMO

待翻譯:Fine-tune on your LangSmith traces with Baseten Loops

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:News Fine-tune on your LangSmith traces with Baseten Loops LangSmith Fine-Tuning is in public beta today, and its open-source CLI, smithtune, trains on Baseten Loops. Authors Mudith Jayasekara Aaron Ellis-Bloor Last upd…

Baseten Blog站內正文待翻譯:Fine-tune on your LangSmith traces with Baseten Loops

待翻譯:NVIDIA Nemotron 3 Diarization: real-time speaker labels at a cent per audio hour

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:AI models NVIDIA Nemotron 3 Diarization: real-time speaker labels at a cent per audio hour NVIDIA Nemotron 3 Diarization answers who spoke when, 300ms–1s behind live, alongside any ASR. Authors Ansel Erol Last updated S…

Baseten Blog站內正文待翻譯:NVIDIA Nemotron 3 Diarization: real-time speaker labels at a cent per audio hour

待翻譯:Introducing Baseten Hosted Tools

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Product Introducing Baseten Hosted Tools Hosted tool execution on the inference backend augments model capabilities and lowers latency. Authors Sai Maddali Marius Killinger Marylise Tauzia Last updated September 16, 202…

Baseten Blog站內正文待翻譯:Introducing Baseten Hosted Tools

待翻譯:LangChain trains custom models for LangSmith Engine with Baseten Loops

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:News LangChain trains custom models for LangSmith Engine with Baseten Loops Loops provides managed infrastructure for fine-tuning models through an API, with support for SFT, RL, and long-context workloads. Authors Aaro…

Baseten Blog站內正文待翻譯:LangChain trains custom models for LangSmith Engine with Baseten Loops

待翻譯:DeepSeek-V4.1-Flash: more efficient prefill for coding agents

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:AI models DeepSeek-V4.1-Flash: more efficient prefill for coding agents Discover DeepSeek-V4.1-Flash, a 552B multimodal MoE model featuring a Causal Encoder-Decoder architecture and efficient KV caching. Authors Albert…

Baseten Blog站內正文待翻譯:DeepSeek-V4.1-Flash: more efficient prefill for coding agents

待翻譯:Blaxel is joining Baseten to build the future of agentic infrastructure

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:News Blaxel is joining Baseten to build the future of agentic infrastructure Baseten has acquired Blaxel Authors Amir Haghighat Tuhin Srivastava Paul Sinai Last updated September 10, 2026 Share Today, Blaxel is joining…

Baseten Blog站內正文待翻譯:Blaxel is joining Baseten to build the future of agentic infrastructure

待翻譯:How Baseten makes pyannote’s diarization models 9.6x faster

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Model performance How Baseten makes pyannote’s diarization models 9.6x faster Using quality-aware quantization, index-based clustering, and scheduling optimizations Authors Matte Lim Ansel Erol Last updated September 9,…

Baseten Blog站內正文待翻譯:How Baseten makes pyannote’s diarization models 9.6x faster

面向編碼代理的新 MCP 與技能,助力使用 Baseten

Baseten 推出遠端 MCP 伺服器和技能,使 Claude Code、Codex、Cursor 等編碼代理能更高效地操作其平臺。基準測試顯示,MCP 使任務平均耗時與成本降低 7.5%,在偏重後端的任務中降幅可達約 50%。

Baseten Blog站內正文面向編碼代理的新 MCP 與技能,助力使用 Baseten

待翻譯:Best open-source models for post-training

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:AI models Best open-source models for post-training Compare top open-source models for post-training. Learn what drives cost and which model fits use case. Authors Chloe Florit Last updated September 2, 2026 Share TL;DR…

Baseten Blog站內正文待翻譯:Best open-source models for post-training

待翻譯:The efficient frontier of LLM inference

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Model performance The efficient frontier of LLM inference Inference techniques either move a deployment along the latency–throughput frontier or push the entire frontier out, creating more efficiency to allocate. Author…

Baseten Blog站內正文待翻譯:The efficient frontier of LLM inference

待翻譯:Agentic kernels in production

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Model performance Agentic kernels in production Baseten's agentic kernel optimization framework cuts latency by 42.3% on Qwen-Image and by 15.2% on FLUX.2. Authors Brian Li Faraz Shahsavan Pankaj Gupta Last updated Augu…

Baseten Blog站內正文待翻譯:Agentic kernels in production

待翻譯:GLM 5.3: Scaling with post-training, intuitively explained

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:AI models GLM 5.3: Scaling with post-training, intuitively explained GLM 5.3's gains came entirely from post-training. An intuitive look at the environment design, RL algorithms, and infrastructure behind the leap. Auth…

Baseten Blog站內正文待翻譯:GLM 5.3: Scaling with post-training, intuitively explained

待翻譯:The two AI gateway patterns in production inference

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:AI engineering The two AI gateway patterns in production inference Access gateways and serving gateways explained, and how they handle identity, tenancy, limits, and metering for teams serving AI models. Authors Amit Ga…

Baseten Blog站內正文待翻譯:The two AI gateway patterns in production inference

待翻譯:How to run any open model inside DeepSeek Harness

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:AI engineering How to run any open model inside DeepSeek Harness Learn how to power DeepSeek Harness with Baseten Model APIs and run open models like Kimi K3, GLM 5.2, and DeepSeek V4 Pro in under 5 minutes. Authors Ale…

Baseten Blog站內正文待翻譯:How to run any open model inside DeepSeek Harness

待翻譯:How leading platforms ensure observability for LLM inference

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Infrastructure How leading platforms ensure observability for LLM inference Learn how metrics, logs, and traces work together to catch slow responses, errors, and failed deployments before users do. Authors Chloe Florit…

Baseten Blog站內正文待翻譯:How leading platforms ensure observability for LLM inference

待翻譯:Inference engineering for DeepSeek V4 Pro 0813

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Model performance Inference engineering for DeepSeek V4 Pro 0813 DeepSeek V4 Pro 0813 is a 1.7T-parameter open frontier model under the MIT license and is available for inference today on Baseten model APIs. Authors Mod…

Baseten Blog站內正文待翻譯:Inference engineering for DeepSeek V4 Pro 0813

待翻譯:Baseten delivers open-source inference for You.com

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Community Baseten delivers open-source inference for You.com Baseten is proud to power the inference behind You.com's search and answer stack. Authors Marylise Tauzia Last updated August 13, 2026 Share TL;DR You.com bui…

Baseten Blog站內正文待翻譯:Baseten delivers open-source inference for You.com

待翻譯:Qwen3.8-Max: Alibaba's new frontier reasoning model

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:AI models Qwen3.8-Max: Alibaba's new frontier reasoning model Qwen3.8-Max is Alibaba's new open-weight frontier model, featuring a 1M token context window and multimodal input. Authors Albert Lee Last updated August 12,…

Baseten Blog站內正文待翻譯:Qwen3.8-Max: Alibaba's new frontier reasoning model

待翻譯:Introducing NVIDIA Nemotron 3.5 Lightning

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:AI models Introducing NVIDIA Nemotron 3.5 Lightning NVIDIA’s Nemotron 3.5 Lightning, now on Baseten, delivers high-throughput, efficient reasoning for faster and more accurate agentic workflows. Authors Marylise Tauzia…

Baseten Blog站內正文待翻譯:Introducing NVIDIA Nemotron 3.5 Lightning

待翻譯:Introducing NVIDIA Nemotron 3.5 ASR Streaming

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:AI models Introducing NVIDIA Nemotron 3.5 ASR Streaming Deploy NVIDIA Nemotron 3.5 ASR for low-latency, production-ready speech recognition with 6x higher throughput and multilingual support. Authors Ansel Erol Ian Carr…

Baseten Blog站內正文待翻譯:Introducing NVIDIA Nemotron 3.5 ASR Streaming

待翻譯:Laguna S 2.1 goes Greek: a repository-scale game transformation

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:AI models Laguna S 2.1 goes Greek: a repository-scale game transformation We put Poolside’s new Laguna S 2.1 model to the test, tasking it with a repository-scale transformation of the open-source game Hypersomnia. Auth…

Baseten Blog站內正文待翻譯:Laguna S 2.1 goes Greek: a repository-scale game transformation

微調Qwen3-TTS實現高質量語音克隆

本文介紹如何微調Qwen3-TTS實現高質量語音克隆,對比ICL、僅說話人嵌入和微調三種方式,並給出完整訓練流程。基於約1.5小時LJ Speech資料在單張H100上訓練約1小時,微調模型TTFA約130ms,比ICL快約16%,並省去推理時的參考音訊處理。

Baseten Blog站內正文微調Qwen3-TTS實現高質量語音克隆

22,580:從GPT-2到Kimi K3,解讀其發展歷程

本文追溯了從GPT-2(1.24億引數)到Kimi K3(2.8萬億引數)的架構演進,七年間規模增長22580倍。文章解釋了關鍵創新,包括KV快取、線性注意力和DeltaNet,揭示了基礎機制如何演化以支撐巨大規模。

Baseten Blog站內正文22,580:從GPT-2到Kimi K3,解讀其發展歷程

如何在任何框架中執行 Kimi K3:使用 Baseten Switch 進行路由

Baseten Switch 是一款本地 Mac 應用,可讓您在流行的人工智慧框架(如 Claude Code 和 Codex)中無縫混合執行開源和閉源模型,只需切換開關即可在 Kimi K3、GLM 5.2 等模型之間切換,並跟蹤成本、效能和使用情況。

Baseten Blog站內正文如何在任何框架中執行 Kimi K3:使用 Baseten Switch 進行路由

Baseten宣佈推出Model Labs平臺

Baseten釋出了專為閉權模型實驗室設計的新平臺Baseten for Model Labs,提供推理基礎設施、模型分發、智慧財產權保護和市場推廣支援,幫助實驗室快速實現模型貨幣化。已有15個合作伙伴如Cartesia、Gradium、Inception、NVIDIA等參與。平臺旨在促進多樣化的模型生態系統發展。

Baseten Blog站內正文Baseten宣佈推出Model Labs平臺

將Kimi K3分詞速度提升18倍,用於百萬令牌的代理工作負載

Baseten推出了新的分詞器(Basetenkenizer),將Kimi K3的分詞速度提升多達18倍,適用於百萬令牌的代理工作負載。它結合了基於Rust的最佳化技術,包括專用預分詞、BPE合併、多核分塊和零複製NumPy傳輸,同時保持與tiktoken完全一致的令牌ID。這一改進顯著降低了長輸入序列的首令牌時間,尤其是在字首快取命中率高的情況下。

Baseten Blog站內正文將Kimi K3分詞速度提升18倍,用於百萬令牌的代理工作負載

如何為Kimi K3構建Day-0 API

Baseten 為 Kimi K3 提供了 Day-0 支援,這是一個 2.8T 引數的新前沿開放模型。本文介紹了在釋出日前執行該模型所需的技術工作,包括硬體配置、權重載入、推理引擎整合、質量驗證、效能最佳化和規模化部署。

Baseten Blog站內正文如何為Kimi K3構建Day-0 API

推出GLM 5.2 Fast

我們推出了一個新的模型API服務層級——GLM-5.2 Fast,專為即時代理工作負載最佳化,提供與標準GLM-5.2相同的權重,但基礎設施針對每使用者吞吐量進行了調整,以實現低延遲和穩定的效能。

Baseten Blog站內正文推出GLM 5.2 Fast

全部來源