本文にスキップ
AI News HubLIVE
公開記事 58収集記事 71信頼度 82更新頻度 120 分
稼働状態 正常ソース種別 公式全文利用権限 公式全文最終取り込み 2026-09-28ID baseten-blog状態 有効

Official AI inference and deployment platform blog; confirm reuse terms before full body display.

最新公開記事

翻訳待ち:Securing the open frontier with NVIDIA OpenShell and Blaxel sandboxes

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:News Securing the open frontier with NVIDIA OpenShell and Blaxel sandboxes Baseten is a launch partner for the NVIDIA Agent Safety Platform and supports OpenShell in the latest generation of Blaxel Sandboxes. Authors Ni…

Baseten Blogサイト内本文翻訳待ち:Securing the open frontier with NVIDIA OpenShell and Blaxel sandboxes

翻訳待ち:Sheila Vashee joins Baseten as CMO

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:News Sheila Vashee joins Baseten as CMO Welcome Sheila Vashee! Authors Tuhin Srivastava Last updated September 25, 2026 Share The next era of AI will be defined by what companies can build with intelligence that is open…

Baseten Blogサイト内本文翻訳待ち:Sheila Vashee joins Baseten as CMO

翻訳待ち:Fine-tune on your LangSmith traces with Baseten Loops

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:News Fine-tune on your LangSmith traces with Baseten Loops LangSmith Fine-Tuning is in public beta today, and its open-source CLI, smithtune, trains on Baseten Loops. Authors Mudith Jayasekara Aaron Ellis-Bloor Last upd…

Baseten Blogサイト内本文翻訳待ち:Fine-tune on your LangSmith traces with Baseten Loops

翻訳待ち:NVIDIA Nemotron 3 Diarization: real-time speaker labels at a cent per audio hour

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:AI models NVIDIA Nemotron 3 Diarization: real-time speaker labels at a cent per audio hour NVIDIA Nemotron 3 Diarization answers who spoke when, 300ms–1s behind live, alongside any ASR. Authors Ansel Erol Last updated S…

Baseten Blogサイト内本文翻訳待ち:NVIDIA Nemotron 3 Diarization: real-time speaker labels at a cent per audio hour

翻訳待ち:Introducing Baseten Hosted Tools

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Product Introducing Baseten Hosted Tools Hosted tool execution on the inference backend augments model capabilities and lowers latency. Authors Sai Maddali Marius Killinger Marylise Tauzia Last updated September 16, 202…

Baseten Blogサイト内本文翻訳待ち:Introducing Baseten Hosted Tools

翻訳待ち:LangChain trains custom models for LangSmith Engine with Baseten Loops

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:News LangChain trains custom models for LangSmith Engine with Baseten Loops Loops provides managed infrastructure for fine-tuning models through an API, with support for SFT, RL, and long-context workloads. Authors Aaro…

Baseten Blogサイト内本文翻訳待ち:LangChain trains custom models for LangSmith Engine with Baseten Loops

翻訳待ち:DeepSeek-V4.1-Flash: more efficient prefill for coding agents

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:AI models DeepSeek-V4.1-Flash: more efficient prefill for coding agents Discover DeepSeek-V4.1-Flash, a 552B multimodal MoE model featuring a Causal Encoder-Decoder architecture and efficient KV caching. Authors Albert…

Baseten Blogサイト内本文翻訳待ち:DeepSeek-V4.1-Flash: more efficient prefill for coding agents

翻訳待ち:Blaxel is joining Baseten to build the future of agentic infrastructure

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:News Blaxel is joining Baseten to build the future of agentic infrastructure Baseten has acquired Blaxel Authors Amir Haghighat Tuhin Srivastava Paul Sinai Last updated September 10, 2026 Share Today, Blaxel is joining…

Baseten Blogサイト内本文翻訳待ち:Blaxel is joining Baseten to build the future of agentic infrastructure

翻訳待ち:How Baseten makes pyannote’s diarization models 9.6x faster

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Model performance How Baseten makes pyannote’s diarization models 9.6x faster Using quality-aware quantization, index-based clustering, and scheduling optimizations Authors Matte Lim Ansel Erol Last updated September 9,…

Baseten Blogサイト内本文翻訳待ち:How Baseten makes pyannote’s diarization models 9.6x faster

コーディングエージェントがBasetenを使うための新しいMCPとスキル

Basetenは、Claude Code、Codex、Cursorなどのコーディングエージェントが同プラットフォームをより高速かつ効率的に操作できるよう、リモートMCPサーバーとスキルを公開した。ベンチマークでは、タスクの所要時間とコストが平均7.5%削減され、バックエンド中心のタスクでは約50%削減された。

Baseten Blogサイト内本文コーディングエージェントがBasetenを使うための新しいMCPとスキル

翻訳待ち:Best open-source models for post-training

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:AI models Best open-source models for post-training Compare top open-source models for post-training. Learn what drives cost and which model fits use case. Authors Chloe Florit Last updated September 2, 2026 Share TL;DR…

Baseten Blogサイト内本文翻訳待ち:Best open-source models for post-training

翻訳待ち:The efficient frontier of LLM inference

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Model performance The efficient frontier of LLM inference Inference techniques either move a deployment along the latency–throughput frontier or push the entire frontier out, creating more efficiency to allocate. Author…

Baseten Blogサイト内本文翻訳待ち:The efficient frontier of LLM inference

翻訳待ち:Agentic kernels in production

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Model performance Agentic kernels in production Baseten's agentic kernel optimization framework cuts latency by 42.3% on Qwen-Image and by 15.2% on FLUX.2. Authors Brian Li Faraz Shahsavan Pankaj Gupta Last updated Augu…

Baseten Blogサイト内本文翻訳待ち:Agentic kernels in production

翻訳待ち:GLM 5.3: Scaling with post-training, intuitively explained

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:AI models GLM 5.3: Scaling with post-training, intuitively explained GLM 5.3's gains came entirely from post-training. An intuitive look at the environment design, RL algorithms, and infrastructure behind the leap. Auth…

Baseten Blogサイト内本文翻訳待ち:GLM 5.3: Scaling with post-training, intuitively explained

翻訳待ち:The two AI gateway patterns in production inference

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:AI engineering The two AI gateway patterns in production inference Access gateways and serving gateways explained, and how they handle identity, tenancy, limits, and metering for teams serving AI models. Authors Amit Ga…

Baseten Blogサイト内本文翻訳待ち:The two AI gateway patterns in production inference

翻訳待ち:How to run any open model inside DeepSeek Harness

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:AI engineering How to run any open model inside DeepSeek Harness Learn how to power DeepSeek Harness with Baseten Model APIs and run open models like Kimi K3, GLM 5.2, and DeepSeek V4 Pro in under 5 minutes. Authors Ale…

Baseten Blogサイト内本文翻訳待ち:How to run any open model inside DeepSeek Harness

翻訳待ち:How leading platforms ensure observability for LLM inference

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Infrastructure How leading platforms ensure observability for LLM inference Learn how metrics, logs, and traces work together to catch slow responses, errors, and failed deployments before users do. Authors Chloe Florit…

Baseten Blogサイト内本文翻訳待ち:How leading platforms ensure observability for LLM inference

翻訳待ち:Inference engineering for DeepSeek V4 Pro 0813

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Model performance Inference engineering for DeepSeek V4 Pro 0813 DeepSeek V4 Pro 0813 is a 1.7T-parameter open frontier model under the MIT license and is available for inference today on Baseten model APIs. Authors Mod…

Baseten Blogサイト内本文翻訳待ち:Inference engineering for DeepSeek V4 Pro 0813

翻訳待ち:Baseten delivers open-source inference for You.com

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Community Baseten delivers open-source inference for You.com Baseten is proud to power the inference behind You.com's search and answer stack. Authors Marylise Tauzia Last updated August 13, 2026 Share TL;DR You.com bui…

Baseten Blogサイト内本文翻訳待ち:Baseten delivers open-source inference for You.com

翻訳待ち:Qwen3.8-Max: Alibaba's new frontier reasoning model

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:AI models Qwen3.8-Max: Alibaba's new frontier reasoning model Qwen3.8-Max is Alibaba's new open-weight frontier model, featuring a 1M token context window and multimodal input. Authors Albert Lee Last updated August 12,…

Baseten Blogサイト内本文翻訳待ち:Qwen3.8-Max: Alibaba's new frontier reasoning model

翻訳待ち:Introducing NVIDIA Nemotron 3.5 Lightning

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:AI models Introducing NVIDIA Nemotron 3.5 Lightning NVIDIA’s Nemotron 3.5 Lightning, now on Baseten, delivers high-throughput, efficient reasoning for faster and more accurate agentic workflows. Authors Marylise Tauzia…

Baseten Blogサイト内本文翻訳待ち:Introducing NVIDIA Nemotron 3.5 Lightning

翻訳待ち:Introducing NVIDIA Nemotron 3.5 ASR Streaming

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:AI models Introducing NVIDIA Nemotron 3.5 ASR Streaming Deploy NVIDIA Nemotron 3.5 ASR for low-latency, production-ready speech recognition with 6x higher throughput and multilingual support. Authors Ansel Erol Ian Carr…

Baseten Blogサイト内本文翻訳待ち:Introducing NVIDIA Nemotron 3.5 ASR Streaming

翻訳待ち:Laguna S 2.1 goes Greek: a repository-scale game transformation

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:AI models Laguna S 2.1 goes Greek: a repository-scale game transformation We put Poolside’s new Laguna S 2.1 model to the test, tasking it with a repository-scale transformation of the open-source game Hypersomnia. Auth…

Baseten Blogサイト内本文翻訳待ち:Laguna S 2.1 goes Greek: a repository-scale game transformation

Qwen3-TTSをファインチューニングして高品質な声質クローンを実現

Qwen3-TTSをファインチューニングして高品質な音声クローンを実現する方法を解説し、ICL、話者埋め込みのみ、ファインチューニングの3手法を比較しながら完全なトレーニング手法を紹介します。LJ Speechデータ約1.5時間を使用し、単一のH100で約1時間のトレーニングで、TTFAは約130msとなり、ICLより約16%高速化し、推論時の参照音声処理が不要になります。

Baseten Blogサイト内本文Qwen3-TTSをファインチューニングして高品質な声質クローンを実現

22,580:GPT-2からKimi K3へ、その発展を解説

本記事は、GPT-2(1.24億パラメータ)からKimi K3(2.8兆パラメータ)へのアーキテクチャ進化を追跡し、7年間で22,580倍の規模拡大を実現した経緯を解説します。KVキャッシュ、線形注意、DeltaNetなどの主要な革新を説明し、巨大な規模に対応するために基礎メカニズムがどのように進化したかを明らかにします。

Baseten Blogサイト内本文22,580:GPT-2からKimi K3へ、その発展を解説

任意のハーネスでKimi K3を実行する方法:Baseten Switchによるルーティング

Baseten Switch はローカルのMacアプリで、Claude CodeやCodexなどの人気ハーネスでオープンモデルとクローズドモデルをシームレスに混在させ、Kimi K3やGLM 5.2などのモデルを切り替えながら、コスト、パフォーマンス、使用状況を追跡できます。

Baseten Blogサイト内本文任意のハーネスでKimi K3を実行する方法:Baseten Switchによるルーティング

BasetenがModel Labs向けプラットフォームを発表

Basetenは、クローズドウェイトモデルラボ向けの新プラットフォーム「Baseten for Model Labs」を発表。推論インフラ、モデル配布、知的財産保護、およびマーケット投入支援を提供し、ラボのモデル収益化を迅速化する。Cartesia、Gradium、Inception、NVIDIAなど15のラボパートナーが参加している。

Baseten Blogサイト内本文BasetenがModel Labs向けプラットフォームを発表

Kimi K3のトークン化を18倍高速化、百万トークンのエージェントワークロード向け

Basetenは新しいトークナイザー(Basetenkenizer)を導入し、Kimi K3のトークン化を百万トークンのエージェントワークロードで最大18倍高速化します。Rustベースの最適化(特殊化された事前トークン化、BPEマージ、マルチコアチャンク処理、ゼロコピーNumPy転送など)を組み合わせ、tiktokenと完全に同じトークンIDを維持します。この改善により、長い入力シーケンスの最初のトークンまでの時間が大幅に短縮され、特にプレフィックスキャッシュヒット率が高い場合に効果的です。

Baseten Blogサイト内本文Kimi K3のトークン化を18倍高速化、百万トークンのエージェントワークロード向け

Kimi K3のDay-0 APIを構築する方法

Basetenは、2.8Tパラメータの新しいオープンフロンティアモデルKimi K3に対してDay-0サポートを提供します。この記事では、ハードウェア構成、重みのロード、推論エンジンの統合、品質検証、パフォーマンス最適化、および大規模展開を含む、モデルを起動日までに実行するために必要な技術的作業について説明します。

Baseten Blogサイト内本文Kimi K3のDay-0 APIを構築する方法

GLM 5.2 Fastのご紹介

リアルタイムのエージェントワークロード向けに最適化された新しいモデルAPI階層、GLM-5.2 Fastを発表します。標準のGLM-5.2と同じ重みを持ちながら、ユーザーあたりのスループットに調整されたインフラ上で動作し、低レイテンシと安定したパフォーマンスを提供します。

Baseten Blogサイト内本文GLM 5.2 Fastのご紹介

全ソース