本文にスキップ
AI News HubLIVE
公開記事 50収集記事 51信頼度 88更新頻度 5 分
稼働状態 正常ソース種別 公式全文利用権限 公式全文最終取り込み 2026-09-23ID together-ai-blog状態 有効

Official source; confirm reuse terms before enabling full body display.

最新公開記事

翻訳待ち:How to train your own Jev for $17

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:We just launched our own Jev-like classifier, together/Tev1-4B-experimental, on top of Qwen3.5 4B on Together’s serverless platform. In this blog post we’ll show you how to fine-tune your own version!

Together AI Blogサイト内本文翻訳待ち:How to train your own Jev for $17

翻訳待ち:Canary rollouts: upgrade models in production without downtime

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:A hard model swap exposes every user at once, and rolling back means cold-starting the old deployment under pressure. Here's how staged traffic ramps, metric gates, and automatic rollback work on dedicated inference.

Together AI Blogサイト内本文翻訳待ち:Canary rollouts: upgrade models in production without downtime

翻訳待ち:How a global fintech scaled coding agent traffic with Dedicated Model Inference

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Inside a global bank's shift to self-serve dedicated inference: how Together's DMI gave engineering teams direct control over scaling, models, and testing.

Together AI Blogサイト内本文翻訳待ち:How a global fintech scaled coding agent traffic with Dedicated Model Inference

翻訳待ち:Migrating from closed to open source models, Together

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Moving from closed to open source models can take weeks, not years. A five-stage playbook: discover, evaluate, adapt, decide, and production.

Together AI Blogサイト内本文翻訳待ち:Migrating from closed to open source models, Together

翻訳待ち:Together AI expands fine-tuning service with more models, live metrics, and finer controls

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Together Fine-Tuning adds the latest open-weight models, live experiment tracking, Expert LoRA, early stopping, tokenized dataset previews, pre-flight validation, and lower training prices on selected models.

Together AI Blogサイト内本文翻訳待ち:Together AI expands fine-tuning service with more models, live metrics, and finer controls

翻訳待ち:Introducing preemptible compute: the same compute, half the price

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Together GPU Clusters now supports preemptible compute: the same GPU capacity at a flat 50% of the on-demand rate, with a five-minute drain window.

Together AI Blogサイト内本文翻訳待ち:Introducing preemptible compute: the same compute, half the price

翻訳待ち:To Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72!

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:We ported ThunderKittens to NVIDIA's Vera Rubin NVL72 and rebuilt our NVFP4 GEMM around the new hardware, taking it from 42% of roofline to over 22 PFLOPS — competitive with cuBLAS and CuTe DSL. Here is what changed in the ISA and how we used it.

Together AI Blogサイト内本文翻訳待ち:To Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72!

翻訳待ち:The Open Source AI Stack

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:All blog posts Inference Published 9/9/2026 The Open Source AI Stack Authors Hassan El Mghari Table of contents 40+ Models Chosen for Production...40+ Models Chosen for Production...40+ Models Chosen for Production... A…

Together AI Blogサイト内本文翻訳待ち:The Open Source AI Stack

翻訳待ち:GLM-5.3 vs. GLM-5.3 Flash on DeepSWE: Cost, Coding, and Routing

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:We ran 900 DeepSWE rollouts on GLM-5.3 and GLM-5.3 Flash. Flash gives up 5.6 points of pass@1 at 17x lower cost, and only 2.6 points at pass@4.

Together AI Blogサイト内本文翻訳待ち:GLM-5.3 vs. GLM-5.3 Flash on DeepSWE: Cost, Coding, and Routing

翻訳待ち:GLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:We ran 904 DeepSWE rollouts on GLM-5.3 and GPT-5.6 Sol. Sol leads pass@1 by 3.7 points; GLM-5.3 wins pass@4 at half the cost, and a GLM-first cascade hits 85.9%.

Together AI Blogサイト内本文翻訳待ち:GLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

翻訳待ち:GLM-5.3 vs. Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:We ran 904 DeepSWE rollouts on GLM-5.3 and Claude Fable 5. A tie on pass@1, but GLM-5.3 wins pass@4 and costs 5.4x less: \$3.99 per rollout vs. \$21.63.

Together AI Blogサイト内本文翻訳待ち:GLM-5.3 vs. Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

翻訳待ち:DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and GPT-5.6 Sol. Sol leads pass@1 by 10 points at 35x the cost; Pro wins pass@4, and a Pro-first cascade hits 83.0%.

Together AI Blogサイト内本文翻訳待ち:DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

翻訳待ち:DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and Claude Fable 5. Fable leads pass@1 at 90x the cost; Pro wins pass@4, and a Pro-first cascade hits 82.7%.

Together AI Blogサイト内本文翻訳待ち:DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

翻訳待ち:A/B test models in production

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Shadow traffic proves a candidate is operationally sound. It can't tell you if users like it better. Run the split at the endpoint instead of in your app code.

Together AI Blogサイト内本文翻訳待ち:A/B test models in production

翻訳待ち:DeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and Coding

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:We ran 900 DeepSWE rollouts on DeepSeek-V4 Flash and GPT-5.6 Luna. Luna leads pass@1 by 14 points; DeepSeek delivers 4.8x the solves per dollar.

Together AI Blogサイト内本文翻訳待ち:DeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and Coding

Kimi K3: 完全開発者ガイド

Kimi K3 は、Moonshot AI が開発した 2.8 兆パラメータのオープンウェイトモデルで、世界初の 3 兆パラメータ級オープンソースモデルです。本記事では、アーキテクチャの特徴、Together AI での呼び出し方、推論強度、視覚入力、構造化出力、ツール呼び出し、動的ツールロードなどを解説します。

Together AI Blogサイト内本文Kimi K3: 完全開発者ガイド

ThunderAgent:大規模合成データ生成のための2倍高速なエージェント推論

ThunderAgentは、エージェント推論のためのプログラム認識スケジューラです。各エージェントワークフローをスケジュール可能なプログラムとして扱うことで、KVキャッシュスラッシングを排除し、2倍以上のシングルノードスループットとほぼ線形のマルチノードスケーリングを実現します。

Together AI Blogサイト内本文ThunderAgent:大規模合成データ生成のための2倍高速なエージェント推論

Together AI、Moonshot AIと戦略的パートナーシップを発表:Kimiモデルをネイティブ提供

Together AIはMoonshot AIと戦略的提携を結び、Moonshotのオープンモデルのローンチプラットフォームとなり、最初のモデルとして2.8兆パラメータのKimi K3を提供開始。

Together AI Blogサイト内本文Together AI、Moonshot AIと戦略的パートナーシップを発表:Kimiモデルをネイティブ提供

専用モデル推論の設定

Together AIの専用モデル推論アーキテクチャの詳細ガイド:エンドポイント、デプロイメント、設定、容量ベースのルーティング。

Together AI Blogサイト内本文専用モデル推論の設定

Kimi K3 vs Claude Fable 5:DeepSWEにおけるコストとコーディング性能の比較

DeepSWEベンチマークでKimi K3とClaude Fable 5の452回のロールアウトを実行しました。Fableはpass@1で1.4ポイントリードするものの、Kimi K3はpass@4で勝利し、1ドルあたりの解決タスク数で2.8倍の効率を達成。Kimi K3はオープンモデルとして、Fableの3分の1のコストで高い性能を提供します。

Together AI Blogサイト内本文Kimi K3 vs Claude Fable 5:DeepSWEにおけるコストとコーディング性能の比較

オープンウェイトAI推論の本番プラットフォーム

Together AIは推論プラットフォームの大幅なアップデートを発表し、パフォーマンス、コスト、品質を完全に制御できるようにしました。カナリアデプロイ、A/Bテスト、オートスケーリング、そして強化学習とファインチューニングを含むカスタムトレーニングのクローズドベータ版など、新機能が追加されています。

Together AI Blogサイト内本文オープンウェイトAI推論の本番プラットフォーム

Together AI と Y Combinator が提携、YC コミュニティ向け初の専用 GPU クラスタを提供

Together AI と Y Combinator は、YC スタートアップ向けに専用 GPU クラスタを提供するパートナーシップを発表。計算ボトルネックを解決し、スタートアップは2年契約ではなく数週間単位でコミット可能。セルフサービスポータルを通じて直接 GPU を管理でき、YC の関与は不要。クラスタはフル稼働中で、単一ノードから大規模拡張まで対応。

Together AI Blogサイト内本文Together AI と Y Combinator が提携、YC コミュニティ向け初の専用 GPU クラスタを提供

推論における99.9%のアップタイムの意味とは?

この記事では、AI推論サービスの信頼性指標(99%、99.9%、99.99%)の実際の意味を分解し、各レベルが耐えなければならない障害ドメインと必要なアーキテクチャ要件を説明します。Together AIの著者らは、信頼性の高い推論インフラ構築の経験を共有し、プロバイダーを選ぶ前に尋ねるべき重要な質問を提供します。

Together AI Blogサイト内本文推論における99.9%のアップタイムの意味とは?

Together GPUクラスターの新機能:プロダクションGPUクラスターの信頼性と制御

Together AIが、受動的健康チェック、ノード修復、強化されたSlurmの信頼性、OIDC、起動スクリプトにより、プロダクションGPUクラスターをどのように改善しているかをご紹介します。

Together AI Blogサイト内本文Together GPUクラスターの新機能:プロダクションGPUクラスターの信頼性と制御

Together AI、Thinking Machines Labの新モデルInklingを初日から提供

Thinking Machines Labが、トークン効率的な推論、ネイティブなマルチモーダル理解、幅広いタスクの汎用性を備えたマルチモーダル混合専門家モデルInklingをリリースしました。Together AIは推論プラットフォームでこのモデルを提供し、制御可能な推論努力、テキスト/画像/音声入力、1Mコンテキストウィンドウをサポートします。

Together AI Blogサイト内本文Together AI、Thinking Machines Labの新モデルInklingを初日から提供

オープンで便利、予測可能:Provisioned Throughputのご紹介

Together AIは、MiniMax M3やGLM-5.2などのフロンティアオープンモデル向けに、トークンベースの料金設定と99%のアップタイムSLAを備えた予約推論容量「Provisioned Throughput」を発表。専有APIと比較して最大90%のコスト削減を実現します。

Together AI Blogサイト内本文オープンで便利、予測可能:Provisioned Throughputのご紹介

オープンソースAIへの移行を加速するための8億ドルのシリーズCを発表

Together AIは、オープンソースAIへの移行を加速するため、8億ドルのシリーズC資金調達を完了した。同社は、クローズドモデルの経済性はスケールせず、フルスタック最適化と組み合わせたオープンモデルで6〜20倍のコスト削減が可能だと主張する。Together AIは、FlashAttention-4やTogether Megakernelなどの革新を立ち上げ、世界最大のAIトークンプロデューサーの1つとなった。

Together AI Blogサイト内本文オープンソースAIへの移行を加速するための8億ドルのシリーズCを発表

Together AI、ICML 2026で全スタックにわたる最先端研究

Together AIはICML 2026で8本の論文が採択され、エージェントからGPUカーネルまでの全スタックの研究を発表。これらの研究はTogetherプラットフォームに統合され、本番環境ですでに活用されている。

Together AI Blogサイト内本文Together AI、ICML 2026で全スタックにわたる最先端研究

ParallelKernelBench:最先端LLMはまだ高速マルチGPUカーネルを書けない

ParallelKernelBenchは、LLMが87の実ワークロードに対して高速なマルチGPU CUDAカーネルを書けるかをテストするベンチマークです。最高のモデルでも3分の1未満しか解けず、ベースラインを上回ったのはさらに少ないですが、生成されたカーネルの中には既存の公開実装を凌ぐものもあります。

Together AI Blogサイト内本文ParallelKernelBench:最先端LLMはまだ高速マルチGPUカーネルを書けない

全ソース