跳到主要內容
AI News HubLIVE
公開文章 50採集文章 51可信度 88刷新頻率 5 分鐘
健康狀態 健康來源類型 官方原文權限 官方原文最近入庫 2026-09-23ID together-ai-blog運行狀態 已啟用

Official source; confirm reuse terms before enabling full body display.

最新公開文章

待翻譯:How to train your own Jev for $17

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:We just launched our own Jev-like classifier, together/Tev1-4B-experimental, on top of Qwen3.5 4B on Together’s serverless platform. In this blog post we’ll show you how to fine-tune your own version!

Together AI Blog站內正文待翻譯:How to train your own Jev for $17

待翻譯:Canary rollouts: upgrade models in production without downtime

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:A hard model swap exposes every user at once, and rolling back means cold-starting the old deployment under pressure. Here's how staged traffic ramps, metric gates, and automatic rollback work on dedicated inference.

Together AI Blog站內正文待翻譯:Canary rollouts: upgrade models in production without downtime

待翻譯:How a global fintech scaled coding agent traffic with Dedicated Model Inference

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Inside a global bank's shift to self-serve dedicated inference: how Together's DMI gave engineering teams direct control over scaling, models, and testing.

Together AI Blog站內正文待翻譯:How a global fintech scaled coding agent traffic with Dedicated Model Inference

待翻譯:Migrating from closed to open source models, Together

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Moving from closed to open source models can take weeks, not years. A five-stage playbook: discover, evaluate, adapt, decide, and production.

Together AI Blog站內正文待翻譯:Migrating from closed to open source models, Together

待翻譯:Together AI expands fine-tuning service with more models, live metrics, and finer controls

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Together Fine-Tuning adds the latest open-weight models, live experiment tracking, Expert LoRA, early stopping, tokenized dataset previews, pre-flight validation, and lower training prices on selected models.

Together AI Blog站內正文待翻譯:Together AI expands fine-tuning service with more models, live metrics, and finer controls

待翻譯:Introducing preemptible compute: the same compute, half the price

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Together GPU Clusters now supports preemptible compute: the same GPU capacity at a flat 50% of the on-demand rate, with a five-minute drain window.

Together AI Blog站內正文待翻譯:Introducing preemptible compute: the same compute, half the price

待翻譯:To Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72!

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:We ported ThunderKittens to NVIDIA's Vera Rubin NVL72 and rebuilt our NVFP4 GEMM around the new hardware, taking it from 42% of roofline to over 22 PFLOPS — competitive with cuBLAS and CuTe DSL. Here is what changed in the ISA and how we used it.

Together AI Blog站內正文待翻譯:To Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72!

待翻譯:The Open Source AI Stack

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:All blog posts Inference Published 9/9/2026 The Open Source AI Stack Authors Hassan El Mghari Table of contents 40+ Models Chosen for Production...40+ Models Chosen for Production...40+ Models Chosen for Production... A…

Together AI Blog站內正文待翻譯:The Open Source AI Stack

待翻譯:GLM-5.3 vs. GLM-5.3 Flash on DeepSWE: Cost, Coding, and Routing

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:We ran 900 DeepSWE rollouts on GLM-5.3 and GLM-5.3 Flash. Flash gives up 5.6 points of pass@1 at 17x lower cost, and only 2.6 points at pass@4.

Together AI Blog站內正文待翻譯:GLM-5.3 vs. GLM-5.3 Flash on DeepSWE: Cost, Coding, and Routing

待翻譯:GLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:We ran 904 DeepSWE rollouts on GLM-5.3 and GPT-5.6 Sol. Sol leads pass@1 by 3.7 points; GLM-5.3 wins pass@4 at half the cost, and a GLM-first cascade hits 85.9%.

Together AI Blog站內正文待翻譯:GLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

待翻譯:GLM-5.3 vs. Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:We ran 904 DeepSWE rollouts on GLM-5.3 and Claude Fable 5. A tie on pass@1, but GLM-5.3 wins pass@4 and costs 5.4x less: \$3.99 per rollout vs. \$21.63.

Together AI Blog站內正文待翻譯:GLM-5.3 vs. Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

待翻譯:DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and GPT-5.6 Sol. Sol leads pass@1 by 10 points at 35x the cost; Pro wins pass@4, and a Pro-first cascade hits 83.0%.

Together AI Blog站內正文待翻譯:DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

待翻譯:DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and Claude Fable 5. Fable leads pass@1 at 90x the cost; Pro wins pass@4, and a Pro-first cascade hits 82.7%.

Together AI Blog站內正文待翻譯:DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

待翻譯:A/B test models in production

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Shadow traffic proves a candidate is operationally sound. It can't tell you if users like it better. Run the split at the endpoint instead of in your app code.

Together AI Blog站內正文待翻譯:A/B test models in production

待翻譯:DeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and Coding

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:We ran 900 DeepSWE rollouts on DeepSeek-V4 Flash and GPT-5.6 Luna. Luna leads pass@1 by 14 points; DeepSeek delivers 4.8x the solves per dollar.

Together AI Blog站內正文待翻譯:DeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and Coding

Kimi K3 完全開發者指南

Kimi K3 是 Moonshot AI 推出的 2.8 萬億引數開源權重模型,也是首個 3 萬億引數級別的開源模型。本文介紹其架構亮點、Together AI 上的呼叫方式、推理強度、視覺輸入、結構化輸出、工具呼叫與動態工具載入等關鍵開發要點。

Together AI Blog站內正文Kimi K3 完全開發者指南

ThunderAgent:大規模合成資料生成的2倍加速智慧體推理

ThunderAgent是一種程式感知的智慧體推理排程器。透過將每個智慧體工作流視為可排程的程式,它消除了KV快取顛簸,實現了超過2倍的單節點吞吐量和近線性的多節點擴充套件。

Together AI Blog站內正文ThunderAgent:大規模合成資料生成的2倍加速智慧體推理

配置專用模型推理

Together AI 的專用模型推理架構詳解:端點、部署、配置三部分模型,以及基於容量的路由機制。

Together AI Blog站內正文配置專用模型推理

Kimi K3 vs GPT-5.6 Sol 在 DeepSWE 上的對比:成本、編碼和路由

我們在 DeepSWE 上對 Kimi K3 和 GPT-5.6 Sol 進行了 904 次測試。Sol 在 pass@1 上領先;Kimi K3 在 pass@4 上以每美元解決任務數 2.8 倍的優勢勝出,兩者之間的路由可達約 85.6%。

Together AI Blog站內正文Kimi K3 vs GPT-5.6 Sol 在 DeepSWE 上的對比:成本、編碼和路由

Kimi K3 vs Claude Fable 5:DeepSWE上的成本與編碼能力對比

我們在DeepSWE基準上對Kimi K3和Claude Fable 5進行了452次實驗。Fable在pass@1上領先1.4個百分點,但Kimi K3在pass@4上勝出,且每美元解決問題的數量是Fable的2.8倍。Kimi K3作為開源模型,成本僅為Fable的三分之一,是追求價效比的理性選擇。

Together AI Blog站內正文Kimi K3 vs Claude Fable 5:DeepSWE上的成本與編碼能力對比

開源權重AI推理的生產平臺

Together AI釋出了其推理平臺的重大更新,為使用者提供對效能、成本和質量的完全控制。新功能包括金絲雀部署、A/B測試、自動縮放以及自定義訓練的封閉測試版,包含強化學習和微調。

Together AI Blog站內正文開源權重AI推理的生產平臺

Together AI 與 Y Combinator 合作推出首個專為 YC 社群打造的 GPU 叢集

Together AI 和 Y Combinator 宣佈合作,為 YC 初創公司提供首個專用 GPU 叢集,解決計算瓶頸。初創公司可靈活承諾數週而非兩年合同,透過自助服務平臺直接管理 GPU,無需經過 YC。該叢集已滿負荷執行,並支援從單節點到大規模擴充套件的需求。

Together AI Blog站內正文Together AI 與 Y Combinator 合作推出首個專為 YC 社群打造的 GPU 叢集

99.9% 的正常執行時間對推理意味著什麼?

本文深入解析了推理服務中 99%、99.9% 和 99.99% 正常執行時間所對應的不同故障域,以及每個層級所需的架構設計。作者分享了 Together AI 在構建可靠推理基礎設施方面的實踐經驗,並提供了在選擇推理提供商時應提出的關鍵問題。

Together AI Blog站內正文99.9% 的正常執行時間對推理意味著什麼?

Together GPU叢集新功能:為生產級GPU叢集提供可靠性與控制力

Together AI釋出一系列更新,旨在提升生產環境GPU叢集的可靠性與操作控制。新功能包括被動健康檢查、自動節點修復、強化版Slurm-on-K8s棧、叢集詳情檢視、外部OIDC認證、啟動指令碼以及可選的驗收測試。這些改進幫助團隊更快發現硬體故障、自動化恢復流程,並提供更細粒度的訪問控制和自定義能力。

Together AI Blog站內正文Together GPU叢集新功能:為生產級GPU叢集提供可靠性與控制力

Together AI 在首日即引入 Thinking Machines Lab 的新模型 Inkling

Thinking Machines Lab 釋出了 Inkling,一個多模態混合專家模型,專注於高效推理、原生多模態理解和廣泛任務適用性。Together AI 在其推理平臺上提供該模型,支援可控推理努力、文本/影像/音訊輸入,以及 1M 上下文視窗。

Together AI Blog站內正文Together AI 在首日即引入 Thinking Machines Lab 的新模型 Inkling

開放、便捷且可預測:推出預留吞吐量功能

Together AI 推出預留吞吐量功能,為 MiniMax M3 和 GLM-5.2 等前沿開放模型提供保留推理容量,採用基於 Token 的定價和 99% 正常執行時間 SLA,成本比專有 API 降低高達 90%。

Together AI Blog站內正文開放、便捷且可預測:推出預留吞吐量功能

宣佈8億美元C輪融資:加速向開源AI的轉變

Together AI完成8億美元C輪融資,由Aramco Ventures、NVIDIA、Vista Equity等領投,旨在加速開源AI的普及。公司強調,閉源模型的成本無法規模化,而開源模型結合全棧最佳化可實現6-20倍成本降低。Together AI已推出FlashAttention-4、Together Megakernel等創新,成為全球最大的AI token生產商之一。

Together AI Blog站內正文宣佈8億美元C輪融資:加速向開源AI的轉變

Together AI在ICML 2026:涵蓋全棧的前沿研究

Together AI在ICML 2026上有八篇論文被接收,覆蓋從智慧代理到GPU核心的整個堆疊。這些研究已整合到Together平臺中,並在生產環境中應用。

Together AI Blog站內正文Together AI在ICML 2026:涵蓋全棧的前沿研究

ParallelKernelBench:前沿LLM尚無法編寫快速的多GPU核心

ParallelKernelBench是一個新的基準測試,評估LLM編寫多GPU CUDA核心的能力。在87個真實問題中,最佳模型僅能正確解決不到三分之一,且只有不到四分之一的解決方案優於基線。文章分析了模型失敗的原因,並展示了幾個意外生成的高效能核心案例。

Together AI Blog站內正文ParallelKernelBench:前沿LLM尚無法編寫快速的多GPU核心

全部來源