跳到主要內容
AI News HubLIVE
公開文章 50採集文章 51可信度 88刷新頻率 5 分鐘
健康狀態 健康來源類型 官方原文權限 官方原文最近入庫 2026-09-23ID together-ai-blog運行狀態 已啟用

Official source; confirm reuse terms before enabling full body display.

最新公開文章

待翻譯:How to train your own Jev for $17

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:We just launched our own Jev-like classifier, together/Tev1-4B-experimental, on top of Qwen3.5 4B on Together’s serverless platform. In this blog post we’ll show you how to fine-tune your own version!

Together AI Blog站內正文待翻譯:How to train your own Jev for $17

待翻譯:Canary rollouts: upgrade models in production without downtime

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:A hard model swap exposes every user at once, and rolling back means cold-starting the old deployment under pressure. Here's how staged traffic ramps, metric gates, and automatic rollback work on dedicated inference.

Together AI Blog站內正文待翻譯:Canary rollouts: upgrade models in production without downtime

待翻譯:How a global fintech scaled coding agent traffic with Dedicated Model Inference

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Inside a global bank's shift to self-serve dedicated inference: how Together's DMI gave engineering teams direct control over scaling, models, and testing.

Together AI Blog站內正文待翻譯:How a global fintech scaled coding agent traffic with Dedicated Model Inference

待翻譯:Migrating from closed to open source models, Together

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Moving from closed to open source models can take weeks, not years. A five-stage playbook: discover, evaluate, adapt, decide, and production.

Together AI Blog站內正文待翻譯:Migrating from closed to open source models, Together

待翻譯:Together AI expands fine-tuning service with more models, live metrics, and finer controls

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Together Fine-Tuning adds the latest open-weight models, live experiment tracking, Expert LoRA, early stopping, tokenized dataset previews, pre-flight validation, and lower training prices on selected models.

Together AI Blog站內正文待翻譯:Together AI expands fine-tuning service with more models, live metrics, and finer controls

待翻譯:Introducing preemptible compute: the same compute, half the price

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Together GPU Clusters now supports preemptible compute: the same GPU capacity at a flat 50% of the on-demand rate, with a five-minute drain window.

Together AI Blog站內正文待翻譯:Introducing preemptible compute: the same compute, half the price

待翻譯:To Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72!

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:We ported ThunderKittens to NVIDIA's Vera Rubin NVL72 and rebuilt our NVFP4 GEMM around the new hardware, taking it from 42% of roofline to over 22 PFLOPS — competitive with cuBLAS and CuTe DSL. Here is what changed in the ISA and how we used it.

Together AI Blog站內正文待翻譯:To Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72!

待翻譯:The Open Source AI Stack

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:All blog posts Inference Published 9/9/2026 The Open Source AI Stack Authors Hassan El Mghari Table of contents 40+ Models Chosen for Production...40+ Models Chosen for Production...40+ Models Chosen for Production... A…

Together AI Blog站內正文待翻譯:The Open Source AI Stack

待翻譯:GLM-5.3 vs. GLM-5.3 Flash on DeepSWE: Cost, Coding, and Routing

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:We ran 900 DeepSWE rollouts on GLM-5.3 and GLM-5.3 Flash. Flash gives up 5.6 points of pass@1 at 17x lower cost, and only 2.6 points at pass@4.

Together AI Blog站內正文待翻譯:GLM-5.3 vs. GLM-5.3 Flash on DeepSWE: Cost, Coding, and Routing

待翻譯:GLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:We ran 904 DeepSWE rollouts on GLM-5.3 and GPT-5.6 Sol. Sol leads pass@1 by 3.7 points; GLM-5.3 wins pass@4 at half the cost, and a GLM-first cascade hits 85.9%.

Together AI Blog站內正文待翻譯:GLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

待翻譯:GLM-5.3 vs. Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:We ran 904 DeepSWE rollouts on GLM-5.3 and Claude Fable 5. A tie on pass@1, but GLM-5.3 wins pass@4 and costs 5.4x less: \$3.99 per rollout vs. \$21.63.

Together AI Blog站內正文待翻譯:GLM-5.3 vs. Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

待翻譯:DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and GPT-5.6 Sol. Sol leads pass@1 by 10 points at 35x the cost; Pro wins pass@4, and a Pro-first cascade hits 83.0%.

Together AI Blog站內正文待翻譯:DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

待翻譯:DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and Claude Fable 5. Fable leads pass@1 at 90x the cost; Pro wins pass@4, and a Pro-first cascade hits 82.7%.

Together AI Blog站內正文待翻譯:DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

待翻譯:A/B test models in production

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Shadow traffic proves a candidate is operationally sound. It can't tell you if users like it better. Run the split at the endpoint instead of in your app code.

Together AI Blog站內正文待翻譯:A/B test models in production

待翻譯:DeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and Coding

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:We ran 900 DeepSWE rollouts on DeepSeek-V4 Flash and GPT-5.6 Luna. Luna leads pass@1 by 14 points; DeepSeek delivers 4.8x the solves per dollar.

Together AI Blog站內正文待翻譯:DeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and Coding

Kimi K3 完全開發者指南

Kimi K3 是 Moonshot AI 推出的 2.8 萬億參數開源權重模型,也是首個 3 萬億參數級別的開源模型。本文介紹其架構亮點、Together AI 上的調用方式、推理強度、視覺輸入、結構化輸出、工具調用與動態工具加載等關鍵開發要點。

Together AI Blog站內正文Kimi K3 完全開發者指南

ThunderAgent:大規模合成數據生成的2倍加速智能體推理

ThunderAgent是一種程序感知的智能體推理調度器。通過將每個智能體工作流視為可調度的程序,它消除了KV緩存顛簸,實現了超過2倍的單節點吞吐量和近線性的多節點擴展。

Together AI Blog站內正文ThunderAgent:大規模合成數據生成的2倍加速智能體推理

配置專用模型推理

Together AI 的專用模型推理架構詳解:端點、部署、配置三部分模型,以及基於容量的路由機制。

Together AI Blog站內正文配置專用模型推理

Kimi K3 vs GPT-5.6 Sol 在 DeepSWE 上的對比:成本、編碼和路由

我們在 DeepSWE 上對 Kimi K3 和 GPT-5.6 Sol 進行了 904 次測試。Sol 在 pass@1 上領先;Kimi K3 在 pass@4 上以每美元解決任務數 2.8 倍的優勢勝出,兩者之間的路由可達約 85.6%。

Together AI Blog站內正文Kimi K3 vs GPT-5.6 Sol 在 DeepSWE 上的對比:成本、編碼和路由

Kimi K3 vs Claude Fable 5:DeepSWE上的成本與編碼能力對比

我們在DeepSWE基準上對Kimi K3和Claude Fable 5進行了452次實驗。Fable在pass@1上領先1.4個百分點,但Kimi K3在pass@4上勝出,且每美元解決問題的數量是Fable的2.8倍。Kimi K3作為開源模型,成本僅為Fable的三分之一,是追求性價比的理性選擇。

Together AI Blog站內正文Kimi K3 vs Claude Fable 5:DeepSWE上的成本與編碼能力對比

開源權重AI推理的生產平台

Together AI發佈了其推理平台的重大更新,為用户提供對性能、成本和質量的完全控制。新功能包括金絲雀部署、A/B測試、自動縮放以及自定義訓練的封閉測試版,包含強化學習和微調。

Together AI Blog站內正文開源權重AI推理的生產平台

Together AI 與 Y Combinator 合作推出首個專為 YC 社區打造的 GPU 集羣

Together AI 和 Y Combinator 宣佈合作,為 YC 初創公司提供首個專用 GPU 集羣,解決計算瓶頸。初創公司可靈活承諾數週而非兩年合同,通過自助服務平台直接管理 GPU,無需經過 YC。該集羣已滿負荷運行,並支持從單節點到大規模擴展的需求。

Together AI Blog站內正文Together AI 與 Y Combinator 合作推出首個專為 YC 社區打造的 GPU 集羣

99.9% 的正常運行時間對推理意味着什麼?

本文深入解析了推理服務中 99%、99.9% 和 99.99% 正常運行時間所對應的不同故障域,以及每個層級所需的架構設計。作者分享了 Together AI 在構建可靠推理基礎設施方面的實踐經驗,並提供了在選擇推理提供商時應提出的關鍵問題。

Together AI Blog站內正文99.9% 的正常運行時間對推理意味着什麼?

Together GPU集羣新功能:為生產級GPU集羣提供可靠性與控制力

Together AI發佈一系列更新,旨在提升生產環境GPU集羣的可靠性與操作控制。新功能包括被動健康檢查、自動節點修復、強化版Slurm-on-K8s棧、集羣詳情視圖、外部OIDC認證、啓動腳本以及可選的驗收測試。這些改進幫助團隊更快發現硬件故障、自動化恢復流程,並提供更細粒度的訪問控制和自定義能力。

Together AI Blog站內正文Together GPU集羣新功能:為生產級GPU集羣提供可靠性與控制力

Together AI 在首日即引入 Thinking Machines Lab 的新模型 Inkling

Thinking Machines Lab 發佈了 Inkling,一個多模態混合專家模型,專注於高效推理、原生多模態理解和廣泛任務適用性。Together AI 在其推理平台上提供該模型,支持可控推理努力、文本/圖像/音頻輸入,以及 1M 上下文窗口。

Together AI Blog站內正文Together AI 在首日即引入 Thinking Machines Lab 的新模型 Inkling

開放、便捷且可預測:推出預留吞吐量功能

Together AI 推出預留吞吐量功能,為 MiniMax M3 和 GLM-5.2 等前沿開放模型提供保留推理容量,採用基於 Token 的定價和 99% 正常運行時間 SLA,成本比專有 API 降低高達 90%。

Together AI Blog站內正文開放、便捷且可預測:推出預留吞吐量功能

宣佈8億美元C輪融資:加速向開源AI的轉變

Together AI完成8億美元C輪融資,由Aramco Ventures、NVIDIA、Vista Equity等領投,旨在加速開源AI的普及。公司強調,閉源模型的成本無法規模化,而開源模型結合全棧優化可實現6-20倍成本降低。Together AI已推出FlashAttention-4、Together Megakernel等創新,成為全球最大的AI token生產商之一。

Together AI Blog站內正文宣佈8億美元C輪融資:加速向開源AI的轉變

Together AI在ICML 2026:涵蓋全棧的前沿研究

Together AI在ICML 2026上有八篇論文被接收,覆蓋從智能代理到GPU內核的整個堆棧。這些研究已集成到Together平台中,並在生產環境中應用。

Together AI Blog站內正文Together AI在ICML 2026:涵蓋全棧的前沿研究

ParallelKernelBench:前沿LLM尚無法編寫快速的多GPU內核

ParallelKernelBench是一個新的基準測試,評估LLM編寫多GPU CUDA內核的能力。在87個真實問題中,最佳模型僅能正確解決不到三分之一,且只有不到四分之一的解決方案優於基線。文章分析了模型失敗的原因,並展示了幾個意外生成的高性能內核案例。

Together AI Blog站內正文ParallelKernelBench:前沿LLM尚無法編寫快速的多GPU內核

全部來源