跳到主要内容
AI News HubLIVE
公开文章 50采集文章 51可信度 88刷新频率 5 分钟
健康状态 健康来源类型 官方原文权限 官方原文最近入库 2026-09-23ID together-ai-blog运行状态 已启用

Official source; confirm reuse terms before enabling full body display.

最新公开文章

待翻译:How to train your own Jev for $17

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:We just launched our own Jev-like classifier, together/Tev1-4B-experimental, on top of Qwen3.5 4B on Together’s serverless platform. In this blog post we’ll show you how to fine-tune your own version!

Together AI Blog站内正文待翻译:How to train your own Jev for $17

待翻译:Canary rollouts: upgrade models in production without downtime

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:A hard model swap exposes every user at once, and rolling back means cold-starting the old deployment under pressure. Here's how staged traffic ramps, metric gates, and automatic rollback work on dedicated inference.

Together AI Blog站内正文待翻译:Canary rollouts: upgrade models in production without downtime

待翻译:How a global fintech scaled coding agent traffic with Dedicated Model Inference

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Inside a global bank's shift to self-serve dedicated inference: how Together's DMI gave engineering teams direct control over scaling, models, and testing.

Together AI Blog站内正文待翻译:How a global fintech scaled coding agent traffic with Dedicated Model Inference

待翻译:Migrating from closed to open source models, Together

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Moving from closed to open source models can take weeks, not years. A five-stage playbook: discover, evaluate, adapt, decide, and production.

Together AI Blog站内正文待翻译:Migrating from closed to open source models, Together

待翻译:Together AI expands fine-tuning service with more models, live metrics, and finer controls

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Together Fine-Tuning adds the latest open-weight models, live experiment tracking, Expert LoRA, early stopping, tokenized dataset previews, pre-flight validation, and lower training prices on selected models.

Together AI Blog站内正文待翻译:Together AI expands fine-tuning service with more models, live metrics, and finer controls

待翻译:Introducing preemptible compute: the same compute, half the price

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Together GPU Clusters now supports preemptible compute: the same GPU capacity at a flat 50% of the on-demand rate, with a five-minute drain window.

Together AI Blog站内正文待翻译:Introducing preemptible compute: the same compute, half the price

待翻译:To Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72!

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:We ported ThunderKittens to NVIDIA's Vera Rubin NVL72 and rebuilt our NVFP4 GEMM around the new hardware, taking it from 42% of roofline to over 22 PFLOPS — competitive with cuBLAS and CuTe DSL. Here is what changed in the ISA and how we used it.

Together AI Blog站内正文待翻译:To Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72!

待翻译:The Open Source AI Stack

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:All blog posts Inference Published 9/9/2026 The Open Source AI Stack Authors Hassan El Mghari Table of contents 40+ Models Chosen for Production...40+ Models Chosen for Production...40+ Models Chosen for Production... A…

Together AI Blog站内正文待翻译:The Open Source AI Stack

待翻译:GLM-5.3 vs. GLM-5.3 Flash on DeepSWE: Cost, Coding, and Routing

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:We ran 900 DeepSWE rollouts on GLM-5.3 and GLM-5.3 Flash. Flash gives up 5.6 points of pass@1 at 17x lower cost, and only 2.6 points at pass@4.

Together AI Blog站内正文待翻译:GLM-5.3 vs. GLM-5.3 Flash on DeepSWE: Cost, Coding, and Routing

待翻译:GLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:We ran 904 DeepSWE rollouts on GLM-5.3 and GPT-5.6 Sol. Sol leads pass@1 by 3.7 points; GLM-5.3 wins pass@4 at half the cost, and a GLM-first cascade hits 85.9%.

Together AI Blog站内正文待翻译:GLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

待翻译:GLM-5.3 vs. Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:We ran 904 DeepSWE rollouts on GLM-5.3 and Claude Fable 5. A tie on pass@1, but GLM-5.3 wins pass@4 and costs 5.4x less: \$3.99 per rollout vs. \$21.63.

Together AI Blog站内正文待翻译:GLM-5.3 vs. Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

待翻译:DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and GPT-5.6 Sol. Sol leads pass@1 by 10 points at 35x the cost; Pro wins pass@4, and a Pro-first cascade hits 83.0%.

Together AI Blog站内正文待翻译:DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

待翻译:DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and Claude Fable 5. Fable leads pass@1 at 90x the cost; Pro wins pass@4, and a Pro-first cascade hits 82.7%.

Together AI Blog站内正文待翻译:DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

待翻译:A/B test models in production

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Shadow traffic proves a candidate is operationally sound. It can't tell you if users like it better. Run the split at the endpoint instead of in your app code.

Together AI Blog站内正文待翻译:A/B test models in production

待翻译:DeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and Coding

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:We ran 900 DeepSWE rollouts on DeepSeek-V4 Flash and GPT-5.6 Luna. Luna leads pass@1 by 14 points; DeepSeek delivers 4.8x the solves per dollar.

Together AI Blog站内正文待翻译:DeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and Coding

Kimi K3 完全开发者指南

Kimi K3 是 Moonshot AI 推出的 2.8 万亿参数开源权重模型,也是首个 3 万亿参数级别的开源模型。本文介绍其架构亮点、Together AI 上的调用方式、推理强度、视觉输入、结构化输出、工具调用与动态工具加载等关键开发要点。

Together AI Blog站内正文Kimi K3 完全开发者指南

ThunderAgent:大规模合成数据生成的2倍加速智能体推理

ThunderAgent是一种程序感知的智能体推理调度器。通过将每个智能体工作流视为可调度的程序,它消除了KV缓存颠簸,实现了超过2倍的单节点吞吐量和近线性的多节点扩展。

Together AI Blog站内正文ThunderAgent:大规模合成数据生成的2倍加速智能体推理

配置专用模型推理

Together AI 的专用模型推理架构详解:端点、部署、配置三部分模型,以及基于容量的路由机制。

Together AI Blog站内正文配置专用模型推理

Kimi K3 vs GPT-5.6 Sol 在 DeepSWE 上的对比:成本、编码和路由

我们在 DeepSWE 上对 Kimi K3 和 GPT-5.6 Sol 进行了 904 次测试。Sol 在 pass@1 上领先;Kimi K3 在 pass@4 上以每美元解决任务数 2.8 倍的优势胜出,两者之间的路由可达约 85.6%。

Together AI Blog站内正文Kimi K3 vs GPT-5.6 Sol 在 DeepSWE 上的对比:成本、编码和路由

Kimi K3 vs Claude Fable 5:DeepSWE上的成本与编码能力对比

我们在DeepSWE基准上对Kimi K3和Claude Fable 5进行了452次实验。Fable在pass@1上领先1.4个百分点,但Kimi K3在pass@4上胜出,且每美元解决问题的数量是Fable的2.8倍。Kimi K3作为开源模型,成本仅为Fable的三分之一,是追求性价比的理性选择。

Together AI Blog站内正文Kimi K3 vs Claude Fable 5:DeepSWE上的成本与编码能力对比

开源权重AI推理的生产平台

Together AI发布了其推理平台的重大更新,为用户提供对性能、成本和质量的完全控制。新功能包括金丝雀部署、A/B测试、自动缩放以及自定义训练的封闭测试版,包含强化学习和微调。

Together AI Blog站内正文开源权重AI推理的生产平台

Together AI 与 Y Combinator 合作推出首个专为 YC 社区打造的 GPU 集群

Together AI 和 Y Combinator 宣布合作,为 YC 初创公司提供首个专用 GPU 集群,解决计算瓶颈。初创公司可灵活承诺数周而非两年合同,通过自助服务平台直接管理 GPU,无需经过 YC。该集群已满负荷运行,并支持从单节点到大规模扩展的需求。

Together AI Blog站内正文Together AI 与 Y Combinator 合作推出首个专为 YC 社区打造的 GPU 集群

99.9% 的正常运行时间对推理意味着什么?

本文深入解析了推理服务中 99%、99.9% 和 99.99% 正常运行时间所对应的不同故障域,以及每个层级所需的架构设计。作者分享了 Together AI 在构建可靠推理基础设施方面的实践经验,并提供了在选择推理提供商时应提出的关键问题。

Together AI Blog站内正文99.9% 的正常运行时间对推理意味着什么?

Together GPU集群新功能:为生产级GPU集群提供可靠性与控制力

Together AI发布一系列更新,旨在提升生产环境GPU集群的可靠性与操作控制。新功能包括被动健康检查、自动节点修复、强化版Slurm-on-K8s栈、集群详情视图、外部OIDC认证、启动脚本以及可选的验收测试。这些改进帮助团队更快发现硬件故障、自动化恢复流程,并提供更细粒度的访问控制和自定义能力。

Together AI Blog站内正文Together GPU集群新功能:为生产级GPU集群提供可靠性与控制力

Together AI 在首日即引入 Thinking Machines Lab 的新模型 Inkling

Thinking Machines Lab 发布了 Inkling,一个多模态混合专家模型,专注于高效推理、原生多模态理解和广泛任务适用性。Together AI 在其推理平台上提供该模型,支持可控推理努力、文本/图像/音频输入,以及 1M 上下文窗口。

Together AI Blog站内正文Together AI 在首日即引入 Thinking Machines Lab 的新模型 Inkling

开放、便捷且可预测:推出预留吞吐量功能

Together AI 推出预留吞吐量功能,为 MiniMax M3 和 GLM-5.2 等前沿开放模型提供保留推理容量,采用基于 Token 的定价和 99% 正常运行时间 SLA,成本比专有 API 降低高达 90%。

Together AI Blog站内正文开放、便捷且可预测:推出预留吞吐量功能

宣布8亿美元C轮融资:加速向开源AI的转变

Together AI完成8亿美元C轮融资,由Aramco Ventures、NVIDIA、Vista Equity等领投,旨在加速开源AI的普及。公司强调,闭源模型的成本无法规模化,而开源模型结合全栈优化可实现6-20倍成本降低。Together AI已推出FlashAttention-4、Together Megakernel等创新,成为全球最大的AI token生产商之一。

Together AI Blog站内正文宣布8亿美元C轮融资:加速向开源AI的转变

Together AI在ICML 2026:涵盖全栈的前沿研究

Together AI在ICML 2026上有八篇论文被接收,覆盖从智能代理到GPU内核的整个堆栈。这些研究已集成到Together平台中,并在生产环境中应用。

Together AI Blog站内正文Together AI在ICML 2026:涵盖全栈的前沿研究

ParallelKernelBench:前沿LLM尚无法编写快速的多GPU内核

ParallelKernelBench是一个新的基准测试,评估LLM编写多GPU CUDA内核的能力。在87个真实问题中,最佳模型仅能正确解决不到三分之一,且只有不到四分之一的解决方案优于基线。文章分析了模型失败的原因,并展示了几个意外生成的高性能内核案例。

Together AI Blog站内正文ParallelKernelBench:前沿LLM尚无法编写快速的多GPU内核

全部来源