跳到主要內容
AI News HubLIVE

Agent 框架動態

待翻譯:Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Sakana AI has released Fugu Max and Fugu Ultra v2, 2 models built on the same learned orchestration architecture. Fugu Max routes tasks to lean open and specialized models, including NVIDIA Nemotron, at $2/$6 per 1M tokens. Fugu Ultra v2 targets peak capability, scoring 48.3 on Chartography and 74.3 on DeepSWE. The post Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration appeared first on MarkTechPost.

MarkTechPost站內正文待翻譯:Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration

待翻譯:How Credit Genie keeps codebase docs fresh with OpenWiki

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:See how Credit Genie uses OpenWiki to automate repo documentation, reduce tribal knowledge, and give engineers and coding agents searchable codebase context.

LangChain Blog站內正文待翻譯:How Credit Genie keeps codebase docs fresh with OpenWiki

待翻譯:Agentic AI-enabled Semantic Commissioning of a Cognitive Digital Twin for Reconfigurable Manufacturing

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09503v1 Announce Type: new Abstract: Rapid bespoke commissioning of the Cognitive Digital Twin (CDT) is a major challenge in reconfigurable manufacturing. Traditional digital twin (DT) construction methods primarily focus on geometric reconstruction, often neglecting the deep semantic integration and functional interoperability necessary for autonomous reasoning. This paper proposes an agent-based, AI-driven workflow to automate end-to-end CDT debugging. The system utilises LangGraph as a multi-agent orchestration engine to achieve dual-path synthesis: the semantic path extracts technical specifications from unstructured documents using Retrieval Augmented Generation (RAG), while the functional path autonomously discovers and binds to real-time indus…

arXiv Robotics站內正文待翻譯:Agentic AI-enabled Semantic Commissioning of a Cognitive Digital Twin for Reconfigurable Manufacturing

待翻譯:Introducing the Agents API

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Build and launch cloud agents with the Agents API, a managed service powered by the Codex harness for orchestration, long-running sessions, and tool use.

OpenAI News站內正文待翻譯:Introducing the Agents API

待翻譯:Connections: managed credentials and per-caller identity for Managed Deep Agents

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Learn how Connections in Managed Deep Agents securely manage credentials, support per-user OAuth, and let agents act with each caller’s identity.

LangChain Blog站內正文待翻譯:Connections: managed credentials and per-caller identity for Managed Deep Agents

多智慧體框架中的上下文組織

deepagents 為子代理引入“上下文模式”(isolated / fork),讓主管代理既可以委派任務,又可以決定子代理繼承多少上下文。fork 模式繼承主管的會話歷史,可複用 prompt caching、減少重複工作;isolated 模式則讓子代理在全新上下文中獨立完成任務。文章還透過 worker、verifier、researcher、memory 四類子代理說明如何選擇。

LangChain Blog站內正文多智慧體框架中的上下文組織

GitHub 推出 Project HydraFusion:在 Copilot CLI 中按程式設計任務動態編排多模型執行時工作流

Project HydraFusion 是 GitHub 釋出的研究預覽,它不再將模型選擇視為一次性設定,而是針對每個請求構建並動態編排一個執行工作流,可在不同提供商的模型間進行起草、批判、升級等操作。目前僅在 GitHub Copilot CLI 中以研究預覽形式提供,並按實際呼叫的各模型標準令牌費率計費。

MarkTechPost站內正文GitHub 推出 Project HydraFusion:在 Copilot CLI 中按程式設計任務動態編排多模型執行時工作流

Project HydraFusion:透過多模型編排實現前沿品質

GitHub 釋出 Project HydraFusion 研究預覽,透過自動在多個提供商的模型之間編排“單模型、級聯、批判”等工作流,為編碼任務平衡質量、成本和延遲。離線評估顯示,在保持前沿質量的同時可將估算工作流成本降低約 36%–67%。

GitHub AI & ML站內正文Project HydraFusion:透過多模型編排實現前沿品質

Perplexity 在 Mac 上釋出混合計算:雲代理編排至本地模型,並由裝置端門控

Perplexity 在 Mac 上推出混合計算,將任務在雲端前沿模型與本地小型模型之間拆分,並透過裝置端隱私門控在將敏感資料傳送到雲端前進行保護。該功能已向 Pro、Max 和企業使用者開放,支援 Apple 晶片 Mac。Perplexity 還開源了用於門控的分類器 PII-Tracer,並在基準測試中表現出色。

MarkTechPost站內正文Perplexity 在 Mac 上釋出混合計算:雲代理編排至本地模型,並由裝置端門控

待翻譯:Multi-Agent Self-Improving Reinforcement Learning for Video Reasoning

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.28675v1 Announce Type: new Abstract: Video reasoning tasks such as grounded video question answering and temporal grounding require selecting temporal evidence that supports the query. In many current training setups, temporal supervision is applied through local objectives such as boundary regression or span generation, while verification is used mainly to rerank candidate segments at inference time. We study whether a frozen verifier can also guide training. Our multi-agent framework couples a trainable \emph{Grounder} with a frozen \emph{Verifier}: the Grounder samples candidate trajectories and evidence segments, the Verifier assigns query-conditioned segment scores, a group-relative policy-gradient objective favors trajectories that outperform t…

arXiv Computer Vision站內正文待翻譯:Multi-Agent Self-Improving Reinforcement Learning for Video Reasoning

待翻譯:The Race between Agentic AI Capabilities and Data Quality Control in Online Surveys

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.28597v1 Announce Type: new Abstract: Online surveys are a foundational data collection instrument in a variety of fields, with attention checks serving as critical guardians of response quality. However, the rapid emergence of agentic AI (goal directed systems powered by a large language model (LLM) brain and/or a multimodal processing unit with tool-augmented capabilities) raises new questions about the robustness of these safeguards. We investigate how well agentic AI architectures can complete web-based surveys and pass standard attention checks. We evaluate a single-agent architecture capable of multimodal input processing and tool-based web interaction on a controlled survey sandbox. We analyze the problem from two perspectives. From an attack p…

arXiv AI站內正文待翻譯:The Race between Agentic AI Capabilities and Data Quality Control in Online Surveys

待翻譯:90 days of attacks on AI infrastructure

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Wiz PricingGet a demo Get a demo Wiz Threat Research operates honeypots across AI and ML services including LiteLLM, Flowise, LangChain, Langflow, ChromaDB, Ollama, and others. Over 90 days of telemetry, we observed sus…

Hacker News AI站內正文待翻譯:90 days of attacks on AI infrastructure

待翻譯:Show HN: ShevtoneAudio Orchestrator – Turning MIDI into Full Orchestration

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Professional MIDI orchestration plugin Turn your MIDI into a full orchestra. ShevtoneAudio Orchestrator is a professional local MIDI orchestration tool built for composers working in film, trailer, television, games and…

Hacker News AI站內正文待翻譯:Show HN: ShevtoneAudio Orchestrator – Turning MIDI into Full Orchestration

待翻譯:Procedura: Agentic 3D Modeling with Procedural Control

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.26238v1 Announce Type: new Abstract: Native 3D generators now recover impressive mesh geometry from a single image. However, a dense mesh stays soft where a machined object should be sharp, it carries no part decomposition, and it exposes no parameter a user could edit. To address this, we explore the paradigm of 3D shape as code, leveraging and scaling the coding ability of an LLM for 3D modeling. We introduce Procedura, a novel 3D modeling agent framework that writes an object as a procedural assembly, a parametric program whose named parts are joined by typed, machine-checkable mates. From a text prompt, the agent plans the object as an assembly graph and writes the program part by part, solving each placement from the mated frames rather than gue…

arXiv Computer Vision站內正文待翻譯:Procedura: Agentic 3D Modeling with Procedural Control

待翻譯:Pushing GPT-5.6 Luna from 0% to 56% on ARC-AGI-3 Public

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Research August 27, 2026 For long-horizon agents, model capability alone does not determine system capability. Infrastructure orchestration multiplies what models can do. INT21’s SwarmOS tested this hypothesis on the AR…

Hacker News AI站內正文待翻譯:Pushing GPT-5.6 Luna from 0% to 56% on ARC-AGI-3 Public

待翻譯:Fusing Perceptual Vision Experts with Multimodal Large Language Models for Explainable Plant Disease Diagnosis: From Benchmark Imagery to Real-World Robotic Field Validation

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.24934v1 Announce Type: new Abstract: Accurate field plant disease diagnosis requires reliable fusion of uncertain and conflicting perceptual evidence. We present the Hybrid Hierarchical Multi-Agent Framework (H$^{2}$MAF), combining decision-level fusion of EfficientNet-B3 and ConvNeXt-Tiny with semantic arbitration by open-weight multimodal large language models (MLLMs), Gemma 4 E4B and Qwen3.5 4B, using structured JSON evidence to generate explainable diagnoses, risk levels, treatment urgency, and financial exposure. (H$^{2}$MAF) is evaluated on 14,364 images (1,370 test images) across PlantDoc (2,922 images, 27 classes) and two non-public, continuously captured Cornell robot-acquired field datasets: Stage 2 (20 GB; 4,215 images) and Stage 4 (40 GB;…

arXiv Computer Vision站內正文待翻譯:Fusing Perceptual Vision Experts with Multimodal Large Language Models for Explainable Plant Disease Diagnosis: From Benchmark Imagery to Real-World Robotic Field Validation

待翻譯:August 2026: LangChain Newsletter — Managed Deep Agents, LLM Gateway, and More

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Managed Deep Agents and LLM Gateway hit public beta, plus Deep Agents v0.7, Tuned Evaluators, Bring Your Own Cloud on AWS, and LangSmith Engine upgrades.

LangChain Blog站內正文待翻譯:August 2026: LangChain Newsletter — Managed Deep Agents, LLM Gateway, and More

待翻譯:Evaluate any agent framework with Amazon Bedrock AgentCore Evaluations

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Amazon Bedrock AgentCore Evaluations decouples agent evaluation from the framework you build on. As long as your agent emits OpenTelemetry telemetry, the service can score it, whether you use LangGraph, LlamaIndex, the OpenAI Agents SDK, Google ADK, the Claude Agent SDK, or Strands Agents. This post explains how the framework-agnostic contract works.

AWS Machine Learning Blog站內正文待翻譯:Evaluate any agent framework with Amazon Bedrock AgentCore Evaluations

待翻譯:Best TypeScript AI Agent Frameworks for Next.js

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:You’ve shipped a working chat feature. Now comes the hard part: figuring out which framework actually fits your production architecture. The tooling landscape has fractured, and comparing AI agent frameworks in a vacuum…

Hacker News AI站內正文待翻譯:Best TypeScript AI Agent Frameworks for Next.js

待翻譯:LangChain State of AI 2024 Report

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Dive into LangSmith product usage patterns that show how the AI ecosystem and the way people are building LLM apps is evolving.

LangChain Blog站內正文待翻譯:LangChain State of AI 2024 Report

待翻譯:LangChain's Second Birthday

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Reflections on how LangChain has evolved — including our products, ecosystem, and community — over the past two years, and where we're headed next.

LangChain Blog站內正文待翻譯:LangChain's Second Birthday

待翻譯:How Podium optimized agent behavior and reduced engineering intervention by 90% with LangSmith

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:See how Podium tests across the lifecycle development of their AI employee agent, using LangSmith for dataset curation and finetuning. They improved agent F1 response quality to 98% and reduced the need for engineering intervention by 90%.

LangChain Blog站內正文待翻譯:How Podium optimized agent behavior and reduced engineering intervention by 90% with LangSmith

待翻譯:Announcing LangGraph v0.1 & LangGraph Cloud: Running agents at scale, reliably

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Our new infrastructure for running agents at scale, LangGraph Cloud, is available in beta. We also have a new stable release of LangGraph.

LangChain Blog站內正文待翻譯:Announcing LangGraph v0.1 & LangGraph Cloud: Running agents at scale, reliably

待翻譯:LangChain Announces Enterprise Agentic AI Platform Built with NVIDIA

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Build, deploy, and monitor production-grade AI agents at scale with LangChain's enterprise agentic AI platform integrated with NVIDIA.

LangChain Blog站內正文待翻譯:LangChain Announces Enterprise Agentic AI Platform Built with NVIDIA

待翻譯:AI Agent Latency 101: How do I speed up my AI agent?

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Learn proven strategies to speed up your AI agent: reduce latency, optimize LLM calls, enable parallelism, and improve UX. Expert tips from LangChain.

LangChain Blog站內正文待翻譯:AI Agent Latency 101: How do I speed up my AI agent?

待翻譯:Connect Amazon Bedrock AgentCore to cross-account knowledge bases

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Learn how Amazon Bedrock AgentCore agents in one account can generate answers from an Amazon Bedrock knowledge base backed by Amazon Redshift Serverless in another account, without copying source data. This post covers the architecture, security boundary, and two orchestration models: a code-based Strands agent and a declarative AgentCore harness.

AWS Machine Learning Blog站內正文待翻譯:Connect Amazon Bedrock AgentCore to cross-account knowledge bases

待翻譯:Evaluating OpenWiki with WikiBench

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:We built WikiBench to test whether generated wikis help coding agents. Pairing a wiki with source code scored higher than source alone, at lower cost.

LangChain Blog站內正文待翻譯:Evaluating OpenWiki with WikiBench

待翻譯:OpenAI's Bet on a Cognitive Architecture

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Why LangChain believes in open, customizable cognitive architectures over closed systems. Build reliable LLM agents with OpenGPTs and LangSmith.

LangChain Blog站內正文待翻譯:OpenAI's Bet on a Cognitive Architecture

待翻譯:Orchestration is the new challenge for CX in the age of AI agents

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Presented by Tata Communications Enterprises are deploying AI agents, voice AI, and automation across messaging, voice, and digital channels faster than the architecture meant to support it. Most of that deployment has involved attaching conversational AI to legacy systems never built for it, says Gaurav Anand, global head of the Customer Interaction Suite at Tata Communications. "In the rush to deploy AI, organizations have largely bolted conversational AI onto legacy systems," Anand says. "As a result, while many enterprises have adopted digital tools, very few have platforms that are truly integrated, scaled, and capable of seamless orchestration." That gap creates a heavy cognitive load for human agents who must piece together context across disjointed tool…

VentureBeat AI站內正文待翻譯:Orchestration is the new challenge for CX in the age of AI agents

待翻譯:Meet Connery: An Open-Source Plugin Infrastructure for OpenGPTs and LLM apps

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discover Connery: open-source plugin infrastructure for LLM apps. Secure integrations, personalization, and human-in-the-loop control for AI agents.

LangChain Blog站內正文待翻譯:Meet Connery: An Open-Source Plugin Infrastructure for OpenGPTs and LLM apps

待翻譯:Qdrant x LangChain: Endgame Performance

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Qdrant and LangChain deliver production-ready RAG performance with async support, optimized resource usage, and scalable vector search for LLM apps.

LangChain Blog站內正文待翻譯:Qdrant x LangChain: Endgame Performance

待翻譯:Data-Driven Characters

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Data-driven-characters is a repo for creating, debugging, and interacting your own chatbots conditioned on your own story corpora.

LangChain Blog站內正文待翻譯:Data-Driven Characters

待翻譯:Eden AI x LangChain: Harnessing LLMs, Embeddings, and AI

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Access multiple LLMs, embeddings, and AI tools through Eden AI's LangChain integration. Unified API for text generation, OCR, speech-to-text, and more.

LangChain Blog站內正文待翻譯:Eden AI x LangChain: Harnessing LLMs, Embeddings, and AI

待翻譯:LangChain State of AI 2023

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discover how developers build LLM applications in 2023. Insights on popular models, vectorstores, retrieval strategies, and testing methods from LangSmith.

LangChain Blog站內正文待翻譯:LangChain State of AI 2023

待翻譯:How to design an Agent for Production

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Build production-ready AI agents with LangChain. Technical guide covering OpenAI functions, tools, prompts, and architecture for Cal.ai's scheduling assistant.

LangChain Blog站內正文待翻譯:How to design an Agent for Production

待翻譯:Auto-Evaluator Opportunities

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Auto-evaluate LLM question-answer chains with LangChain's free tool. Generate test sets, grade answers, and optimize chain performance.

LangChain Blog站內正文待翻譯:Auto-Evaluator Opportunities

待翻譯:Applying OpenAI's RAG Strategies

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Implement OpenAI's proven RAG strategies with LangChain. Explore query transformations, routing, post-processing, and evaluation methods for optimal retrieval.

LangChain Blog站內正文待翻譯:Applying OpenAI's RAG Strategies

待翻譯:Autonomous Agents & Agent Simulations

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Explore how LangChain implements autonomous agents like AutoGPT and BabyAGI. Learn about planning techniques, memory systems, and agent simulations.

LangChain Blog站內正文待翻譯:Autonomous Agents & Agent Simulations

待翻譯:LLM Agents Perform Controlled Experiments Using Simulation Models

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23622v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong capabilities in reasoning, planning, and tool use, but many scientific and engineering tasks require more than plausible text and code generation. They require understanding how a system responds to intervention, which in practice depends on controlled experimentation. In this work, we propose a multi-agent framework that enables LLM agents to conduct controlled experiments with scientific simulation models for pharmaceutical process design. Given a user query and a baseline configuration, the system constructs a structured task representation, designs experiments, executes comparative simulation, interprets the resulting outcomes, and synthesizes evidence-based recom…

arXiv AI站內正文待翻譯:LLM Agents Perform Controlled Experiments Using Simulation Models

待翻譯:Using LangSmith to Support Fine-tuning

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Learn how to fine-tune and evaluate LLMs with LangSmith for dataset management. Complete guide covers LLaMA2 and GPT-3.5 fine-tuning with practical examples.

LangChain Blog站內正文待翻譯:Using LangSmith to Support Fine-tuning

待翻譯:Benchmarking Question/Answering Over CSV Data

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Build better Q&A systems for CSV data using LangChain agents, retrieval, and LLM evaluation. Includes benchmarks, debugging insights, and open-source code.

LangChain Blog站內正文待翻譯:Benchmarking Question/Answering Over CSV Data

待翻譯:Announcing our $10M seed round led by Benchmark

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:LangChain secures $10M seed round from Benchmark to empower developers building AI apps with our open-source framework for data-aware, agentic LLMs.

LangChain Blog站內正文待翻譯:Announcing our $10M seed round led by Benchmark

待翻譯:Making Data Ingestion Production Ready: a LangChain-Powered Airbyte Destination

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Scale retrieval apps to production with LangChain's Airbyte integration. Automate data ingestion with scheduling, text splitting, and 50+ embeddings.

LangChain Blog站內正文待翻譯:Making Data Ingestion Production Ready: a LangChain-Powered Airbyte Destination

待翻譯:LangServe Playground and Configurability

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Deploy LangChain apps with LangServe's playground UI and configurable parameters. Experiment with models, share with teams, stream in real-time.

LangChain Blog站內正文待翻譯:LangServe Playground and Configurability

待翻譯:Retrieval

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Build better AI apps with flexible retrieval methods in LangChain. Use any retriever—from semantic to hybrid—to create personalized ChatGPT for your data.

LangChain Blog站內正文待翻譯:Retrieval

待翻譯:Cube x LangChain: Building AI experiences with LLMs and the semantic layer

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Build AI-powered data experiences with Cube's semantic layer and LangChain. Prevent hallucinations, query in natural language, create conversational interfaces.

LangChain Blog站內正文待翻譯:Cube x LangChain: Building AI experiences with LLMs and the semantic layer

更多增長標籤

Agent 框架 AI News | AI News Hub