跳到主要内容
AI News HubLIVE

Agent 框架动态

待翻译:The SaaSpocalypse that wasn’t, with Atlassian CEO Mike Cannon-Brookes

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Today, I’m talking with Mike Cannon-Brookes, who is cofounder and CEO of Atlassian. Atlassian is one of those companies that every other company runs on — it makes important platform tools like Jira and Trello that allow people to organize and manage big teams, create shared databases of company information, and generally allow work to happen. As you’ll hear Mike say, all of Atlassian’s products are actually different expressions of a single core platform, which really shapes how Atlassian itself is structured and how those products are built. All of this means Atlassian is also right in the middle of the way AI is changing how all these companies work — AI tools might be able to look at all the different tools and systems you have and just read them for you, m…

The Verge AI站内正文待翻译:The SaaSpocalypse that wasn’t, with Atlassian CEO Mike Cannon-Brookes

待翻译:LensDesigner: A Self-Improving Agent for Optical Lens Design

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2609.30450v1 Announce Type: new Abstract: Optical lens design is a complex, non-convex optimization challenge that relies heavily on human experience and intuition. Existing optimized-based automatic lens design methods struggle to navigate this vast parameter space without meticulous manual tuning. In this paper, we present LensDesigner, an autonomous agent framework that mirrors the problem-solving workflow of expert opticians. To overcome the initial cold start problem, we construct LensLib100K, an extensive optical lens library, and employ Optics-Aware Retrieval to supply physically valid structural seeds. Within an interactive physical simulation environment, the agent executes macroscopic orchestration while receiving immediate optical feedback. Fur…

arXiv Computer Vision站内正文待翻译:LensDesigner: A Self-Improving Agent for Optical Lens Design

待翻译:SlideLab: Audience-Centered Scientific Slide Generation and Evaluation

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2609.30294v1 Announce Type: new Abstract: Scientific presentations are more than summaries of research papers. They need to present the work in a coherent sequence, explain the main ideas clearly, and help the audience follow the presentation. We present SlideLab, a training-free multi-agent framework for generating scientific presentations from research papers. SlideLab first plans the presentation narrative, then builds and iteratively refines a shared slide deck using agents for content planning, visual generation, layout refinement, and grounding verification. In a blind human preference study, SlideLab was preferred over both open-source and commercial systems on 77% of papers while using roughly 4 times fewer inference tokens than the strongest open…

arXiv Computational Linguistics站内正文待翻译:SlideLab: Audience-Centered Scientific Slide Generation and Evaluation

待翻译:Building Production Agents with Jev and LangGraph

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:See how LangGraph orchestrates Jev, TypeSafe AI's decision model, to build faster, cheaper production agents.

LangChain Blog站内正文待翻译:Building Production Agents with Jev and LangGraph

待翻译:NarrateAI: production-ready LLM quality assurance on Amazon Bedrock

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:NarrateAI delivers production-ready LLM quality assurance on Amazon Bedrock. This post details five techniques—adaptive pipeline orchestration, cross-account multi-model failover, real-time streaming evaluation, composite evaluation, and data accuracy verification—that reach about 99% numerical accuracy while streaming responses in real time.

AWS Machine Learning Blog站内正文待翻译:NarrateAI: production-ready LLM quality assurance on Amazon Bedrock

待翻译:New in LangSmith: Engine v2, Managed Deep Agents, Fine-Tuning, and more

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:LangChain announced new updates to LangSmith. Updates include Engine v2 with red teaming and automatic testing, a new version of Managed Deep Agents, trajectories and more.

LangChain Blog站内正文待翻译:New in LangSmith: Engine v2, Managed Deep Agents, Fine-Tuning, and more

待翻译:BaseCamp --- An Agentic AI Framework for Automating DNA Sequencing Data Pipelines

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2609.28557v1 Announce Type: new Abstract: DNA sequencing pipelines, spanning quality control, alignment, variant calling, and annotation, are now reliably executed by workflow management systems that orchestrate established bioinformatics tools at scale. What remains manual is the decision layer surrounding that execution: selecting quality thresholds appropriate to a sample and platform, adjudicating borderline variant calls, diagnosing anomalies, and determining which findings warrant expert review. These decisions are repetitive, judgment-intensive, inconsistent across operators, and frequently undocumented. This paper introduces BaseCamp, a novel agentic AI framework for automating the decision layer of DNA sequencing pipelines. The framework decompos…

arXiv AI站内正文待翻译:BaseCamp --- An Agentic AI Framework for Automating DNA Sequencing Data Pipelines

待翻译:LangSmith Custom Apps: Build custom interfaces around your agent data

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:LangSmith Custom Apps lets you build the interface you want with your LangSmith data, publish it into your workspace, and skip the hosting, auth, and permissions work. Learn more.

LangChain Blog站内正文待翻译:LangSmith Custom Apps: Build custom interfaces around your agent data

待翻译:Introducing LangSmith Fine-Tuning

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:LangChain introduces LangSmith Fine-Tuning and SmithTune, a CLI built for post-training models. Train specialized models without building data pipelines by hand.

LangChain Blog站内正文待翻译:Introducing LangSmith Fine-Tuning

待翻译:Trajectories now in LangSmith: A readable view of every agent session

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Trajectories in LangSmith provide a conversational view of an agent session. Trajectories make trace data easy to navigate and speed up debugging for long-running agents.

LangChain Blog站内正文待翻译:Trajectories now in LangSmith: A readable view of every agent session

待翻译:New in LangSmith Engine: red teaming and automated testing

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:LangSmith Engine now includes Red Teaming to proactively detect agent issues and automated agent testing. Learn more about the Engine v2 release.

LangChain Blog站内正文待翻译:New in LangSmith Engine: red teaming and automated testing

待翻译:Managed Deep Agents delivers a better user experience for agents in production

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Managed Deep Agents is the simplest way to build, deploy, and run agents in production. The 0.8 release adds support for user-owned credentials, user-level memory, HTTP channels, file transfer in Slack and a pre-built tool for web search.

LangChain Blog站内正文待翻译:Managed Deep Agents delivers a better user experience for agents in production

待翻译:Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2609.26891v1 Announce Type: new Abstract: Modern language-model agents are built around the \textit{agent loop}, where the LLM is placed in an environment exposing a set of tools, and the LLM has full control over the workflow by alternating between tool calls and observing their output. However, certain workflows currently require additional engineering beyond the agent loop itself, such as memory systems and self-improving systems. We built an LLM agent framework, JAZ, to explore the extent to which a minimal harness that is little more than the agent loop itself can accomplish tasks these specialized systems are built for. JAZ exposes a single LLM-based primitive invoke and provides a set of built-in hooks that allow the programmer to apply constraints…

arXiv AI站内正文待翻译:Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity

待翻译:Agentic conversational video intelligence built on AWS

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Learn how to build a conversational video intelligence solution on AWS using an agentic architecture. A single Strands Agents SDK agent orchestrates Amazon Bedrock, Amazon Rekognition, and Amazon Transcribe at runtime, deciding which service to call so you can ask natural language questions about your videos and get answers in seconds.

AWS Machine Learning Blog站内正文待翻译:Agentic conversational video intelligence built on AWS

待翻译:The Reliability Layer for Healthcare AI: Common LangSmith Use Cases

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:See how LangSmith helps healthcare AI teams turn clinical review into reusable evaluators, datasets, and release gates for safer AI in production.

LangChain Blog站内正文待翻译:The Reliability Layer for Healthcare AI: Common LangSmith Use Cases

待翻译:Jev-as-a-Judge Is Now Available in LangSmith

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Use Jev as a judge for LangSmith evals to evaluate agent traces with faster, cheaper structured feedback across production runs, datasets, and regression tests.

LangChain Blog站内正文待翻译:Jev-as-a-Judge Is Now Available in LangSmith

待翻译:Python Workers are now generally available

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Python Workers allow developers to run Python web frameworks and AI orchestration libraries natively in the Cloudflare Workers runtime. You can seamlessly integrate with Cloudflare's ecosystem including D1, R2, and Workers AI without writing any JavaScript glue code.

Cloudflare AI Blog站内正文待翻译:Python Workers are now generally available

待翻译:Can Jev Be a Better Agent Evaluator?

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:We tested Jev against LLM judges on accuracy, repeatability, latency, and cost to see whether System One models could offer a new approach to agent evaluation.

LangChain Blog站内正文待翻译:Can Jev Be a Better Agent Evaluator?

待翻译:Migrating multi-model AI agents to Amazon Bedrock AgentCore runtime

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Migrate a multi-model healthcare AI agent from self-managed Amazon ECS with AWS Fargate to Amazon Bedrock AgentCore runtime, preserving triple-model orchestration and vector-enhanced knowledge retrieval while reducing infrastructure management. The framework-agnostic pattern applies across healthcare, financial services, and manufacturing.

AWS Machine Learning Blog站内正文待翻译:Migrating multi-model AI agents to Amazon Bedrock AgentCore runtime

待翻译:Salesforce Agentforce: Bridging the Enterprise AI Gap from ‘Vibe Coding’ to Battle-Tested Orchestration

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Building an AI prototype is easy, but operating autonomous agents at scale requires production-grade tooling. Salesforce Agentforce bridges the gap from "vibe coding" to enterprise reliability by combining synthetic stress-testing, real-time optimization, dynamic agentic UIs, and deterministic guardrails—as proven by Southwest Airlines' 7x ROI. The post Salesforce Agentforce: Bridging the Enterprise AI Gap from ‘Vibe Coding’ to Battle-Tested Orchestration appeared first on MarkTechPost.

MarkTechPost站内正文待翻译:Salesforce Agentforce: Bridging the Enterprise AI Gap from ‘Vibe Coding’ to Battle-Tested Orchestration

待翻译:MAGS: Multi-agent Auto-formalization Guarantees Safety for Agentic Outputs

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2609.19391v1 Announce Type: new Abstract: LLM coding agents now generate complex programs at a scale that makes thorough human review increasingly difficult, raising the risk of safety and security failures. Common approaches, including fuzz testing, static analysis, and LLM-as-a-Verifier, can detect many failures but struggle to cover all possible edge cases. Formal verification addresses this by providing machine-checkable guarantees over specified properties, but traditionally demands substantial manual specification and proof engineering. We introduce a unified multi-agent framework, MAGS, that generates executable programs with formal safety guarantees, using Dafny as a verification-aware intermediate representation where safety properties can be mec…

arXiv AI站内正文待翻译:MAGS: Multi-agent Auto-formalization Guarantees Safety for Agentic Outputs

待翻译:Building a Harness with Jev

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Open Source Agent Architecture LangChain Building a Harness with Jev September 17, 2026 5 min Go back to blog Create agents Agents run in a loop: an LLM decides what to do, a tool executes, a model evaluates the results…

LangChain Blog站内正文待翻译:Building a Harness with Jev

待翻译:Building an Agent Harness for Life Sciences: Introducing Deep Life Sci

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Deep Life Sci is LangChain's open source agentic assistant for clinical and lab scientists. It pulls from 600K+ ClinicalTrials.gov studies, 29M PubMed abstracts, and 12M PubMed Central full-text articles, with sandboxed sub-agents for real data analysis.

LangChain Blog站内正文待翻译:Building an Agent Harness for Life Sciences: Introducing Deep Life Sci

待翻译:How Included Health Built Federated Healthcare Agents with LangGraph and Deep Agents

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:See how Included Health used Deep Agents, LangGraph, and LangSmith to build Dot, a federated healthcare navigation agent with human handoff and clinical oversight.

LangChain Blog站内正文待翻译:How Included Health Built Federated Healthcare Agents with LangGraph and Deep Agents

待翻译:OmniHarness: Harnessing Generalizable Visual Generation via Symbolic Policy Learning

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2609.16057v1 Announce Type: new Abstract: Unified multimodal large language models (MLLMs) and multi-agent systems have advanced visual generation. However, three limitations remain. (1) Existing methods often distill task-specific experience with limited generalizability. (2) Reflection is often deferred until task completion. (3) Knowledge is often acquired only in response to downstream task demands. To address these limitations, we introduce OmniHarness, a framework for generalizable visual generation via symbolic policy learning. OmniHarness abstracts verified executions into symbolic policies for visual generation task families, capturing shared procedures and applicability conditions while removing instance-specific inputs. The harness instantiates…

arXiv Machine Learning站内正文待翻译:OmniHarness: Harnessing Generalizable Visual Generation via Symbolic Policy Learning

待翻译:Optimizing cost and latency with Amazon Bedrock prompt caching

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Prompt caching in Amazon Bedrock can cut input token costs by up to 90% when you repeatedly send the same context to foundation models. This post walks through six practical prompt caching scenarios using the Converse API: message content, system prompt, tool definition, mixed TTL, tenant isolation, and LangChain integration.

AWS Machine Learning Blog站内正文待翻译:Optimizing cost and latency with Amazon Bedrock prompt caching

待翻译:Scaling Agents in Healthcare & Life Sciences: Lessons from Madrigal Pharmaceuticals, Abridge, and Vizient

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Agent programs in healthcare and life sciences are being built under a different set of constraints than those in most industries. There’s plenty of upside if the constraints can be resolved. Success can mean hours of manual review compressed into minutes, data spread across a dozen systems finally queryable in one place, and clinicians getting time back from documentation. At the same time, the cost of a wrong answer can be higher here than almost anywhere else, which changes how teams build.

LangChain Blog站内正文待翻译:Scaling Agents in Healthcare & Life Sciences: Lessons from Madrigal Pharmaceuticals, Abridge, and Vizient

待翻译:Agent-net Open Sources Webagent: A Go Harness That Turns Any Website into a Guarded AI Agent

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Agent-net, the team building an agent-to-agent marketplace where AI agents discover, trust, and pay each other, has released Webagent, an open source harness for standing up public-facing business agents. So, basically you give it your website, get an agent, and let it talk to other agents. Instead of writing orchestration code, a business fills in […] The post Agent-net Open Sources Webagent: A Go Harness That Turns Any Website into a Guarded AI Agent appeared first on MarkTechPost.

MarkTechPost站内正文待翻译:Agent-net Open Sources Webagent: A Go Harness That Turns Any Website into a Guarded AI Agent

待翻译:ArtSociety: Multi-Agent Multimodal Collaboration for Art Emotion Understanding

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2609.13240v1 Announce Type: new Abstract: The AffectiveArt Multidimensional Art Emotion Understanding task asks to jointly predict an artwork's fine-grained emotion (12 classes, 1549:1 head-to-tail ratio), binary valence/arousal, and five attribute-grounded descriptions -- sub-tasks that exhibit strong empirical trade-offs, so the single-model solutions we tried do not jointly optimize all of them well. We present ArtSociety, a multi-agent framework that assembles heterogeneous multimodal experts -- a DINOv2-Giant vision agent (A1), a scene-grounded CoT fine-tuned MLLM (A2), and three closed-source reasoning agents (A3-A5) -- and coordinates them with two training-free controllers: (i) a rare-class-aware voting arbiter that lowers the agreement threshold…

arXiv Computer Vision站内正文待翻译:ArtSociety: Multi-Agent Multimodal Collaboration for Art Emotion Understanding

待翻译:OrchSLM: Probing the Dynamics of Small Language Model Orchestration

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2609.13470v1 Announce Type: new Abstract: Although large language models (LLMs) have demonstrated remarkable capabilities, their reliance on cloud-scale infrastructure poses fundamental challenges for deployment in agentic pipelines, including latency, privacy, connectivity, and substantial computational cost. Small language models (SLMs) offer a compelling alternative: recent studies suggest that many repetitive and narrowly scoped subtasks in agentic workloads may be better served by specialized SLMs than by monolithic LLMs. However, the limited capacity and context windows of SLMs can constrain long-horizon reasoning and interaction-heavy orchestration strategies such as iterative verification and debate. This motivates a complementary, non-interactive…

arXiv AI站内正文待翻译:OrchSLM: Probing the Dynamics of Small Language Model Orchestration

待翻译:Toward Self-Adaptive Physical AI: Can LLM Agents Manage Long-Horizon Physical Tasks?

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2609.13436v1 Announce Type: new Abstract: Large Language Model (LLM) agents offer a promising path toward autonomously managing long-term physical tasks without human intervention. However, physical tasks require agents to continuously observe the environment, make consequential actions, and remain effective as the environment changes. Existing approaches either require substantial data and retraining, or primarily focus on agents operating in the virtual world. In this work, we explore the feasibility of building a self-adaptive physical AI agent that manages long-term physical tasks in a zero-shot manner and adapts to environmental changes without human intervention. We design a multi-agent framework that integrates planning, tool calling, observation,…

arXiv AI站内正文待翻译:Toward Self-Adaptive Physical AI: Can LLM Agents Manage Long-Horizon Physical Tasks?

待翻译:Agent Harness vs Agent Framework vs MCP: Which Layer Owns the Loop, State, Tools, Permissions, and Recovery

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:A practitioner's map of the 3 layers in a modern agent stack, with verified sources and an overlap analysis. The post Agent Harness vs Agent Framework vs MCP: Which Layer Owns the Loop, State, Tools, Permissions, and Recovery appeared first on MarkTechPost.

MarkTechPost站内正文待翻译:Agent Harness vs Agent Framework vs MCP: Which Layer Owns the Loop, State, Tools, Permissions, and Recovery

待翻译:Enterprise Analytics Beyond Dashboards: Intelligent Data Orchestration with LLMs

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:In 17 years of building enterprise data platforms, I’ve watched every organization eventually ask the same question: “Can I ask one question and get one answer across everything my company knows?” A finance analyst wants actual revenue from the warehouse, pipeline data from the CRM, commentary from planning documents, and market signals from external providers. […]

O'Reilly AI & ML Radar站内正文待翻译:Enterprise Analytics Beyond Dashboards: Intelligent Data Orchestration with LLMs

待翻译:How We Built LangChain’s Paid Media Agent

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:How LangChain built a paid media agent to analyze campaign performance, optimize ads, propose changes, and turn marketing data into action.

LangChain Blog站内正文待翻译:How We Built LangChain’s Paid Media Agent

待翻译:Context Engineering Inside the Harness: 4 Mechanisms That Beat Context Overflow and Goal Loss on Long-Horizon Tasks

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:A shallow agent is an LLM calling tools in a loop, and on long tasks it fails in 2 ways: context overflow and goal loss. This article opens the harness layer that fixes both, with the actual thresholds shipped by LangChain Deep Agents, Claude Code, Manus, OpenAI Codex and Amazon Bedrock AgentCore, plus an interactive simulator that shows a 200K window filling up. The post Context Engineering Inside the Harness: 4 Mechanisms That Beat Context Overflow and Goal Loss on Long-Horizon Tasks appeared first on MarkTechPost.

MarkTechPost站内正文待翻译:Context Engineering Inside the Harness: 4 Mechanisms That Beat Context Overflow and Goal Loss on Long-Horizon Tasks

待翻译:Automating Quadratic Unconstrained Binary Optimization (QUBO) Formulation Generation from Natural Language

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2609.10629v1 Announce Type: new Abstract: Quadratic Unconstrained Binary Optimization (QUBO) is a central formulation for combinatorial optimization and has gained increasing attention due to its compatibility with quantum, hybrid quantum-classical, and quantum-inspired solvers. However, translating natural-language problem descriptions into correct QUBO formulations remains difficult, requiring the identification of binary variables, constraints, objective functions, penalty terms, and suitable penalty weights. This process is time-consuming and often demands substantial domain expertise. To address this challenge, we propose an end-to-end multi-agent framework that automatically generates QUBO formulations from natural-language problem descriptions, sup…

arXiv AI站内正文待翻译:Automating Quadratic Unconstrained Binary Optimization (QUBO) Formulation Generation from Natural Language

待翻译:Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Sakana AI has released Fugu Max and Fugu Ultra v2, 2 models built on the same learned orchestration architecture. Fugu Max routes tasks to lean open and specialized models, including NVIDIA Nemotron, at $2/$6 per 1M tokens. Fugu Ultra v2 targets peak capability, scoring 48.3 on Chartography and 74.3 on DeepSWE. The post Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration appeared first on MarkTechPost.

MarkTechPost站内正文待翻译:Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration

待翻译:How Credit Genie keeps codebase docs fresh with OpenWiki

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:See how Credit Genie uses OpenWiki to automate repo documentation, reduce tribal knowledge, and give engineers and coding agents searchable codebase context.

LangChain Blog站内正文待翻译:How Credit Genie keeps codebase docs fresh with OpenWiki

待翻译:Agentic AI-enabled Semantic Commissioning of a Cognitive Digital Twin for Reconfigurable Manufacturing

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2609.09503v1 Announce Type: new Abstract: Rapid bespoke commissioning of the Cognitive Digital Twin (CDT) is a major challenge in reconfigurable manufacturing. Traditional digital twin (DT) construction methods primarily focus on geometric reconstruction, often neglecting the deep semantic integration and functional interoperability necessary for autonomous reasoning. This paper proposes an agent-based, AI-driven workflow to automate end-to-end CDT debugging. The system utilises LangGraph as a multi-agent orchestration engine to achieve dual-path synthesis: the semantic path extracts technical specifications from unstructured documents using Retrieval Augmented Generation (RAG), while the functional path autonomously discovers and binds to real-time indus…

arXiv Robotics站内正文待翻译:Agentic AI-enabled Semantic Commissioning of a Cognitive Digital Twin for Reconfigurable Manufacturing

待翻译:Introducing the Agents API

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Build and launch cloud agents with the Agents API, a managed service powered by the Codex harness for orchestration, long-running sessions, and tool use.

OpenAI News站内正文待翻译:Introducing the Agents API

待翻译:Connections: managed credentials and per-caller identity for Managed Deep Agents

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Learn how Connections in Managed Deep Agents securely manage credentials, support per-user OAuth, and let agents act with each caller’s identity.

LangChain Blog站内正文待翻译:Connections: managed credentials and per-caller identity for Managed Deep Agents

多智能体框架中的上下文组织

deepagents 为子代理引入“上下文模式”(isolated / fork),让主管代理既可以委派任务,又可以决定子代理继承多少上下文。fork 模式继承主管的会话历史,可复用 prompt caching、减少重复工作;isolated 模式则让子代理在全新上下文中独立完成任务。文章还通过 worker、verifier、researcher、memory 四类子代理说明如何选择。

LangChain Blog站内正文多智能体框架中的上下文组织

GitHub 推出 Project HydraFusion:在 Copilot CLI 中按编程任务动态编排多模型运行时工作流

Project HydraFusion 是 GitHub 发布的研究预览,它不再将模型选择视为一次性设置,而是针对每个请求构建并动态编排一个执行工作流,可在不同提供商的模型间进行起草、批判、升级等操作。目前仅在 GitHub Copilot CLI 中以研究预览形式提供,并按实际调用的各模型标准令牌费率计费。

MarkTechPost站内正文GitHub 推出 Project HydraFusion:在 Copilot CLI 中按编程任务动态编排多模型运行时工作流

Project HydraFusion:通过多模型编排实现前沿品质

GitHub 发布 Project HydraFusion 研究预览,通过自动在多个提供商的模型之间编排“单模型、级联、批判”等工作流,为编码任务平衡质量、成本和延迟。离线评估显示,在保持前沿质量的同时可将估算工作流成本降低约 36%–67%。

GitHub AI & ML站内正文Project HydraFusion:通过多模型编排实现前沿品质

Perplexity 在 Mac 上发布混合计算:云代理编排至本地模型,并由设备端门控

Perplexity 在 Mac 上推出混合计算,将任务在云端前沿模型与本地小型模型之间拆分,并通过设备端隐私门控在将敏感数据发送到云端前进行保护。该功能已向 Pro、Max 和企业用户开放,支持 Apple 芯片 Mac。Perplexity 还开源了用于门控的分类器 PII-Tracer,并在基准测试中表现出色。

MarkTechPost站内正文Perplexity 在 Mac 上发布混合计算:云代理编排至本地模型,并由设备端门控

待翻译:Multi-Agent Self-Improving Reinforcement Learning for Video Reasoning

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.28675v1 Announce Type: new Abstract: Video reasoning tasks such as grounded video question answering and temporal grounding require selecting temporal evidence that supports the query. In many current training setups, temporal supervision is applied through local objectives such as boundary regression or span generation, while verification is used mainly to rerank candidate segments at inference time. We study whether a frozen verifier can also guide training. Our multi-agent framework couples a trainable \emph{Grounder} with a frozen \emph{Verifier}: the Grounder samples candidate trajectories and evidence segments, the Verifier assigns query-conditioned segment scores, a group-relative policy-gradient objective favors trajectories that outperform t…

arXiv Computer Vision站内正文待翻译:Multi-Agent Self-Improving Reinforcement Learning for Video Reasoning

待翻译:The Race between Agentic AI Capabilities and Data Quality Control in Online Surveys

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.28597v1 Announce Type: new Abstract: Online surveys are a foundational data collection instrument in a variety of fields, with attention checks serving as critical guardians of response quality. However, the rapid emergence of agentic AI (goal directed systems powered by a large language model (LLM) brain and/or a multimodal processing unit with tool-augmented capabilities) raises new questions about the robustness of these safeguards. We investigate how well agentic AI architectures can complete web-based surveys and pass standard attention checks. We evaluate a single-agent architecture capable of multimodal input processing and tool-based web interaction on a controlled survey sandbox. We analyze the problem from two perspectives. From an attack p…

arXiv AI站内正文待翻译:The Race between Agentic AI Capabilities and Data Quality Control in Online Surveys

待翻译:90 days of attacks on AI infrastructure

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Wiz PricingGet a demo Get a demo Wiz Threat Research operates honeypots across AI and ML services including LiteLLM, Flowise, LangChain, Langflow, ChromaDB, Ollama, and others. Over 90 days of telemetry, we observed sus…

Hacker News AI站内正文待翻译:90 days of attacks on AI infrastructure

待翻译:Show HN: ShevtoneAudio Orchestrator – Turning MIDI into Full Orchestration

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Professional MIDI orchestration plugin Turn your MIDI into a full orchestra. ShevtoneAudio Orchestrator is a professional local MIDI orchestration tool built for composers working in film, trailer, television, games and…

Hacker News AI站内正文待翻译:Show HN: ShevtoneAudio Orchestrator – Turning MIDI into Full Orchestration

待翻译:Procedura: Agentic 3D Modeling with Procedural Control

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26238v1 Announce Type: new Abstract: Native 3D generators now recover impressive mesh geometry from a single image. However, a dense mesh stays soft where a machined object should be sharp, it carries no part decomposition, and it exposes no parameter a user could edit. To address this, we explore the paradigm of 3D shape as code, leveraging and scaling the coding ability of an LLM for 3D modeling. We introduce Procedura, a novel 3D modeling agent framework that writes an object as a procedural assembly, a parametric program whose named parts are joined by typed, machine-checkable mates. From a text prompt, the agent plans the object as an assembly graph and writes the program part by part, solving each placement from the mated frames rather than gue…

arXiv Computer Vision站内正文待翻译:Procedura: Agentic 3D Modeling with Procedural Control

更多增长标签

Agent 框架 AI News | AI News Hub