跳到主要內容
AI News HubLIVE

Agent 框架動態

待翻譯:The SaaSpocalypse that wasn’t, with Atlassian CEO Mike Cannon-Brookes

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Today, I’m talking with Mike Cannon-Brookes, who is cofounder and CEO of Atlassian. Atlassian is one of those companies that every other company runs on — it makes important platform tools like Jira and Trello that allow people to organize and manage big teams, create shared databases of company information, and generally allow work to happen. As you’ll hear Mike say, all of Atlassian’s products are actually different expressions of a single core platform, which really shapes how Atlassian itself is structured and how those products are built. All of this means Atlassian is also right in the middle of the way AI is changing how all these companies work — AI tools might be able to look at all the different tools and systems you have and just read them for you, m…

The Verge AI站內正文待翻譯:The SaaSpocalypse that wasn’t, with Atlassian CEO Mike Cannon-Brookes

待翻譯:LensDesigner: A Self-Improving Agent for Optical Lens Design

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.30450v1 Announce Type: new Abstract: Optical lens design is a complex, non-convex optimization challenge that relies heavily on human experience and intuition. Existing optimized-based automatic lens design methods struggle to navigate this vast parameter space without meticulous manual tuning. In this paper, we present LensDesigner, an autonomous agent framework that mirrors the problem-solving workflow of expert opticians. To overcome the initial cold start problem, we construct LensLib100K, an extensive optical lens library, and employ Optics-Aware Retrieval to supply physically valid structural seeds. Within an interactive physical simulation environment, the agent executes macroscopic orchestration while receiving immediate optical feedback. Fur…

arXiv Computer Vision站內正文待翻譯:LensDesigner: A Self-Improving Agent for Optical Lens Design

待翻譯:SlideLab: Audience-Centered Scientific Slide Generation and Evaluation

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.30294v1 Announce Type: new Abstract: Scientific presentations are more than summaries of research papers. They need to present the work in a coherent sequence, explain the main ideas clearly, and help the audience follow the presentation. We present SlideLab, a training-free multi-agent framework for generating scientific presentations from research papers. SlideLab first plans the presentation narrative, then builds and iteratively refines a shared slide deck using agents for content planning, visual generation, layout refinement, and grounding verification. In a blind human preference study, SlideLab was preferred over both open-source and commercial systems on 77% of papers while using roughly 4 times fewer inference tokens than the strongest open…

arXiv Computational Linguistics站內正文待翻譯:SlideLab: Audience-Centered Scientific Slide Generation and Evaluation

待翻譯:Building Production Agents with Jev and LangGraph

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:See how LangGraph orchestrates Jev, TypeSafe AI's decision model, to build faster, cheaper production agents.

LangChain Blog站內正文待翻譯:Building Production Agents with Jev and LangGraph

待翻譯:NarrateAI: production-ready LLM quality assurance on Amazon Bedrock

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:NarrateAI delivers production-ready LLM quality assurance on Amazon Bedrock. This post details five techniques—adaptive pipeline orchestration, cross-account multi-model failover, real-time streaming evaluation, composite evaluation, and data accuracy verification—that reach about 99% numerical accuracy while streaming responses in real time.

AWS Machine Learning Blog站內正文待翻譯:NarrateAI: production-ready LLM quality assurance on Amazon Bedrock

待翻譯:New in LangSmith: Engine v2, Managed Deep Agents, Fine-Tuning, and more

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:LangChain announced new updates to LangSmith. Updates include Engine v2 with red teaming and automatic testing, a new version of Managed Deep Agents, trajectories and more.

LangChain Blog站內正文待翻譯:New in LangSmith: Engine v2, Managed Deep Agents, Fine-Tuning, and more

待翻譯:BaseCamp --- An Agentic AI Framework for Automating DNA Sequencing Data Pipelines

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.28557v1 Announce Type: new Abstract: DNA sequencing pipelines, spanning quality control, alignment, variant calling, and annotation, are now reliably executed by workflow management systems that orchestrate established bioinformatics tools at scale. What remains manual is the decision layer surrounding that execution: selecting quality thresholds appropriate to a sample and platform, adjudicating borderline variant calls, diagnosing anomalies, and determining which findings warrant expert review. These decisions are repetitive, judgment-intensive, inconsistent across operators, and frequently undocumented. This paper introduces BaseCamp, a novel agentic AI framework for automating the decision layer of DNA sequencing pipelines. The framework decompos…

arXiv AI站內正文待翻譯:BaseCamp --- An Agentic AI Framework for Automating DNA Sequencing Data Pipelines

待翻譯:LangSmith Custom Apps: Build custom interfaces around your agent data

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:LangSmith Custom Apps lets you build the interface you want with your LangSmith data, publish it into your workspace, and skip the hosting, auth, and permissions work. Learn more.

LangChain Blog站內正文待翻譯:LangSmith Custom Apps: Build custom interfaces around your agent data

待翻譯:Introducing LangSmith Fine-Tuning

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:LangChain introduces LangSmith Fine-Tuning and SmithTune, a CLI built for post-training models. Train specialized models without building data pipelines by hand.

LangChain Blog站內正文待翻譯:Introducing LangSmith Fine-Tuning

待翻譯:Trajectories now in LangSmith: A readable view of every agent session

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Trajectories in LangSmith provide a conversational view of an agent session. Trajectories make trace data easy to navigate and speed up debugging for long-running agents.

LangChain Blog站內正文待翻譯:Trajectories now in LangSmith: A readable view of every agent session

待翻譯:New in LangSmith Engine: red teaming and automated testing

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:LangSmith Engine now includes Red Teaming to proactively detect agent issues and automated agent testing. Learn more about the Engine v2 release.

LangChain Blog站內正文待翻譯:New in LangSmith Engine: red teaming and automated testing

待翻譯:Managed Deep Agents delivers a better user experience for agents in production

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Managed Deep Agents is the simplest way to build, deploy, and run agents in production. The 0.8 release adds support for user-owned credentials, user-level memory, HTTP channels, file transfer in Slack and a pre-built tool for web search.

LangChain Blog站內正文待翻譯:Managed Deep Agents delivers a better user experience for agents in production

待翻譯:Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.26891v1 Announce Type: new Abstract: Modern language-model agents are built around the \textit{agent loop}, where the LLM is placed in an environment exposing a set of tools, and the LLM has full control over the workflow by alternating between tool calls and observing their output. However, certain workflows currently require additional engineering beyond the agent loop itself, such as memory systems and self-improving systems. We built an LLM agent framework, JAZ, to explore the extent to which a minimal harness that is little more than the agent loop itself can accomplish tasks these specialized systems are built for. JAZ exposes a single LLM-based primitive invoke and provides a set of built-in hooks that allow the programmer to apply constraints…

arXiv AI站內正文待翻譯:Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity

待翻譯:Agentic conversational video intelligence built on AWS

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Learn how to build a conversational video intelligence solution on AWS using an agentic architecture. A single Strands Agents SDK agent orchestrates Amazon Bedrock, Amazon Rekognition, and Amazon Transcribe at runtime, deciding which service to call so you can ask natural language questions about your videos and get answers in seconds.

AWS Machine Learning Blog站內正文待翻譯:Agentic conversational video intelligence built on AWS

待翻譯:The Reliability Layer for Healthcare AI: Common LangSmith Use Cases

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:See how LangSmith helps healthcare AI teams turn clinical review into reusable evaluators, datasets, and release gates for safer AI in production.

LangChain Blog站內正文待翻譯:The Reliability Layer for Healthcare AI: Common LangSmith Use Cases

待翻譯:Jev-as-a-Judge Is Now Available in LangSmith

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Use Jev as a judge for LangSmith evals to evaluate agent traces with faster, cheaper structured feedback across production runs, datasets, and regression tests.

LangChain Blog站內正文待翻譯:Jev-as-a-Judge Is Now Available in LangSmith

待翻譯:Python Workers are now generally available

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Python Workers allow developers to run Python web frameworks and AI orchestration libraries natively in the Cloudflare Workers runtime. You can seamlessly integrate with Cloudflare's ecosystem including D1, R2, and Workers AI without writing any JavaScript glue code.

Cloudflare AI Blog站內正文待翻譯:Python Workers are now generally available

待翻譯:Can Jev Be a Better Agent Evaluator?

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:We tested Jev against LLM judges on accuracy, repeatability, latency, and cost to see whether System One models could offer a new approach to agent evaluation.

LangChain Blog站內正文待翻譯:Can Jev Be a Better Agent Evaluator?

待翻譯:Migrating multi-model AI agents to Amazon Bedrock AgentCore runtime

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Migrate a multi-model healthcare AI agent from self-managed Amazon ECS with AWS Fargate to Amazon Bedrock AgentCore runtime, preserving triple-model orchestration and vector-enhanced knowledge retrieval while reducing infrastructure management. The framework-agnostic pattern applies across healthcare, financial services, and manufacturing.

AWS Machine Learning Blog站內正文待翻譯:Migrating multi-model AI agents to Amazon Bedrock AgentCore runtime

待翻譯:Salesforce Agentforce: Bridging the Enterprise AI Gap from ‘Vibe Coding’ to Battle-Tested Orchestration

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Building an AI prototype is easy, but operating autonomous agents at scale requires production-grade tooling. Salesforce Agentforce bridges the gap from "vibe coding" to enterprise reliability by combining synthetic stress-testing, real-time optimization, dynamic agentic UIs, and deterministic guardrails—as proven by Southwest Airlines' 7x ROI. The post Salesforce Agentforce: Bridging the Enterprise AI Gap from ‘Vibe Coding’ to Battle-Tested Orchestration appeared first on MarkTechPost.

MarkTechPost站內正文待翻譯:Salesforce Agentforce: Bridging the Enterprise AI Gap from ‘Vibe Coding’ to Battle-Tested Orchestration

待翻譯:MAGS: Multi-agent Auto-formalization Guarantees Safety for Agentic Outputs

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.19391v1 Announce Type: new Abstract: LLM coding agents now generate complex programs at a scale that makes thorough human review increasingly difficult, raising the risk of safety and security failures. Common approaches, including fuzz testing, static analysis, and LLM-as-a-Verifier, can detect many failures but struggle to cover all possible edge cases. Formal verification addresses this by providing machine-checkable guarantees over specified properties, but traditionally demands substantial manual specification and proof engineering. We introduce a unified multi-agent framework, MAGS, that generates executable programs with formal safety guarantees, using Dafny as a verification-aware intermediate representation where safety properties can be mec…

arXiv AI站內正文待翻譯:MAGS: Multi-agent Auto-formalization Guarantees Safety for Agentic Outputs

待翻譯:Building a Harness with Jev

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Open Source Agent Architecture LangChain Building a Harness with Jev September 17, 2026 5 min Go back to blog Create agents Agents run in a loop: an LLM decides what to do, a tool executes, a model evaluates the results…

LangChain Blog站內正文待翻譯:Building a Harness with Jev

待翻譯:Building an Agent Harness for Life Sciences: Introducing Deep Life Sci

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Deep Life Sci is LangChain's open source agentic assistant for clinical and lab scientists. It pulls from 600K+ ClinicalTrials.gov studies, 29M PubMed abstracts, and 12M PubMed Central full-text articles, with sandboxed sub-agents for real data analysis.

LangChain Blog站內正文待翻譯:Building an Agent Harness for Life Sciences: Introducing Deep Life Sci

待翻譯:How Included Health Built Federated Healthcare Agents with LangGraph and Deep Agents

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:See how Included Health used Deep Agents, LangGraph, and LangSmith to build Dot, a federated healthcare navigation agent with human handoff and clinical oversight.

LangChain Blog站內正文待翻譯:How Included Health Built Federated Healthcare Agents with LangGraph and Deep Agents

待翻譯:OmniHarness: Harnessing Generalizable Visual Generation via Symbolic Policy Learning

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.16057v1 Announce Type: new Abstract: Unified multimodal large language models (MLLMs) and multi-agent systems have advanced visual generation. However, three limitations remain. (1) Existing methods often distill task-specific experience with limited generalizability. (2) Reflection is often deferred until task completion. (3) Knowledge is often acquired only in response to downstream task demands. To address these limitations, we introduce OmniHarness, a framework for generalizable visual generation via symbolic policy learning. OmniHarness abstracts verified executions into symbolic policies for visual generation task families, capturing shared procedures and applicability conditions while removing instance-specific inputs. The harness instantiates…

arXiv Machine Learning站內正文待翻譯:OmniHarness: Harnessing Generalizable Visual Generation via Symbolic Policy Learning

待翻譯:Optimizing cost and latency with Amazon Bedrock prompt caching

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Prompt caching in Amazon Bedrock can cut input token costs by up to 90% when you repeatedly send the same context to foundation models. This post walks through six practical prompt caching scenarios using the Converse API: message content, system prompt, tool definition, mixed TTL, tenant isolation, and LangChain integration.

AWS Machine Learning Blog站內正文待翻譯:Optimizing cost and latency with Amazon Bedrock prompt caching

待翻譯:Scaling Agents in Healthcare & Life Sciences: Lessons from Madrigal Pharmaceuticals, Abridge, and Vizient

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Agent programs in healthcare and life sciences are being built under a different set of constraints than those in most industries. There’s plenty of upside if the constraints can be resolved. Success can mean hours of manual review compressed into minutes, data spread across a dozen systems finally queryable in one place, and clinicians getting time back from documentation. At the same time, the cost of a wrong answer can be higher here than almost anywhere else, which changes how teams build.

LangChain Blog站內正文待翻譯:Scaling Agents in Healthcare & Life Sciences: Lessons from Madrigal Pharmaceuticals, Abridge, and Vizient

待翻譯:Agent-net Open Sources Webagent: A Go Harness That Turns Any Website into a Guarded AI Agent

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Agent-net, the team building an agent-to-agent marketplace where AI agents discover, trust, and pay each other, has released Webagent, an open source harness for standing up public-facing business agents. So, basically you give it your website, get an agent, and let it talk to other agents. Instead of writing orchestration code, a business fills in […] The post Agent-net Open Sources Webagent: A Go Harness That Turns Any Website into a Guarded AI Agent appeared first on MarkTechPost.

MarkTechPost站內正文待翻譯:Agent-net Open Sources Webagent: A Go Harness That Turns Any Website into a Guarded AI Agent

待翻譯:ArtSociety: Multi-Agent Multimodal Collaboration for Art Emotion Understanding

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.13240v1 Announce Type: new Abstract: The AffectiveArt Multidimensional Art Emotion Understanding task asks to jointly predict an artwork's fine-grained emotion (12 classes, 1549:1 head-to-tail ratio), binary valence/arousal, and five attribute-grounded descriptions -- sub-tasks that exhibit strong empirical trade-offs, so the single-model solutions we tried do not jointly optimize all of them well. We present ArtSociety, a multi-agent framework that assembles heterogeneous multimodal experts -- a DINOv2-Giant vision agent (A1), a scene-grounded CoT fine-tuned MLLM (A2), and three closed-source reasoning agents (A3-A5) -- and coordinates them with two training-free controllers: (i) a rare-class-aware voting arbiter that lowers the agreement threshold…

arXiv Computer Vision站內正文待翻譯:ArtSociety: Multi-Agent Multimodal Collaboration for Art Emotion Understanding

待翻譯:OrchSLM: Probing the Dynamics of Small Language Model Orchestration

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.13470v1 Announce Type: new Abstract: Although large language models (LLMs) have demonstrated remarkable capabilities, their reliance on cloud-scale infrastructure poses fundamental challenges for deployment in agentic pipelines, including latency, privacy, connectivity, and substantial computational cost. Small language models (SLMs) offer a compelling alternative: recent studies suggest that many repetitive and narrowly scoped subtasks in agentic workloads may be better served by specialized SLMs than by monolithic LLMs. However, the limited capacity and context windows of SLMs can constrain long-horizon reasoning and interaction-heavy orchestration strategies such as iterative verification and debate. This motivates a complementary, non-interactive…

arXiv AI站內正文待翻譯:OrchSLM: Probing the Dynamics of Small Language Model Orchestration

待翻譯:Toward Self-Adaptive Physical AI: Can LLM Agents Manage Long-Horizon Physical Tasks?

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.13436v1 Announce Type: new Abstract: Large Language Model (LLM) agents offer a promising path toward autonomously managing long-term physical tasks without human intervention. However, physical tasks require agents to continuously observe the environment, make consequential actions, and remain effective as the environment changes. Existing approaches either require substantial data and retraining, or primarily focus on agents operating in the virtual world. In this work, we explore the feasibility of building a self-adaptive physical AI agent that manages long-term physical tasks in a zero-shot manner and adapts to environmental changes without human intervention. We design a multi-agent framework that integrates planning, tool calling, observation,…

arXiv AI站內正文待翻譯:Toward Self-Adaptive Physical AI: Can LLM Agents Manage Long-Horizon Physical Tasks?

待翻譯:Agent Harness vs Agent Framework vs MCP: Which Layer Owns the Loop, State, Tools, Permissions, and Recovery

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:A practitioner's map of the 3 layers in a modern agent stack, with verified sources and an overlap analysis. The post Agent Harness vs Agent Framework vs MCP: Which Layer Owns the Loop, State, Tools, Permissions, and Recovery appeared first on MarkTechPost.

MarkTechPost站內正文待翻譯:Agent Harness vs Agent Framework vs MCP: Which Layer Owns the Loop, State, Tools, Permissions, and Recovery

待翻譯:Enterprise Analytics Beyond Dashboards: Intelligent Data Orchestration with LLMs

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:In 17 years of building enterprise data platforms, I’ve watched every organization eventually ask the same question: “Can I ask one question and get one answer across everything my company knows?” A finance analyst wants actual revenue from the warehouse, pipeline data from the CRM, commentary from planning documents, and market signals from external providers. […]

O'Reilly AI & ML Radar站內正文待翻譯:Enterprise Analytics Beyond Dashboards: Intelligent Data Orchestration with LLMs

待翻譯:How We Built LangChain’s Paid Media Agent

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:How LangChain built a paid media agent to analyze campaign performance, optimize ads, propose changes, and turn marketing data into action.

LangChain Blog站內正文待翻譯:How We Built LangChain’s Paid Media Agent

待翻譯:Context Engineering Inside the Harness: 4 Mechanisms That Beat Context Overflow and Goal Loss on Long-Horizon Tasks

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:A shallow agent is an LLM calling tools in a loop, and on long tasks it fails in 2 ways: context overflow and goal loss. This article opens the harness layer that fixes both, with the actual thresholds shipped by LangChain Deep Agents, Claude Code, Manus, OpenAI Codex and Amazon Bedrock AgentCore, plus an interactive simulator that shows a 200K window filling up. The post Context Engineering Inside the Harness: 4 Mechanisms That Beat Context Overflow and Goal Loss on Long-Horizon Tasks appeared first on MarkTechPost.

MarkTechPost站內正文待翻譯:Context Engineering Inside the Harness: 4 Mechanisms That Beat Context Overflow and Goal Loss on Long-Horizon Tasks

待翻譯:Automating Quadratic Unconstrained Binary Optimization (QUBO) Formulation Generation from Natural Language

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10629v1 Announce Type: new Abstract: Quadratic Unconstrained Binary Optimization (QUBO) is a central formulation for combinatorial optimization and has gained increasing attention due to its compatibility with quantum, hybrid quantum-classical, and quantum-inspired solvers. However, translating natural-language problem descriptions into correct QUBO formulations remains difficult, requiring the identification of binary variables, constraints, objective functions, penalty terms, and suitable penalty weights. This process is time-consuming and often demands substantial domain expertise. To address this challenge, we propose an end-to-end multi-agent framework that automatically generates QUBO formulations from natural-language problem descriptions, sup…

arXiv AI站內正文待翻譯:Automating Quadratic Unconstrained Binary Optimization (QUBO) Formulation Generation from Natural Language

待翻譯:Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Sakana AI has released Fugu Max and Fugu Ultra v2, 2 models built on the same learned orchestration architecture. Fugu Max routes tasks to lean open and specialized models, including NVIDIA Nemotron, at $2/$6 per 1M tokens. Fugu Ultra v2 targets peak capability, scoring 48.3 on Chartography and 74.3 on DeepSWE. The post Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration appeared first on MarkTechPost.

MarkTechPost站內正文待翻譯:Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration

待翻譯:How Credit Genie keeps codebase docs fresh with OpenWiki

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:See how Credit Genie uses OpenWiki to automate repo documentation, reduce tribal knowledge, and give engineers and coding agents searchable codebase context.

LangChain Blog站內正文待翻譯:How Credit Genie keeps codebase docs fresh with OpenWiki

待翻譯:Agentic AI-enabled Semantic Commissioning of a Cognitive Digital Twin for Reconfigurable Manufacturing

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09503v1 Announce Type: new Abstract: Rapid bespoke commissioning of the Cognitive Digital Twin (CDT) is a major challenge in reconfigurable manufacturing. Traditional digital twin (DT) construction methods primarily focus on geometric reconstruction, often neglecting the deep semantic integration and functional interoperability necessary for autonomous reasoning. This paper proposes an agent-based, AI-driven workflow to automate end-to-end CDT debugging. The system utilises LangGraph as a multi-agent orchestration engine to achieve dual-path synthesis: the semantic path extracts technical specifications from unstructured documents using Retrieval Augmented Generation (RAG), while the functional path autonomously discovers and binds to real-time indus…

arXiv Robotics站內正文待翻譯:Agentic AI-enabled Semantic Commissioning of a Cognitive Digital Twin for Reconfigurable Manufacturing

待翻譯:Introducing the Agents API

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Build and launch cloud agents with the Agents API, a managed service powered by the Codex harness for orchestration, long-running sessions, and tool use.

OpenAI News站內正文待翻譯:Introducing the Agents API

待翻譯:Connections: managed credentials and per-caller identity for Managed Deep Agents

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Learn how Connections in Managed Deep Agents securely manage credentials, support per-user OAuth, and let agents act with each caller’s identity.

LangChain Blog站內正文待翻譯:Connections: managed credentials and per-caller identity for Managed Deep Agents

多智能體框架中的上下文組織

deepagents 為子代理引入“上下文模式”(isolated / fork),讓主管代理既可以委派任務,又可以決定子代理繼承多少上下文。fork 模式繼承主管的會話歷史,可複用 prompt caching、減少重複工作;isolated 模式則讓子代理在全新上下文中獨立完成任務。文章還通過 worker、verifier、researcher、memory 四類子代理説明如何選擇。

LangChain Blog站內正文多智能體框架中的上下文組織

GitHub 推出 Project HydraFusion:在 Copilot CLI 中按編程任務動態編排多模型運行時工作流

Project HydraFusion 是 GitHub 發佈的研究預覽,它不再將模型選擇視為一次性設置,而是針對每個請求構建並動態編排一個執行工作流,可在不同提供商的模型間進行起草、批判、升級等操作。目前僅在 GitHub Copilot CLI 中以研究預覽形式提供,並按實際調用的各模型標準令牌費率計費。

MarkTechPost站內正文GitHub 推出 Project HydraFusion:在 Copilot CLI 中按編程任務動態編排多模型運行時工作流

Project HydraFusion:通過多模型編排實現前沿品質

GitHub 發佈 Project HydraFusion 研究預覽,通過自動在多個提供商的模型之間編排“單模型、級聯、批判”等工作流,為編碼任務平衡質量、成本和延遲。離線評估顯示,在保持前沿質量的同時可將估算工作流成本降低約 36%–67%。

GitHub AI & ML站內正文Project HydraFusion:通過多模型編排實現前沿品質

Perplexity 在 Mac 上發佈混合計算:雲代理編排至本地模型,並由設備端門控

Perplexity 在 Mac 上推出混合計算,將任務在雲端前沿模型與本地小型模型之間拆分,並通過設備端隱私門控在將敏感數據發送到雲端前進行保護。該功能已向 Pro、Max 和企業用户開放,支持 Apple 芯片 Mac。Perplexity 還開源了用於門控的分類器 PII-Tracer,並在基準測試中表現出色。

MarkTechPost站內正文Perplexity 在 Mac 上發佈混合計算:雲代理編排至本地模型,並由設備端門控

待翻譯:Multi-Agent Self-Improving Reinforcement Learning for Video Reasoning

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.28675v1 Announce Type: new Abstract: Video reasoning tasks such as grounded video question answering and temporal grounding require selecting temporal evidence that supports the query. In many current training setups, temporal supervision is applied through local objectives such as boundary regression or span generation, while verification is used mainly to rerank candidate segments at inference time. We study whether a frozen verifier can also guide training. Our multi-agent framework couples a trainable \emph{Grounder} with a frozen \emph{Verifier}: the Grounder samples candidate trajectories and evidence segments, the Verifier assigns query-conditioned segment scores, a group-relative policy-gradient objective favors trajectories that outperform t…

arXiv Computer Vision站內正文待翻譯:Multi-Agent Self-Improving Reinforcement Learning for Video Reasoning

待翻譯:The Race between Agentic AI Capabilities and Data Quality Control in Online Surveys

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.28597v1 Announce Type: new Abstract: Online surveys are a foundational data collection instrument in a variety of fields, with attention checks serving as critical guardians of response quality. However, the rapid emergence of agentic AI (goal directed systems powered by a large language model (LLM) brain and/or a multimodal processing unit with tool-augmented capabilities) raises new questions about the robustness of these safeguards. We investigate how well agentic AI architectures can complete web-based surveys and pass standard attention checks. We evaluate a single-agent architecture capable of multimodal input processing and tool-based web interaction on a controlled survey sandbox. We analyze the problem from two perspectives. From an attack p…

arXiv AI站內正文待翻譯:The Race between Agentic AI Capabilities and Data Quality Control in Online Surveys

待翻譯:90 days of attacks on AI infrastructure

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Wiz PricingGet a demo Get a demo Wiz Threat Research operates honeypots across AI and ML services including LiteLLM, Flowise, LangChain, Langflow, ChromaDB, Ollama, and others. Over 90 days of telemetry, we observed sus…

Hacker News AI站內正文待翻譯:90 days of attacks on AI infrastructure

待翻譯:Show HN: ShevtoneAudio Orchestrator – Turning MIDI into Full Orchestration

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Professional MIDI orchestration plugin Turn your MIDI into a full orchestra. ShevtoneAudio Orchestrator is a professional local MIDI orchestration tool built for composers working in film, trailer, television, games and…

Hacker News AI站內正文待翻譯:Show HN: ShevtoneAudio Orchestrator – Turning MIDI into Full Orchestration

待翻譯:Procedura: Agentic 3D Modeling with Procedural Control

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.26238v1 Announce Type: new Abstract: Native 3D generators now recover impressive mesh geometry from a single image. However, a dense mesh stays soft where a machined object should be sharp, it carries no part decomposition, and it exposes no parameter a user could edit. To address this, we explore the paradigm of 3D shape as code, leveraging and scaling the coding ability of an LLM for 3D modeling. We introduce Procedura, a novel 3D modeling agent framework that writes an object as a procedural assembly, a parametric program whose named parts are joined by typed, machine-checkable mates. From a text prompt, the agent plans the object as an assembly graph and writes the program part by part, solving each placement from the mated frames rather than gue…

arXiv Computer Vision站內正文待翻譯:Procedura: Agentic 3D Modeling with Procedural Control

更多增長標籤

Agent 框架 AI News | AI News Hub