Skip to content
AI News HubLIVE

Agent Frameworks updates

Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration

Sakana AI has released Fugu Max and Fugu Ultra v2, 2 models built on the same learned orchestration architecture. Fugu Max routes tasks to lean open and specialized models, including NVIDIA Nemotron, at $2/$6 per 1M tokens. Fugu Ultra v2 targets peak capability, scoring 48.3 on Chartography and 74.3 on DeepSWE. The post Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration appeared first on MarkTechPost.

MarkTechPostIn-site articleSakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration

How Credit Genie keeps codebase docs fresh with OpenWiki

See how Credit Genie uses OpenWiki to automate repo documentation, reduce tribal knowledge, and give engineers and coding agents searchable codebase context.

LangChain BlogIn-site articleHow Credit Genie keeps codebase docs fresh with OpenWiki

Agentic AI-enabled Semantic Commissioning of a Cognitive Digital Twin for Reconfigurable Manufacturing

arXiv:2609.09503v1 Announce Type: new Abstract: Rapid bespoke commissioning of the Cognitive Digital Twin (CDT) is a major challenge in reconfigurable manufacturing. Traditional digital twin (DT) construction methods primarily focus on geometric reconstruction, often neglecting the deep semantic integration and functional interoperability necessary for autonomous reasoning. This paper proposes an agent-based, AI-driven workflow to automate end-to-end CDT debugging. The system utilises LangGraph as a multi-agent orchestration engine to achieve dual-path synthesis: the semantic path extracts technical specifications from unstructured documents using Retrieval Augmented Generation (RAG), while the functional path autonomously discovers and binds to real-time industrial telemetry data using M…

arXiv RoboticsIn-site articleAgentic AI-enabled Semantic Commissioning of a Cognitive Digital Twin for Reconfigurable Manufacturing

Introducing the Agents API

Build and launch cloud agents with the Agents API, a managed service powered by the Codex harness for orchestration, long-running sessions, and tool use.

OpenAI NewsIn-site articleIntroducing the Agents API

Organizing Context in a Multi-Agent Harness

deepagents introduces context modes that decide whether subagents inherit a supervisor’s conversation (fork) or start clean (isolated). Fork can be faster and cheaper by reusing prompt caching, while isolated is ideal for independent review or parallel research. The post maps worker, verifier, researcher, and memory agents to the right mode.

LangChain BlogIn-site articleOrganizing Context in a Multi-Agent Harness

How HPE Zerto built an agentic troubleshooting system with Amazon Bedrock

HPE Zerto partnered with AWS to build an Amazon Bedrock-powered agentic troubleshooting system that runs on-premises as part of the Zerto product. Using the Strands Agents framework, Guardrails, Knowledge Bases, and multi-agent orchestration, it grounds AI answers in live disaster-recovery data. Since launch in Q2 2026, more than 20% of HPE Zerto customers have adopted it, and supported workflows saw a 10% reduction in support cases.

AWS Machine Learning BlogIn-site articleHow HPE Zerto built an agentic troubleshooting system with Amazon Bedrock

GitHub Introduces Project HydraFusion: Runtime Multi-Model Orchestration That Builds a Workflow Per Coding Task in Copilot CLI

Project HydraFusion is a research preview from GitHub that reframes model selection as an optimization problem. Instead of picking one model, it dynamically chooses a workflow for each request — drafting, critiquing, or escalating across models from multiple providers. Currently available in GitHub Copilot CLI only, it bills per token at each model's standard rate.

MarkTechPostIn-site articleGitHub Introduces Project HydraFusion: Runtime Multi-Model Orchestration That Builds a Workflow Per Coding Task in Copilot CLI

Project HydraFusion: Frontier quality via multi-model orchestration

GitHub is introducing Project HydraFusion, a research preview in GitHub Copilot that automatically orchestrates models from multiple providers—using single, cascade, or critique workflows—to balance quality, cost, and latency for each coding task. In offline agentic benchmark evaluations, HydraFusion matched or exceeded the Claude Opus 5 baseline on verified quality while cutting estimated workflow cost by 36–67%.

GitHub AI & MLIn-site articleProject HydraFusion: Frontier quality via multi-model orchestration

Antigravity Teamwork Multi-Agent Framework Tackles Math and Engineering Challenges

Antigravity's Teamwork multi-agent framework has achieved breakthroughs in mathematics, hardware simulation, and open-source optimization, solving seven open math problems and creating a RISC-V simulator that boots an OS from scratch. This post details Teamwork's updates, how it works, and its various collaboration patterns.

Hacker News AIIn-site articleAntigravity Teamwork Multi-Agent Framework Tackles Math and Engineering Challenges

Multi-Agent Self-Improving Reinforcement Learning for Video Reasoning

arXiv:2608.28675v1 Announce Type: new Abstract: Video reasoning tasks such as grounded video question answering and temporal grounding require selecting temporal evidence that supports the query. In many current training setups, temporal supervision is applied through local objectives such as boundary regression or span generation, while verification is used mainly to rerank candidate segments at inference time. We study whether a frozen verifier can also guide training. Our multi-agent framework couples a trainable \emph{Grounder} with a frozen \emph{Verifier}: the Grounder samples candidate trajectories and evidence segments, the Verifier assigns query-conditioned segment scores, a group-relative policy-gradient objective favors trajectories that outperform their within-input peers, and…

arXiv Computer VisionIn-site articleMulti-Agent Self-Improving Reinforcement Learning for Video Reasoning

The Race between Agentic AI Capabilities and Data Quality Control in Online Surveys

arXiv:2608.28597v1 Announce Type: new Abstract: Online surveys are a foundational data collection instrument in a variety of fields, with attention checks serving as critical guardians of response quality. However, the rapid emergence of agentic AI (goal directed systems powered by a large language model (LLM) brain and/or a multimodal processing unit with tool-augmented capabilities) raises new questions about the robustness of these safeguards. We investigate how well agentic AI architectures can complete web-based surveys and pass standard attention checks. We evaluate a single-agent architecture capable of multimodal input processing and tool-based web interaction on a controlled survey sandbox. We analyze the problem from two perspectives. From an attack perspective, we demonstrate h…

arXiv AIIn-site articleThe Race between Agentic AI Capabilities and Data Quality Control in Online Surveys

90 days of attacks on AI infrastructure

Wiz PricingGet a demo Get a demo Wiz Threat Research operates honeypots across AI and ML services including LiteLLM, Flowise, LangChain, Langflow, ChromaDB, Ollama, and others. Over 90 days of telemetry, we observed sus…

Hacker News AIIn-site article90 days of attacks on AI infrastructure

Show HN: ShevtoneAudio Orchestrator – Turning MIDI into Full Orchestration

Professional MIDI orchestration plugin Turn your MIDI into a full orchestra. ShevtoneAudio Orchestrator is a professional local MIDI orchestration tool built for composers working in film, trailer, television, games and…

Hacker News AIIn-site articleShow HN: ShevtoneAudio Orchestrator – Turning MIDI into Full Orchestration

Procedura: Agentic 3D Modeling with Procedural Control

arXiv:2608.26238v1 Announce Type: new Abstract: Native 3D generators now recover impressive mesh geometry from a single image. However, a dense mesh stays soft where a machined object should be sharp, it carries no part decomposition, and it exposes no parameter a user could edit. To address this, we explore the paradigm of 3D shape as code, leveraging and scaling the coding ability of an LLM for 3D modeling. We introduce Procedura, a novel 3D modeling agent framework that writes an object as a procedural assembly, a parametric program whose named parts are joined by typed, machine-checkable mates. From a text prompt, the agent plans the object as an assembly graph and writes the program part by part, solving each placement from the mated frames rather than guessing it, and admitting a pa…

arXiv Computer VisionIn-site articleProcedura: Agentic 3D Modeling with Procedural Control

Pushing GPT-5.6 Luna from 0% to 56% on ARC-AGI-3 Public

Research August 27, 2026 For long-horizon agents, model capability alone does not determine system capability. Infrastructure orchestration multiplies what models can do. INT21’s SwarmOS tested this hypothesis on the AR…

Hacker News AIIn-site articlePushing GPT-5.6 Luna from 0% to 56% on ARC-AGI-3 Public

Fusing Perceptual Vision Experts with Multimodal Large Language Models for Explainable Plant Disease Diagnosis: From Benchmark Imagery to Real-World Robotic Field Validation

arXiv:2608.24934v1 Announce Type: new Abstract: Accurate field plant disease diagnosis requires reliable fusion of uncertain and conflicting perceptual evidence. We present the Hybrid Hierarchical Multi-Agent Framework (H$^{2}$MAF), combining decision-level fusion of EfficientNet-B3 and ConvNeXt-Tiny with semantic arbitration by open-weight multimodal large language models (MLLMs), Gemma 4 E4B and Qwen3.5 4B, using structured JSON evidence to generate explainable diagnoses, risk levels, treatment urgency, and financial exposure. (H$^{2}$MAF) is evaluated on 14,364 images (1,370 test images) across PlantDoc (2,922 images, 27 classes) and two non-public, continuously captured Cornell robot-acquired field datasets: Stage 2 (20 GB; 4,215 images) and Stage 4 (40 GB; 7,227 images), covering Ear…

arXiv Computer VisionIn-site articleFusing Perceptual Vision Experts with Multimodal Large Language Models for Explainable Plant Disease Diagnosis: From Benchmark Imagery to Real-World Robotic Field Validation

Evaluate any agent framework with Amazon Bedrock AgentCore Evaluations

Amazon Bedrock AgentCore Evaluations decouples agent evaluation from the framework you build on. As long as your agent emits OpenTelemetry telemetry, the service can score it, whether you use LangGraph, LlamaIndex, the OpenAI Agents SDK, Google ADK, the Claude Agent SDK, or Strands Agents. This post explains how the framework-agnostic contract works.

AWS Machine Learning BlogIn-site articleEvaluate any agent framework with Amazon Bedrock AgentCore Evaluations

Best TypeScript AI Agent Frameworks for Next.js

You’ve shipped a working chat feature. Now comes the hard part: figuring out which framework actually fits your production architecture. The tooling landscape has fractured, and comparing AI agent frameworks in a vacuum…

Hacker News AIIn-site articleBest TypeScript AI Agent Frameworks for Next.js

LangChain State of AI 2024 Report

Dive into LangSmith product usage patterns that show how the AI ecosystem and the way people are building LLM apps is evolving.

LangChain BlogIn-site articleLangChain State of AI 2024 Report

LangChain's Second Birthday

Reflections on how LangChain has evolved — including our products, ecosystem, and community — over the past two years, and where we're headed next.

LangChain BlogIn-site articleLangChain's Second Birthday

How Podium optimized agent behavior and reduced engineering intervention by 90% with LangSmith

See how Podium tests across the lifecycle development of their AI employee agent, using LangSmith for dataset curation and finetuning. They improved agent F1 response quality to 98% and reduced the need for engineering intervention by 90%.

LangChain BlogIn-site articleHow Podium optimized agent behavior and reduced engineering intervention by 90% with LangSmith

AI Agent Latency 101: How do I speed up my AI agent?

Learn proven strategies to speed up your AI agent: reduce latency, optimize LLM calls, enable parallelism, and improve UX. Expert tips from LangChain.

LangChain BlogIn-site articleAI Agent Latency 101: How do I speed up my AI agent?

Connect Amazon Bedrock AgentCore to cross-account knowledge bases

Learn how Amazon Bedrock AgentCore agents in one account can generate answers from an Amazon Bedrock knowledge base backed by Amazon Redshift Serverless in another account, without copying source data. This post covers the architecture, security boundary, and two orchestration models: a code-based Strands agent and a declarative AgentCore harness.

AWS Machine Learning BlogIn-site articleConnect Amazon Bedrock AgentCore to cross-account knowledge bases

Evaluating OpenWiki with WikiBench

We built WikiBench to test whether generated wikis help coding agents. Pairing a wiki with source code scored higher than source alone, at lower cost.

LangChain BlogIn-site articleEvaluating OpenWiki with WikiBench

OpenAI's Bet on a Cognitive Architecture

Why LangChain believes in open, customizable cognitive architectures over closed systems. Build reliable LLM agents with OpenGPTs and LangSmith.

LangChain BlogIn-site articleOpenAI's Bet on a Cognitive Architecture

Orchestration is the new challenge for CX in the age of AI agents

Presented by Tata Communications Enterprises are deploying AI agents, voice AI, and automation across messaging, voice, and digital channels faster than the architecture meant to support it. Most of that deployment has involved attaching conversational AI to legacy systems never built for it, says Gaurav Anand, global head of the Customer Interaction Suite at Tata Communications. "In the rush to deploy AI, organizations have largely bolted conversational AI onto legacy systems," Anand says. "As a result, while many enterprises have adopted digital tools, very few have platforms that are truly integrated, scaled, and capable of seamless orchestration." That gap creates a heavy cognitive load for human agents who must piece together context across disjointed tools to understand what an AI s…

VentureBeat AIIn-site articleOrchestration is the new challenge for CX in the age of AI agents

Qdrant x LangChain: Endgame Performance

Qdrant and LangChain deliver production-ready RAG performance with async support, optimized resource usage, and scalable vector search for LLM apps.

LangChain BlogIn-site articleQdrant x LangChain: Endgame Performance

Data-Driven Characters

Data-driven-characters is a repo for creating, debugging, and interacting your own chatbots conditioned on your own story corpora.

LangChain BlogIn-site articleData-Driven Characters

Eden AI x LangChain: Harnessing LLMs, Embeddings, and AI

Access multiple LLMs, embeddings, and AI tools through Eden AI's LangChain integration. Unified API for text generation, OCR, speech-to-text, and more.

LangChain BlogIn-site articleEden AI x LangChain: Harnessing LLMs, Embeddings, and AI

LangChain State of AI 2023

Discover how developers build LLM applications in 2023. Insights on popular models, vectorstores, retrieval strategies, and testing methods from LangSmith.

LangChain BlogIn-site articleLangChain State of AI 2023

How to design an Agent for Production

Build production-ready AI agents with LangChain. Technical guide covering OpenAI functions, tools, prompts, and architecture for Cal.ai's scheduling assistant.

LangChain BlogIn-site articleHow to design an Agent for Production

Auto-Evaluator Opportunities

Auto-evaluate LLM question-answer chains with LangChain's free tool. Generate test sets, grade answers, and optimize chain performance.

LangChain BlogIn-site articleAuto-Evaluator Opportunities

Applying OpenAI's RAG Strategies

Implement OpenAI's proven RAG strategies with LangChain. Explore query transformations, routing, post-processing, and evaluation methods for optimal retrieval.

LangChain BlogIn-site articleApplying OpenAI's RAG Strategies

Autonomous Agents & Agent Simulations

Explore how LangChain implements autonomous agents like AutoGPT and BabyAGI. Learn about planning techniques, memory systems, and agent simulations.

LangChain BlogIn-site articleAutonomous Agents & Agent Simulations

LLM Agents Perform Controlled Experiments Using Simulation Models

arXiv:2608.23622v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong capabilities in reasoning, planning, and tool use, but many scientific and engineering tasks require more than plausible text and code generation. They require understanding how a system responds to intervention, which in practice depends on controlled experimentation. In this work, we propose a multi-agent framework that enables LLM agents to conduct controlled experiments with scientific simulation models for pharmaceutical process design. Given a user query and a baseline configuration, the system constructs a structured task representation, designs experiments, executes comparative simulation, interprets the resulting outcomes, and synthesizes evidence-based recommendations for process param…

arXiv AIIn-site articleLLM Agents Perform Controlled Experiments Using Simulation Models

Using LangSmith to Support Fine-tuning

Learn how to fine-tune and evaluate LLMs with LangSmith for dataset management. Complete guide covers LLaMA2 and GPT-3.5 fine-tuning with practical examples.

LangChain BlogIn-site articleUsing LangSmith to Support Fine-tuning

Benchmarking Question/Answering Over CSV Data

Build better Q&A systems for CSV data using LangChain agents, retrieval, and LLM evaluation. Includes benchmarks, debugging insights, and open-source code.

LangChain BlogIn-site articleBenchmarking Question/Answering Over CSV Data

Announcing our $10M seed round led by Benchmark

LangChain secures $10M seed round from Benchmark to empower developers building AI apps with our open-source framework for data-aware, agentic LLMs.

LangChain BlogIn-site articleAnnouncing our $10M seed round led by Benchmark

LangServe Playground and Configurability

Deploy LangChain apps with LangServe's playground UI and configurable parameters. Experiment with models, share with teams, stream in real-time.

LangChain BlogIn-site articleLangServe Playground and Configurability

Retrieval

Build better AI apps with flexible retrieval methods in LangChain. Use any retriever—from semantic to hybrid—to create personalized ChatGPT for your data.

LangChain BlogIn-site articleRetrieval

More growth tags

Agent Frameworks AI News | AI News Hub