AI News HubLIVE

来源分布

  • LangChain Blog41
  • Hacker News AI3
  • MarkTechPost3
  • AWS Machine Learning Blog2
  • Analytics Vidhya1

主题分布

  • Agent50
  • 研究27
  • 模型15
  • 政策7
  • 芯片6
  • 创业融资3

日期线

  • 2026-08-2622
  • 2026-08-253
  • 2026-08-052
  • 2026-08-062
  • 2026-08-112
  • 2026-08-122
  • 2026-08-132
  • 2026-08-182

最新动态

待翻译:Using LangSmith to Support Fine-tuning

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Learn how to fine-tune and evaluate LLMs with LangSmith for dataset management. Complete guide covers LLaMA2 and GPT-3.5 fine-tuning with practical examples.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Learn how to fine-tune and evaluate LLMs with LangSmith for dataset management. Complete guide covers LLaMA2 and GPT-3.5 fine-tuning with practical examples.
站内正文

待翻译:Benchmarking Question/Answering Over CSV Data

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Build better Q&A systems for CSV data using LangChain agents, retrieval, and LLM evaluation. Includes benchmarks, debugging insights, and open-source code.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Build better Q&A systems for CSV data using LangChain agents, retrieval, and LLM evaluation. Includes benchmarks, debugging insights, and open-source code.
站内正文

待翻译:Timescale Vector x LangChain: Making PostgreSQL A Better Vector Database for AI Applications

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Build faster AI apps with Timescale Vector for LangChain. Get 243% faster similarity search, time-based RAG, and PostgreSQL simplicity. Free 90-day trial.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Build faster AI apps with Timescale Vector for LangChain. Get 243% faster similarity search, time-based RAG, and PostgreSQL simplicity. Free 90-day trial.
站内正文

待翻译:Announcing our $10M seed round led by Benchmark

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:LangChain secures $10M seed round from Benchmark to empower developers building AI apps with our open-source framework for data-aware, agentic LLMs.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • LangChain secures $10M seed round from Benchmark to empower developers building AI apps with our open-source framework for data-aware, agentic LLMs.
站内正文

待翻译:Making Data Ingestion Production Ready: a LangChain-Powered Airbyte Destination

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Scale retrieval apps to production with LangChain's Airbyte integration. Automate data ingestion with scheduling, text splitting, and 50+ embeddings.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Scale retrieval apps to production with LangChain's Airbyte integration. Automate data ingestion with scheduling, text splitting, and 50+ embeddings.
站内正文

待翻译:LangServe Playground and Configurability

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Deploy LangChain apps with LangServe's playground UI and configurable parameters. Experiment with models, share with teams, stream in real-time.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Deploy LangChain apps with LangServe's playground UI and configurable parameters. Experiment with models, share with teams, stream in real-time.
站内正文

待翻译:Retrieval

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Build better AI apps with flexible retrieval methods in LangChain. Use any retriever—from semantic to hybrid—to create personalized ChatGPT for your data.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Build better AI apps with flexible retrieval methods in LangChain. Use any retriever—from semantic to hybrid—to create personalized ChatGPT for your data.
站内正文

待翻译:Introducing Pytest and Vitest integrations for LangSmith Evaluations

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Introducing a new way to run evals using LangSmith’s Pytest and Vitest/Jest integrations.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Introducing a new way to run evals using LangSmith’s Pytest and Vitest/Jest integrations.
站内正文

待翻译:Cube x LangChain: Building AI experiences with LLMs and the semantic layer

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Build AI-powered data experiences with Cube's semantic layer and LangChain. Prevent hallucinations, query in natural language, create conversational interfaces.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Build AI-powered data experiences with Cube's semantic layer and LangChain. Prevent hallucinations, query in natural language, create conversational interfaces.
站内正文

待翻译:Role Based Access Control (RBAC) for LangSmith

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:LangSmith's Role Based Access Control (RBAC) helps enterprises manage resource access with custom roles and API keys.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • LangSmith's Role Based Access Control (RBAC) helps enterprises manage resource access with custom roles and API keys.
站内正文

待翻译:Automating Web Research

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Automate web research with LangChain's retriever. Run parallel searches, scrape pages, and synthesize information with LLMs—locally or in the cloud.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Automate web research with LangChain's retriever. Run parallel searches, scrape pages, and synthesize information with LLMs—locally or in the cloud.
站内正文

待翻译:Plan-and-Execute Agents

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Build reliable AI agents with Plan-and-Execute framework. Separate planning from execution for complex tasks with fewer errors.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Build reliable AI agents with Plan-and-Execute framework. Separate planning from execution for complex tasks with fewer errors.
站内正文

待翻译:Workspaces in LangSmith for improved collaboration and organization

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Workspaces in LangSmith lets enterprises separate resources between different teams, business units, or deployment environments.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Workspaces in LangSmith lets enterprises separate resources between different teams, business units, or deployment environments.
站内正文

待翻译:LLMs and SQL

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Query SQL databases using natural language with LLMs. Learn techniques to reduce hallucinations and build reliable text-to-SQL solutions with LangChain.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Query SQL databases using natural language with LLMs. Learn techniques to reduce hallucinations and build reliable text-to-SQL solutions with LangChain.
站内正文

待翻译:Multi-modal RAG on slide decks

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Build multi-modal RAG apps for slide decks using GPT-4V. Compare approaches, evaluate with benchmarks, and deploy with LangChain templates for visual Q&A.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Build multi-modal RAG apps for slide decks using GPT-4V. Compare approaches, evaluate with benchmarks, and deploy with LangChain templates for visual Q&A.
站内正文

待翻译:The rise of "context engineering"

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Learn context engineering: building dynamic systems that provide LLMs the right information, tools, and format to reliably accomplish tasks.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Learn context engineering: building dynamic systems that provide LLMs the right information, tools, and format to reliably accomplish tasks.
站内正文

待翻译:Debugging Deep Agents with LangSmith

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Debug deep agents with LangSmith's tracing and analysis. Analyze complex traces, optimize prompts with Polly, and improve performance.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Debug deep agents with LangSmith's tracing and analysis. Analyze complex traces, optimize prompts with Polly, and improve performance.
站内正文

待翻译:Structured Tools

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Build powerful LangChain agents with Structured Tools. Accept multiple inputs, create complex tool schemas, and unlock advanced AI agent capabilities.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Build powerful LangChain agents with Structured Tools. Accept multiple inputs, create complex tool schemas, and unlock advanced AI agent capabilities.
站内正文

待翻译:Multi-Vector Retriever for RAG on tables, text, and images

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Learn how to implement multi-vector retriever for RAG across tables, text, and images. Explore cookbooks for semi-structured and multi-modal data retrieval.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Learn how to implement multi-vector retriever for RAG across tables, text, and images. Explore cookbooks for semi-structured and multi-modal data retrieval.
站内正文

待翻译:LangSmith Engine Improves Agent Issue Detection by 2x

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:LangSmith Engine now detects agent issues over 2x better, proposes stronger fixes, supports Slack and Linear workflows, and is available for self-hosted deployments.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • LangSmith Engine now detects agent issues over 2x better, proposes stronger fixes, supports Slack and Linear workflows, and is available for self-hosted deployments.
站内正文

待翻译:Using skills with Deep Agents

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Learn how to use agent skills with Deep Agents CLI to build token-efficient AI agents. Discover, load, and execute skills dynamically.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Learn how to use agent skills with Deep Agents CLI to build token-efficient AI agents. Discover, load, and execute skills dynamically.
站内正文

待翻译:Building Production Agentic AI at IBM: Architecture, Decisions, and Lessons

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Building Production Agentic AI at IBM: Architecture, Decisions, and What We Learned TL;DR — IBM’s Technology Lifecycle Services built a multi-agent system from scratch — the agents themselves in Python with LangGraph. I…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Building Production Agentic AI at IBM: Architecture, Decisions, and What We Learned TL;DR — IBM’s Technology Lifecycle Services built a multi-agent system from scratch — the agent…
站内正文

待翻译:Building Self-Correcting Memory in OpenWiki

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Learn how OpenWiki uses evidence-backed claims to detect stale knowledge, reduce hallucinations, and build self-correcting memory for evolving codebases.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Learn how OpenWiki uses evidence-backed claims to detect stale knowledge, reduce hallucinations, and build self-correcting memory for evolving codebases.
站内正文

待翻译:How We Build Agent Environments & Tasks

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:How we create synthetic agent environments and tasks: a spec generation step, a spec-to-task step, and a world spec that holds shared knowledge.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • How we create synthetic agent environments and tasks: a spec generation step, a spec-to-task step, and a world spec that holds shared knowledge.
站内正文

待翻译:Toyota Scales Enterprise AI with Deep Agents and LangSmith

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:See how Toyota North America uses Deep Agents and LangSmith to run 50+ production agents, cut delivery from 6 months to 4 days, and track AI ROI.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • See how Toyota North America uses Deep Agents and LangSmith to run 50+ production agents, cut delivery from 6 months to 4 days, and track AI ROI.
站内正文

待翻译:Prism Reviewer – Multi-agent AI code reviewer built with LangGraph and LiteLLM

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:🌈 Prism Reviewer Developed by Vyoman Labs Prism Reviewer is an agentic, AI-driven multi-agent code review system developed by Vyoman Labs and orchestrated via LangGraph and LiteLLM. It acts as an autonomous gatekeeper…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • 🌈 Prism Reviewer Developed by Vyoman Labs Prism Reviewer is an agentic, AI-driven multi-agent code review system developed by Vyoman Labs and orchestrated via LangGraph and LiteL…
站内正文

待翻译:Decoding AI’s Open-Source Course Maps Three Ways to Run an Agent Loop and the Provider Economics Behind Each

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Most teams treat ‘which model’ as the important decision. The harness engineering literature keeps pointing somewhere else. In LangChain’s Terminal-Bench experiment, changing only the harness—same model throughout—moved a coding agent from roughly 30th place into the top 5. That result reframes the question. If the harness decides quality, then how you run the loop becomes […] The post Decoding AI’s Open-Source Course Maps Three Ways to Run an Agent Loop and the Provider Economics Behind Each appeared first on MarkTechPost.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Most teams treat ‘which model’ as the important decision. The harness engineering literature keeps pointing somewhere else. In LangChain’s Terminal-Bench experiment, changing only…
站内正文

待翻译:Test Agent Changes with LangSmith Preview Builds

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Preview Builds let teams test pull request branches in temporary, production-like LangSmith deployments before merging agent changes.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Preview Builds let teams test pull request branches in temporary, production-like LangSmith deployments before merging agent changes.
站内正文

待翻译:Introducing LangSmith Tuned Evaluators

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:LangSmith Tuned Evaluators attach quality feedback to production traces, starting with Perceived Error, to help teams find and fix agent mistakes.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • LangSmith Tuned Evaluators attach quality feedback to production traces, starting with Perceived Error, to help teams find and fix agent mistakes.
站内正文

待翻译:How to Add Skills in Agents using LangChain

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Ever wondered how ChatGPT, Gemini, and other chat interfaces generate PDFs, PowerPoints, and more when all they have under the hood is an LLM? The trick isn’t a smarter model. It’s something simpler: skills which are instructions an agent loads only when needed. Next, let’s explore how skills work using LangChain and how they can make […] The post How to Add Skills in Agents using LangChain appeared first on Analytics Vidhya.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Ever wondered how ChatGPT, Gemini, and other chat interfaces generate PDFs, PowerPoints, and more when all they have under the hood is an LLM? The trick isn’t a smarter model. It’…
站内正文

待翻译:AgentCore Payments middleware for LangChain agents

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Let your LangChain agents pay for APIs with deterministic session budgets. AgentCore Payments middleware signs x402 payments; LangSmith traces every one.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Let your LangChain agents pay for APIs with deterministic session budgets. AgentCore Payments middleware signs x402 payments; LangSmith traces every one.
站内正文

待翻译:Why managed agents are the next big thing in agent building

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Managed Deep Agents gives developers a managed way to build, run, and deploy Deep Agents with built-in runtime, streaming, sandboxes, evals, memory, and auth.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Managed Deep Agents gives developers a managed way to build, run, and deploy Deep Agents with built-in runtime, streaming, sandboxes, evals, memory, and auth.
站内正文

待翻译:LangSmith BYOC on AWS is generally available

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:LangSmith Bring Your Own Cloud is now generally available on AWS, giving Enterprise teams managed observability, evaluation, and deployment inside their own VPC.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • LangSmith Bring Your Own Cloud is now generally available on AWS, giving Enterprise teams managed observability, evaluation, and deployment inside their own VPC.
站内正文

待翻译:How to Debug AI Agents

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Learn how agent observability enables effective evaluation of AI agents. Understand tracing, debugging reasoning, and performance insights to iterate and improve agent behavior.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Learn how agent observability enables effective evaluation of AI agents. Understand tracing, debugging reasoning, and performance insights to iterate and improve agent behavior.
站内正文

待翻译:Building Monday Com Sidekick Why Capable Agents Need More Than Just Tools

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Building monday.com Sidekick: why capable agents need more than just tools August 11, 2026 14 min Go back to blog Create agents This is a guest post from Omri Bruchim, AI Engineering Group Lead, monday.com In early test…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Building monday.com Sidekick: why capable agents need more than just tools August 11, 2026 14 min Go back to blog Create agents This is a guest post from Omri Bruchim, AI Engineer…
站内正文

待翻译:How many of your agent's calls actually need a frontier model?

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:We benchmarked NVIDIA NeMo Switchyard on 145 agent tasks. Only 7% of turns needed a frontier model, and routing cut cost 74% for six points of accuracy.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • We benchmarked NVIDIA NeMo Switchyard on 145 agent tasks. Only 7% of turns needed a frontier model, and routing cut cost 74% for six points of accuracy.
站内正文

待翻译:How nOps shipped FinOps agents 75% faster with Amazon Bedrock AgentCore

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:nOps rebuilt its Clara FinOps AI agent on Amazon Bedrock AgentCore, replacing a self-managed Amazon EKS stack running LangChain and LangGraph. The move cut time-to-production by 75% (from 10-12 months to 4 months), improved response quality, and reduced operational overhead while keeping analytics governed through Databricks Lakehouse Metric Views.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • nOps rebuilt its Clara FinOps AI agent on Amazon Bedrock AgentCore, replacing a self-managed Amazon EKS stack running LangChain and LangGraph. The move cut time-to-production by 7…
站内正文

待翻译:Top LLM Observability and Evaluation Platforms in 2026: Langfuse, LangSmith, Braintrust, Arize, and More Compared

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:A verified 2026 comparison of LLM observability platforms covering tracing depth, evaluation capability, production monitoring, and pricing. The post Top LLM Observability and Evaluation Platforms in 2026: Langfuse, LangSmith, Braintrust, Arize, and More Compared appeared first on MarkTechPost.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • A verified 2026 comparison of LLM observability platforms covering tracing depth, evaluation capability, production monitoring, and pricing. The post Top LLM Observability and Eva…
站内正文

待翻译:Show HN: Aidress – LangChain integration for cross-agent discovery and trust

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Tools are utilities designed to be called by a model: their inputs are designed to be generated by models, and their outputs are designed to be passed back to models. A toolkit is a collection of tools meant to be used…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Tools are utilities designed to be called by a model: their inputs are designed to be generated by models, and their outputs are designed to be passed back to models. A toolkit is…
站内正文

待翻译:Managed Deep Agents is now in public beta

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Deploy Deep Agents to a managed LangSmith runtime with durable execution, memory, sandboxes, channels, evals, and production-ready infrastructure.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Deploy Deep Agents to a managed LangSmith runtime with durable execution, memory, sandboxes, channels, evals, and production-ready infrastructure.
站内正文

待翻译:Deep Agents vs LangChain vs LangGraph

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Deep Agents, LangChain, and LangGraph each offer distinct approaches to building agents. In this post, we cover the key distinctions between our open source frameworks and when you should reach for each one.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Deep Agents, LangChain, and LangGraph each offer distinct approaches to building agents. In this post, we cover the key distinctions between our open source frameworks and when yo…
站内正文

待翻译:How LendingTree built a multi-agent mortgage assistant on Amazon Bedrock

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Learn how LendingTree built a production multi-agent mortgage assistant on Amazon Bedrock. Three coordinated agents use LangGraph, the Model Context Protocol, and Amazon Nova models with built-in guardrails to deliver 24/7 personalized mortgage guidance while meeting strict financial-services compliance.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Learn how LendingTree built a production multi-agent mortgage assistant on Amazon Bedrock. Three coordinated agents use LangGraph, the Model Context Protocol, and Amazon Nova mode…
站内正文

待翻译:How we built an autonomous SRE agent for Kubernetes

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Learn how LangChain built an autonomous SRE agent for Kubernetes deployments with Deep Agents, human approval for changes, LangSmith tracing, and evals.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Learn how LangChain built an autonomous SRE agent for Kubernetes deployments with Deep Agents, human approval for changes, LangSmith tracing, and evals.
站内正文

待翻译:How to Evaluate Voice Agents with LangSmith

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Learn how to evaluate voice agents across execution, outcomes, and caller experience using LangSmith traces, code evaluators, LLM judges, and human review.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Learn how to evaluate voice agents across execution, outcomes, and caller experience using LangSmith traces, code evaluators, LLM judges, and human review.
站内正文

待翻译:Building an Advanced AI Skill Security Auditing Pipeline with NVIDIA SkillSpector, LangGraph, YARA Rules, SARIF, and CI Policy Gates

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Learn how to build an end-to-end security assessment pipeline for AI agent skills using NVIDIA SkillSpector and LangGraph. In this tutorial, we construct a synthetic skill marketplace, scan for malicious prompt injection, credential access, and risky dependencies, and implement custom YARA rules, baseline suppressions, and CI deployment gates. The post Building an Advanced AI Skill Security Auditing Pipeline with NVIDIA SkillSpector, LangGraph, YARA Rules, SARIF, and CI Policy Gates appeared first on MarkTechPost.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Learn how to build an end-to-end security assessment pipeline for AI agent skills using NVIDIA SkillSpector and LangGraph. In this tutorial, we construct a synthetic skill marketp…
站内正文

待翻译:How Stripe Built Kai on Deep Agents in 1 Week

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Learn how Stripe built Kai, a company-wide AI agent on LangChain, LangGraph, and Deep Agents, reaching 5,000 users in roughly 4 weeks.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Learn how Stripe built Kai, a company-wide AI agent on LangChain, LangGraph, and Deep Agents, reaching 5,000 users in roughly 4 weeks.
站内正文

使用 ReviewBench 评估代码审查代理

LangChain 构建了 ReviewBench,一个基于真实拉取请求反馈的基准,用于评估代码审查代理。文章介绍了从真实审查评论中筛选任务、运行方式、评分指标、初步结果以及未来方向。

  • ReviewBench 基于 LangSmith 仓库中受信任审查者的真实 PR 评论构建。
  • 通过 LLM 门控和人工筛选将原始评论转化为可验证的基准任务。
站内正文

LangSmith LLM网关:为生产环境中的AI代理提供运行时控制

LangSmith LLM网关现已公开测试,为生产环境中的代理模型调用提供集中治理层,包括成本控制、速率限制、模型回退和敏感数据保护,帮助团队避免供应商锁定并有效管理模型使用。

  • LangSmith LLM网关作为代理与模型之间的集中治理层,提供运行时控制。
  • 支持成本上限、速率限制、模型回退和敏感数据编辑等关键控制。
站内正文

Similarweb如何使用LangSmith评估Agent报告

了解Similarweb如何使用LangSmith通过评分标准、忠实度检查、追踪和基线比较来评估长篇Agent研究报告。

  • 根据输出类型匹配评估方法:标准答案适用于聚焦问题,而长篇报告需要评分标准、忠实度检查和基线比较。
  • 将分数视为信号而非答案:Similarweb使用LangSmith将每个分数与评估者评论、追踪和A/B比较关联起来。
站内正文

公司导航

LangChain — AI 公司追踪 | AI News Hub