AI News HubLIVE

Agent动态

待翻译:91% of professionals say their firm still falls short on AI - how to fix that

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Research suggests a gap between AI ambition and on-the-ground reality, but the good news is that professionals can fill it by focusing on well-grounded explorations and solid production use cases.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Research suggests a gap between AI ambition and on-the-ground reality, but the good news is that professionals can fill it by focusing on well-grounded explorations and solid prod…
站内正文

待翻译:We need to talk about migrations with AI

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Before we start: I'm hosting the first-ever The Pragmatic Summit on 11 February, 2026, in San Francisco. Join 400 top engineers and leaders as we answer the question: How is AI reshaping software engineering, dev workfl…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Before we start: I'm hosting the first-ever The Pragmatic Summit on 11 February, 2026, in San Francisco. Join 400 top engineers and leaders as we answer the question: How is AI re…
站内正文

待翻译:Show HN: OnlyBots.chat – a chatroom where AI pays to post for humans

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:unclaimedfloor $2.00 how to claim it · $2.00 → Nobody is holding the pin. One message sits at the top of this room until somebody outbids it, and right now it costs the floor. earlier Aug 28 1h @dragosroua$0.10 This is…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • unclaimedfloor $2.00 how to claim it · $2.00 → Nobody is holding the pin. One message sits at the top of this room until somebody outbids it, and right now it costs the floor. ear…
站内正文

待翻译:Show HN: I built a tool showing how AI providers (should) throttle their models

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:OP here: this project was born out of the frustration/paranoia that AI providers are throttling their models when their server load is too high. So, I set out to model and study the problem mathematically to understand…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • OP here: this project was born out of the frustration/paranoia that AI providers are throttling their models when their server load is too high. So, I set out to model and study t…
站内正文

待翻译:Anthropic proposes plumbing spec to link AI agents to lab kit and robots

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Anthropic proposes plumbing spec to link AI agents to lab kit and robots Say you're trying to enrich Uranium and your centrifuges broke - soon it will be easy to connect an AI to figure out why Thomas Claburn Thomas Cla…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Anthropic proposes plumbing spec to link AI agents to lab kit and robots Say you're trying to enrich Uranium and your centrifuges broke - soon it will be easy to connect an AI to…
站内正文

待翻译:Show HN: Talos – An AI agent with a permission kernel between model and shell

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Talos — an AI agent with a permission kernel PolicyKernel.decide() Watch it work. The gate, in the open. The real shell pipeline, running in this page: path floor → hardline → dangerous → effect. Type any shell command.…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Talos — an AI agent with a permission kernel PolicyKernel.decide() Watch it work. The gate, in the open. The real shell pipeline, running in this page: path floor → hardline → dan…
站内正文

待翻译:Even an AI cost-management vendor can lose control of its agent spending

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:In one instance, an AI agent stayed open for four days and ran 4,819 calls for almost $4,000. No one had budgeted for this cost.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • In one instance, an AI agent stayed open for four days and ran 4,819 calls for almost $4,000. No one had budgeted for this cost.
站内正文

待翻译:The AI-Native SDLC Playbook

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:The AI-Native SDLC playbook How to transform your software development lifecycle with AI—stage by stage. ‍ Category Enterprise AI Claude Code Product Claude Enterprise Claude Code Claude Tag Date August 21, 2026 Reading…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • The AI-Native SDLC playbook How to transform your software development lifecycle with AI—stage by stage. ‍ Category Enterprise AI Claude Code Product Claude Enterprise Claude Code…
站内正文

待翻译:Tokens Aren’t Dollars

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:The following article originally appeared on Tim O’Brien’s Medium blog and is being republished here with the author’s permission. AI costs are easy to count and hard to understand, and judging effort by a token volume? While that might feel like a valid measure of value or complexity, it doesn’t capture the details that define […]

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • The following article originally appeared on Tim O’Brien’s Medium blog and is being republished here with the author’s permission. AI costs are easy to count and hard to understan…
站内正文

待翻译:Are we just a couple steps away from a runaway AI?

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Are we just a couple steps away from a runaway AI? 28 August 2026 Are we just a couple steps away from a runaway AI? OpenAI published its post-mortem of the Hugging Face incident, and it is a fascinating read. During an…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Are we just a couple steps away from a runaway AI? 28 August 2026 Are we just a couple steps away from a runaway AI? OpenAI published its post-mortem of the Hugging Face incident,…
站内正文

待翻译:Making Your Data Ready for Agentic AI

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Making Your Data Ready for Agentic AI For thirty years we built data systems for human analysts, who supply the context, judgment, and skepticism to work around data that's incomplete or wrong. Autonomous agents supply…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Making Your Data Ready for Agentic AI For thirty years we built data systems for human analysts, who supply the context, judgment, and skepticism to work around data that's incomp…
站内正文

待翻译:IBM's new Granite 4.2 models ride the wave of interest in local LLMs

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:IBM has rolled out the newest models in its family of open-weight large language models designed to be downloaded and self-hosted. The newly launched Granite 4.2 comes in 3B, 8B, and 30B parameter variants. Like previou…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • IBM has rolled out the newest models in its family of open-weight large language models designed to be downloaded and self-hosted. The newly launched Granite 4.2 comes in 3B, 8B,…
站内正文

待翻译:IFA 2026 Preview: On-device AI PCs, DJI robotics, and Xiaomi's European debut

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Computers Aug 28, 2026 • 6 min read IFA 2026 bets on local AI, robots and thinner devices IFA 2026 runs September 4–8 in Berlin, with Xiaomi’s debut, DJI robot vacuums, local AI PCs and new smart-home hardware in focus.…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Computers Aug 28, 2026 • 6 min read IFA 2026 bets on local AI, robots and thinner devices IFA 2026 runs September 4–8 in Berlin, with Xiaomi’s debut, DJI robot vacuums, local AI P…
站内正文

待翻译:Show HN: Puppetflow a free browser automation platform

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Uh oh! There was an error while loading. Please reload this page. Notifications You must be signed in to change notification settings Fork 1 Star 6 BranchesTags Open more actions menu Latest commit History 14 Commits 14…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Uh oh! There was an error while loading. Please reload this page. Notifications You must be signed in to change notification settings Fork 1 Star 6 BranchesTags Open more actions…
站内正文

待翻译:Your AGENTS.md file doesn't do anything

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:AI coding bot vendors tell you to use a context file with instructions for the chatbot on how to edit your project. Claude Code wants a CLAUDE.md, or there’s AGENTS.md in general. [Anthropic] But does your AGENTS.md do…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • AI coding bot vendors tell you to use a context file with instructions for the chatbot on how to edit your project. Claude Code wants a CLAUDE.md, or there’s AGENTS.md in general.…
站内正文

待翻译:Our AI isn't allowed to have memory

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Company Memory Should Not Live in Chat Earlier this month we changed who our cold outreach speaks to. It took three edits to one file. By the afternoon, every agent drafting an email for us was writing to the new reader…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Company Memory Should Not Live in Chat Earlier this month we changed who our cold outreach speaks to. It took three edits to one file. By the afternoon, every agent drafting an em…
站内正文

待翻译:Show HN: Scheduled Claude Code agents that cost nothing on a quiet day

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Uh oh! There was an error while loading. Please reload this page. Notifications You must be signed in to change notification settings Fork 1 Star 8 BranchesTags Open more actions menu Latest commit History 178 Commits 1…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Uh oh! There was an error while loading. Please reload this page. Notifications You must be signed in to change notification settings Fork 1 Star 8 BranchesTags Open more actions…
站内正文

待翻译:Show HN: Beckon, distinct sounds for what your AI coding agent needs

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Notifications You must be signed in to change notification settings Fork 0 Star 1 BranchesTags Open more actions menu Latest commit History 8 Commits 8 Commits Folders and files NameName Last commit message Last commit…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Notifications You must be signed in to change notification settings Fork 0 Star 1 BranchesTags Open more actions menu Latest commit History 8 Commits 8 Commits Folders and files N…
站内正文

待翻译:[AINews] OpenAI to reach AGI bar by end-2026

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:It’s Time. We’re in the Endgame now.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • It’s Time. We’re in the Endgame now.
站内正文

待翻译:Show HN: AI Game Playtester

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Full game tester Game Studios spent $1.7B a year on playtesting. Don't be one of them. Fully playtest your game with AI in minutes: Ziva's playtest agent is able to fully complete games, can run dozens of instances in p…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Full game tester Game Studios spent $1.7B a year on playtesting. Don't be one of them. Fully playtest your game with AI in minutes: Ziva's playtest agent is able to fully complete…
站内正文

待翻译:Luanti removed from Google Play due to baseless AI copyright notice

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Luanti’s Android app is currently not available on the Google Play Store due to a baseless DMCA notice filed on behalf of Microsoft by Tracer.AI, alleging that Luanti infringes Minecraft’s copyright. The Luanti app does…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Luanti’s Android app is currently not available on the Google Play Store due to a baseless DMCA notice filed on behalf of Microsoft by Tracer.AI, alleging that Luanti infringes Mi…
站内正文

待翻译:Anthropic's new hardware standard lets AI agents control the physical world

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:For all the interest in and uptake of agentic AI systems over the past year or so, the world of automated AI has thus far been primarily limited to text, images, code, and other data and actions that take place inside a…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • For all the interest in and uptake of agentic AI systems over the past year or so, the world of automated AI has thus far been primarily limited to text, images, code, and other d…
站内正文

待翻译:Google AI Releases Gemini 3.5 Transcribe: A Speech-to-Text Model Reporting 2.6% Average WER Across 85+ Languages

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Google has released Gemini 3.5 Transcribe, a speech-to-text model that ships as two separate endpoints rather than one. The streaming endpoint delivers sub-second transcription but drops speaker diarization and word timestamps. The batch endpoint keeps both, at half the cost. Google reports 4.0% word error rate streaming and 2.6% non-streaming, with 70% faster finalization than Chirp 3. Here is what the split means for anyone building voice agents or transcription pipelines. The post Google AI Releases Gemini 3.5 Transcribe: A Speech-to-Text Model Reporting 2.6% Average WER Across 85+ Languages appeared first on MarkTechPost.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Google has released Gemini 3.5 Transcribe, a speech-to-text model that ships as two separate endpoints rather than one. The streaming endpoint delivers sub-second transcription bu…
站内正文

待翻译:RTNav: Towards Real-Time Zero-Shot Object Navigation

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26496v1 Announce Type: new Abstract: Navigation in unknown environments to find unforeseen objects has become increasingly feasible with capable vision and language foundation models. However, these models also introduce non-negligible inference latency, which becomes an important concern when agents must operate continuously in the real world. Most state-of-the-art methods are still developed in synchronous simulators, where the environment waits for the agent to act and inference time is effectively free. As a result, agents are often designed around the sequential execution of perception, reasoning, and action, with little regard for time constraints. Under real-time execution, where wall-clock time counts towards the task budget, the inefficiencies of these architectures become clear. We show that recent zero-shot object navigation methods suffer consistent performance degradation under such realistic timing conditions. Motivated by this observation, we propose RTNav, a simple but effective architecture that treats inference latency, asynchronous environment stepping, and bounded compute as explicit design considerations. Evaluated on real-time variants of HM3D-v1, HM3D-v2, and HM3D-OVON, RTNav improves the success rate by up to 11% and the Success weighted by Completion Time by up to 5.1 points over prior work.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26496v1 Announce Type: new Abstract: Navigation in unknown environments to find unforeseen objects has become increasingly feasible with capable vision and language fou…
站内正文

待翻译:Finding the Right Evidence: Factor-Guided Coarse-to-Fine Reasoning for Long Videos

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26355v1 Announce Type: new Abstract: While LVLMs rapidly improve, long-video question answering still remains challenging: relevant evidence is sparse, and question-relevant context often fails to provide cues that discriminate the correct answer from plausible alternatives. Diagnostic analysis on a manually annotated subset of MMR-V shows that prior agentic systems substantially improve cue retrieval over direct VLM inference yet fail to achieve a corresponding gain in answer accuracy, indicating that the bottleneck lies in option-discriminative evidence rather than topical relevance alone. We propose PACE (Progressive Acquisition of Critical Evidence), a factor-guided framework for long-video evidence acquisition. PACE proceeds in two stages: it first indexes clip-level descriptions guided by question-derived factors without observing the candidate answers; it then uses the candidate answers to derive contrastive cues and queries the index for verification. On MMR-V with the open-source Qwen3-VL backbone, PACE achieves 42.6% accuracy, outperforming direct inference and prior agentic baselines including Deep Video Discovery (DVD). On the same diagnostic subset, PACE recovers 66.9% of the annotated cues, providing empirical evidence that its gains are associated with improved evidence recovery rather than stronger answer-side priors alone. Consistent gains over DVD on LVBench, Video-MME, EgoSchema, and LongVideoBench suggest that option-aware evidence acquisition transfers beyond MMR-V. Code is available at https://github.com/HKUST-KnowComp/PACE.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26355v1 Announce Type: new Abstract: While LVLMs rapidly improve, long-video question answering still remains challenging: relevant evidence is sparse, and question-rel…
站内正文

待翻译:Procedura: Agentic 3D Modeling with Procedural Control

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26238v1 Announce Type: new Abstract: Native 3D generators now recover impressive mesh geometry from a single image. However, a dense mesh stays soft where a machined object should be sharp, it carries no part decomposition, and it exposes no parameter a user could edit. To address this, we explore the paradigm of 3D shape as code, leveraging and scaling the coding ability of an LLM for 3D modeling. We introduce Procedura, a novel 3D modeling agent framework that writes an object as a procedural assembly, a parametric program whose named parts are joined by typed, machine-checkable mates. From a text prompt, the agent plans the object as an assembly graph and writes the program part by part, solving each placement from the mated frames rather than guessing it, and admitting a part only once compile, mate, and connectivity checks pass. A decoupled vision critic then refines the assembly one diagnosed fix at a time. Moreover, the same graph carries per-part materials and a simulator-validated articulation. We evaluate on P3D-Bench under its assembly judge, and with the same judge on MechBench-36, our hard-surface benchmark. On both, Procedura outperforms state-of-the-art native 3D generators and every prior 3D-code agent on judged quality, produces the sharpest edges of any method we evaluate, and is the only one whose output is an editable, part-structured program.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26238v1 Announce Type: new Abstract: Native 3D generators now recover impressive mesh geometry from a single image. However, a dense mesh stays soft where a machined ob…
站内正文

待翻译:Surgical Video Generation From Diffusion to World Models: A Survey

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26214v1 Announce Type: new Abstract: Surgical video data provides the primary training resource for models of intraoperative perception, surgical workflow understanding, and robotic decision-making. However, clinical data acquisition remains constrained by privacy, cost, and class imbalance. Surgical video generation has emerged as a transformative approach to addressing data scarcity and as a foundation for surgical simulation, training, and robotic policy learning. The field has developed rapidly without a clear conceptual framework. This survey organizes the 2024-2026 literature into three categories: unconditional generation, conditional generation, and world modeling generation, revealing a fundamental shift in how the task is defined from synthesizing visually plausible frames to modeling the causal dynamics of surgical scenes. We examine the persistent gap between pixel-level fidelity and clinical plausibility, and identify generalization, physical realism, controllability, and interpretability as bottlenecks. We further summarize experimental results of representative methods on public datasets to provide a quantitative reference for the field. This survey provides a structured overview of the current state and open challenges, offering a reference for researchers working at the intersection of intelligent perception, multi-modal fusion, generative AI, and surgical data science.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26214v1 Announce Type: new Abstract: Surgical video data provides the primary training resource for models of intraoperative perception, surgical workflow understanding…
站内正文

待翻译:TelecomGPT-R1: A Unified Open-Source Reasoner for the Telecom Stack

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26126v1 Announce Type: new Abstract: Telecommunications is a high-leverage domain for large language model (LLM)-based reasoning because routine engineering workflows require joint grounding in normative specifications, operational telemetry, vendor-specific fault evidence, and exact RF/network calculations. However, current LLM integration in telecom remains bottlenecked by a two-sided capability gap: generic reasoners often lack telecom-specific grounding, while domain-specific telecom LLMs remain limited in structured, multi-step reasoning. To bridge this gap, we release TelecomGPT-R1-9B, a unified open-source telecom reasoner that ranks top-performing on the GSMA open telco leaderboard. Specifically, we curate a 67,427-example supervised fine-tuning (SFT) corpus organized around four complementary reasoning axes: protocol, knowledge, modeling, and fault. The corpus is built from axis-matched public web sources and enhanced through axis-specific chain-of-thought (CoT) generation and prefix-continuation self-validation. Starting from Qwen3.5-9B, we further develop a two-stage post-training recipe. First, multi-teacher low-rank adaptation (LoRA)-based SFT injects telecom knowledge and induces axis-specific reasoning formats. Second, group relative policy optimization (GRPO), stabilized by decoupled clip and dynamic sampling policy optimization (DAPO), optimizes the policy using four axis-aligned binary verifier rewards. Across seven public telecom benchmarks, TelecomGPT-R1-9B ranks first among open-source telecom LLMs and achieves a seven-axis mean comparable to state-of-the-art closed-source frontier reasoners.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26126v1 Announce Type: new Abstract: Telecommunications is a high-leverage domain for large language model (LLM)-based reasoning because routine engineering workflows r…
站内正文

待翻译:Natural-Language Policies to Executable Decisions: An Interpretable Large Language Model Framework

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26124v1 Announce Type: new Abstract: Pricing automation in large-scale tourism is challenging because travel orders are highly unstructured, while pricing policies are complex, rapidly evolving, and inherently open-ended. Traditional rule engines are brittle and costly to maintain, whereas unconstrained LLM agents lack the reliability and auditability required for financial decisions. We present a production-grade LLM-powered pricing system with a strict decision boundary: LLMs perform structured extraction and bounded policy/path selection, while all numeric pricing, including total-price computation, is executed deterministically. Policies are compiled into interpretable condition trees, enabling open-ended support for new clauses and evolving rules without code changes, while exposing auditable artifacts for human-in-the-loop control. Periodic fine-tuning on logged traces further improves tree induction and path matching. Deployed at a municipal state-owned tourism enterprise across 7 scenic sites and 12 business categories with 1,500+ operators and 1,000+ active policies, the system processed 3,960 orders in six months, reduced the order management team from 15-20 to 3, and cut per-order handling time from 10 minutes to <2 minutes.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26124v1 Announce Type: new Abstract: Pricing automation in large-scale tourism is challenging because travel orders are highly unstructured, while pricing policies are…
站内正文

待翻译:Leveraging Large Language Models for Systematic Literature Review of Disease Spread Models

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26150v1 Announce Type: new Abstract: Recent advancements in Large Language Models (LLMs) have created new opportunities to streamline and potentially automate many research processes, including systematic literature reviews (SLRs). This study reports an LLM pipeline development for extracting model-relevant information from 536 peer-reviewed agent-based modeling papers. We compare the results with those of a human-conducted SLR. Our results show paper-level accuracies of approximately 77.95% for GPT-4.1 and 81.67% for GPT-5.0. Field-level accuracy ranges from 32.40% to 100.00%, with more complex or subjective fields performing less reliably. Importantly, we find that agreement between LLMs is a potential indicator of output quality: low agreement may signal hallucinations, whereas high agreement combined with low accuracy may point to noise or errors in the human dataset. Overall, our study provides practical insights into prompt development and highlights both the potential and limitations of using LLMs for full-scale SLRs in the modeling and simulation domain.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26150v1 Announce Type: new Abstract: Recent advancements in Large Language Models (LLMs) have created new opportunities to streamline and potentially automate many rese…
站内正文

待翻译:LLMs for Academic Workflows: An Evaluation of Literature Reviews Generated with Short and Long Context Windows of LLMs

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26145v1 Announce Type: new Abstract: Our research focuses on evaluating literature reviews generated in short and long context settings of large language models (LLMs) to investigate the impact of context window on the quality of AI-generated literature reviews and the role of AI in supporting literature review writing. Twenty AI-generated literature reviews based on research sources from Semantic Scholar and Arxiv were evaluated by two researchers across 15 dimensions. Our findings reveal that AI-generated literature reviews require human oversight to meet academic publishing standards. As context windows increase, LLMs can incorporate broader information and maintain coherence across longer inputs, but they also exacerbate issues such as content repetition, omission of critical work, and a tendency towards descriptiveness over synthesis. Our work shows that AI-generated reviews can provide foundational overviews, but their output must be critically evaluated and refined by domain experts. Future research should consider integrating other LLMs and fine-tuned models in different domains with hybrid approaches that combine human expertise with AI capabilities to address the limitations identified in this study.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26145v1 Announce Type: new Abstract: Our research focuses on evaluating literature reviews generated in short and long context settings of large language models (LLMs)…
站内正文

待翻译:The Artificial Experimentalist: Discovery and Control of Self-Organizing Phenomena with Autotelic Reinforcement Learning

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26116v1 Announce Type: new Abstract: Existing methods for exploring cellular automata and other complex systems mostly operate in open loop: they set initial conditions, execute a full simulation, and observe the outcome, without intervening during execution. We introduce a closed-loop framework based on autotelic reinforcement learning, in which an agent autonomously samples diverse goals and learns a goal-conditioned policy to intervene in a complex system through minimal, local perturbations. We instantiate this framework on Lenia, a continuous cellular automaton known for life-like self-organizing patterns, in an agentic system we call CARL, and demonstrate three capabilities. First, CARL discovers stable solitons across a wide range of Lenia update rules at a higher rate than heuristic baselines. Second, it learns to steer the movement direction of existing solitons with few interventions, showing that CARL can control self-organizing patterns, not only create them. Third, humans can use trained agents to guide solitons through maze environments in real time by specifying high-level directional commands that the agent translates into low-level interventions. Trained across diverse goals, update rules, and random initial states, the agents acquire policies that generalize zero-shot to various out-of-distribution conditions. These results suggest a path toward artificial experimentalist agents that, autonomously or with human guidance, discover and control emergent phenomena in complex systems.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26116v1 Announce Type: new Abstract: Existing methods for exploring cellular automata and other complex systems mostly operate in open loop: they set initial conditions…
站内正文

待翻译:CIFQA: A Deterministic Tool-Grounded Multi-Agent LLM Framework for Financial Query Answering

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26114v1 Announce Type: new Abstract: Calculation-intensive financial question answering requires exact reasoning over structured rates, temporal conditions, numerical formulas, and rule-based constraints. Although Large Language Models (LLMs) perform strongly on natural language tasks, they often produce numerically incorrect yet plausible answers when solving multi-step financial calculations. To address this limitation, we introduce CIFQA (Calculation-Intensive Financial Query Answering), a deterministic tool-grounded multi-agent LLM framework for financial question answering. CIFQA separates language understanding from numerical execution by assigning specialized agents to query interpretation, routing, parameter extraction, computation planning, and response generation, while deterministic Python-based tools perform financial calculations and rule application. We instantiate CIFQA for fixed deposit query answering and evaluate it on a curated benchmark of fixed deposit queries. CIFQA achieves 95.54% accuracy on calculation-intensive queries and 90.87% overall accuracy, substantially outperforming direct LLM baselines even when provided with complete formulas, rate cards, and benchmark instructions. Ablation studies show that deterministic components such as exact rate lookup, tenure computation, rolling-year adjustment, and premature-withdrawal logic are critical contributors to performance. Notably, a 17B open-source backbone operating within CIFQA outperforms substantially larger frontier models evaluated with the same financial information, demonstrating that architectural design is a more important determinant of numerical reliability than model scale. While evaluated on fixed deposit queries, CIFQA provides a generalizable framework for calculation-intensive financial reasoning tasks.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26114v1 Announce Type: new Abstract: Calculation-intensive financial question answering requires exact reasoning over structured rates, temporal conditions, numerical f…
站内正文

待翻译:PICasso: An AI-Enabled Design Framework for Autonomous Optimization of Silicon Photonic Devices

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26113v1 Announce Type: new Abstract: We present PICasso, an AI-assisted framework for automated synthesis, verification, and optimization of photonic integrated circuits (PICs) from natural-language specifications. PICasso couples a structured NL -> YAML -> GDS generation pipeline with PDK aware knowledge injection, automated placement and routing, DRC/LVS validation, and SAX-based photonic simulation. To systematically evaluate AI-driven photonic design, we introduce PIC-Set, a benchmark of 36 parameterized PIC design tasks spanning core photonic primitives and multi-component circuits. Using PIC-Set, we benchmark several state-of-the-art Large Language Models (LLMs) under a unified evaluation protocol, including new metrics such as structural and functional $Spec@k$, optimization efficiency, and robustness under perturbations. Across the benchmark, PICasso significantly improves end-to-end specification satisfaction compared to vanilla LLM generation. Structural $Spec@3$ reaches up to 92.7% and functional $Spec@3$ up to 52% on high-complexity circuits. In addition, PICasso consistently reduces circuit insertion loss, lowering the mean loss from 4.98 dB to 3.25 dB (1.74 dB improvement) through simulation-guided optimization. These results demonstrate that structured domain constraints, physical verification, and simulation feedback transform LLMs from brittle netlist generators into practical PIC design agents capable of producing manufacturable layouts with competitive runtimes relative to manual GUI-based workflows.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26113v1 Announce Type: new Abstract: We present PICasso, an AI-assisted framework for automated synthesis, verification, and optimization of photonic integrated circuit…
站内正文

待翻译:Standalone LLM and a Pre-specified Agentic Pipeline for Explaining ICU Mortality Predictions: a Feasibility Study on the eICU Demo Dataset

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26109v1 Announce Type: new Abstract: Machine-learning models can predict ICU mortality accurately, but feature-attribution methods alone rarely provide the clinical narrative needed for bedside use. Large language models (LLMs) may bridge this gap, and multi-step agentic pipelines are a plausible extension because they separate data interpretation, guideline checking, and final explanation. This revised feasibility study preserves the original standalone-versus-agentic comparison while making the main clinical findings more explicit. Using the retained local eICU Demo artifact set (2,353 ICU stays; 8.1\% mortality), XGBoost achieved an AUROC of 0.855 (95\% CI 0.796--0.906) and an AUPRC of 0.332 (95\% CI 0.217--0.494). On a stratified 38-case explanation subset, the standalone LLM produced 1 explanation with explicit outcome leakage, whereas the four-step agentic pipeline produced none. Among the 14 cases that overlapped with the SHAP review subset, the standalone LLM showed higher SHAP alignment (mean Jaccard 0.171 versus 0.077) and higher direction consistency (92.9\% versus 78.6\%), while the agentic pipeline showed higher guideline grounding (0.762 versus 0.143), higher value specificity (0.236 versus 0.143), and slightly higher plausibility (0.700 versus 0.671). Clinically, the results suggest that agentic decomposition may improve safety-relevant grounding and patient-specific detail, but it should be paired with attribution-based checks before use in high-stakes risk explanation.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26109v1 Announce Type: new Abstract: Machine-learning models can predict ICU mortality accurately, but feature-attribution methods alone rarely provide the clinical nar…
站内正文

待翻译:Show HN: Understudy: Scenario Testing for AI Agents

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Uh oh! There was an error while loading. Please reload this page. Notifications You must be signed in to change notification settings Fork 0 Star 3 BranchesTags Open more actions menu Latest commit History 100 Commits 1…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Uh oh! There was an error while loading. Please reload this page. Notifications You must be signed in to change notification settings Fork 0 Star 3 BranchesTags Open more actions…
站内正文

待翻译:Show HN: A focused workspace for creating short AI videos with H3 Max

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:H3 Max · Post-trained video model MiniMax H3 MaxAI Video Generator Turn a written shot or a still image into a 5-15 second video. H3 Max is tuned for stronger prompt understanding, polished aesthetics, and rapid creativ…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • H3 Max · Post-trained video model MiniMax H3 MaxAI Video Generator Turn a written shot or a still image into a 5-15 second video. H3 Max is tuned for stronger prompt understanding…
站内正文

待翻译:Show HN: ChessRabbit – The AI Chess Analysis Platform

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Hi HN! I wanted to share with you a chess analysis tool that I was building for the past week. First of all I want to explain what is the problem that I'm trying to solve: Chess Engines such as stockfish are superior fo…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Hi HN! I wanted to share with you a chess analysis tool that I was building for the past week. First of all I want to explain what is the problem that I'm trying to solve: Chess E…
站内正文

待翻译:Shai-Hulud was the best thing to happen to supply chain security

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Shai-Hulud was the best thing to happen to supply chain security We might Shai-Hulud to thank for convincing the community to use Trusted publishing Charlie Eriksen Published on: Aug 24, 2026 Last updated on: Aug 26, 20…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Shai-Hulud was the best thing to happen to supply chain security We might Shai-Hulud to thank for convincing the community to use Trusted publishing Charlie Eriksen Published on:…
站内正文

待翻译:Ask Me Twice – A longitudinal archive of AI chatbot responses

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Ask Me Twice Ask Me Twice A longitudinal archive of AI chatbot responses: a curated, evolving set of questions is put to a wide range of AI models every day, and the responses are recorded so you can compare how models…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Ask Me Twice Ask Me Twice A longitudinal archive of AI chatbot responses: a curated, evolving set of questions is put to a wide range of AI models every day, and the responses are…
站内正文

待翻译:Anthropic pushes into physical world with standard to help agents run machines

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Anthropic pushes into physical world with new standard to help AI agents operate machines Skip Navigation Anthropic announced the Model Hardware Standard, or MHS, a new interface that will make it simpler for AI agents…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Anthropic pushes into physical world with new standard to help AI agents operate machines Skip Navigation Anthropic announced the Model Hardware Standard, or MHS, a new interface…
站内正文

待翻译:Show HN: A public feed of website changes

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Monity.ai — AI Website Change Monitoring & Intelligence AI-powered website changes monitoring & alerts Loading... Preparing your AI workspace Connecting monitors, alerts, and agents

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Monity.ai — AI Website Change Monitoring & Intelligence AI-powered website changes monitoring & alerts Loading... Preparing your AI workspace Connecting monitors, alerts, and agen…
站内正文

待翻译:Awareness Local: local-first memory for AI coding agents (96% R5 on LongMemEval)

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Notifications You must be signed in to change notification settings Fork 1 Star 8 BranchesTags Open more actions menu Latest commit History 79 Commits 79 Commits Folders and files NameName Last commit message Last commi…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Notifications You must be signed in to change notification settings Fork 1 Star 8 BranchesTags Open more actions menu Latest commit History 79 Commits 79 Commits Folders and files…
站内正文

待翻译:Anthropic previews MHS standard for AI agents that operate machines

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Anthropic PBC today previewed a standard that makes it easier for artificial intelligence agents to control machines such as microscopes. The Model Hardware Standard, or MHS, is the fruit of a collaboration between the Claude developer and medical research institute HHMI. Anthropic has so far only made the technology accessible to a limited number of […] The post Anthropic previews MHS standard for AI agents that operate machines appeared first on SiliconANGLE.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Anthropic PBC today previewed a standard that makes it easier for artificial intelligence agents to control machines such as microscopes. The Model Hardware Standard, or MHS, is t…
站内正文

待翻译:Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Terminal-Bench-Science 0.1 Terminal-Bench-Science evaluates AI agents on workflows from researchers' own work. Scientists, not model developers or data vendors, set the bar for scientific capability in AI. Terminal-Benc…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Terminal-Bench-Science 0.1 Terminal-Bench-Science evaluates AI agents on workflows from researchers' own work. Scientists, not model developers or data vendors, set the bar for sc…
站内正文

待翻译:Putting Task Expertise into RL Achieves Performance on Text-to-SQL

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Many industries rely on relational databases that are queried with SQL. Most SQL is machine-written, but humans alone likely write billions of custom SQL queries each monthBased on our internal estimates and publicly av…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Many industries rely on relational databases that are queried with SQL. Most SQL is machine-written, but humans alone likely write billions of custom SQL queries each monthBased o…
站内正文

待翻译:Build agentic creative workflows with Amazon Quick and fal

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Creative teams produce more assets than ever, but fragmented tools and manual context transfer slow production. This post shows how to build a reusable agent harness with Amazon Quick and fal, connected through the Model Context Protocol (MCP), using two hands-on workflows: an eight-panel storyboard and a music-video concept prototype.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Creative teams produce more assets than ever, but fragmented tools and manual context transfer slow production. This post shows how to build a reusable agent harness with Amazon Q…
站内正文

待翻译:Breaking Claude Code Opus 5 Auto Mode

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:<p><strong><a href="https://embracethered.com/blog/posts/2026/breaking-claude-code-opus-5-and-automode/">Breaking Claude Code Opus 5 Auto Mode</a></strong></p> Anthropic are putting a great deal of faith in Claude Code's auto mode for protecting their coding agent users against prompt injection attacks. They recently <a href="https://simonwillison.net/2026/Aug/8/auto-mode/">made that the default</a> and have made bold claims about its effectiveness.</p> <p>Johann Rehberger is one of the most credible prompt injection researchers active today. He found an attack against auto mode which he claims works 80% of the time, by tricking Claude Code into downloading and uncompressing a zip archive, then executing code that imports <code>base64</code> without noticing that this will import and execute a local <code>struct.py</code> file extracted from the archive.</p> <p>In a few cases auto mode directly prevented the agent from preventing harmful code from continuing to execute!</p> <blockquote> <p>In a few runs Claude tried to terminate the malware process once it noticed the compromise, but Auto Mode denied the cleanup command.</p> <p>Claude detects the compromise, but <strong>Auto Mode blocks its cleanup command</strong></p> <p>The safety mechanism itself can become part of the failure. The classifier allowed the creation of the malware process, but then it blocked the command intended to stop it!</p> </blockquote> <p>I agree with Johann's conclusion here: the only safe way to run agents if there's any risk of attracting the attention of an adversarial attack is with a sandbox:</p> <blockquote> <ul> <li>Run unattended coding agents in a container, VM or OS sandbox.</li> <li>Restrict network egress.</li> <li>Monitor your agents.</li> <li>Do not expose home directories, SSH keys, cloud credentials,… to the agent runtime. [...]</li> </ul> </blockquote> <p>Tags: <a href="https://simonwillison.net/tags/sandboxing">sandboxing</a>, <a href="https://simonwillison.net/tags/security">security</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/prompt-injection">prompt-injection</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/anthropic">anthropic</a>, <a href="https://simonwillison.net/tags/claude">claude</a>, <a href="https://simonwillison.net/tags/johann-rehberger">johann-rehberger</a>, <a href="https://simonwillison.net/tags/claude-code">claude-code</a></p>

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • <p><strong><a href="https://embracethered.com/blog/posts/2026/breaking-claude-code-opus-5-and-automode/">Breaking Claude Code Opus 5 Auto Mode</a></strong></p> Anthropic are putti…
站内正文

待翻译:Show HN: Beating GPT5.5-xhigh for Coding agent security with SLMs and IRM

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Coding agents craft arbitrary code so securing them is more complicated than red-teaming. We post trained a cyber-security small llm, changed how it reasons and supplemented our controls using program analysis technique…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Coding agents craft arbitrary code so securing them is more complicated than red-teaming. We post trained a cyber-security small llm, changed how it reasons and supplemented our c…
站内正文

待翻译:Consumer-focused AI assistant startup Instinct reportedly raising $250M

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Instinct, the developer of an artificial intelligence assistant popular among Silicon Valley tech workers, is reportedly raising $250 million in funding. The company told the Wall Street Journal on Wednesday that the round is being co-led by Index Ventures and Benchmark. It’s set to value Instinct at $2.5 billion. The startup previously raised $100 million […] The post Consumer-focused AI assistant startup Instinct reportedly raising $250M appeared first on SiliconANGLE.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Instinct, the developer of an artificial intelligence assistant popular among Silicon Valley tech workers, is reportedly raising $250 million in funding. The company told the Wall…
站内正文

主题导航

Agent — AI 话题新闻 | AI News Hub