AI News HubLIVE

今日の必読ニュース

ロボット

翻訳待ち:I've tested dozens of robot vacuums - this new $599 Roborock has all the features I need

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:The Roborock Qrevo 2 Pro is a midrange robot vacuum and mop with a hands-free cleaning experience.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • The Roborock Qrevo 2 Pro is a midrange robot vacuum and mop with a hands-free cleaning experience.
サイト内本文
Agent

翻訳待ち:Show HN: I built a tool showing how AI providers (should) throttle their models

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:OP here: this project was born out of the frustration/paranoia that AI providers are throttling their models when their server load is too high. So, I set out to model and study the problem mathematically to understand…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • OP here: this project was born out of the frustration/paranoia that AI providers are throttling their models when their server load is too high. So, I set out to model and study t…
サイト内本文

翻訳待ち:Anthropic proposes plumbing spec to link AI agents to lab kit and robots

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Anthropic proposes plumbing spec to link AI agents to lab kit and robots Say you're trying to enrich Uranium and your centrifuges broke - soon it will be easy to connect an AI to figure out why Thomas Claburn Thomas Cla…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Anthropic proposes plumbing spec to link AI agents to lab kit and robots Say you're trying to enrich Uranium and your centrifuges broke - soon it will be easy to connect an AI to…
サイト内本文

翻訳待ち:Show HN: Talos – An AI agent with a permission kernel between model and shell

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Talos — an AI agent with a permission kernel PolicyKernel.decide() Watch it work. The gate, in the open. The real shell pipeline, running in this page: path floor → hardline → dangerous → effect. Type any shell command.…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Talos — an AI agent with a permission kernel PolicyKernel.decide() Watch it work. The gate, in the open. The real shell pipeline, running in this page: path floor → hardline → dan…
サイト内本文

翻訳待ち:Even an AI cost-management vendor can lose control of its agent spending

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:In one instance, an AI agent stayed open for four days and ran 4,819 calls for almost $4,000. No one had budgeted for this cost.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • In one instance, an AI agent stayed open for four days and ran 4,819 calls for almost $4,000. No one had budgeted for this cost.
サイト内本文

翻訳待ち:The AI-Native SDLC Playbook

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:The AI-Native SDLC playbook How to transform your software development lifecycle with AI—stage by stage. ‍ Category Enterprise AI Claude Code Product Claude Enterprise Claude Code Claude Tag Date August 21, 2026 Reading…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • The AI-Native SDLC playbook How to transform your software development lifecycle with AI—stage by stage. ‍ Category Enterprise AI Claude Code Product Claude Enterprise Claude Code…
サイト内本文
研究

翻訳待ち:Why AI doesn't make companies more productive

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:In 1987, Nobel Prize–winning economist Robert Solow wrote, "You can see the computer age everywhere but in the productivity statistics." The same could be said about AI today. Gartner projects worldwide AI spending of $…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • In 1987, Nobel Prize–winning economist Robert Solow wrote, "You can see the computer age everywhere but in the productivity statistics." The same could be said about AI today. Gar…
サイト内本文
チップ

翻訳待ち:Quantization and Pruning Methods to Make Your LLM Leaner

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:This article walks through what each technique actually does, why skipping them costs real money and real latency, and then gets hands-on with five specific methods people are running in production right now.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • This article walks through what each technique actually does, why skipping them costs real money and real latency, and then gets hands-on with five specific methods people are run…
サイト内本文
ツール

翻訳待ち:How to get free Google AI Pro for an entire year - and save $240: 3 ways

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Google is offering its premium AI subscription plan for free. It usually costs $20 a month - and comes with many benefits.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Google is offering its premium AI subscription plan for free. It usually costs $20 a month - and comes with many benefits.
サイト内本文

翻訳待ち:Show HN: I made a manga critique site with professional Japanese manga editors

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Hey guys, I’m Sean founder & CEO of M2W. We made an AI editor for manga creators worldwide. Manga creators work solo or duo and can’t get critical feedback from public, but editors are too busy and expensive to hire. So…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Hey guys, I’m Sean founder & CEO of M2W. We made an AI editor for manga creators worldwide. Manga creators work solo or duo and can’t get critical feedback from public, but editor…
サイト内本文
その他の更新(158件)
Agent

翻訳待ち:Tokens Aren’t Dollars

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:The following article originally appeared on Tim O’Brien’s Medium blog and is being republished here with the author’s permission. AI costs are easy to count and hard to understand, and judging effort by a token volume? While that might feel like a valid measure of value or complexity, it doesn’t capture the details that define […]

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • The following article originally appeared on Tim O’Brien’s Medium blog and is being republished here with the author’s permission. AI costs are easy to count and hard to understan…
サイト内本文

翻訳待ち:Are we just a couple steps away from a runaway AI?

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Are we just a couple steps away from a runaway AI? 28 August 2026 Are we just a couple steps away from a runaway AI? OpenAI published its post-mortem of the Hugging Face incident, and it is a fascinating read. During an…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Are we just a couple steps away from a runaway AI? 28 August 2026 Are we just a couple steps away from a runaway AI? OpenAI published its post-mortem of the Hugging Face incident,…
サイト内本文

翻訳待ち:Making Your Data Ready for Agentic AI

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Making Your Data Ready for Agentic AI For thirty years we built data systems for human analysts, who supply the context, judgment, and skepticism to work around data that's incomplete or wrong. Autonomous agents supply…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Making Your Data Ready for Agentic AI For thirty years we built data systems for human analysts, who supply the context, judgment, and skepticism to work around data that's incomp…
サイト内本文

翻訳待ち:IBM's new Granite 4.2 models ride the wave of interest in local LLMs

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:IBM has rolled out the newest models in its family of open-weight large language models designed to be downloaded and self-hosted. The newly launched Granite 4.2 comes in 3B, 8B, and 30B parameter variants. Like previou…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • IBM has rolled out the newest models in its family of open-weight large language models designed to be downloaded and self-hosted. The newly launched Granite 4.2 comes in 3B, 8B,…
サイト内本文

翻訳待ち:Show HN: Puppetflow a free browser automation platform

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Uh oh! There was an error while loading. Please reload this page. Notifications You must be signed in to change notification settings Fork 1 Star 6 BranchesTags Open more actions menu Latest commit History 14 Commits 14…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Uh oh! There was an error while loading. Please reload this page. Notifications You must be signed in to change notification settings Fork 1 Star 6 BranchesTags Open more actions…
サイト内本文

翻訳待ち:Your AGENTS.md file doesn't do anything

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:AI coding bot vendors tell you to use a context file with instructions for the chatbot on how to edit your project. Claude Code wants a CLAUDE.md, or there’s AGENTS.md in general. [Anthropic] But does your AGENTS.md do…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • AI coding bot vendors tell you to use a context file with instructions for the chatbot on how to edit your project. Claude Code wants a CLAUDE.md, or there’s AGENTS.md in general.…
サイト内本文

翻訳待ち:Our AI isn't allowed to have memory

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Company Memory Should Not Live in Chat Earlier this month we changed who our cold outreach speaks to. It took three edits to one file. By the afternoon, every agent drafting an email for us was writing to the new reader…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Company Memory Should Not Live in Chat Earlier this month we changed who our cold outreach speaks to. It took three edits to one file. By the afternoon, every agent drafting an em…
サイト内本文

翻訳待ち:Show HN: Scheduled Claude Code agents that cost nothing on a quiet day

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Uh oh! There was an error while loading. Please reload this page. Notifications You must be signed in to change notification settings Fork 1 Star 8 BranchesTags Open more actions menu Latest commit History 178 Commits 1…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Uh oh! There was an error while loading. Please reload this page. Notifications You must be signed in to change notification settings Fork 1 Star 8 BranchesTags Open more actions…
サイト内本文

翻訳待ち:Show HN: Beckon, distinct sounds for what your AI coding agent needs

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Notifications You must be signed in to change notification settings Fork 0 Star 1 BranchesTags Open more actions menu Latest commit History 8 Commits 8 Commits Folders and files NameName Last commit message Last commit…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Notifications You must be signed in to change notification settings Fork 0 Star 1 BranchesTags Open more actions menu Latest commit History 8 Commits 8 Commits Folders and files N…
サイト内本文

翻訳待ち:Show HN: AI Game Playtester

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Full game tester Game Studios spent $1.7B a year on playtesting. Don't be one of them. Fully playtest your game with AI in minutes: Ziva's playtest agent is able to fully complete games, can run dozens of instances in p…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Full game tester Game Studios spent $1.7B a year on playtesting. Don't be one of them. Fully playtest your game with AI in minutes: Ziva's playtest agent is able to fully complete…
サイト内本文

翻訳待ち:Luanti removed from Google Play due to baseless AI copyright notice

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Luanti’s Android app is currently not available on the Google Play Store due to a baseless DMCA notice filed on behalf of Microsoft by Tracer.AI, alleging that Luanti infringes Minecraft’s copyright. The Luanti app does…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Luanti’s Android app is currently not available on the Google Play Store due to a baseless DMCA notice filed on behalf of Microsoft by Tracer.AI, alleging that Luanti infringes Mi…
サイト内本文

翻訳待ち:Anthropic's new hardware standard lets AI agents control the physical world

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:For all the interest in and uptake of agentic AI systems over the past year or so, the world of automated AI has thus far been primarily limited to text, images, code, and other data and actions that take place inside a…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • For all the interest in and uptake of agentic AI systems over the past year or so, the world of automated AI has thus far been primarily limited to text, images, code, and other d…
サイト内本文

翻訳待ち:Google AI Releases Gemini 3.5 Transcribe: A Speech-to-Text Model Reporting 2.6% Average WER Across 85+ Languages

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Google has released Gemini 3.5 Transcribe, a speech-to-text model that ships as two separate endpoints rather than one. The streaming endpoint delivers sub-second transcription but drops speaker diarization and word timestamps. The batch endpoint keeps both, at half the cost. Google reports 4.0% word error rate streaming and 2.6% non-streaming, with 70% faster finalization than Chirp 3. Here is what the split means for anyone building voice agents or transcription pipelines. The post Google AI Releases Gemini 3.5 Transcribe: A Speech-to-Text Model Reporting 2.6% Average WER Across 85+ Languages appeared first on MarkTechPost.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Google has released Gemini 3.5 Transcribe, a speech-to-text model that ships as two separate endpoints rather than one. The streaming endpoint delivers sub-second transcription bu…
サイト内本文

翻訳待ち:RTNav: Towards Real-Time Zero-Shot Object Navigation

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26496v1 Announce Type: new Abstract: Navigation in unknown environments to find unforeseen objects has become increasingly feasible with capable vision and language foundation models. However, these models also introduce non-negligible inference latency, which becomes an important concern when agents must operate continuously in the real world. Most state-of-the-art methods are still developed in synchronous simulators, where the environment waits for the agent to act and inference time is effectively free. As a result, agents are often designed around the sequential execution of perception, reasoning, and action, with little regard for time constraints. Under real-time execution, where wall-clock time counts towards the task budget, the inefficiencies of these architectures become clear. We show that recent zero-shot object navigation methods suffer consistent performance degradation under such realistic timing conditions. Motivated by this observation, we propose RTNav, a simple but effective architecture that treats inference latency, asynchronous environment stepping, and bounded compute as explicit design considerations. Evaluated on real-time variants of HM3D-v1, HM3D-v2, and HM3D-OVON, RTNav improves the success rate by up to 11% and the Success weighted by Completion Time by up to 5.1 points over prior work.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26496v1 Announce Type: new Abstract: Navigation in unknown environments to find unforeseen objects has become increasingly feasible with capable vision and language fou…
サイト内本文

翻訳待ち:Finding the Right Evidence: Factor-Guided Coarse-to-Fine Reasoning for Long Videos

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26355v1 Announce Type: new Abstract: While LVLMs rapidly improve, long-video question answering still remains challenging: relevant evidence is sparse, and question-relevant context often fails to provide cues that discriminate the correct answer from plausible alternatives. Diagnostic analysis on a manually annotated subset of MMR-V shows that prior agentic systems substantially improve cue retrieval over direct VLM inference yet fail to achieve a corresponding gain in answer accuracy, indicating that the bottleneck lies in option-discriminative evidence rather than topical relevance alone. We propose PACE (Progressive Acquisition of Critical Evidence), a factor-guided framework for long-video evidence acquisition. PACE proceeds in two stages: it first indexes clip-level descriptions guided by question-derived factors without observing the candidate answers; it then uses the candidate answers to derive contrastive cues and queries the index for verification. On MMR-V with the open-source Qwen3-VL backbone, PACE achieves 42.6% accuracy, outperforming direct inference and prior agentic baselines including Deep Video Discovery (DVD). On the same diagnostic subset, PACE recovers 66.9% of the annotated cues, providing empirical evidence that its gains are associated with improved evidence recovery rather than stronger answer-side priors alone. Consistent gains over DVD on LVBench, Video-MME, EgoSchema, and LongVideoBench suggest that option-aware evidence acquisition transfers beyond MMR-V. Code is available at https://github.com/HKUST-KnowComp/PACE.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26355v1 Announce Type: new Abstract: While LVLMs rapidly improve, long-video question answering still remains challenging: relevant evidence is sparse, and question-rel…
サイト内本文

翻訳待ち:Procedura: Agentic 3D Modeling with Procedural Control

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26238v1 Announce Type: new Abstract: Native 3D generators now recover impressive mesh geometry from a single image. However, a dense mesh stays soft where a machined object should be sharp, it carries no part decomposition, and it exposes no parameter a user could edit. To address this, we explore the paradigm of 3D shape as code, leveraging and scaling the coding ability of an LLM for 3D modeling. We introduce Procedura, a novel 3D modeling agent framework that writes an object as a procedural assembly, a parametric program whose named parts are joined by typed, machine-checkable mates. From a text prompt, the agent plans the object as an assembly graph and writes the program part by part, solving each placement from the mated frames rather than guessing it, and admitting a part only once compile, mate, and connectivity checks pass. A decoupled vision critic then refines the assembly one diagnosed fix at a time. Moreover, the same graph carries per-part materials and a simulator-validated articulation. We evaluate on P3D-Bench under its assembly judge, and with the same judge on MechBench-36, our hard-surface benchmark. On both, Procedura outperforms state-of-the-art native 3D generators and every prior 3D-code agent on judged quality, produces the sharpest edges of any method we evaluate, and is the only one whose output is an editable, part-structured program.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26238v1 Announce Type: new Abstract: Native 3D generators now recover impressive mesh geometry from a single image. However, a dense mesh stays soft where a machined ob…
サイト内本文

翻訳待ち:Surgical Video Generation From Diffusion to World Models: A Survey

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26214v1 Announce Type: new Abstract: Surgical video data provides the primary training resource for models of intraoperative perception, surgical workflow understanding, and robotic decision-making. However, clinical data acquisition remains constrained by privacy, cost, and class imbalance. Surgical video generation has emerged as a transformative approach to addressing data scarcity and as a foundation for surgical simulation, training, and robotic policy learning. The field has developed rapidly without a clear conceptual framework. This survey organizes the 2024-2026 literature into three categories: unconditional generation, conditional generation, and world modeling generation, revealing a fundamental shift in how the task is defined from synthesizing visually plausible frames to modeling the causal dynamics of surgical scenes. We examine the persistent gap between pixel-level fidelity and clinical plausibility, and identify generalization, physical realism, controllability, and interpretability as bottlenecks. We further summarize experimental results of representative methods on public datasets to provide a quantitative reference for the field. This survey provides a structured overview of the current state and open challenges, offering a reference for researchers working at the intersection of intelligent perception, multi-modal fusion, generative AI, and surgical data science.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26214v1 Announce Type: new Abstract: Surgical video data provides the primary training resource for models of intraoperative perception, surgical workflow understanding…
サイト内本文

翻訳待ち:TelecomGPT-R1: A Unified Open-Source Reasoner for the Telecom Stack

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26126v1 Announce Type: new Abstract: Telecommunications is a high-leverage domain for large language model (LLM)-based reasoning because routine engineering workflows require joint grounding in normative specifications, operational telemetry, vendor-specific fault evidence, and exact RF/network calculations. However, current LLM integration in telecom remains bottlenecked by a two-sided capability gap: generic reasoners often lack telecom-specific grounding, while domain-specific telecom LLMs remain limited in structured, multi-step reasoning. To bridge this gap, we release TelecomGPT-R1-9B, a unified open-source telecom reasoner that ranks top-performing on the GSMA open telco leaderboard. Specifically, we curate a 67,427-example supervised fine-tuning (SFT) corpus organized around four complementary reasoning axes: protocol, knowledge, modeling, and fault. The corpus is built from axis-matched public web sources and enhanced through axis-specific chain-of-thought (CoT) generation and prefix-continuation self-validation. Starting from Qwen3.5-9B, we further develop a two-stage post-training recipe. First, multi-teacher low-rank adaptation (LoRA)-based SFT injects telecom knowledge and induces axis-specific reasoning formats. Second, group relative policy optimization (GRPO), stabilized by decoupled clip and dynamic sampling policy optimization (DAPO), optimizes the policy using four axis-aligned binary verifier rewards. Across seven public telecom benchmarks, TelecomGPT-R1-9B ranks first among open-source telecom LLMs and achieves a seven-axis mean comparable to state-of-the-art closed-source frontier reasoners.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26126v1 Announce Type: new Abstract: Telecommunications is a high-leverage domain for large language model (LLM)-based reasoning because routine engineering workflows r…
サイト内本文

翻訳待ち:Natural-Language Policies to Executable Decisions: An Interpretable Large Language Model Framework

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26124v1 Announce Type: new Abstract: Pricing automation in large-scale tourism is challenging because travel orders are highly unstructured, while pricing policies are complex, rapidly evolving, and inherently open-ended. Traditional rule engines are brittle and costly to maintain, whereas unconstrained LLM agents lack the reliability and auditability required for financial decisions. We present a production-grade LLM-powered pricing system with a strict decision boundary: LLMs perform structured extraction and bounded policy/path selection, while all numeric pricing, including total-price computation, is executed deterministically. Policies are compiled into interpretable condition trees, enabling open-ended support for new clauses and evolving rules without code changes, while exposing auditable artifacts for human-in-the-loop control. Periodic fine-tuning on logged traces further improves tree induction and path matching. Deployed at a municipal state-owned tourism enterprise across 7 scenic sites and 12 business categories with 1,500+ operators and 1,000+ active policies, the system processed 3,960 orders in six months, reduced the order management team from 15-20 to 3, and cut per-order handling time from 10 minutes to <2 minutes.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26124v1 Announce Type: new Abstract: Pricing automation in large-scale tourism is challenging because travel orders are highly unstructured, while pricing policies are…
サイト内本文

翻訳待ち:Leveraging Large Language Models for Systematic Literature Review of Disease Spread Models

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26150v1 Announce Type: new Abstract: Recent advancements in Large Language Models (LLMs) have created new opportunities to streamline and potentially automate many research processes, including systematic literature reviews (SLRs). This study reports an LLM pipeline development for extracting model-relevant information from 536 peer-reviewed agent-based modeling papers. We compare the results with those of a human-conducted SLR. Our results show paper-level accuracies of approximately 77.95% for GPT-4.1 and 81.67% for GPT-5.0. Field-level accuracy ranges from 32.40% to 100.00%, with more complex or subjective fields performing less reliably. Importantly, we find that agreement between LLMs is a potential indicator of output quality: low agreement may signal hallucinations, whereas high agreement combined with low accuracy may point to noise or errors in the human dataset. Overall, our study provides practical insights into prompt development and highlights both the potential and limitations of using LLMs for full-scale SLRs in the modeling and simulation domain.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26150v1 Announce Type: new Abstract: Recent advancements in Large Language Models (LLMs) have created new opportunities to streamline and potentially automate many rese…
サイト内本文

翻訳待ち:LLMs for Academic Workflows: An Evaluation of Literature Reviews Generated with Short and Long Context Windows of LLMs

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26145v1 Announce Type: new Abstract: Our research focuses on evaluating literature reviews generated in short and long context settings of large language models (LLMs) to investigate the impact of context window on the quality of AI-generated literature reviews and the role of AI in supporting literature review writing. Twenty AI-generated literature reviews based on research sources from Semantic Scholar and Arxiv were evaluated by two researchers across 15 dimensions. Our findings reveal that AI-generated literature reviews require human oversight to meet academic publishing standards. As context windows increase, LLMs can incorporate broader information and maintain coherence across longer inputs, but they also exacerbate issues such as content repetition, omission of critical work, and a tendency towards descriptiveness over synthesis. Our work shows that AI-generated reviews can provide foundational overviews, but their output must be critically evaluated and refined by domain experts. Future research should consider integrating other LLMs and fine-tuned models in different domains with hybrid approaches that combine human expertise with AI capabilities to address the limitations identified in this study.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26145v1 Announce Type: new Abstract: Our research focuses on evaluating literature reviews generated in short and long context settings of large language models (LLMs)…
サイト内本文

翻訳待ち:The Artificial Experimentalist: Discovery and Control of Self-Organizing Phenomena with Autotelic Reinforcement Learning

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26116v1 Announce Type: new Abstract: Existing methods for exploring cellular automata and other complex systems mostly operate in open loop: they set initial conditions, execute a full simulation, and observe the outcome, without intervening during execution. We introduce a closed-loop framework based on autotelic reinforcement learning, in which an agent autonomously samples diverse goals and learns a goal-conditioned policy to intervene in a complex system through minimal, local perturbations. We instantiate this framework on Lenia, a continuous cellular automaton known for life-like self-organizing patterns, in an agentic system we call CARL, and demonstrate three capabilities. First, CARL discovers stable solitons across a wide range of Lenia update rules at a higher rate than heuristic baselines. Second, it learns to steer the movement direction of existing solitons with few interventions, showing that CARL can control self-organizing patterns, not only create them. Third, humans can use trained agents to guide solitons through maze environments in real time by specifying high-level directional commands that the agent translates into low-level interventions. Trained across diverse goals, update rules, and random initial states, the agents acquire policies that generalize zero-shot to various out-of-distribution conditions. These results suggest a path toward artificial experimentalist agents that, autonomously or with human guidance, discover and control emergent phenomena in complex systems.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26116v1 Announce Type: new Abstract: Existing methods for exploring cellular automata and other complex systems mostly operate in open loop: they set initial conditions…
サイト内本文

翻訳待ち:CIFQA: A Deterministic Tool-Grounded Multi-Agent LLM Framework for Financial Query Answering

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26114v1 Announce Type: new Abstract: Calculation-intensive financial question answering requires exact reasoning over structured rates, temporal conditions, numerical formulas, and rule-based constraints. Although Large Language Models (LLMs) perform strongly on natural language tasks, they often produce numerically incorrect yet plausible answers when solving multi-step financial calculations. To address this limitation, we introduce CIFQA (Calculation-Intensive Financial Query Answering), a deterministic tool-grounded multi-agent LLM framework for financial question answering. CIFQA separates language understanding from numerical execution by assigning specialized agents to query interpretation, routing, parameter extraction, computation planning, and response generation, while deterministic Python-based tools perform financial calculations and rule application. We instantiate CIFQA for fixed deposit query answering and evaluate it on a curated benchmark of fixed deposit queries. CIFQA achieves 95.54% accuracy on calculation-intensive queries and 90.87% overall accuracy, substantially outperforming direct LLM baselines even when provided with complete formulas, rate cards, and benchmark instructions. Ablation studies show that deterministic components such as exact rate lookup, tenure computation, rolling-year adjustment, and premature-withdrawal logic are critical contributors to performance. Notably, a 17B open-source backbone operating within CIFQA outperforms substantially larger frontier models evaluated with the same financial information, demonstrating that architectural design is a more important determinant of numerical reliability than model scale. While evaluated on fixed deposit queries, CIFQA provides a generalizable framework for calculation-intensive financial reasoning tasks.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26114v1 Announce Type: new Abstract: Calculation-intensive financial question answering requires exact reasoning over structured rates, temporal conditions, numerical f…
サイト内本文

翻訳待ち:PICasso: An AI-Enabled Design Framework for Autonomous Optimization of Silicon Photonic Devices

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26113v1 Announce Type: new Abstract: We present PICasso, an AI-assisted framework for automated synthesis, verification, and optimization of photonic integrated circuits (PICs) from natural-language specifications. PICasso couples a structured NL -> YAML -> GDS generation pipeline with PDK aware knowledge injection, automated placement and routing, DRC/LVS validation, and SAX-based photonic simulation. To systematically evaluate AI-driven photonic design, we introduce PIC-Set, a benchmark of 36 parameterized PIC design tasks spanning core photonic primitives and multi-component circuits. Using PIC-Set, we benchmark several state-of-the-art Large Language Models (LLMs) under a unified evaluation protocol, including new metrics such as structural and functional $Spec@k$, optimization efficiency, and robustness under perturbations. Across the benchmark, PICasso significantly improves end-to-end specification satisfaction compared to vanilla LLM generation. Structural $Spec@3$ reaches up to 92.7% and functional $Spec@3$ up to 52% on high-complexity circuits. In addition, PICasso consistently reduces circuit insertion loss, lowering the mean loss from 4.98 dB to 3.25 dB (1.74 dB improvement) through simulation-guided optimization. These results demonstrate that structured domain constraints, physical verification, and simulation feedback transform LLMs from brittle netlist generators into practical PIC design agents capable of producing manufacturable layouts with competitive runtimes relative to manual GUI-based workflows.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26113v1 Announce Type: new Abstract: We present PICasso, an AI-assisted framework for automated synthesis, verification, and optimization of photonic integrated circuit…
サイト内本文

翻訳待ち:Standalone LLM and a Pre-specified Agentic Pipeline for Explaining ICU Mortality Predictions: a Feasibility Study on the eICU Demo Dataset

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26109v1 Announce Type: new Abstract: Machine-learning models can predict ICU mortality accurately, but feature-attribution methods alone rarely provide the clinical narrative needed for bedside use. Large language models (LLMs) may bridge this gap, and multi-step agentic pipelines are a plausible extension because they separate data interpretation, guideline checking, and final explanation. This revised feasibility study preserves the original standalone-versus-agentic comparison while making the main clinical findings more explicit. Using the retained local eICU Demo artifact set (2,353 ICU stays; 8.1\% mortality), XGBoost achieved an AUROC of 0.855 (95\% CI 0.796--0.906) and an AUPRC of 0.332 (95\% CI 0.217--0.494). On a stratified 38-case explanation subset, the standalone LLM produced 1 explanation with explicit outcome leakage, whereas the four-step agentic pipeline produced none. Among the 14 cases that overlapped with the SHAP review subset, the standalone LLM showed higher SHAP alignment (mean Jaccard 0.171 versus 0.077) and higher direction consistency (92.9\% versus 78.6\%), while the agentic pipeline showed higher guideline grounding (0.762 versus 0.143), higher value specificity (0.236 versus 0.143), and slightly higher plausibility (0.700 versus 0.671). Clinically, the results suggest that agentic decomposition may improve safety-relevant grounding and patient-specific detail, but it should be paired with attribution-based checks before use in high-stakes risk explanation.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26109v1 Announce Type: new Abstract: Machine-learning models can predict ICU mortality accurately, but feature-attribution methods alone rarely provide the clinical nar…
サイト内本文

翻訳待ち:Show HN: Understudy: Scenario Testing for AI Agents

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Uh oh! There was an error while loading. Please reload this page. Notifications You must be signed in to change notification settings Fork 0 Star 3 BranchesTags Open more actions menu Latest commit History 100 Commits 1…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Uh oh! There was an error while loading. Please reload this page. Notifications You must be signed in to change notification settings Fork 0 Star 3 BranchesTags Open more actions…
サイト内本文

翻訳待ち:Show HN: A focused workspace for creating short AI videos with H3 Max

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:H3 Max · Post-trained video model MiniMax H3 MaxAI Video Generator Turn a written shot or a still image into a 5-15 second video. H3 Max is tuned for stronger prompt understanding, polished aesthetics, and rapid creativ…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • H3 Max · Post-trained video model MiniMax H3 MaxAI Video Generator Turn a written shot or a still image into a 5-15 second video. H3 Max is tuned for stronger prompt understanding…
サイト内本文

翻訳待ち:Show HN: ChessRabbit – The AI Chess Analysis Platform

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Hi HN! I wanted to share with you a chess analysis tool that I was building for the past week. First of all I want to explain what is the problem that I'm trying to solve: Chess Engines such as stockfish are superior fo…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Hi HN! I wanted to share with you a chess analysis tool that I was building for the past week. First of all I want to explain what is the problem that I'm trying to solve: Chess E…
サイト内本文

翻訳待ち:Shai-Hulud was the best thing to happen to supply chain security

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Shai-Hulud was the best thing to happen to supply chain security We might Shai-Hulud to thank for convincing the community to use Trusted publishing Charlie Eriksen Published on: Aug 24, 2026 Last updated on: Aug 26, 20…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Shai-Hulud was the best thing to happen to supply chain security We might Shai-Hulud to thank for convincing the community to use Trusted publishing Charlie Eriksen Published on:…
サイト内本文

翻訳待ち:Ask Me Twice – A longitudinal archive of AI chatbot responses

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Ask Me Twice Ask Me Twice A longitudinal archive of AI chatbot responses: a curated, evolving set of questions is put to a wide range of AI models every day, and the responses are recorded so you can compare how models…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Ask Me Twice Ask Me Twice A longitudinal archive of AI chatbot responses: a curated, evolving set of questions is put to a wide range of AI models every day, and the responses are…
サイト内本文

翻訳待ち:Show HN: A public feed of website changes

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Monity.ai — AI Website Change Monitoring & Intelligence AI-powered website changes monitoring & alerts Loading... Preparing your AI workspace Connecting monitors, alerts, and agents

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Monity.ai — AI Website Change Monitoring & Intelligence AI-powered website changes monitoring & alerts Loading... Preparing your AI workspace Connecting monitors, alerts, and agen…
サイト内本文

翻訳待ち:Awareness Local: local-first memory for AI coding agents (96% R5 on LongMemEval)

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Notifications You must be signed in to change notification settings Fork 1 Star 8 BranchesTags Open more actions menu Latest commit History 79 Commits 79 Commits Folders and files NameName Last commit message Last commi…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Notifications You must be signed in to change notification settings Fork 1 Star 8 BranchesTags Open more actions menu Latest commit History 79 Commits 79 Commits Folders and files…
サイト内本文

翻訳待ち:Anthropic previews MHS standard for AI agents that operate machines

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Anthropic PBC today previewed a standard that makes it easier for artificial intelligence agents to control machines such as microscopes. The Model Hardware Standard, or MHS, is the fruit of a collaboration between the Claude developer and medical research institute HHMI. Anthropic has so far only made the technology accessible to a limited number of […] The post Anthropic previews MHS standard for AI agents that operate machines appeared first on SiliconANGLE.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Anthropic PBC today previewed a standard that makes it easier for artificial intelligence agents to control machines such as microscopes. The Model Hardware Standard, or MHS, is t…
サイト内本文

翻訳待ち:Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Terminal-Bench-Science 0.1 Terminal-Bench-Science evaluates AI agents on workflows from researchers' own work. Scientists, not model developers or data vendors, set the bar for scientific capability in AI. Terminal-Benc…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Terminal-Bench-Science 0.1 Terminal-Bench-Science evaluates AI agents on workflows from researchers' own work. Scientists, not model developers or data vendors, set the bar for sc…
サイト内本文

翻訳待ち:Putting Task Expertise into RL Achieves Performance on Text-to-SQL

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Many industries rely on relational databases that are queried with SQL. Most SQL is machine-written, but humans alone likely write billions of custom SQL queries each monthBased on our internal estimates and publicly av…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Many industries rely on relational databases that are queried with SQL. Most SQL is machine-written, but humans alone likely write billions of custom SQL queries each monthBased o…
サイト内本文

翻訳待ち:Build agentic creative workflows with Amazon Quick and fal

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Creative teams produce more assets than ever, but fragmented tools and manual context transfer slow production. This post shows how to build a reusable agent harness with Amazon Quick and fal, connected through the Model Context Protocol (MCP), using two hands-on workflows: an eight-panel storyboard and a music-video concept prototype.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Creative teams produce more assets than ever, but fragmented tools and manual context transfer slow production. This post shows how to build a reusable agent harness with Amazon Q…
サイト内本文

翻訳待ち:Breaking Claude Code Opus 5 Auto Mode

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:<p><strong><a href="https://embracethered.com/blog/posts/2026/breaking-claude-code-opus-5-and-automode/">Breaking Claude Code Opus 5 Auto Mode</a></strong></p> Anthropic are putting a great deal of faith in Claude Code's auto mode for protecting their coding agent users against prompt injection attacks. They recently <a href="https://simonwillison.net/2026/Aug/8/auto-mode/">made that the default</a> and have made bold claims about its effectiveness.</p> <p>Johann Rehberger is one of the most credible prompt injection researchers active today. He found an attack against auto mode which he claims works 80% of the time, by tricking Claude Code into downloading and uncompressing a zip archive, then executing code that imports <code>base64</code> without noticing that this will import and execute a local <code>struct.py</code> file extracted from the archive.</p> <p>In a few cases auto mode directly prevented the agent from preventing harmful code from continuing to execute!</p> <blockquote> <p>In a few runs Claude tried to terminate the malware process once it noticed the compromise, but Auto Mode denied the cleanup command.</p> <p>Claude detects the compromise, but <strong>Auto Mode blocks its cleanup command</strong></p> <p>The safety mechanism itself can become part of the failure. The classifier allowed the creation of the malware process, but then it blocked the command intended to stop it!</p> </blockquote> <p>I agree with Johann's conclusion here: the only safe way to run agents if there's any risk of attracting the attention of an adversarial attack is with a sandbox:</p> <blockquote> <ul> <li>Run unattended coding agents in a container, VM or OS sandbox.</li> <li>Restrict network egress.</li> <li>Monitor your agents.</li> <li>Do not expose home directories, SSH keys, cloud credentials,… to the agent runtime. [...]</li> </ul> </blockquote> <p>Tags: <a href="https://simonwillison.net/tags/sandboxing">sandboxing</a>, <a href="https://simonwillison.net/tags/security">security</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/prompt-injection">prompt-injection</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/anthropic">anthropic</a>, <a href="https://simonwillison.net/tags/claude">claude</a>, <a href="https://simonwillison.net/tags/johann-rehberger">johann-rehberger</a>, <a href="https://simonwillison.net/tags/claude-code">claude-code</a></p>

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • <p><strong><a href="https://embracethered.com/blog/posts/2026/breaking-claude-code-opus-5-and-automode/">Breaking Claude Code Opus 5 Auto Mode</a></strong></p> Anthropic are putti…
サイト内本文

翻訳待ち:Show HN: Beating GPT5.5-xhigh for Coding agent security with SLMs and IRM

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Coding agents craft arbitrary code so securing them is more complicated than red-teaming. We post trained a cyber-security small llm, changed how it reasons and supplemented our controls using program analysis technique…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Coding agents craft arbitrary code so securing them is more complicated than red-teaming. We post trained a cyber-security small llm, changed how it reasons and supplemented our c…
サイト内本文

翻訳待ち:Consumer-focused AI assistant startup Instinct reportedly raising $250M

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Instinct, the developer of an artificial intelligence assistant popular among Silicon Valley tech workers, is reportedly raising $250 million in funding. The company told the Wall Street Journal on Wednesday that the round is being co-led by Index Ventures and Benchmark. It’s set to value Instinct at $2.5 billion. The startup previously raised $100 million […] The post Consumer-focused AI assistant startup Instinct reportedly raising $250M appeared first on SiliconANGLE.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Instinct, the developer of an artificial intelligence assistant popular among Silicon Valley tech workers, is reportedly raising $250 million in funding. The company told the Wall…
サイト内本文

翻訳待ち:Show HN: Make apps in seconds inside of sandbox and share them with a link

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:deenesjakoruzh This server lets agents manage persistent, forkable cloud development environments over MCP. Check compute credits – see the account's remaining compute-credit balance. Create environments – spin up a per…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • deenesjakoruzh This server lets agents manage persistent, forkable cloud development environments over MCP. Check compute credits – see the account's remaining compute-credit bala…
サイト内本文

翻訳待ち:OpenAI Is Developing a 'Persistent' AI Agent

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:In recent days, OpenAI has started adding code for a new “Persistent mode” setting to its command line version of Codex, according to changes made to the product’s code base reviewed by WIRED. Changes to the Codex comma…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • In recent days, OpenAI has started adding code for a new “Persistent mode” setting to its command line version of Codex, according to changes made to the product’s code base revie…
サイト内本文

翻訳待ち:Show HN: ChessRabbit – The AI Chess Analysis Platform

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Hi HN! I wanted to share with you a chess analysis tool that I was building for the past week. First of all I want to explain what is the problem that I'm trying to solve: Chess Engines such as stockfish are superior fo…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Hi HN! I wanted to share with you a chess analysis tool that I was building for the past week. First of all I want to explain what is the problem that I'm trying to solve: Chess E…
サイト内本文

翻訳待ち:The enterprise AI payoff shifts beyond models to mission-critical workflows

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Enterprise AI capabilities are improving almost everywhere, yet the returns still trail the spending. The technology is reaching production, but it often stops short of the business process where revenue, innovation and risk actually live — a gap that is now reshaping how enterprises measure AI success. That gap is widest in industries where a […] The post The enterprise AI payoff shifts beyond models to mission-critical workflows appeared first on SiliconANGLE.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Enterprise AI capabilities are improving almost everywhere, yet the returns still trail the spending. The technology is reaching production, but it often stops short of the busine…
サイト内本文

翻訳待ち:Almanac

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Discussion | Link

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Discussion | Link
サイト内本文

翻訳待ち:Meta memo reveals what its new 'Hatch' AI agent can do

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:It can DJ. It can order food. It can book you a table at a restaurant. And of course, it can access your Instagram. These are some things Meta says its coming AI agent can do, according to an internal memo seen by Busin…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • It can DJ. It can order food. It can book you a table at a restaurant. And of course, it can access your Instagram. These are some things Meta says its coming AI agent can do, acc…
サイト内本文

翻訳待ち:Cohere Releases Parse 5 (parse-v5.0): A 2.3B Vision Language Model That Turns Enterprise Documents Into Markdown

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Cohere has released Parse (parse-v5.0), a 2.3B-parameter vision language model that converts PDFs, slides and images into Markdown with HTML tables, bounding boxes and image descriptions. It runs at $1.50 per 1,000 pages through the API, or on dedicated Model Vault instances from $2,500 a month. Cohere reports a ParseBench score of 79.2, ahead of Mistral OCR 4, Azure Document Intelligence and Databricks AI Parse — but that figure averages three of the benchmark's five dimensions and drops charts and visual grounding entirely. The post Cohere Releases Parse 5 (parse-v5.0): A 2.3B Vision Language Model That Turns Enterprise Documents Into Markdown appeared first on MarkTechPost.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Cohere has released Parse (parse-v5.0), a 2.3B-parameter vision language model that converts PDFs, slides and images into Markdown with HTML tables, bounding boxes and image descr…
サイト内本文

翻訳待ち:Show HN: Apronagents – give each AI coding agent a disposable Git remote

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Notifications You must be signed in to change notification settings Fork 0 Star 1 BranchesTags Open more actions menu Latest commit History 177 Commits 177 Commits Folders and files NameName Last commit message Last com…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Notifications You must be signed in to change notification settings Fork 0 Star 1 BranchesTags Open more actions menu Latest commit History 177 Commits 177 Commits Folders and fil…
サイト内本文

翻訳待ち:I built a long-horizon AI harness that doesn't live in the chat

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:The product is the crane A long-horizon harness you run. Not a plugin pack inside someone else’s. Most things branded “harness engineering” are skills, agents, and slash commands that sit inside Claude Code or Copilot.…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • The product is the crane A long-horizon harness you run. Not a plugin pack inside someone else’s. Most things branded “harness engineering” are skills, agents, and slash commands…
サイト内本文

翻訳待ち:Introducing OpenAI models on Amazon Bedrock for in-country inferencing in India

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Amazon Bedrock now supports the OpenAI GPT-5.6 models, Terra and Luna, in India with India geographic cross-Region inference. If you have local data processing requirements, you can now use these models at scale while Amazon Bedrock keeps inference requests and data within India.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Amazon Bedrock now supports the OpenAI GPT-5.6 models, Terra and Luna, in India with India geographic cross-Region inference. If you have local data processing requirements, you c…
サイト内本文

翻訳待ち:Introducing India cross-Region inference for OpenAI GPT-5.6 models on Amazon Bedrock

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Amazon Bedrock now supports the OpenAI GPT-5.6 models, Terra and Luna, in India with India geographic cross-Region inference. If you have local data processing requirements, you can now use these models at scale while Amazon Bedrock keeps inference requests and data within India.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Amazon Bedrock now supports the OpenAI GPT-5.6 models, Terra and Luna, in India with India geographic cross-Region inference. If you have local data processing requirements, you c…
サイト内本文

翻訳待ち:The AI 'Ghosts' Contaminating Academic Publishing

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Advertisement &bull; Go ad free · Aug 27, 2026 at 2:14 PM “The academic record is being quietly haunted” by researchers with names like Elena Vasquez and Marcus Chen. “Elena Vasquez and Marcus Chen have appeared as volc…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Advertisement &bull; Go ad free · Aug 27, 2026 at 2:14 PM “The academic record is being quietly haunted” by researchers with names like Elena Vasquez and Marcus Chen. “Elena Vasqu…
サイト内本文

翻訳待ち:10 Essential Agentic AI Concepts Explained Simply

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:AI agents are everywhere right now. You hear terms like tool calling, agent loops, MCP, guardrails thrown around as if its common language… it isn’t! But that is about to change. Agentic AI isn’t nearly as complicated as it sounds once you understand the few core ideas that actually matter. Here are 10 agentic AI concepts […] The post 10 Essential Agentic AI Concepts Explained Simply appeared first on Analytics Vidhya.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • AI agents are everywhere right now. You hear terms like tool calling, agent loops, MCP, guardrails thrown around as if its common language… it isn’t! But that is about to change.…
サイト内本文

翻訳待ち:Runable raises $21M to realize small businesses’ growth vision using AI agents

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Runable Inc., a platform that uses artificial intelligence to help businesses build, run and grow, announced Wednesday that it raised $21 million in early-stage funding to scale its operations and reach more enterprise outfits. Susquehanna Venture Capital and Nexus Venture Partners co-led the Series A funding round, alongside continued support from existing investors Together Fund […] The post Runable raises $21M to realize small businesses’ growth vision using AI agents appeared first on SiliconANGLE.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Runable Inc., a platform that uses artificial intelligence to help businesses build, run and grow, announced Wednesday that it raised $21 million in early-stage funding to scale i…
サイト内本文

翻訳待ち:CMS with AI, Not AI CMS: Wagtail 8.0's New API

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Wagtail’s just-released 8.0 release notes are very unusual. Zero admin UI improvements in the highlights, even though UX is one of Wagtail’s biggest strengths. We made a strategic choice to focus on a shiny new API inst…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Wagtail’s just-released 8.0 release notes are very unusual. Zero admin UI improvements in the highlights, even though UX is one of Wagtail’s biggest strengths. We made a strategic…
サイト内本文

翻訳待ち:DHH: Future of Programming, AI, Agentic Engineering, Vibe Coding and Linux [video]

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:- YouTube AboutPressCopyrightContact usCreatorsAdvertiseDevelopersTermsPrivacyPolicy & SafetyHow YouTube worksTest new features

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • - YouTube AboutPressCopyrightContact usCreatorsAdvertiseDevelopersTermsPrivacyPolicy & SafetyHow YouTube worksTest new features
サイト内本文

翻訳待ち:How AI-armed script kiddies will soon wield the power of state-sponsored threat actors

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Low-skill hacktivists are now 'enabled with the same tooling and sophistication as a state-sponsored group,' according to cybersecurity consultant Unit 42.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Low-skill hacktivists are now 'enabled with the same tooling and sophistication as a state-sponsored group,' according to cybersecurity consultant Unit 42.
サイト内本文

翻訳待ち:Replit’s new default: Auto mode picks the best model for each task

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:AI coding company Replit is throwing its weight behind the model-routing trend by making its “intelligent model routing” system the The post Replit’s new default: Auto mode picks the best model for each task appeared first on The New Stack.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • AI coding company Replit is throwing its weight behind the model-routing trend by making its “intelligent model routing” system the The post Replit’s new default: Auto mode picks…
サイト内本文
ツール

翻訳待ち:Microsoft's virtual intern Teams Facilitator will be late for the meeting

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Teams question detection bot postpones debut, will get 2-month extension to practice interrupting you

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Teams question detection bot postpones debut, will get 2-month extension to practice interrupting you
サイト内本文

翻訳待ち:New Studio M5 Ultra ready for local home AI lab? Compared to OpenRouter prices

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Post Log inSign up Post Gabor Herget on X: "New Studio M5 Ultra ready for local home AI lab? I compared it to OpenRouter prices. 512GB will be available from October, and I guess that the price will be around 15-17k. Le…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Post Log inSign up Post Gabor Herget on X: "New Studio M5 Ultra ready for local home AI lab? I compared it to OpenRouter prices. 512GB will be available from October, and I guess…
サイト内本文

翻訳待ち:Nicola Coughlan and Matt Lucas among stars backing campaign against AI voice cloning

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:About 80 people sign open letter to Andy Burnham calling for legislation to protect voice ownership Nicola Coughlan, Hugh Bonneville and Matt Lucas are among a raft of actors backing a campaign against artificial intelligence (AI) voice cloning. Save Our Voices Now, which is also being supported by Luke Evans, Jen Brister, Siobhán McSweeney and Pearl Mackie, aims to stop the practice in which AI technology is used to replicate a real person’s voice to say anything it is prompted to. Continue reading...

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • About 80 people sign open letter to Andy Burnham calling for legislation to protect voice ownership Nicola Coughlan, Hugh Bonneville and Matt Lucas are among a raft of actors back…
サイト内本文

翻訳待ち:Game over for Unauthorized AI Performances: The Sag-Aftra Video Game Strike

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:By: Quintin DiStefano Video game characters may be fictional, but the actors portraying them are real people. Every line of dialogue and every swing of the sword is crafted by voice and motion-capture actors whose perfo…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • By: Quintin DiStefano Video game characters may be fictional, but the actors portraying them are real people. Every line of dialogue and every swing of the sword is crafted by voi…
サイト内本文

翻訳待ち:Revalvo

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Discussion | Link

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Discussion | Link
サイト内本文

翻訳待ち:Hy4 Preview

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Tencent Hy

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Tencent Hy
サイト内本文

翻訳待ち:Pentagon’s blacklisting of Anthropic was unlawful, US judge rules

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Anthropic ​argued designation as ‘supply chain risk’ could cost the company billions ‌of dollars in lost business ‌and reputational harm A US federal judge ruled Thursday that sanctions imposed in February by the Trump administration against AI giant Anthropic were illegal, finding that the government had punished the artificial intelligence company for publicly criticising the Pentagon. “The empty invocation of national security is not a blank check to punish and retaliate against government critics,” Judge Rita Lin said in a 59-page decision. Continue reading...

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Anthropic ​argued designation as ‘supply chain risk’ could cost the company billions ‌of dollars in lost business ‌and reputational harm A US federal judge ruled Thursday that san…
サイト内本文

翻訳待ち:Anthropic was illegally blacklisted by the Trump administration, court rules

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:On Thursday, a judge ruled that the Pentagon's blacklisting of Anthropic earlier this year was unconstitutional, delivering the AI lab a win in a monthslong rollercoaster of a battle with the Trump administration. The lawsuit, filed in March in a California district court, accused the Trump administration of unlawfully retaliating against Anthropic for setting "red lines," or unacceptable military use cases of its AI technology. "The empty invocation of national security is not a blank check to punish and retaliate against government critics," Judge Rita F. Lin, a district judge in the northern district of California, wrote in the ruling … Read the full story at The Verge.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • On Thursday, a judge ruled that the Pentagon's blacklisting of Anthropic earlier this year was unconstitutional, delivering the AI lab a win in a monthslong rollercoaster of a bat…
サイト内本文

翻訳待ち:Alphabet stock sheds $700B as AI bills climb

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Alphabet’s stock is down more than 15% from its May peak, wiping out roughly $700 billion in value, as concerns rise over its mounting AI bills and failure to produce a decisive technological breakthrough. After leading…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Alphabet’s stock is down more than 15% from its May peak, wiping out roughly $700 billion in value, as concerns rise over its mounting AI bills and failure to produce a decisive t…
サイト内本文

翻訳待ち:Black Box: episode 5 – The white mask – podcast

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Revisited: Guardian journalist Michael Safi looks into the world of artificial intelligence, exploring the dangers and promises it holds for society Today in Focus is on a summer break and will be back with new episodes from 1 September. In the meantime, we are bringing you season one of Black Box, before the launch of season two in early September. This episode was first broadcast on 18 March 2024. In January 2020, Robert Williams was arrested by Detroit police for a crime he had not committed. The officers were acting on a tipoff, but not from a witness or informant. In fact, not from a person at all. Continue reading...

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Revisited: Guardian journalist Michael Safi looks into the world of artificial intelligence, exploring the dangers and promises it holds for society Today in Focus is on a summer…
サイト内本文

翻訳待ち:AI-Powered Tilichu VST/AU Synth Released by Butterlamp Audio

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Butterlamp Audio has introduced a new wavetable synthesizer built for classic, knob-by-knob sound design with depth on every layer. Tilichu features two morphing wavetable oscillators, granular and spectral sample engin…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Butterlamp Audio has introduced a new wavetable synthesizer built for classic, knob-by-knob sound design with depth on every layer. Tilichu features two morphing wavetable oscilla…
サイト内本文

翻訳待ち:AI Finds Critical Flaw in Bitcoin Lightning, Devs Issue Emergency Warning

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:In brief Core Lightning confirmed that several AI-generated security reports identified real flaws. The project told operators to verify and install its forthcoming update promptly. Operators who cannot upgrade should u…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • In brief Core Lightning confirmed that several AI-generated security reports identified real flaws. The project told operators to verify and install its forthcoming update promptl…
サイト内本文

翻訳待ち:Unlimited AI calls, done sustainably

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Unlimited, sustainably. Switch once. Keep building. 0code is sustainably unlimited and will always be free. Why switch to 0code? One request, no SDK. Call GET /api/chat/ from any language, script, or prototype. During h…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Unlimited, sustainably. Switch once. Keep building. 0code is sustainably unlimited and will always be free. Why switch to 0code? One request, no SDK. Call GET /api/chat/ from any…
サイト内本文

翻訳待ち:Three UK airports hit by cyber-attack with data of 8.7M customers accessed

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Manchester, London Stansted and East Midlands airports have been hit by a cyber-attack in which hackers accessed the data of about 8.7 million customers. The incident involved data related to “car park, lounge and fast-…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Manchester, London Stansted and East Midlands airports have been hit by a cyber-attack in which hackers accessed the data of about 8.7 million customers. The incident involved dat…
サイト内本文

翻訳待ち:The advent of capable AI tools has highighted a variant of Simpson's paradox

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Terence Tao: "The advent of capable AI tools has highighted a v…" - Mathstodon

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Terence Tao: "The advent of capable AI tools has highighted a v…" - Mathstodon
サイト内本文

翻訳待ち:OpenTag

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Discussion | Link

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Discussion | Link
サイト内本文

翻訳待ち:Gemini Omni 1.1 Flash

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Discussion | Link

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Discussion | Link
サイト内本文

翻訳待ち:AI Weeds Out Job Applicants with Disabilities, Lawsuit Claims

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:AI Weeds Out Job Applicants With Disabilities, Lawsuit Claims by Jean Marbella, The Baltimore Sun/TNS | August 25, 2026 Facebook Twitter Google+ --> LinkedIn Pinterest --> Email BALTIMORE — A lawsuit claiming that Workd…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • AI Weeds Out Job Applicants With Disabilities, Lawsuit Claims by Jean Marbella, The Baltimore Sun/TNS | August 25, 2026 Facebook Twitter Google+ --> LinkedIn Pinterest --> Email B…
サイト内本文

翻訳待ち:Galaxy Z Fold 8 camera trouble? 4 settings I recommend changing now

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:The cameras on the Galaxy Z Fold 8 work best when they're pushed to their limits. Here are the settings I changed to get the best results.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • The cameras on the Galaxy Z Fold 8 work best when they're pushed to their limits. Here are the settings I changed to get the best results.
サイト内本文

翻訳待ち:This excellent HP laptop is nearly 50% off at Best Buy - and comes with a cheap TV

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Snag the HP OmniBook 3 16-inch laptop for under $700, plus get a 43-inch Toshiba smart TV for just $99 when you bundle at Best Buy.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Snag the HP OmniBook 3 16-inch laptop for under $700, plus get a 43-inch Toshiba smart TV for just $99 when you bundle at Best Buy.
サイト内本文

翻訳待ち:Surprising AI breakthroughs raise soul-searching questions for mathematicians | Letter

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Whether mathematics thrives or perishes in the age of AI depends on what it is that society values in the human intellect, writes Dr Henry Bradford I share Kasra Rafi and Bruce Schneier’s impression that recent mathematical breakthroughs by AI consist in clever recombination of existing ideas, not development of truly novel theory (No, AI doesn’t mean the end of mathematics – at least not yet, 25 August). The question is: what happens to mathematics if this changes? Like many mathematicians, I have done much soul-searching in recent weeks, especially since a key problem in my own field of group theory (the existence of non-sofic groups) was solved this month by OpenAI’s Astra model. Astra’s proof consists largely in a slight twist on theorems by my colleagues Gabor Kun and Andreas Thom. That said, even a few months ago I would have found the idea that AI was capable of such a breakthrough incredible, so it seems foolhardy now to bet against AIs achieving superhuman capabilities in all areas of mathematical thought in the coming years. Continue reading...

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Whether mathematics thrives or perishes in the age of AI depends on what it is that society values in the human intellect, writes Dr Henry Bradford I share Kasra Rafi and Bruce Sc…
サイト内本文

翻訳待ち:BRIN is 1/4570th the size of a B-tree, until 5% of rows are updated

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:All posts Oracle vs PostgresAug 27, 202611 min read BRIN is 1/4570th the size of a B-tree, until 5% of rows are updated I co-owned the Zonemaps module at Oracle, and BRIN is the same design in Postgres. It stops pruning…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • All posts Oracle vs PostgresAug 27, 202611 min read BRIN is 1/4570th the size of a B-tree, until 5% of rows are updated I co-owned the Zonemaps module at Oracle, and BRIN is the s…
サイト内本文

翻訳待ち:Sony's new Bravia 6 OLED TV aims to bring flagship visuals to the $1,300 price point

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:The Bravia 6 OLED TV is a new entry-level option in a variety of screen sizes.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • The Bravia 6 OLED TV is a new entry-level option in a variety of screen sizes.
サイト内本文
政策

翻訳待ち:Terence Tao – How the top mathematician uses AI [video]

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:- YouTube AboutPressCopyrightContact usCreatorsAdvertiseDevelopersTermsPrivacyPolicy & SafetyHow YouTube worksTest new features

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • - YouTube AboutPressCopyrightContact usCreatorsAdvertiseDevelopersTermsPrivacyPolicy & SafetyHow YouTube worksTest new features
サイト内本文

翻訳待ち:Judge rules Trump administration illegally punished AI firm Anthropic

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:A judge late Thursday permanently barred the Trump administration from enforcing a set of rules aimed at cutting off Anthropic from the federal government, ruling the administration had unconstitutionally punished the a…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • A judge late Thursday permanently barred the Trump administration from enforcing a set of rules aimed at cutting off Anthropic from the federal government, ruling the administrati…
サイト内本文

翻訳待ち:MI: Measure Price Calculator – AI Pricing for Shopify

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Pricing $9.99/month. Free trial available. Free trial available Rating 0.0 (0 Reviews) Developer MoonImpact.co Install View demo store Featured images gallery Build product price calculators for size, area, length, weig…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Pricing $9.99/month. Free trial available. Free trial available Rating 0.0 (0 Reviews) Developer MoonImpact.co Install View demo store Featured images gallery Build product price…
サイト内本文

翻訳待ち:Who actually benefits from new datacentres in Australia? - podcast

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:National cabinet met on Wednesday to discuss the federal government’s plan for new AI datacentre development. Queensland and the Northern Territory have won a carve-out so that theirs can be powered by coal and gas – rather than renewables. Meanwhile, the fight between the West Australian premier Roger Cook and the Productivity Commission over his state’s generous GST deal persists. And federally, migration arrival numbers continue to spur debate across the political parties. To discuss this week’s mix of politics and policy, political editor Tom McIlroy speaks to the CEO of the Grattan Institute, Dr Aruna Sathanapally, and economics editor Patrick Commins Read more: Albanese backs down on states powering AI datacentres using renewable energy In this hour of need, contemplating a cut to Australia’s refugee intake is cruel populism | Ben Doherty Continue reading...

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • National cabinet met on Wednesday to discuss the federal government’s plan for new AI datacentre development. Queensland and the Northern Territory have won a carve-out so that th…
サイト内本文

翻訳待ち:US Aims to Revive Civil War-Era Court to Claim Iran Oil as Prize

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:The Justice Department is preparing to activate a long-dormant maritime war court to streamline military capture of Iranian oil tankers as US prizes, according to three people familiar with the plans. Reviving prize cou…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • The Justice Department is preparing to activate a long-dormant maritime war court to streamline military capture of Iranian oil tankers as US prizes, according to three people fam…
サイト内本文

翻訳待ち:Apple escaped Android's 'toxic hellstew' - now Siri AI is creating a new one

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Apple's new toxic hellstew is a battle over regulation, litigation, and keeping your data safe.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Apple's new toxic hellstew is a battle over regulation, litigation, and keeping your data safe.
サイト内本文
ロボット

翻訳待ち:These Mammotion robot lawn mowers can climb hills - and they're up to $560 off right now

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:The Mammotion LUBA 3 AWD 3000H and the LUBA Mini 2 mowers are both heavily discounted - here's which you should get.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • The Mammotion LUBA 3 AWD 3000H and the LUBA Mini 2 mowers are both heavily discounted - here's which you should get.
サイト内本文

翻訳待ち:Please stop flooding our projects with AI slop to furnish your CV

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Please stop flooding our projects with AI slop to furnish your CV | neilalexander.dev Neil Home Blog GitHub Please stop flooding our projects with AI slop to furnish your CV 30 June 2026 by Neil Successful contributions…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Please stop flooding our projects with AI slop to furnish your CV | neilalexander.dev Neil Home Blog GitHub Please stop flooding our projects with AI slop to furnish your CV 30 Ju…
サイト内本文
チップ

翻訳待ち:IFA 2026 Preview: On-device AI PCs, DJI robotics, and Xiaomi's European debut

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Computers Aug 28, 2026 • 6 min read IFA 2026 bets on local AI, robots and thinner devices IFA 2026 runs September 4–8 in Berlin, with Xiaomi’s debut, DJI robot vacuums, local AI PCs and new smart-home hardware in focus.…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Computers Aug 28, 2026 • 6 min read IFA 2026 bets on local AI, robots and thinner devices IFA 2026 runs September 4–8 in Berlin, with Xiaomi’s debut, DJI robot vacuums, local AI P…
サイト内本文

翻訳待ち:Anthropic pushes into physical world with standard to help agents run machines

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Anthropic pushes into physical world with new standard to help AI agents operate machines Skip Navigation Anthropic announced the Model Hardware Standard, or MHS, a new interface that will make it simpler for AI agents…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Anthropic pushes into physical world with new standard to help AI agents operate machines Skip Navigation Anthropic announced the Model Hardware Standard, or MHS, a new interface…
サイト内本文

翻訳待ち:Fast, fault-tolerant PyTorch training on AI Runtime

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:At scale, your training efficiency is determined by a single metric: "goodput", the...

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • At scale, your training efficiency is determined by a single metric: "goodput", the...
サイト内本文

翻訳待ち:Yet Another AI Builder But

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:wanno: AI web & app builder Describe your idea. It's live. Only for iPhone Free · In‑App Purchases · Designed for iPhone. Not verified for macOS. Age Rating 4+ Years Category Productivity Developer FEDERICO GUILLERMO CA…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • wanno: AI web & app builder Describe your idea. It's live. Only for iPhone Free · In‑App Purchases · Designed for iPhone. Not verified for macOS. Age Rating 4+ Years Category Prod…
サイト内本文

翻訳待ち:Nvidia reportedly acquires AI project hosting platform Hugging Face for $12.9B

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Nvidia Corp. has reportedly bought Hugging Face Inc., a startup with a popular platform for hosting open-source artificial intelligence projects. Rumors that an acquisition was in the cards first leaked on Monday. Business Insider broke the news that Hugging Face had received interest from multiple prospective buyers. On late Wednesday, The Information reported that Nvidia […] The post Nvidia reportedly acquires AI project hosting platform Hugging Face for $12.9B appeared first on SiliconANGLE.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Nvidia Corp. has reportedly bought Hugging Face Inc., a startup with a popular platform for hosting open-source artificial intelligence projects. Rumors that an acquisition was in…
サイト内本文

翻訳待ち:Sopro V2: SOTA voice cloning TTS model that runs on your CPU

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Sopro V2 Turbo Today we are presenting a new family of TTS models called Sopro V2, and open-sourcing our fastest one: sopro-v2-turbo. Sopro V2 Turbo is a 120M-parameter voice-cloning text-to-speech model that streams, r…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Sopro V2 Turbo Today we are presenting a new family of TTS models called Sopro V2, and open-sourcing our fastest one: sopro-v2-turbo. Sopro V2 Turbo is a 120M-parameter voice-clon…
サイト内本文

翻訳待ち:The AI storage stack gets an inference-era rethink

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Artificial intelligence is changing what storage and data management platforms look like. In collaboration with Super Micro Computer Inc. and Solidigm, DataDirect Networks Inc. has introduced DDN Enterprise AI HyperPOD, built on Nvidia Corp.’s AI Data Platform. The goal is to simplify the storage, scaling and deployment of AI inference for enterprise workloads. “Customers are […] The post The AI storage stack gets an inference-era rethink appeared first on SiliconANGLE.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Artificial intelligence is changing what storage and data management platforms look like. In collaboration with Super Micro Computer Inc. and Solidigm, DataDirect Networks Inc. ha…
サイト内本文

翻訳待ち:Different hats I wear as an AI Engineer

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Hitika Aug 17, 2026 The more I worked with AI systems, the more I noticed a familiar pattern in my own learning: exposure, feedback, mistakes, adjustment, and repetition. Over the past month, I have been interning at a…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Hitika Aug 17, 2026 The more I worked with AI systems, the more I noticed a familiar pattern in my own learning: exposure, feedback, mistakes, adjustment, and repetition. Over the…
サイト内本文

翻訳待ち:Signia MaX hearing aids: four neural networks running simultaneously

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Hearing Aid Technology Signia MaX Hearing Aids in San Mateo & San Carlos Signia MaX, the full name being Multi-adaptive Xperience, is Signia's new flagship platform, announced on August 24, 2026 and available in the Uni…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Hearing Aid Technology Signia MaX Hearing Aids in San Mateo & San Carlos Signia MaX, the full name being Multi-adaptive Xperience, is Signia's new flagship platform, announced on…
サイト内本文

翻訳待ち:Z.AI's Use of Chinese Chips for New Model is About Optimization

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Using domestic chips shows that Chinese vendors are becoming more self-reliant and improving inference performance on AI chips.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Using domestic chips shows that Chinese vendors are becoming more self-reliant and improving inference performance on AI chips.
サイト内本文

翻訳待ち:Pushing GPT-5.6 Luna from 0% to 56% on ARC-AGI-3 Public

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Research August 27, 2026 For long-horizon agents, model capability alone does not determine system capability. Infrastructure orchestration multiplies what models can do. INT21’s SwarmOS tested this hypothesis on the AR…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Research August 27, 2026 For long-horizon agents, model capability alone does not determine system capability. Infrastructure orchestration multiplies what models can do. INT21’s…
サイト内本文

翻訳待ち:AMD Jumps from ROCm 7.14 to ROCm 10.0 with Rocm.ai

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:AMD Jumps From ROCm 7.14 To ROCm 10.0 With ROCm.AI Back in July ROCm 7.14 was announced as their new production release built atop TheRock build system and introducing Ryzen AI 400 series support. The versioning choice…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • AMD Jumps From ROCm 7.14 To ROCm 10.0 With ROCm.AI Back in July ROCm 7.14 was announced as their new production release built atop TheRock build system and introducing Ryzen AI 40…
サイト内本文

翻訳待ち:Best Agent Sandboxes in 2026: Cold Start, Per-Second Pricing, and Network Policy Across E2B, Daytona, Modal, Cloudflare, and Vercel

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Every agent that writes code needs somewhere to run it, and no two vendors quote the same units. This comparison measures burst cold start across E2B, Daytona, Modal, Cloudflare, and Vercel, normalizes per-second rates to cost per 1,000 executions, and maps filesystem persistence, idle billing, and egress policy against primary sources verified August 27, 2026. The post Best Agent Sandboxes in 2026: Cold Start, Per-Second Pricing, and Network Policy Across E2B, Daytona, Modal, Cloudflare, and Vercel appeared first on MarkTechPost.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Every agent that writes code needs somewhere to run it, and no two vendors quote the same units. This comparison measures burst cold start across E2B, Daytona, Modal, Cloudflare,…
サイト内本文

翻訳待ち:Chinese AI Models Overtake American Rivals in Popularity

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Take that, OpenAI! Anthropic! Chinese AI models have surpassed their U.S. counterparts in token consumption on OpenRouter. You might think U.S. AI companies dictate the AI economy. You&rsquo;d be wrong. According to dat…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Take that, OpenAI! Anthropic! Chinese AI models have surpassed their U.S. counterparts in token consumption on OpenRouter. You might think U.S. AI companies dictate the AI economy…
サイト内本文

翻訳待ち:This duck will teach you reinforcement learning — and pick up your socks

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:You could have a mechanical duck waddling through your home before Christmas. Hugging Face‘s Pollen Robotics on Thursday opened pre-orders The post This duck will teach you reinforcement learning — and pick up your socks appeared first on The New Stack.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • You could have a mechanical duck waddling through your home before Christmas. Hugging Face‘s Pollen Robotics on Thursday opened pre-orders The post This duck will teach you reinfo…
サイト内本文

翻訳待ち:Jensen Huang says Nvidia achieved AGI, again — not that it matters

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:On Nvidia's earnings call Wednesday, CEO Jensen Huang casually announced the company had "achieved AGI," one of the tech industry's ultimate goals some of its biggest players have spent years chasing. Almost immediately, Huang dismissed the coveted milestone as "senseless." He's right. For the supposed finish line of the AI race, there is no consensus on what artificial general intelligence means, let alone how we'll know when we've actually got there, which makes achieving it equally arbitrary. Asked about OpenAI's pursuit of AGI, Huang said that when it comes to Nvidia, "for many tasks, we could say that we've already achieved AGI." … Read the full story at The Verge.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • On Nvidia's earnings call Wednesday, CEO Jensen Huang casually announced the company had "achieved AGI," one of the tech industry's ultimate goals some of its biggest players have…
サイト内本文

翻訳待ち:Deepgram deepens Amazon SageMaker AI observability with Enhanced Metrics

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Self-hosted speech AI carries an observability trade-off: the numbers that drive capacity planning and cost management stay locked inside the vendor container. Deepgram closes that gap on Amazon SageMaker AI with two capabilities that land billing, usage, and per-GPU metrics directly in your own Amazon CloudWatch account.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Self-hosted speech AI carries an observability trade-off: the numbers that drive capacity planning and cost management stay locked inside the vendor container. Deepgram closes tha…
サイト内本文

翻訳待ち:Reduce ASR inference costs by 75% with NVIDIA MPS on Amazon EC2

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Serving automatic speech recognition (ASR) models at scale is costly when each request uses only a fraction of a GPU. Learn how NVIDIA CUDA Multi-Process Service (MPS) with NVIDIA Triton Inference Server on Amazon EC2 GPU instances cuts GPU infrastructure by 75% while holding sub-second latency at 92.1 requests per second per GPU.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Serving automatic speech recognition (ASR) models at scale is costly when each request uses only a fraction of a GPU. Learn how NVIDIA CUDA Multi-Process Service (MPS) with NVIDIA…
サイト内本文
研究

翻訳待ち:AI writing has begun to appear on the opinion pages

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Exclusive / AI writing has already begun to appear on the opinion pages Aug 26, 2026, 9:48pm EDT TechnologyMedia Illustration/Jake Angelo/Semafor PostEmailWhatsapp The Scoop Humans are still writing the vast majority of…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Exclusive / AI writing has already begun to appear on the opinion pages Aug 26, 2026, 9:48pm EDT TechnologyMedia Illustration/Jake Angelo/Semafor PostEmailWhatsapp The Scoop Human…
サイト内本文

翻訳待ち:SOLO: Stable Omni-terrain Long-Horizon Perceptive Humanoid Locomotion

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26583v1 Announce Type: new Abstract: Humans traverse complex terrain over long distances without losing balance, whereas perceptive humanoid policies become fragile as perception and control errors accumulate. We present SOLO, a unified framework addressing two compounding causes of this long-horizon fragility: dense terrain reconstruction smooths action-critical details, and pointwise imitation lacks temporal credit assignment. Its Query Reconstructor (QR) uses Fourier-encoded cell queries to retrieve spatially specific evidence from depth-proprioception tokens, preserving sharp terrain boundaries. Trajectory-Aware MSE (TA-MSE) Distillation adds next-state teacher-student disagreement to the PPO reward, enabling Generalized Advantage Estimation to propagate future disagreement penalties to preceding actions. In simulation, QR reduces height-map L1 error by factors of 3.3-4.0, while TA-MSE surpasses PPO and MSE+PPO in curriculum progression. On stress-test terrains, SOLO achieves 97.5% mean traversal success and 96% stepping-stone success, versus 75.0-75.6% and 0-3% for dense-reconstructor variants. Deployed zero-shot with only a chest-mounted depth camera and proprioception, SOLO completes a continuous 1.5-km outdoor route and an indoor mixed-terrain course. Project page: https://sunpihai-up.github.io/solo/

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26583v1 Announce Type: new Abstract: Humans traverse complex terrain over long distances without losing balance, whereas perceptive humanoid policies become fragile as…
サイト内本文

翻訳待ち:Memory Anchors for Continual Robot Learning

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26545v1 Announce Type: new Abstract: Robot policies deployed in the wild should have the capability to continually learn new tasks without forgetting existing behaviors. A common approach to combat such catastrophic forgetting is to train on new task data with a replay buffer of previously learned task data. Although this buffer is commonly sampled randomly from all prior experiences, we show that a small set of these experiences contributes greatly in anchoring past performance. We call these experiences Memory Anchors. We identify Memory Anchors in regions where representations of new-task observations collapse onto those of old-task observations even though the tasks require conflicting actions, like when a familiar object must be manipulated in a new way. Rehearsing old data in this region plays a key role in preventing destructive overwriting of past task knowledge, serving as this critical Memory Anchor role. Excluding only 10% Memory Anchors before sampling the buffer leads to more than a 4.5x increase in catastrophic forgetting on the LIBERO benchmark suites. Conversely, enriching the replay buffer with Memory Anchors can decrease high-conflict task forgetting by 63% and enables successful continual learning of two task sequences on a real robot. Videos and additional visualizations can be found at https://robot-adaptation.github.io/MemoryAnchors

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26545v1 Announce Type: new Abstract: Robot policies deployed in the wild should have the capability to continually learn new tasks without forgetting existing behaviors…
サイト内本文

翻訳待ち:Closing the Loop on the Poppy Humanoid: Bipedal Locomotion with Linear-Quadratic Control and Learned Cost Functions

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26505v1 Announce Type: new Abstract: The Poppy Humanoid is an open-source, low-cost robot suitable for research and education in artificial intelligence. However, we are unaware of any published methodology that achieves reliable, unassisted bipedal locomotion on the standard Poppy hardware. This paper contributes a functional closed-loop walking controller for Poppy, based on the linear-quadratic regulator (LQR) framework for trajectory tracking. Starting with data collected from open-loop playback of a nominal walking trajectory, our proposed method learns a quadratic cost function for an LQR controller that substantially improves the reliability of the motion. The closed-loop controller is validated empirically, demonstrating statistically significant improvements in walking performance compared to open-loop trajectory playback.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26505v1 Announce Type: new Abstract: The Poppy Humanoid is an open-source, low-cost robot suitable for research and education in artificial intelligence. However, we ar…
サイト内本文

翻訳待ち:Dispersive Forward Tree Search for Optimal Control: Coverage, Complexity, and Computation

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26314v1 Announce Type: new Abstract: Steering-based planners require solutions to state-to-state boundary value problems, which can be inaccessible for nonlinear platforms. Forward propagation evades the steering requirement, but the finite-sample behavior of the associated planners remains uncharacterized and their implementations underperform in practice. This paper develops a propagation-based kinodynamic planner with deterministic finite-sample near-optimality guarantees. We work within the large class of differentially flat nonlinear systems and show that a forward tree of locally dispersive control commands contains a near-optimal trajectory at a certified tree size. We provide a general mechanism to construct dispersive command sets for control-affine systems, which are necessary to implement the search algorithm prescribed by the theory. We show that covering the certified trajectory class irrespective of cost provably demands a tree exponentially sized in the problem horizon, and present a cost-conditioned dominance pruning procedure that retains near-optimality at a tree size polynomial in the horizon. We implement the resulting search algorithm, Dispersive Forward Tree search (DFT*), as breadth-first expansion of the forward tree, which maps naturally onto parallel hardware. We design efficient dispersive samplers for the unicycle, the trailer car, and the quadrotor and evaluate challenging planning tasks for these platforms. DFT* delivers consistently competitive and often substantially better solution quality than state-of-the-art kinodynamic planners at comparable solution times on embedded-tier processors, accelerating further as parallel compute is scaled. We also implement DFT* in a receding-horizon loop to demonstrate real-time planning in dynamic environments at embedded-tier compute budgets.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26314v1 Announce Type: new Abstract: Steering-based planners require solutions to state-to-state boundary value problems, which can be inaccessible for nonlinear platfo…
サイト内本文

翻訳待ち:Learning Woody Clearing With Loss Alignment for Zero-Shot Regrowth and Woody Segmentation

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26489v1 Announce Type: new Abstract: Detecting woody clearing is vital for managing biodiversity. Deep learning models can detect change in woody vegetation from bitemporal remote sensing imagery, however generated products may not meet end-user specifications due to unaligned loss definitions. Further limitations of deep learning models are the reliance on large datasets which can be difficult to attain for spatially rare and ambiguous events such as regrowth detection. In this work we train a model to detect woody change using bitemporal Sentinel-2 imagery consisting of 7 years' worth of annual imagery across the state of New South Wales, Australia. To align the objective of the model with end-user metrics, we introduce the loss scaling coefficient $\alpha$ which transforms the objective to optimize for specific $F_{\beta}$ scores. Introducing $\alpha$ was found to increase precision by 1.85x or recall by 1.12x. We propose input imagery augmentation and generation techniques that allow the woody change detection model to zero-shot transfer to regrowth and woody segmentation tasks. For woody segmentation, image generation techniques using activation maximization with low $\alpha$ values for stability and image generation techniques derived from handcrafted features utilizing a mosaic of clearing patches and artificial trees for contextual grounding were found to outperform prior woody segmentation works of the study area, reducing the overall error by up to 18.2%. For zero-shot woody regrowth, creating pseudo-post and prior images resulted in the model achieving an F1 score of 0.845, creating a foundation for future regrowth detection work.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26489v1 Announce Type: new Abstract: Detecting woody clearing is vital for managing biodiversity. Deep learning models can detect change in woody vegetation from bitemp…
サイト内本文

翻訳待ち:Mapping Woody Vegetation from Multi-Source Imagery and Prediction Fusion for Enhanced Data Efficiency and Accuracy

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26471v1 Announce Type: new Abstract: Tree cover maps are a fundamental remote sensing product, used to derive ecological insights about the landscape and are essential to change detection, vegetation mapping and fire monitoring programs. However, comprehensive tree cover mapping requires reliable and high-quality imagery, free of cloud and weather defects to ensure accurate model outputs. Deep learning approaches can generate high quality maps with minimal human intervention but require large amounts of human annotated data to be successful. In this work we propose a framework consisting of methods that aim to improve the data efficiency and robustness of deep learning models using data fusion techniques to segment woody vegetation defined as vegetation over the height of 2m across the state of New South Wales, Australia. To improve robustness against varying image quality, we propose an image composition method that normalizes the imagery and removes defects, whilst also minimizing the reliance on individual image quality by proposing a prediction fusion method. The two methods resulted in an error reduction of 38.2% and 53.6% respectively compared to single-source imagery. To address deep learning approaches' limitation of requiring large amounts of data, we apply label transfer to multiple sources of imagery as a form of data augmentation to improve data efficiency. Learning from multiple image sources was shown to be the biggest improvement in performance, resulting in an error reduction between 28.1% to 76.2% across the different validation experiments, whilst reducing the standard deviation of performance across image dates by a factor of 13.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26471v1 Announce Type: new Abstract: Tree cover maps are a fundamental remote sensing product, used to derive ecological insights about the landscape and are essential…
サイト内本文

翻訳待ち:FIRSTPASS: A Multi-Domain, Multi-Round Peer Review Dataset Grounded in Real Editorial Outcomes

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26129v1 Announce Type: new Abstract: Scientific peer review datasets have trained AI systems exclusively on Computer Science and Machine Learning venues, producing models that critique ablation studies yet have never seen a biology reviewer demand contamination controls or a chemist question Nuclear Magnetic Resonance (NMR) spectral assignments. We introduce FIRSTPASS, the first large-scale peer review dataset built on complete multi-round editorial dialogues from a multidisciplinary high-impact journal. Curated from Nature Communications mandatory transparent peer review (instituted November 2022), FIRSTPASS comprises 3,668 records spanning five scientific domains (biology, chemistry, neuroscience, physics, and earth science), capturing the full iterative structure of scientific validation: initial referee reports, author point-by-point responses, and updated reviewer assessments. Each record carries an outcome label derived directly from editorial decisions (STANDARD for two-round review; EXTENDED for three or more rounds), providing ground truth absent in all prior corpora. An automated audit confirms 100% content integrity. Expert reviews average 2,155 words, substantially denser than conference venue reviews. All data, parsing pipelines, and evaluation scripts are released to enable reproducible benchmarking of AI scientific judgment across disciplines.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26129v1 Announce Type: new Abstract: Scientific peer review datasets have trained AI systems exclusively on Computer Science and Machine Learning venues, producing mode…
サイト内本文

翻訳待ち:Training-Time Explainability for Multilingual Hate Speech Detection: Aligning Model Reasoning with Human Rationales

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26125v1 Announce Type: new Abstract: Online hate against Muslim communities often appears in culturally coded, multilingual forms that evade conventional AI moderation. Such systems, though accurate, remain opaque and risk bias, over-censorship, or under-moderation, particularly when detached from sociocultural context. We propose a \emph{training-time} explainability framework that aligns model reasoning with human-annotated rationales, improving both classification performance and interpretability. Our approach is evaluated on HateXplain (English) and BullySent (Hinglish), reflecting the prevalence of anti-Muslim hate across both languages. Using LIME, Integrated Gradients, Grad X Input, and attention, we assess accuracy, explanation quality, and cross-method agreement. Results show that gradient- and attention-based regularization improve F-scores, enhance plausibility and faithfulness, and capture culturally specific cues for detecting implicit anti-Muslim hate, offering a path toward multilingual, culturally aware content moderation.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26125v1 Announce Type: new Abstract: Online hate against Muslim communities often appears in culturally coded, multilingual forms that evade conventional AI moderation.…
サイト内本文

翻訳待ち:ElementCheck: Complexity-Aware Long-Form Text Factuality Evaluation via Sentence Elements

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26118v1 Announce Type: new Abstract: Existing long-form factuality evaluation relies on the decompose-retrieve-verify pipeline. However, the pipeline suffers from noise from claim decomposition and fixed verification granularity, resulting in unreliable results. We propose ElementCheck, a complexity-aware framework that verifies long-form outputs via sentence elements. Instead of uniformly decomposing sentences into atomic sub-claims, ElementCheck extracts entity pairs that are explicitly linked through verifiable connections in the original sentence as elements, and organizes these into an element graph. The graph topology provides a structural signal for estimating sentence complexity, enabling direct verification for simple sentences and targeted element-level refinement and verification for complex ones. To support fine-grained evaluation, we construct a new benchmark FastFact-Sent by mapping isolated claims from FastFact-Bench back to their source sentences. Experiments on FastFact-Sent and two domain-specific benchmarks show ElementCheck consistently improves factuality verification across five backbone models while maintaining a favorable accuracy-cost trade-off. Further analyses demonstrate that complexity-aware verification reduces unnecessary re-verification and maintains stability across different backbones.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26118v1 Announce Type: new Abstract: Existing long-form factuality evaluation relies on the decompose-retrieve-verify pipeline. However, the pipeline suffers from noise…
サイト内本文

翻訳待ち:FedCMAPSS: A Benchmark for Federated Learning in Remaining Useful Life Estimation

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26433v1 Announce Type: new Abstract: Data-driven prognostics and health management has emerged as a key enabler for Industry 4.0, yet the development of robust remaining useful life (RUL) estimation models is often limited by the scarcity of run-to-failure data. While federated learning offers a promising paradigm to collaboratively train predictive models without sharing sensor data, research efforts have operated so far in the absence of a common evaluation framework. To address this gap, this paper introduces FedCMAPSS, a benchmark for federated RUL estimation based on the commonly-used NASA C-MAPSS dataset. We define a set of five standardized tasks designed to simulate real-world industrial challenges, ranging from ideal IID settings to extreme statistical heterogeneity, and conduct a systematic evaluation of state-of-the-art federated optimization algorithms across multiple neural architectures. By establishing reproducible baselines and making the source code and data splits publicly available, this work aims to provide a standard foundation for developing and comparing federated predictive maintenance solutions.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26433v1 Announce Type: new Abstract: Data-driven prognostics and health management has emerged as a key enabler for Industry 4.0, yet the development of robust remainin…
サイト内本文

翻訳待ち:Algebraic Multigrid Acceleration for Efficient Label Spreading

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26309v1 Announce Type: new Abstract: Modern machine learning models rely on large amounts of labeled data. However, manual annotation of large-scale datasets is expensive and time-consuming. Label spreading is a semi-supervised learning technique that addresses this challenge by propagating information from a few labeled examples to a larger pool of unlabeled data. Despite its effectiveness, its application to large-scale, high-dimensional datasets is limited by computational costs and memory constraints. To address these limitations, we propose Algebraic Multigrid Acceleration for Efficient Label Spreading (AMELS), an efficient label spreading framework that improves scalability by fast construction of neighborhood graphs and the incorporation of algebraic multigrid solvers. The latter is an iterative solver that replaces the ordinary random walk iteration typically performed in label spreading. Due to the multilevel nature of algebraic multigrid solvers, AMELS spreads given label information across a graph of any size in a single multigrid cycle. We demonstrate that AMELS achieves significant runtime reductions compared to existing implementations while also being more robust to hyperparameter choices in terms of both runtime and classification accuracy. Our framework therefore enables efficient label spreading on large-scale image datasets and produces accurate labels even when only a few labeled samples are available.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26309v1 Announce Type: new Abstract: Modern machine learning models rely on large amounts of labeled data. However, manual annotation of large-scale datasets is expensi…
サイト内本文

翻訳待ち:Pruning Binarized Neural Networks: A Dedicated Framework and Globally Weighted Algorithms

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26233v1 Announce Type: new Abstract: Extreme compression of deep neural networks, up to full binarization, dramatically reduces memory footprint and arithmetic complexity, facilitating deployment on constrained edge hardware with field-programmable gate arrays (FPGAs) and microcontrollers. Although combining binarization with pruning promises additional efficiency gains, existing pruning strategies are ill-suited to binarized representations and rarely translate into meaningful hardware savings. We introduce a PyTorch-based, research-oriented framework that incorporates freezing and pruning mechanisms for designing and optimizing binarized neural networks. The framework enables rapid and reproducible evaluation of state-of-the-art approaches and the fast prototyping of new ones. Leveraging this framework, we propose a novel pruning method that accounts for the relative importance of learned parameters across abstraction levels. Such a global weighting mechanism consistently achieves a superior trade-off between model accuracy and pruning rate, achieving a 70% pruning rate on VGG11 with constant accuracy, while state-of-the-art results reach only 41% in the binarized setting.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26233v1 Announce Type: new Abstract: Extreme compression of deep neural networks, up to full binarization, dramatically reduces memory footprint and arithmetic complexi…
サイト内本文

翻訳待ち:The Accuracy-Efficiency Paradox Quantifying Net Energy Loss in on-Device Energy Forecasting

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26134v1 Announce Type: new Abstract: Energy forecasting aims to maximize accuracy to ensure energy efficiency by reducing energy waste, an objective that applies equally to on-device forecasting for mission-critical edge environments, including military systems. However, this paper identifies the Accuracy-Efficiency Paradox: high-precision energy forecasting models can ironically trigger a net energy deficit. This stems from both edge AI's inference energy consumption and battery aging. We propose a Total Cost of Ownership (TCO) framework for energy forecasting, designed to minimize net energy loss. This framework treats not only inference energy consumption but also battery aging as a unified form of energy loss, as degradation represents a physical dissipation of the system's future energy-carrying capacity. We demonstrate that in thermally sensitive edge environments, energy saved by the superior precision of complex architectures is often outweighed by the total energy lost through their high operational intensity.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26134v1 Announce Type: new Abstract: Energy forecasting aims to maximize accuracy to ensure energy efficiency by reducing energy waste, an objective that applies equall…
サイト内本文

翻訳待ち:My $120 GE smart scale's body fat is just 1.5 × BMI − 17.5

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:GE scales are a scam and math theater I exported 3 months of readings from my GE CS10H smart scale. The body fat percentage comes from a formula: fat % = 0.434 × weight − 17.5. Every one of 37 readings fits within 0.1 p…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • GE scales are a scam and math theater I exported 3 months of readings from my GE CS10H smart scale. The body fat percentage comes from a formula: fat % = 0.434 × weight − 17.5. Ev…
サイト内本文

翻訳待ち:Tim O'Reilly – Writing with AI

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:O'Reilly and Tim O'Reilly Aug 27, 2026 By Tim O’Reilly A lot of professional writers say they never do it. And for writers who do, the consequences can be serious. Book contracts have been withdrawn. People have lost jo…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • O'Reilly and Tim O'Reilly Aug 27, 2026 By Tim O’Reilly A lot of professional writers say they never do it. And for writers who do, the consequences can be serious. Book contracts…
サイト内本文

翻訳待ち:Lawsuit says Oura sleep tracking has 'a coin flip's chance of being correct'

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:The class action lawsuit claims that Oura's sleep-tracking mechanisms sell 'faulty AI-based inference as reliable science.'

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • The class action lawsuit claims that Oura's sleep-tracking mechanisms sell 'faulty AI-based inference as reliable science.'
サイト内本文

翻訳待ち:Show HN: An open T2I benchmark with all 9k+ generated images published

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:52 models x 192 prompts = 9,984 images for you to see Loading gallery...

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • 52 models x 192 prompts = 9,984 images for you to see Loading gallery...
サイト内本文

翻訳待ち:Google’s AI note-taking app now allows you to interact with books

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Google's AI note-taking app, Gemini Notebook, can now pull information from the books you've purchased. The new "Expert Intelligence" feature allows you to bring titles from Google Play Books directly into Gemini Notebook, which means you can ask questions about the material, as well as generate plans, infographics, AI podcasts, and more based on their contents. During a briefing with The Verge, Google Labs editorial director Steven Johnson showed how you can use the tool to generate a recipe book using the information in Michael Pollan's Food Rules. In another example, Gemini Notebook applied the knowledge from Kim Scott's management-focus … Read the full story at The Verge.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Google's AI note-taking app, Gemini Notebook, can now pull information from the books you've purchased. The new "Expert Intelligence" feature allows you to bring titles from Googl…
サイト内本文

翻訳待ち:Looking beyond natural sequences

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
サイト内本文

翻訳待ち:Reimagining Homework in the Age of AI by Prioritizing Engagement

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:&larr; Back to Blog A recent study of students using AI in China found that they completed homework faster and scored higher on it, while their later exam performance fell. My hypothesis is that the 30% drop in time spe…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • &larr; Back to Blog A recent study of students using AI in China found that they completed homework faster and scored higher on it, while their later exam performance fell. My hyp…
サイト内本文

翻訳待ち:A new start after 60 in an old occupation | Brief letters

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Grocers and greengrocers | Serbelloni group AI warning | Walking over cars | Christmas catalogues | Appropriate names Living with nostalgia as I am, I enjoyed solving the Wordsearch in Monday’s paper featuring old occupations. However, the man described as a “grocer” in your feature (A new start after 60, 24 August) is surely a greengrocer. Could that word have been included in the Wordsearch as an “old occupation”, or just an old word? David Jones Spalding, Lincolnshire • I daresay not many people knew about the Serbelloni group’s findings (Letters, 25 August). However, millions watched the machines crushing humanity in The Terminator a dozen years later, in 1984. We can’t say that we weren’t warned. Cassy Firth Morley, West Yorkshire Continue reading...

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Grocers and greengrocers | Serbelloni group AI warning | Walking over cars | Christmas catalogues | Appropriate names Living with nostalgia as I am, I enjoyed solving the Wordsear…
サイト内本文

翻訳待ち:3 new ways to plan and book travel in Search

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Book hotels and track airfares, plus view miles and rewards with AI Mode in Google Search.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Book hotels and track airfares, plus view miles and rewards with AI Mode in Google Search.
サイト内本文
スタートアップ

翻訳待ち:In a divided America, left and right unite to oppose AI data centers

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:In a divided America, the left and right unite to oppose artificial intelligence data centers 1 of 5 | Protesters gather outside the New Mexico Environment Department to voice opposition to the Project Jupiter data cent…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • In a divided America, the left and right unite to oppose artificial intelligence data centers 1 of 5 | Protesters gather outside the New Mexico Environment Department to voice opp…
サイト内本文

翻訳待ち:Supporting Thailand’s next generation of AI startups

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:OpenAI and Thailand’s MHESI launch an eight-week accelerator helping 10 health, wellness, and education startups turn AI prototypes into trusted products.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • OpenAI and Thailand’s MHESI launch an eight-week accelerator helping 10 health, wellness, and education startups turn AI prototypes into trusted products.
サイト内本文
モデル

翻訳待ち:Relaxation-Aware Multimodal Sensing of Soft Gripper Driven by Structure-Perception-Learning

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26622v1 Announce Type: new Abstract: Achieving stable, sustained grasping with soft robotic hands remains a fundamental challenge. Compliance enables safe and adaptive contact, yet the intrinsic viscoelasticity of soft polymers leads to stress relaxation and a continuous decay of grasping force during holding. Inspired by human grasping, which combines phase-dependent stiffness regulation with continuous sensing and feedback, this paper presents an integrated structure--perception--learning framework. We develop a variable-stiffness soft gripper that uses onboard vision and infrared thermography to track deformation and the temperature field in real time, preserving continuous tracking of the interaction state. To mitigate relaxation-induced force decay, we propose a temperature-coupled viscoelastic force representation, together with a physics-informed learning model, to reconstruct the force trend and provide explicit compensation during holding. Experiments show that, in a 280s force-controlled grasp-and-hold task, the proposed method maintains the desired force with a mean absolute error of 0.066N, outperforming fixed-aperture and instantaneous-only baselines by 80% and 95%, respectively. Overall, the results support a mechanism--AI co-design view: mechanisms shape feasible interactions, while learning compensates remaining uncertainty in viscoelastic dynamics, together enabling stable, sustained grasping.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26622v1 Announce Type: new Abstract: Achieving stable, sustained grasping with soft robotic hands remains a fundamental challenge. Compliance enables safe and adaptive…
サイト内本文

翻訳待ち:TrapVLA: Trapping Vision-Language-Action Models in Configured Failure Modes

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26578v1 Announce Type: new Abstract: This work introduces Configured Failure Trapping, a novel backdoor attack task against Vision-Language-Action (VLA) models, which aims to activate attacks through stealthy textual triggers and induce configured failure modes. Unlike prior backdoor attacks that treat any task failure as a successful attack, Configured Failure Trapping requires the attacker to control how the robot fails (e.g., causing the robot to grasp with a specified positional offset), making it substantially more challenging and hard to detect. To support the new task, we propose an effective data engine for synthesizing high-quality target trajectories and an automated suite for measuring configured-failure fidelity. Then, based on this foundation, we construct two new benchmarks, namely Trap-LIBERO and Trap-RoboTwin, that instantiate Configured Failure Trapping across four representative failure modes. To address this task, we identify sparse action deviation as a critical challenge and accordingly propose a novel method named TrapVLA, which explicitly learns trigger-induced action residuals to steer the policy toward the configured failure behavior. Extensive experiments across simulation benchmarks and real-world robotic settings show that TrapVLA effectively injects configured failure modes into VLA models while largely preserving performance on clean data. Project page: https://john-liua.github.io/TrapVLA/

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26578v1 Announce Type: new Abstract: This work introduces Configured Failure Trapping, a novel backdoor attack task against Vision-Language-Action (VLA) models, which a…
サイト内本文

翻訳待ち:Cross-Platform Benchmark of Neural 3D Reconstruction for Autonomous Laboratory Robots

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26383v1 Announce Type: new Abstract: Autonomous robots performing laboratory tasks depend on 3D reconstruction pipelines that can turn raw camera streams into actionable object representations within the latency budget of a physical control loop. Neural 3D reconstruction methods have demonstrated high-quality view synthesis, but their real-time viability across the compute platforms on which laboratory robots actually run remains poorly characterized. In this work, we present a systematic compute-platform benchmark of neural 3D reconstruction methods, evaluating NeRF and 3D Gaussian Splatting training and rendering on GPU-enabled computing devices ranging from single-board computers to server-class nodes, and place Meta's SAM3D single-image reconstruction on the same axes to quantify its latency and fidelity gap relative to per-scene optimization. Our results show that Gaussian Splatting yields higher rendering quality than NeRF at greater GPU cost, and that onboard compute is insufficient for full per-scene optimization at interactive rates. Our preliminary assessment on SAM3D indicates that it delivers plausible object geometry within seconds, but with detail mismatches that can compromise downstream manipulation. Together, these findings motivate tiered pipelines in which lightweight feed-forward reconstruction sustains the real-time perception-and-tracking loop for laboratory robots, while heavier neural reconstruction is scheduled selectively on suitable compute.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26383v1 Announce Type: new Abstract: Autonomous robots performing laboratory tasks depend on 3D reconstruction pipelines that can turn raw camera streams into actionabl…
サイト内本文

翻訳待ち:Constraint-Aware Physics-Informed Neural Networks for Static Shape Estimation of Co-Manipulative Continuum Robots

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26273v1 Announce Type: new Abstract: Static shape estimation of co-manipulative continuum robots (CCRs) is challenging because the continuum arms and manipulated flexible object form a closed chain that must satisfy both static equilibrium and geometric loop-closure constraints. This paper presents a constraint-aware physics-informed neural network (PINN) for static shape estimation of a tendon-driven CCR modeled using the geometric variable strain formulation. The proposed method incorporates a projected static equilibrium residual and a configuration-level geometric residual to enforce the governing mechanics and closed-chain geometry. In simulation, the PINN is compared with a purely data-driven artificial neural network (ANN) under limited and noisy training data. With 140 samples and 50% label noise, the PINN reduces the relative configuration error, equilibrium residual, and closed-chain residual by 67.88%, 67.35%, and 88.06%, respectively. Using the full dataset, the PINN achieves 0.1597% relative configuration error with an inference time of 0.1773 ms, compared with 17.97 s for an iterative nonlinear solver. Experimental fine-tuning reduces the marker RMSE from 2.657 mm to 0.497 mm and increases R2 from -0.788 to 0.937. These results demonstrate accurate, physically consistent, and computationally efficient static shape estimation of closed-chain CCRs.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26273v1 Announce Type: new Abstract: Static shape estimation of co-manipulative continuum robots (CCRs) is challenging because the continuum arms and manipulated flexib…
サイト内本文

翻訳待ち:WALL-SS: Scaling Long-horizon World Models via Next-Scale Autoregression

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26239v1 Announce Type: new Abstract: Generative world models provide robots with predictive models of how the world evolves under interaction, with growing potential for simulation, planning, policy evaluation, and robot learning. Beyond clip-level future prediction, a unified generative formulation should relate actions to consequences, support flexible horizons and continuous interaction, and enable reward-driven optimization. We introduce WALL-SS, a world model that generates visual futures through Scale-wise autoregressive Scaling, enabling action-controllable and long-horizon robotic simulation. WALL-SS represents embodied trajectories as causal sequences of temporally interleaved observations and actions, making action-dependent state transitions explicit while naturally supporting variable-length generation, streaming extension through reusable causal states, and direct optimization through sequence probabilities. To make this formulation effective over long horizons, we generate each future observation in a coarse-to-fine manner and develop three complementary components within the same hierarchy. Action-conditioned next-scale prediction injects scale-aligned action representations to improve action-future coupling and model both successful and failed behaviors. Scale-compressed long-horizon memory retains recent interactions at fine resolution while compressing distant observations and actions, with scale-wise dream forcing enhancing robustness to self-generated context. Finally, on-policy alignment optimizes autoregressive visual dynamics with action-following and long-term consistency rewards while preserving the pretrained visual distribution. Experiments show that WALL-SS improves action following and trajectory accuracy, supports coherent minute-long streaming rollout under bounded memory, and consistently benefits from on-policy alignment in reducing action drift and long-horizon inconsistency.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26239v1 Announce Type: new Abstract: Generative world models provide robots with predictive models of how the world evolves under interaction, with growing potential fo…
サイト内本文

翻訳待ち:Video-FLAIR: Not Whether to Reason, But How

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26495v1 Announce Type: new Abstract: Multimodal queries can require different types of reasoning. Some can be answered via perceptual reasoning, extracting information directly from the visual signal, while others require compositional reasoning that combines observations or deliberative reasoning that evaluates competing hypotheses. However, many existing methods apply a uniform reasoning strategy across queries, leading to unnecessary computation on simple tasks and insufficient reasoning on complex ones. We introduce Video-FLAIR, a training framework that learns to select the appropriate reasoning mode for each query using reinforcement learning. During training, the model generates responses under all three modes for the same prompt, enabling direct comparison. A composite reward compares these responses to favor the most effective one based on correctness, grounding, and cost, while discouraging unsupported or misaligned deliberation. This yields a supervision signal for learning adaptive reasoning without per-query annotations. Video-FLAIR improves accuracy over the Qwen2.5-VL base model by +5.4 on MathVista, +4.8 on Video-Holmes, and +4.8 on Video-MMMU, while reducing average token usage to 95 compared to 417 for always-thinking baselines.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26495v1 Announce Type: new Abstract: Multimodal queries can require different types of reasoning. Some can be answered via perceptual reasoning, extracting information…
サイト内本文

翻訳待ち:Zero-Shot Video Restoration and Enhancement with Text-to-Image Latent Diffusion Models and Multi-Modal References

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26476v1 Announce Type: new Abstract: Zero-shot image restoration methods with text-to-image latent diffusion models have achieved great success in universal image restoration tasks without training. However, applying them to video restoration will result in severe temporal flickering. In this paper, we propose a novel framework for zero-shot video restoration and enhancement which uses a text-to-image latent diffusion model and multi-modal references. Through the proposed dual prompt tuning inversion and sampling, the inference time can be reduced to nearly 1/3 of the original. The performance and temporal consistency can be also significantly stregthened. By using the proposed texture-aware video token merging, the temporal correlation between frames can be further utilized to improve the temporal consistency. We futher propose the referenced self-attention and referenced token merging to support image reference. Experimental results demonstrate the superiority of the proposed method in restoring and enhancing temporally consistent videos.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26476v1 Announce Type: new Abstract: Zero-shot image restoration methods with text-to-image latent diffusion models have achieved great success in universal image resto…
サイト内本文

翻訳待ち:VIPER: An Expert-Curated Benchmark for Vision-Language Models in Veterinary Pathology

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26382v1 Announce Type: new Abstract: Pathology vision-language models are advancing rapidly, yet existing benchmarks remain focused on human tissue, particularly oncology, leaving non-human pathology largely unaddressed. This gap is especially important in toxicologic pathology, where microscopic tissue examination of laboratory animals is a core component of preclinical drug safety assessment. To address it, we introduce VIPER, the first expert-curated benchmark for vision-language model evaluation in toxicologic pathology. VIPER contains 1,251 questions associated with 419 H&E-stained rat histology images across seven organ systems, covering multiple-choice, KPrim, and free-text formats. All questions were curated and validated by board-certified veterinary pathologists. In total, we benchmarked 16 models, including two newly introduced veterinary-pathology models, seven human pathology-specialized models, and seven general-purpose frontier models. The results identify a substantial domain gap between veterinary and human pathology, expose the risk of over-diagnosis of normal tissue in frontier models, and show that domain-specific training remains critical for visually grounded predictions. VIPER data and evaluation code are available at https://github.com/mahmoodlab/viper.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26382v1 Announce Type: new Abstract: Pathology vision-language models are advancing rapidly, yet existing benchmarks remain focused on human tissue, particularly oncolo…
サイト内本文

翻訳待ち:A Unified Framework for the Mechanics of Information in Convolutional Neural Network Image Space

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26363v1 Announce Type: new Abstract: This paper introduces a unified mathematical framework for modeling information propagation through convolutional neural networks (CNNs), with the aim of connecting descriptions of physical space and information space. A correspondence is presented linking discrete filter symmetry and the relativistic energy--momentum relation under the widely used nonlinear rectified convolution operation. Specifically, symmetric filter components (e.g. the sum $\Sigma = [1,1]$) operate analogously to rest energy $mc^2$ in preserving the image centre of mass (e.g. isotropic diffusion), whereas antisymmetric components (e.g. the gradient $\nabla = [-1,1]$) operate analogously to the momentum term $pc$ in generally inducing a displacement (e.g. vibration or translation). For typical small discrete filters, this displacement is determined by the ratio of antisymmetric to total filter energy, analogously to how the displacement of a relativistic particle relates to a Lorentz transform with beta parameter $\beta = \frac{v}{c}=\frac{pc}{E}$ equal to the ratio of momentum $pc$ to total energy $E$. Repeated filtering leads to the Gaussian scale-space and emergent scale-invariant features. These constructions share a Laplacian-driven structure with the classical heat (diffusion) equation and, via standard mathematical correspondences, with the Schr\"odinger equation and aspects of the Friedmann equations, together with emergent Morse topological structure. Demonstrations in 3D images reveal blob-like, scale-invariant Morse critical points in images spanning a wide range of physical scales, including organic sugar molecules and inorganic silicon crystals, human and primate brains in magnetic resonance images (MRI), galaxies and the cosmic microwave background (CMB).

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26363v1 Announce Type: new Abstract: This paper introduces a unified mathematical framework for modeling information propagation through convolutional neural networks (…
サイト内本文

翻訳待ち:Modality Maturity Index: A benchmark for assessing multimodal capabilities of omni models

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26317v1 Announce Type: new Abstract: Frontier language models are increasingly marketed as omni systems that can perceive and respond across modalities. Existing evaluation frameworks, however, focus almost exclusively on bimodal understanding, typically text plus one other modality. We propose the Modality Maturity Index (MMI), a benchmark designed to evaluate the multimodal capabilities of large language models across five modalities (text, image, audio, video and document) and combinations of up to three modalities in both inputs and outputs. MMI consists of 893 questions, each carefully crafted to require the model to demonstrate its understanding of multiple input modalities and to generate responses that incorporate various output formats. The questions are designed to be self-contained, with clear expectations for the correct modality or mix of modalities required for an accurate response. Every MMI prompt carries human-authored rubric criteria for each output modality expected in the response; a model's MMI Value expresses the average of the per-modality scores for each prompt. Because low scores can reflect either failure to generate a modality (lack of presence) or failure to generate correct content, we introduce also a supplementary Modality Presence Score (MPS), a per-prompt F1 over the expected output modalities. Applying MMI to five frontier multimodal models, we find that the MPS ranges from only 15.6 (Claude Opus 4.6) to 34.9 (GPT-5.4). Given the low availability of returned modalities to even grade, we report MPS as our main result pending model improvements. To assess the viability of judging output correctness with LLM judges and rubrics, we run a separate experiment with custom generation tools. On the assets that generates, we find that an LLM judge applying the rubrics agrees with rubric-blind human annotators (who score the outputs directly and never see the criteria) on 70.8% of judgments.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26317v1 Announce Type: new Abstract: Frontier language models are increasingly marketed as omni systems that can perceive and respond across modalities. Existing evalua…
サイト内本文

翻訳待ち:Which India Survives Translation? Narrative Homogenisation Across Indian Oral Traditions in LLMs

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26123v1 Announce Type: new Abstract: Large language models (LLMs) are trained predominantly on English-language internet text that over-represents certain cultural narratives, raising concerns that models flatten the diversity of non-Western storytelling traditions into a single homogenized archetype. We present a pilot computational study examining this across three maximally distinct Indian regional oral and literary traditions: the Rajasthani Pabuji epic, classical Tamil Sangam poetry, and Bengali folk tales. We collected authentic reference corpora for each tradition (11, 21, and 10 passages respectively) and prompted two LLMs (Claude Sonnet and Gemini) with 54 generation requests spanning three prompt types per tradition - generic, culturally specific, and regional-language. Using Sentence-BERT embeddings and cosine similarity, we measure reference drift (how closely outputs track their own tradition's authentic texts relative to the other two) and cross-tradition convergence (how similar outputs are across traditions). We find that while outputs remain closer to their own tradition's reference than to others, cross-tradition similarity is high (0.52-0.66) relative to what the traditions' genuine distance would predict, indicating partial homogenisation. Unexpectedly, prompting in the regional language (Hindi, Tamil, or Bengali) consistently reduced fidelity to the authentic tradition relative to English prompting, by as much as 27 percentage points for Rajasthani and Bengali traditions. We discuss this against conflicting prior results on multilingual prompting and argue it reflects a difference between eliciting general cultural diversity and simulating one narrow, lesser-documented oral tradition. We position this pilot as a lightweight, scalable complement to recent large-scale human-annotation studies of Indian cultural misrepresentation in LLM-generated stories, as part of a broader doctoral research program.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26123v1 Announce Type: new Abstract: Large language models (LLMs) are trained predominantly on English-language internet text that over-represents certain cultural narr…
サイト内本文

翻訳待ち:Can a Model Catch Its Own Hallucinations for Free?: Label-Free Doubt Signals Hold Their Own Against a Labelled Dataset for Abstention

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26121v1 Announce Type: new Abstract: Large language models state false facts as fluently as true ones, yet a model often "knows" internally when it is on shaky ground: the probability it assigns to its own answer tends to dip on the facts it gets wrong. The usual way to act on this, teaching a model to abstain rather than guess, requires a labelled dataset of right and wrong answers. We ask whether the model's own confidence, which is free and needs no labels, can do that job instead. We fine-tune each model (with LoRA) to answer when its frozen confidence is high and to say "I'm not sure" when it is low, using the signal alone and no correctness labels. Across six open-weights models (1B-8B, two families) on short-form factual question answering, with correctness adjudicated by an independent judge model, this label-free recipe holds its own against label-supervised abstention-tuning: at matched coverage we find no statistically detectable difference between the two. A control that drills hard examples instead of abstaining does not help, indicating the gain comes from calibration, not rote memorization. The signal's one blind spot is confidently wrong facts, which it cannot flag. A model's own doubt is thus a near-free substitute for a labelled dataset when teaching it when to abstain. Code and artifacts are available on request.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26121v1 Announce Type: new Abstract: Large language models state false facts as fluently as true ones, yet a model often "knows" internally when it is on shaky ground:…
サイト内本文

翻訳待ち:Recipes for Steering and Scaling LLMs via Sampling

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26120v1 Announce Type: new Abstract: Large Language Models (LLMs) are probabilistic models, typically defined by an autoregressive factorization. While recent work has begun to study richer target distributions beyond the base model, the sampling strategies remain highly inefficient. In this paper, we present a flexible and theoretically grounded framework for steering and scaling autoregressive LLMs with sampling. Within this framework, we describe two algorithms -- one based on Sequential Monte Carlo (SMC) and one based on Replica Exchange (RE) -- that steer generation toward powering, product or tilting of the base model distribution. We illustrate this framework through scaling the generation quality of LLMs without external supervision or reward models. Experimental results demonstrate our methods scale more favorably than Best-of-N and standard MCMC baselines. Overall, this paper offers a systematic recipe for probabilistic inference with LLMs via sampling.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26120v1 Announce Type: new Abstract: Large Language Models (LLMs) are probabilistic models, typically defined by an autoregressive factorization. While recent work has…
サイト内本文

翻訳待ち:DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMs

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26119v1 Announce Type: new Abstract: Whether large language models can be prompted to generate rhetorical fallacies on demand, and whether current safety post-training constrains this behavior, has received less attention than the related question of detecting fallacies in existing text. We close this gap with DeflectBench, evaluating 23,990 generations from four frontier models across three deflection strategies (whataboutism, ad hominem, red herring), seven prompt framings, and 80 claims spanning four controversy levels. Refusal is governed primarily by request structure rather than claim content. Per claim refusal varies by only 11 percentage points across the 80 claims, while a single prompt frame change can swing within model refusal by nearly 100 percentage points and switching the requested fallacy type can swing it by over 80 percentage points within explicit framings. An educational debate coach prompt framing collapses refusal to near zero across all four model families, but the bypassed behavior is not clean compliance. Models typically produce labeled compliance, naming the requested manipulation in the same response that contains it. The four models distribute differently across refusal, labeled compliance, soft refusal, and clean compliance. The code and dataset are released at https://github.com/ArtKanke/DeflectBench.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26119v1 Announce Type: new Abstract: Whether large language models can be prompted to generate rhetorical fallacies on demand, and whether current safety post-training…
サイト内本文

翻訳待ち:TreeGraft: Adaptive Multi-Drafter Grafting for Tree-Based Speculative Decoding

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26112v1 Announce Type: new Abstract: Speculative decoding accelerates large language model inference through a draft-then-verify paradigm. Building on this, tree-structured methods improve inference by organizing proposals into multiple candidate paths, increasing the accepted length. However, existing tree-structured methods use a single drafter for all drafting steps, creating a dilemma: a smaller drafter is fast but yields lower-quality trees, whereas a larger drafter improves tree quality but suffers from high latency. To address this, we propose TreeGraft, a multi-drafter framework in which drafters of different costs jointly construct a shared draft tree. TreeGraft uses the stronger drafter to rescore candidates by updating scores assigned by the weaker drafter, reselect grafting positions, and recover promising paths left unexplored. It also integrates stronger drafter expansions non-destructively, preserving existing branches that may still be accepted by the target model. Together, these designs improve the quality of the shared draft tree. To control the drafting cost, TreeGraft introduces a lightweight scheduler distilled from an offline value system to decide when to call the stronger drafter. Across 10 model pairs and 6 benchmarks, TreeGraft outperforms the better of the two fixed single-drafter endpoint strategies by 15.1% on average, reaching a maximum gain of 26.6%. Our code is available at https://anonymous.4open.science/r/TreeGraft-E983.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26112v1 Announce Type: new Abstract: Speculative decoding accelerates large language model inference through a draft-then-verify paradigm. Building on this, tree-struct…
サイト内本文

翻訳待ち:The Latent Diagnostic Taxonomy: A Framework for Constructing Classifiers and Diagnosing Their Decisions, Applied to Prompt Injection Detection

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26423v1 Announce Type: new Abstract: This paper proposes a framework for constructing a classifier as a safeguard layer, and for developing a complementary diagnostic that identifies which of the classifier's confident decisions can be trusted. This framework, the Latent Diagnostic Taxonomy, consists of (i) constructing a dimensionality-optimized classifier, in which the embedding dimensionality is empirically selected via cross-validated performance rather than fixed a priori, (ii) locating a relatively small set of latent support vectors (~ 29% of total training examples) representing influential prompts for identifying tokens that alter the classifier's predicted labels, and (iii) utilizing such tokens and their associated attack magnitudes for constructing a diagnostic taxonomy. This diagnostic taxonomy provides an end-to-end guideline for flagging prompts that require different treatments: rely Safely on the classifier's decision; flag Heuristic Bias and Heuristic Override cases; route Insufficient Context cases for further human/safety review. Applying the framework to a classifier trained on a public prompt injection dataset, we find that a substantial fraction of its confident decisions (~ 77%) are not robust to removing a single token, and that this brittleness separates into two distinct failure patterns: a confidence calibration failure and a genuinely exploitable shortcut. For each zone of the taxonomy, we also recommend strategies for remediating diagnosed prompts. We illustrate the framework as a series of steps, demonstrating how each step operates.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26423v1 Announce Type: new Abstract: This paper proposes a framework for constructing a classifier as a safeguard layer, and for developing a complementary diagnostic t…
サイト内本文

翻訳待ち:CG4AI: A Column Generation Framework for Training AI Models Under Constraints

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26375v1 Announce Type: new Abstract: Standard machine-learning training minimizes a loss function over a dataset, but does not guarantee that the resulting model will satisfy predefined rules or constraints on its outputs. In many real-world applications, ranging from autonomous systems to network routing, such guarantees are essential. We propose CG4AI, a framework that builds a convex combination of AI models while enforcing linear constraints on the combined output. A master linear program (LP) determines the optimal mixture weights, while a pricing subproblem generates new models guided by LP dual variables, focusing attention on the most violated constraints. A cutting-plane procedure extends feasibility guarantees beyond the training set. We apply CG4AI to two problems: (i) digit classification on MNIST, where we demonstrate four distinct uses of constraints, learning from constraints alone, improving adversarial robustness, correcting misclassified examples, and enforcing output relabeling; and (ii) the multi-commodity flow problem, where link capacity constraints are enforced on neural-network routing predictors. Experiments on MNIST and standard SNDLIB benchmark networks show that CG4AI reliably produces feasible predictors while achieving better accuracy than single-model baselines.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26375v1 Announce Type: new Abstract: Standard machine-learning training minimizes a loss function over a dataset, but does not guarantee that the resulting model will s…
サイト内本文

翻訳待ち:Beyond Capability Benchmarks: Learning Operational Fingerprints of LLM Cloud Services from Production Incident Metadata

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26332v1 Announce Type: new Abstract: Managed LLM services are now part of real production systems, but model selection and service planning still rely heavily on capability benchmarks that reveal little about operational behavior after deployment. We present Operational Embedding (OpEmbed), a framework for learning compact operational fingerprints of LLM cloud services from structured, privacy-preserving support-case metadata, without using case text. OpEmbed aggregates model--time windows into an eight-channel operational signature and learns a low-dimensional representation via temporal contrastive learning, cross-view reconstruction, and generational-ordinality regularization. Evaluated on more than 33,000 production support cases spanning seven LLM families over 26 months at Google Cloud, OpEmbed recovers interpretable family- and version-level structure, improves leave-one-model-out operational forecasting over non-learned baselines, remains useful under limited early-window data, and supports cross-model fault-type transfer. We report the practical lessons learned from building and evaluating this tool for model onboarding, support readiness assessment, and operational monitoring.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26332v1 Announce Type: new Abstract: Managed LLM services are now part of real production systems, but model selection and service planning still rely heavily on capabi…
サイト内本文

翻訳待ち:Privacy Without Regret: Differentially Private Inference-Time Alignment

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26324v1 Announce Type: new Abstract: Best-of-N (BoN) sampling is the simplest and most widely deployed inference-time alignment strategy, but it suffers from two distinct problems: reward hacking, in which the selected response exploits errors in the proxy reward model, and the absence of any privacy protection for the sensitive human preference data used to train that reward model. We show that a single intervention-adding calibrated noise to reward scores before selection-resolves both. Our first result, Private Best-of-N (PrivBoN), establishes that Gumbel noise at an appropriate scale simultaneously provides $\epsilon$-differential privacy and implements KL-regularized alignment. Whenever the privacy budget exceeds a critical threshold $\epsilon^*$, the privacy-mandated noise is the regret-optimal regularization, and privacy imposes zero additional alignment cost-matching the information-theoretic skyline of Huang et al. (2025). Because $\epsilon^*$ depends on an unknown coverage coefficient, we introduce Private Inference-Time Pessimism (PrivITP), which combines $\chi^2$-regularized rejection sampling with a two-phase Gaussian mechanism. PrivITP achieves ex-post $(\epsilon,\delta)$-DP with a privacy cost independent of the number of responses $n$, cleanly decouples the regularization parameter from the privacy parameter, and attains the skyline up to a noise-inflation term. Experiments across several language models, datasets, and reward models confirm our results: PrivBoN and PrivITP are scaling-monotonic (unlike BoN, which degrades past a critical $n$), and PrivITP matches or outperforms PrivBoN at equivalent privacy levels, with the largest gains in the strong-privacy regime.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26324v1 Announce Type: new Abstract: Best-of-N (BoN) sampling is the simplest and most widely deployed inference-time alignment strategy, but it suffers from two distin…
サイト内本文

翻訳待ち:Muon with Finite Newton-Schulz: The Smoothing Benefit in Nonsmooth Nonconvex Optimization

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26288v1 Announce Type: new Abstract: Muon has emerged as a strong optimizer for the matrix-valued parameters in large language model pretraining, approximately orthogonalizing its momentum with a few Newton-Schulz iterations. Existing theory either replaces this iteration with the exact polar factor it approximates, or treats its finite depth as an approximation error, and thus the iteration Muon actually runs can only hurt the guarantees. We show that finite Newton-Schulz can instead be beneficial for nonsmooth nonconvex optimization. To this end, we analyze Muon through the online-to-nonconvex conversion, which views the update rule as an online learner and converts its regret bound into a stationarity guarantee. The finite Newton-Schulz iteration smooths the discontinuous polar map into a Lipschitz map of the singular values, and Muon with finite Newton-Schulz can be regarded as an online learner with a smoothed spectral potential. This smoothing is exactly what the conversion needs: we prove that a Newton-Schulz depth growing only logarithmically in the target accuracy suffices for convergence to stationary points in nonsmooth nonconvex optimization, whereas Muon with the exact-polar update may fail to converge. The resulting sample complexity bounds match the best-known guarantees for nonsmooth nonconvex optimization and are optimal for smooth nonconvex optimization up to problem-dependent factors. The argument extends beyond Newton-Schulz to general spectral maps with the same smoothing property.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26288v1 Announce Type: new Abstract: Muon has emerged as a strong optimizer for the matrix-valued parameters in large language model pretraining, approximately orthogon…
サイト内本文

翻訳待ち:NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26222v1 Announce Type: new Abstract: Safety evaluation is critical for assessing whether aligned Large Language Models (LLMs) remain robust against jailbreak attacks. Existing automated testing methods, however, largely rely on response-level feedback: each candidate prompt typically requires generating a target-model response to evaluate its attack effectiveness. This process is expensive and, more importantly, provides only sparse guidance on strongly aligned models, where most candidates are rejected with the same failure outcome. This paper presents NeuronFuzz, a white-box fuzzing framework that exploits internal safety neurons as continuous execution feedback for LLM safety evaluation. A SafetyOracle converts safety-neuron activations into a continuous safety alarm score that serves as feedback for fuzzing and can be obtained during prefill, eliminating response generation from the fuzzing loop. To construct the SafetyOracle, NeuronFuzz uses template-invariant harmful and benign inputs and stability-aware selection to identify a compact set of safety neurons whose activations capture harmful-intent recognition. Moreover, since the safety alarm score is differentiable, NeuronFuzz uses its gradients to identify safety-sensitive template positions and a masked language model to generate fluent, context-compatible mutations while preserving original harmful payload and avoiding additional optimization variables. We evaluate NeuronFuzz across 21 text and multimodal models. Across five white-box source models, it achieves a 76-100% jailbreak discovery rate, outperforming baselines by up to 48 percentage points. Its optimized templates further transfer zero-shot to open-weight and six proprietary target models, achieving average ASR and top-5 ensemble ASR (EASR) of 69.6%/92.6% and 44.1%/60.0%, respectively.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26222v1 Announce Type: new Abstract: Safety evaluation is critical for assessing whether aligned Large Language Models (LLMs) remain robust against jailbreak attacks. E…
サイト内本文

翻訳待ち:SLM-Conditioned Hierarchical Relation Routing for Labeled Property Graph Learning

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26132v1 Announce Type: new Abstract: Labeled property graphs combine relational structure with heterogeneous textual and categorical properties attached to both nodes and relationships. Conventional graph neural networks typically represent these properties as static feature vectors, limiting their ability to determine which semantic evidence should influence message propagation for a particular prediction target. We propose SLM-Conditioned Hierarchical Relation Routing, an architecture that integrates a small language model directly into graph message selection. A topology GNN provides a stable structural representation and prediction anchor. For each target node, incident messages combine the neighbor's structural state, node-property encoding, relationship-property encoding, and relationship type. A parameter-efficient SLM processes structured graph soft tokens and produces a target-conditioned routing query. This query first selects relevant messages within each relationship type and subsequently routes information across relation-level summaries. The resulting representation provides a bounded residual update to the topology anchor, preserving structural evidence while allowing contextual semantic information to modify the prediction. The architecture supports interpretable analysis at both the neighbor and relationship-type levels and provides a general mechanism for integrating language-derived semantics into property-rich graph learning.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26132v1 Announce Type: new Abstract: Labeled property graphs combine relational structure with heterogeneous textual and categorical properties attached to both nodes a…
サイト内本文

翻訳待ち:Methodological and Conceptual Framework for 5D Multi-Table Analysis: A Unified Approach for Complex Data Reuse

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26149v1 Announce Type: new Abstract: Multi-table learning remains a major challenge in machine learning for healthcare and other complex information systems. Relational data combine several sources of complexity, including large data volume, high-dimensional variables, high-cardinality categorical features, complex inter-table dependencies, and repeated temporal observations. We introduce the Relational Hypergraph Transformer (RHT), a unified architecture that represents relational databases as hypergraphs, learns pentadimensional embeddings (PentE), and performs sparse relational attention with complexity proportional to the average relational degree rather than the square of the number of entities. We formally define the architecture, derive the complexity of its attention mechanism, and provide an open-source reference implementation. We evaluate RHT on the public Synthea synthetic electronic health record dataset using multi-label prediction of SNOMED CT condition codes per encounter, a task characterized by high categorical cardinality and long-tailed label distributions. Comparisons with tabular, relational, and temporal graph baselines show that RHT produces more semantically coherent embeddings while remaining computationally scalable. In this benchmark, the highest rare-code recall is achieved by XGBoost, whereas RHT attains the strongest embedding semantic coherence. We also report ablation studies quantifying the contribution of each architectural component. Clinical validation on MIMIC-IV is planned following PhysioNet credentialing. Source code and experimental protocols are provided in the accompanying repository.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26149v1 Announce Type: new Abstract: Multi-table learning remains a major challenge in machine learning for healthcare and other complex information systems. Relational…
サイト内本文

翻訳待ち:Large Models for Battery Prognostics and Health Management: A Review and Future Roadmap

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26111v1 Announce Type: new Abstract: Battery Prognostics and Health Management (BPHM) is critical for ensuring the safe, reliable, and cost-effective operation of batteries across electric vehicles, grid storage, and consumer electronics. Conventional BPHM approaches, including physics-based models and task-centric deep learning methods, face challenges in computational efficiency and parameterization, cross-domain generalization, dependence on extensive labeled run-to-failure data, and model interpretability. Recent Large Models (LMs), built upon Transformer architectures and self-supervised pre-training, offer a transformative new paradigm to overcome these long-standing bottlenecks. This review provides the first comprehensive survey of LM applications in BPHM, systematically examining how these models address challenges in the field. We begin by elucidating the foundational technologies enabling LMs, including Transformer architectures, self-supervised learning, large-scale multimodal datasets, and PEFT techniques. We then categorize recent progress along four critical dimensions: mitigating data scarcity, enhancing generalization and robustness, integrating domain knowledge for interpretability, and enabling system-level automation. Despite promising results, significant challenges remain across data accessibility, intelligence validation, trustworthiness, and deployment feasibility. To guide future research, we propose a roadmap focused on building collaborative data ecosystems, validating intelligence for industrial applications, enhancing trustworthiness with physics-informed designs, and enabling efficient on-device deployment. This review establishes a systematic approach to understand and advance LM-driven BPHM, providing researchers and practitioners with essential insights for developing next-generation battery management systems capable of safe, reliable, and autonomous operation throughout battery lifecycles.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26111v1 Announce Type: new Abstract: Battery Prognostics and Health Management (BPHM) is critical for ensuring the safe, reliable, and cost-effective operation of batte…
サイト内本文

翻訳待ち:EduRiskX: A Neuro-Symbolic Framework with F-Logic Reasoning for Early Academic Risk Prediction

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26107v1 Announce Type: new Abstract: Predicting students' academic risk in online education is crucial for enabling timely interventions that can improve retention and learning outcomes. However, existing models often suffer from limited early detection capability and insufficient interpretability, leading to a "black-box" trust crisis that hinders their adoption in real-world pedagogical settings. To address these challenges, we propose EduRiskX, a neuro-symbolic framework that integrates a temporal Transformer-based predictor with F-Logic symbolic reasoning. The neural component models longitudinal student activity sequences using temporal attention, class-weighted loss, and dynamic weekly truncation. Acting as a data-driven expert system, an F-Logic rule base -- grounded in established educational theories (Engagement Theory and Student Integration Model) to mimic the diagnostic logic of human educators -- is constructed exclusively from the training data. The neural risk probability and the symbolic confidence score are then combined through a logistic regression-based fusion mechanism that learns the relative contribution of each signal. Experiments on the Open University Learning Analytics Dataset (OULAD) using a strict 80/10/10 student-level split show that EduRiskX achieves an accuracy of 0.900 and an F1-score of 0.894 at the end of the semester (Week 38), with an average early detection week of 9.32 and a detection rate of 94.30 percent. Compared with state-of-the-art time-series models (PatchTST, iTransformer) and common deep learning baselines (LSTM, CNN), EduRiskX yields improved recall and earlier risk identification under identical conditions. Beyond predictive performance, the F-Logic module provides structured rule-based explanations linking predictions to observable behavioral patterns and educational theories.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.26107v1 Announce Type: new Abstract: Predicting students' academic risk in online education is crucial for enabling timely interventions that can improve retention and…
サイト内本文

翻訳待ち:XPress: Parallel Refinement for Diffusion Drafters in Speculative Decoding

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:XPress: Parallel Refinement for Diffusion Drafters in Speculative Decoding August 2, 2026 12 minutes Case study. A real GSM8K prompt decoded three ways under the same timing setup: autoregressive, the dFlash drafter alo…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • XPress: Parallel Refinement for Diffusion Drafters in Speculative Decoding August 2, 2026 12 minutes Case study. A real GSM8K prompt decoded three ways under the same timing setup…
AI デイリーブリーフィング | AI News Hub