AI News HubLIVE

創業融資動態

待翻譯:Gamescom highlights gaming boom amid AI concerns

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:https://p.dw.com/p/5JOAz AI is bringing significant challenges for the gaming industry, but it could also help significantly reduce costsImage: Political-Moments/IMAGO Earlier this month, gaming giant Electronic Arts wa…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • https://p.dw.com/p/5JOAz AI is bringing significant challenges for the gaming industry, but it could also help significantly reduce costsImage: Political-Moments/IMAGO Earlier thi…
站內正文

待翻譯:Toward Measuring AI's Effects on Skill Formation: The Stock-Formation Gap

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:--> [Submitted on 12 Apr 2026 (v1), last revised 9 Aug 2026 (this version, v3)] Title:Toward Measuring AI's Effects on Skill Formation: The Stock-Formation Gap View a PDF of the paper titled Toward Measuring AI's Effect…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • --> [Submitted on 12 Apr 2026 (v1), last revised 9 Aug 2026 (this version, v3)] Title:Toward Measuring AI's Effects on Skill Formation: The Stock-Formation Gap View a PDF of the p…
站內正文

待翻譯:Evaluate any agent framework with Amazon Bedrock AgentCore Evaluations

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Amazon Bedrock AgentCore Evaluations decouples agent evaluation from the framework you build on. As long as your agent emits OpenTelemetry telemetry, the service can score it, whether you use LangGraph, LlamaIndex, the OpenAI Agents SDK, Google ADK, the Claude Agent SDK, or Strands Agents. This post explains how the framework-agnostic contract works.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Amazon Bedrock AgentCore Evaluations decouples agent evaluation from the framework you build on. As long as your agent emits OpenTelemetry telemetry, the service can score it, whe…
站內正文

待翻譯:Apache DataFusion vs. DuckDB

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Apache DataFusion vs DuckDB Apache DataFusion and DuckDB are both fast, in-process analytical query engines. DataFusion is an embeddable Rust library designed to be extended. DuckDB is a self-contained database designed…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Apache DataFusion vs DuckDB Apache DataFusion and DuckDB are both fast, in-process analytical query engines. DataFusion is an embeddable Rust library designed to be extended. Duck…
站內正文

待翻譯:#1 on BABILong at 10M, our own gaming audit published

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:A deep dive into a living memory's forward pass - Sapience Labs Sapience Labs · technical deep dive · August 2026 A deep dive into a living memory's forward pass What actually happens when one AI session w…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • A deep dive into a living memory's forward pass - Sapience Labs Sapience Labs · technical deep dive · August 2026 A deep dive into a living memory's forward pass Wha…
站內正文

待翻譯:Claude Desktop can now easily run Qwen, DeepSeek and Kimi models — after Ollama’s first effort stalled

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Open-weight model runner Ollama has reintroduced an integration with Claude Desktop that lets users connect Anthropic’s app to models served The post Claude Desktop can now easily run Qwen, DeepSeek and Kimi models — after Ollama’s first effort stalled appeared first on The New Stack.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Open-weight model runner Ollama has reintroduced an integration with Claude Desktop that lets users connect Anthropic’s app to models served The post Claude Desktop can now easily…
站內正文

待翻譯:LangChain State of AI 2023

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discover how developers build LLM applications in 2023. Insights on popular models, vectorstores, retrieval strategies, and testing methods from LangSmith.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Discover how developers build LLM applications in 2023. Insights on popular models, vectorstores, retrieval strategies, and testing methods from LangSmith.
站內正文

待翻譯:Bruin Startup Program: Open-source data stack and AI data analyst for startups

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Bruin for Startups Open-source data stack and AI data analyst for early-stage startups. A Bruin engineer onboards you and sets everything up with open-source tools. Run it locally or self-host it, then start analyzing y…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Bruin for Startups Open-source data stack and AI data analyst for early-stage startups. A Bruin engineer onboards you and sets everything up with open-source tools. Run it locally…
站內正文

待翻譯:Code Documentation Quality, Measured

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Quality scores 100 = best · Line = minimum Three documentation quality scores out of 100. Each bar includes its minimum passing score. Codebase coverage Important systems and workflows are documented. Minimum passing sc…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Quality scores 100 = best · Line = minimum Three documentation quality scores out of 100. Each bar includes its minimum passing score. Codebase coverage Important systems and work…
站內正文

待翻譯:Nvidia Jetson Orin-guided Russian AI drone killed three civilians in Ukraine

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:5 Join the conversation Follow us Add us as a preferred source on Google A Russian Molniya drone carrying an Nvidia Jetson Orin module crashed and killed three civilians at a gas station in Zaporizhzhia last month after…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • 5 Join the conversation Follow us Add us as a preferred source on Google A Russian Molniya drone carrying an Nvidia Jetson Orin module crashed and killed three civilians at a gas…
站內正文

待翻譯:Velocity-coupled Representation Refinement for Satellite Orbit Prediction

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23728v1 Announce Type: new Abstract: Satellite orbit prediction, which aims to forecast future orbital trajectories from historical observations, is important for collision warning and safe space operations. With advances in time-series forecasting, learning-based methods have emerged as a promising solution for satellite prediction. In orbital dynamics, a satellite state is typically described by position and velocity, where position characterizes trajectory geometry and velocity reflects its instantaneous direction and rate of change. However, most existing methods mainly focus on temporal dependencies within position sequences while rarely exploiting the intrinsic coupling between position and velocity, which is essential for modeling satellite motion. To this end, we propose OrbitNet, a velocity-aware representation learning method for accurate satellite orbit prediction. It lifts conventional position-sequence forecasting to a position-velocity coupled representation learning paradigm by exploiting relationships among satellite state variables. Specifically, we develop a velocity-coupled representation refinement strategy to enhance positional representations through cross-variable interactions between position and velocity. We further introduce orbital segment modeling, which partitions historical trajectories into temporal segments and performs segment-level temporal learning to capture local motion variations and long-range evolution patterns. Extensive experiments show that OrbitNet outperforms large time-series foundation models and representative general forecasting methods under both in-domain evaluation on Starlink and zero-shot evaluation across six unseen satellite constellations. We expect this work to encourage further exploration of satellite-aware representation learning for trajectory time-series forecasting.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23728v1 Announce Type: new Abstract: Satellite orbit prediction, which aims to forecast future orbital trajectories from historical observations, is important for colli…
站內正文

待翻譯:GAP-Prompt: Gated Adaptive Prompting for Efficient Continual Learning

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23782v1 Announce Type: new Abstract: Continual learning faces the persistent challenge of catastrophic forgetting, where sequential task updates degrade previously acquired knowledge. While prompt-based methods integrated with pre-trained models offer a compelling solution by freezing the backbone, they often rely on static, task-level prompting strategies that overlook fine-grained intra-task diversity. In this paper, we propose Gated Adaptive Prompting (GAP-Prompt), a novel method that introduces instance-level adaptability to the prompting process. GAP-Prompt consists of three synergistic modules: (1) instance-conditioned gating, which dynamically determines optimal prompt injection layers for each individual image; (2) dynamic knowledge fusion, which performs instance-aware aggregation of current and historical prompts, enabling knowledge integration across tasks; and (3) shared prompt distillation, which anchors foundational knowledge in early shared layers to mitigate forgetting. Extensive evaluations on CIFAR-100, ImageNet-R, and CUB-200 benchmarks demonstrate that GAP-Prompt consistently achieves state-of-the-art performance. Notably, on the fine-grained CUB-200 dataset, GAP-Prompt reaches 87.29% accuracy, approaching the joint training upper bound (88.00%) and outperforming existing methods by a significant margin.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23782v1 Announce Type: new Abstract: Continual learning faces the persistent challenge of catastrophic forgetting, where sequential task updates degrade previously acqu…
站內正文

待翻譯:From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23660v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to provide prior causal knowledge for structural causal discovery, yet whether their direct-edge judgments and confidence can be trusted remains unclear. We systematically evaluate 12 instruction-tuned open-weight models across six benchmark causal graphs, five prompting strategies, and four confidence sources: verbalized, logit-based, cross-prompt agreement, and cross-model agreement. Under our language-only pairwise protocol, our evaluation yields three key findings. (i) LLM-based causal judgments are strongly recall-dominant: models predict overly dense graphs with many false-positive edges, while prompting mainly shifts the precision-recall trade-off rather than resolving overprediction. Gains from model scale diminish on the largest graphs and do not eliminate miscalibration. (ii) LLMs often capture causal relatedness without reliably identifying directness or orientation. Relative to published reference graphs, models misclassify 40.0% of indirect and 36.0% of reversed non-edges as direct edges, versus 28.2% of other non-edges. Moreover, 80.8% and 84.6% of these false positives receive verbalized confidence of at least 80%, revealing substantial overconfidence in structurally incorrect predictions. (iii) Conventional confidence estimates are unreliable, whereas agreement offers a more promising signal. Logit-based confidence frequently collapses near 1.0 regardless of correctness, while cross-prompt and cross-model agreement achieve better mean calibration and discrimination, though their advantages are not statistically significant after Holm correction. A benchmark-familiarity audit further identifies potential familiarity in five model-dataset pairs, all involving AsiaM. Overall, our results suggest LLMs are better viewed as sources of externally validated soft causal priors than as direct evidence of causal structure.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23660v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to provide prior causal knowledge for structural causal discovery, yet whether t…
站內正文

待翻譯:ESQ-Bench: A Multi-Tier Enterprise Oracle Benchmark for Evaluating NL2SQL Dialect Generalization and Silent Semantic Divergence

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23569v1 Announce Type: new Abstract: State-of-the-art Natural Language to SQL (NL2SQL) models report execution accuracy exceeding 89 percent on established benchmarks such as Spider and BIRD. However, these benchmarks rely on simplified academic schemas and open-source SQL dialects that do not reflect the complexity of enterprise database environments. We introduce ESQ-Bench, an Oracle-first NL2SQL benchmark with systematic complexity tiers and silent-divergence evaluation across three enterprise schema complexity tiers. We constructed and released six populated schemas (465 tables, 164,682 rows, zero empty tables) with identical seed data on Oracle, PostgreSQL, MySQL, and SQL Server, a four-metric evaluation harness (EM, EX, SR, SD), and 550 gold-validated question-query pairs (Tier-1: 95; Tier-2: 228; Tier-3: 227). Schema-linked prompting with GPT-4o shows monotonic execution-match degradation across tiers: 79.8, 60.3, and 57.2 percent EX on executed queries (June 2026), versus 75.6, 80.4, and 95.8 percent on an earlier 142-question pilot slice. EM stays below 7 percent tier-wide; operational silent-divergence reaches 73 to 99 percent among EX-passing queries. Failure analysis shows wrong-result semantics dominate at higher tiers. Claude Sonnet 4.6 with schema-linked prompts reaches 87.4, 74.9, and 68.7 percent EX (executed queries), exceeding GPT-4o schema-linked on every tier. GPT-4o zero-shot EX on executed queries (78.7, 73.5, and 77.8 percent) inverts schema-linked at Tiers 2 to 3 due to lower execution rates and survivor bias in the zero-shot versus schema-linked analysis. Local Llama 3.2 schema-linked reaches only 13.3 percent bank-wide EX (73 out of 550), underscoring the gap between closed API models and open-weight baselines on enterprise Oracle schemas.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23569v1 Announce Type: new Abstract: State-of-the-art Natural Language to SQL (NL2SQL) models report execution accuracy exceeding 89 percent on established benchmarks s…
站內正文

待翻譯:RENDER: Controlling Reader-Facing Evidence in LLM Memory Evaluation

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23568v1 Announce Type: new Abstract: Memory and RAG evaluations often treat the answering model's input as an implementation detail, even though systems may render the same history as a memory entry, summary, typed record, or raw excerpt. We introduce RENDER, a benchmark control that fixes the conversation while varying the reader-facing artifact. RENDER combines a five-level packet ladder, localizing when answer-bearing content enters the input, with deterministic templates approximating ChatGPT-style entries, LangChain summaries, MemGPT-style typed records, and raw conversation. On 500 LongMemEval questions and nine models, matched-budget resolved packets beat recency-truncated raw dialogue by 42.4-72.6 points. In deployed-style templates, best-worst spread is 24.6-48.8 points per model; under the primary scorer, ChatGPT-style entries have higher point estimates than raw conversation on 7 of 9 models. Judge rescoring preserves the positive aggregate effect, but model-specific significance is mixed. Three models scoring 0 percent on formal ledger packets answer the same facts from natural-language entries at 45.4-53.4 percent. The effect persists under retrieval noise and transfers to HotpotQA, suggesting that memory/RAG evaluations should report or control the reader-facing artifact.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23568v1 Announce Type: new Abstract: Memory and RAG evaluations often treat the answering model's input as an implementation detail, even though systems may render the…
站內正文

待翻譯:Amazon announced it's shuttering Mechanical Turk on Sept. 30

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Amazon service that Jeff Bezos called artificial AI is shutting down Skip Navigation Amazon announced it's shuttering Mechanical Turk, a platform that matches workers with small digital tasks, on Sept. 30. The platform…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Amazon service that Jeff Bezos called artificial AI is shutting down Skip Navigation Amazon announced it's shuttering Mechanical Turk, a platform that matches workers with small d…
站內正文

待翻譯:Show HN: Hacker News keeps flagging my startup – A multi-Platform search engine

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Bluhe History Sign in to save chats Your conversations will be saved so you can continue them later Toggle Theme Expand Video

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Bluhe History Sign in to save chats Your conversations will be saved so you can continue them later Toggle Theme Expand Video
站內正文

待翻譯:Agentic web search infrastructure startup Keenable raises $26M

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Artificial intelligence startup Keenable.ai Inc. said today it has exited stealth with $26 million in funding to try and revamp a web search infrastructure that was built for humans, rather than the billions of autonomous agents that many believe will soon dominate the internet. The seed round was backed by investors including Accel, Brightwing Capital, […] The post Agentic web search infrastructure startup Keenable raises $26M appeared first on SiliconANGLE.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Artificial intelligence startup Keenable.ai Inc. said today it has exited stealth with $26 million in funding to try and revamp a web search infrastructure that was built for huma…
站內正文

待翻譯:AI and Constitutions (From My Email)

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:“Dear Tyler, I enjoyed reading your notes on visiting Anthropic to advise on Claude’s constitution. Framing AI governance around the common law, case law (“Talmud”), and independent adjudication is a much more adaptive…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • “Dear Tyler, I enjoyed reading your notes on visiting Anthropic to advise on Claude’s constitution. Framing AI governance around the common law, case law (“Talmud”), and independe…
站內正文

待翻譯:Workspaces in LangSmith for improved collaboration and organization

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Workspaces in LangSmith lets enterprises separate resources between different teams, business units, or deployment environments.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Workspaces in LangSmith lets enterprises separate resources between different teams, business units, or deployment environments.
站內正文

待翻譯:Kids outlearn AI–and we still don't know why

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:People have been talking to each other for at least 100,000 years, as best we can tell. And in all that time, there has been only one thing in the world that could learn a human language to perfect fluency: a human chil…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • People have been talking to each other for at least 100,000 years, as best we can tell. And in all that time, there has been only one thing in the world that could learn a human l…
站內正文

待翻譯:The Design System as the Control Plane for AI-Generated UI

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:AI-assisted development has made it easier to generate frontend code quickly. A developer can ask for a form, a dashboard widget, a settings page, or a modal flow and get a working first draft in seconds. That speed is useful, especially when teams are moving through routine UI work. But speed creates a problem that’s […]

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • AI-assisted development has made it easier to generate frontend code quickly. A developer can ask for a form, a dashboard widget, a settings page, or a modal flow and get a workin…
站內正文

待翻譯:Building Self-Correcting Memory in OpenWiki

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Learn how OpenWiki uses evidence-backed claims to detect stale knowledge, reduce hallucinations, and build self-correcting memory for evolving codebases.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Learn how OpenWiki uses evidence-backed claims to detect stale knowledge, reduce hallucinations, and build self-correcting memory for evolving codebases.
站內正文

待翻譯:Ropedia launches next-gen wearable capture device for robotic AI training data

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Singapore-based robotics data infrastructure firm Ropedia Pte. Ltd. today announced the launch of HOMIE Gen2, the next generation of the company’s head-mounted wearable that records human movement data to train robots. Smart robotics requires tremendous amounts of high-quality data taken from various environments. Most of the training data used to build the artificial intelligence foundation […] The post Ropedia launches next-gen wearable capture device for robotic AI training data appeared first on SiliconANGLE.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Singapore-based robotics data infrastructure firm Ropedia Pte. Ltd. today announced the launch of HOMIE Gen2, the next generation of the company’s head-mounted wearable that recor…
站內正文

待翻譯:“You can rent a feature, but you can’t rent a foundation”: why MotherDuck bought the startup already powering its data pipelines

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:What do you do when a technology you’ve become dependent on belongs to someone else? You buy the startup behind The post “You can rent a feature, but you can’t rent a foundation”: why MotherDuck bought the startup already powering its data pipelines appeared first on The New Stack.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • What do you do when a technology you’ve become dependent on belongs to someone else? You buy the startup behind The post “You can rent a feature, but you can’t rent a foundation”:…
站內正文

待翻譯:Show HN: I built a search engine for 800 niche job boards

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Brian's Job Search 🥇 Discover your next job with Brian's Job Search. Optimized for Google, this is a user-friendly and efficient way to explore a wide range of career opportunities. Start your journey towards a fulfill…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Brian's Job Search 🥇 Discover your next job with Brian's Job Search. Optimized for Google, this is a user-friendly and efficient way to explore a wide range of career opportuniti…
站內正文

待翻譯:Jensen Huang: Land power and shell – The next critical resource for AI factories

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Land, power and shell: The next critical resource for AI factories. AI factories are the defining infrastructure of the AI era—where compute transforms energy and data into intelligence that powers every business, indus…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Land, power and shell: The next critical resource for AI factories. AI factories are the defining infrastructure of the AI era—where compute transforms energy and data into intell…
站內正文

待翻譯:Tolerance-Dependent Inspection Disagreement Between a Fixed CMM and a Portable Articulated-Arm CMM

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.21404v1 Announce Type: new Abstract: Fixed coordinate measuring machines (CMMs) and portable articulated-arm CMMs are often assigned to the same inspection task, but their nominal accuracy specifications do not show whether a change of instrument will preserve the disposition of a part. The question is not simply how far the two results differ, but whether that difference crosses the tolerance boundary. We examined this issue with recorded measurements of cylindrical, cubic, and spherical features under nominal 20 {\deg}C and 30 {\deg}C conditions. Repeated records and two roughness profiles without sufficient acquisition information were removed, leaving six dimensional and four form profiles. For each dimensional feature, the distances of the two system means from nominal define the exact tolerance interval in which the systems receive opposite direct labels. The fixed-CMM stream was approximately 11.2 {\mu}m higher than the articulated-arm stream at both conditions. All four form profiles fell on opposite sides of the recorded 10 {\mu}m upper limit. The dimensional disagreement intervals also overlapped strongly; their mean widths were 6.573 {\mu}m at 20 {\deg}C and 4.995 {\mu}m at 30 {\deg}C. The results clarify why an average difference between instruments is not, by itself, a measure of substitution risk. The proposed tolerance map identifies the feature-tolerance combinations for which instrument choice can change the recorded inspection label and, therefore, where a controlled equivalence study and a task-specific uncertainty budget are needed before substitution.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.21404v1 Announce Type: new Abstract: Fixed coordinate measuring machines (CMMs) and portable articulated-arm CMMs are often assigned to the same inspection task, but th…
站內正文

待翻譯:Selective Cross-View Consistency for World Action Models: Held-Out Viewpoint Robustness Without Test-Time Camera Information

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.21402v1 Announce Type: new Abstract: World action models (WAMs) jointly denoise future video frames and robot actions, and the video prior is expected to generalize their control. Camera viewpoint change remains one of their hardest perturbation axes. We study a question specific to this model class: when training with same-state cross-view image pairs, on which output coordinates should a consistency loss be imposed? The WAM denoising target mixes view-covariant coordinates, namely the predicted future scene, with view-invariant coordinates, namely the action chunk, future proprioception, and value. We show that consistency applied to the covariant block is provably harmful, shrinking legitimate view-specific content to a fraction $1/(1+4\lambda)$ of its true value, and we verify this shrinkage law in controlled experiments. Selective cross-view consistency (SCVC) therefore constrains only the invariant block, requires no camera labels, extrinsics, depth, or view synthesis at training or test time, and leaves the deployment interface unchanged. We introduce a carve-and-hold-out evaluation protocol on the LIBERO-Plus camera track that separates a distribution-matched ceiling from genuine interpolation and extrapolation to held-out viewpoints, with a matched pair-trained control isolating the effect of the consistency term from pair exposure. On held-out orbital viewpoints beyond the training envelope, SCVC improves closed-loop success over the matched control by 12.2 points (95% CI [7.4, 17.0]; +15.5, CI [11.7, 19.4], under an independent second seed) -- an effect two further camera axes replicate -- while interpolation within the envelope shows no gain in either seed (-1.2 and -4.3 points) and in-distribution competence is preserved (-0.6, -0.2). We also report a cross-backbone audit showing that published camera-robustness numbers are confounded by wrist-camera pose stability.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.21402v1 Announce Type: new Abstract: World action models (WAMs) jointly denoise future video frames and robot actions, and the video prior is expected to generalize the…
站內正文

待翻譯:Boosting Knowledge-based Visual Question Answering with Structured Context Reasoning

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.21431v1 Announce Type: new Abstract: Knowledge-based Visual Question Answering aims to answer questions about an image by integrating external knowledge with visual and textual information. Recent approaches often rely on in-context learning to prompt Large Language Models (LLMs) with multimodal context in a zero-shot or few-shot manner. However, we observe that directly concatenating heterogeneous visual descriptions and retrieved knowledge into long, unstructured prompts often degrades reasoning performance, due to both excessive irrelevant context and the lack of explicit relational structure. In this paper, we propose an LLM-based Structured Context Reasoning (SCoRe) framework that infers both explicit and implicit relationships for prediction. SCoRe consists of three stages: Context Acquisition, which generates diverse visual notes and retrieves explicit knowledge via an efficient two-stage multimodal retrieval strategy; Context Selection, which filters relevant visual, explicit, and implicit knowledge using LLM-guided selection; and Context Compression, which performs Relational Logic Distillation (RLD) to transform raw text into explicit entity-relation triplets. These relational triplets serve as a concise and structured prompt for final answer prediction. Extensive experiments on the OK-VQA and A-OKVQA benchmarks demonstrate that SCoRe consistently outperforms state-of-the-art methods.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.21431v1 Announce Type: new Abstract: Knowledge-based Visual Question Answering aims to answer questions about an image by integrating external knowledge with visual and…
站內正文

待翻譯:Few-Shot Cross-Dataset Adaptation for Tuberculosis Detection Using DenseNet

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.21427v1 Announce Type: new Abstract: Tuberculosis (TB) is one of the most common and dangerous bacterial ailments. Every year, it causes a large number of deaths worldwide. Although many deep learning models can detect tuberculosis from chest X-rays quite accurately, severe domain shift across datasets makes the task challenging. Different imaging protocols, patient demographics, and equipment across domains make the task of generalization difficult. In real-world settings, a model may perform well on one dataset but show a noticeable drop in performance when tested on another. In this work, we address this domain adaptation challenge through a few-shot scaling study. A controlled cross-dataset evaluation is presented in this paper using TBX11K as the source domain and the Mendeley TB dataset as the target domain. It is investigated how varying the number of target samples affects model performance under three training regimes: frozen backbone adaptation, full fine-tuning of a source-pretrained DenseNet121 model, and training from scratch. The results indicate that the model can perform well even with limited data and can achieve 98.36\% accuracy with just 75 labeled samples per class. The adaptation curves demonstrate how fine-tuning effectively mitigates domain shift. These findings establish full fine-tuning of pretrained models as a highly effective and practical strategy for mitigating domain shift in low-resource clinical deployment scenarios.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.21427v1 Announce Type: new Abstract: Tuberculosis (TB) is one of the most common and dangerous bacterial ailments. Every year, it causes a large number of deaths worldw…
站內正文

待翻譯:Mitigating Bias in Large Vision-Language Models via Counterfactual Ensemble Decoding

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.21415v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) have achieved remarkable performance across a wide range of tasks; however, they often inherit social biases from their training data, resulting in biased behavior when processing portraits from different social groups. Existing debiasing approaches typically compare token probabilities between the original and biased generations during decoding, but they are fundamentally limited by their reliance on a single, stereotyped viewpoint and fail to account for the diversity of social perspectives. Inspired by the social science principle that diversity fosters fairness, we propose Counterfactual Ensemble Decoding (CED), a novel framework that constructs multi-group counterfactual perspectives within the visual representation space and integrates them during decoding to promote equitable model behavior. CED first performs counterfactual steering in the visual space by identifying semantic directions associated with each social group and generating counterfactual representations along these directions, thereby offering diverse perspectives that disrupt stereotypical narratives. During decoding, CED locates the decoder layer exhibiting the greatest divergence among these perspectives and ensembles their token distributions using uncertainty-aware weights, prioritizing high-confidence tokens from different groups to yield a more balanced probability distribution that guides fairer generation. Extensive experiments on three social bias evaluation benchmarks demonstrate that \tool achieves substantial improvements over leading baselines, reducing bias by up to 47.97% across scenarios involving occupations, descriptors, and persona traits. Moreover, CED also preserves the core capabilities of the original model with minimal degradation.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.21415v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) have achieved remarkable performance across a wide range of tasks; however, they often inherit…
站內正文

待翻譯:Wazobia Eval: A Benchmark for Nigerian Pidgin Emotion Understanding, Sarcasm Detection, and Cultural Reasoning

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.21369v1 Announce Type: new Abstract: Nigerian Pidgin is one of Africa's most widely spoken languages, yet remains severely underrepresented in language model evaluation. Existing benchmarks primarily focus on translation, transcription, or generic sentiment analysis, leaving critical aspects of culturally grounded language understanding unmeasured. We introduce Wazobia Eval, a benchmark for evaluating Nigerian Pidgin emotion understanding, sarcasm detection, and cultural reasoning. The benchmark is built on a manually annotated dataset containing over 550 examples and a 16-category emotion taxonomy designed to capture culturally specific emotional registers that are not represented in conventional sentiment frameworks. Wazobia Eval provides standardized evaluation protocols and benchmark tasks for assessing model performance on nuanced Nigerian language understanding. We present the benchmark design, annotation methodology, taxonomy development process, and preliminary pilot evaluation results. Our goal is to provide foundational evaluation infrastructure for Nigerian language AI and establish a reproducible benchmark for future research. The dataset is publicly available at https://huggingface.co/WAZOBIALABS.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.21369v1 Announce Type: new Abstract: Nigerian Pidgin is one of Africa's most widely spoken languages, yet remains severely underrepresented in language model evaluation…
站內正文

待翻譯:There Is No Neutral Harness: Modern LLM Leaderboards Are Manufactured by Config-Fragile Items

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.21382v1 Announce Type: new Abstract: Multiple-choice benchmarks fix the questions and the correct answers, but not the harness: the order of the options, the wording of the prompt, and whether a language model's answer is read from generated text or from per-option likelihoods. Work on this harness sensitivity reports it as aggregate score variance, leaving unexamined which items the variance falls on and whether they are the items that separate one model from the next. We treat the evaluation harness of large language models (LLMs) as an independent variable and resolve its effect to single items. We introduce the \textit{fragility grid}: 12 open-weight instruction-tuned LLMs from 4 families answer the same 3{,}679 items from 4 benchmarks (ARC, HellaSwag, MMLU, TruthfulQA) under 26 equally defensible harness configurations, recording one correctness bit for every model, item, and configuration. The comparison is matched, since the items, the weights, and the greedy decoding stay fixed while only the harness varies. Under the grid a model's score is a band rather than a point: gemma4-31b scores between 31 and 89 percent depending only on the harness. Three results follow. On the items that two adjacent models both answer stably the pair is tied, and config-fragile items carry 95.7 percent of a pair's gap on average. Four of the 12 models reach rank one under some configuration, so the harness selects the winner. Item discrimination, the property that benchmark-compression methods maximize, correlates with fragility at 0.28 (95 percent CI 0.25 to 0.30), so compression keeps the fragile items rather than removing them. The scoring choice, not the option order that protocols usually fix, is the load-bearing axis. We release the per-item records and the analysis script, from which every number regenerates on a CPU in seconds, and we position the fragility grid as a check a leaderboard can run before it reports an order.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.21382v1 Announce Type: new Abstract: Multiple-choice benchmarks fix the questions and the correct answers, but not the harness: the order of the options, the wording of…
站內正文

待翻譯:RIACT: A Responsible AI System for Personalized Study Habit Tracking and Early Burnout Signal Detection in University Students

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.21379v1 Announce Type: new Abstract: Student burnout is highly prevalent in higher education, with reported rates ranging from 12% to over 70% and consistently exceeding those of the working population - yet it is typically identified only retrospectively, after academic decline has already occurred. A contributing factor is that students have little structured visibility into their own study behaviour, and existing productivity tools record activity without interpreting it. This paper presents RIACT (Record, Insight, Analyze, Coach, Track), a web-based application that combines structured study session logging with a hybrid AI architecture to surface personalized insights and early burnout signals. Students log sessions by location and time; the system computes net focus time by accounting for breaks, detects burnout signals through transparent, deterministic rules operating on week-over-week behavioural comparisons, and uses a large language model - constrained to a fixed output schema - to contextualize patterns and generate personalized recommendations. The design embeds responsible AI principles throughout: warnings are governed by auditable rules rather than model judgement, all output is framed as an observation rather than a diagnosis and data collection is limited to self-logged behavioural fields. We describe the system's design rationale, situate it within the literature on student burnout and explainable AI in education and propose an evaluation framework for validating its behavioural signals against established burnout instruments.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.21379v1 Announce Type: new Abstract: Student burnout is highly prevalent in higher education, with reported rates ranging from 12% to over 70% and consistently exceedin…
站內正文

待翻譯:Quintessent bags $40M to develop lasers for AI clusters

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Optical interconnect startup Quintessent Inc. today announced that it has raised $40 million in funding to ramp up production of its hardware. Cycle Capital led the Series A raise. It was joined by Goldman Sachs, Susquehanna International Group, Osage University Partners and several others. The investment follows a $11.4 million seed round in 2024. Data […] The post Quintessent bags $40M to develop lasers for AI clusters appeared first on SiliconANGLE.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Optical interconnect startup Quintessent Inc. today announced that it has raised $40 million in funding to ramp up production of its hardware. Cycle Capital led the Series A raise…
站內正文

待翻譯:AI text watermarking and quality loss

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:After the EU AI Act’s transparency guidelines went into effect, and following the announcement that most major AI labs will soon be watermarking their text output to comply, there has been some debate about how much of…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • After the EU AI Act’s transparency guidelines went into effect, and following the announcement that most major AI labs will soon be watermarking their text output to comply, there…
站內正文

待翻譯:The Teaser Period: Why the AI Boom Is Built to Break

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Groundbreaker Aug 20, 2026 Nothing looked wrong in the summer of 2006. Home prices had risen for the better part of a decade. Delinquencies were near historic lows. Credit spreads were tight, the ratings held, and the s…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Groundbreaker Aug 20, 2026 Nothing looked wrong in the summer of 2006. Home prices had risen for the better part of a decade. Delinquencies were near historic lows. Credit spreads…
站內正文

待翻譯:Mathematical Theories Could Be the Key to Explainable AI Systems

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Kodamai, an enterprise AI startup, is addressing the growing concerns around the explainability and governance of AI systems by applying mathematically grounded theories.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Kodamai, an enterprise AI startup, is addressing the growing concerns around the explainability and governance of AI systems by applying mathematically grounded theories.
站內正文

待翻譯:Ode With Anthropic Makes First Acquisition to Expand Enterprise AI

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The standalone company, backed by Wall Street firms, is looking to boost its AI technology implementation capabilities.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • The standalone company, backed by Wall Street firms, is looking to boost its AI technology implementation capabilities.
站內正文

待翻譯:Canonical backs quest to translate mountains of C into safe Rust with AI

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Bungs banknotes at Bristol boffins to find out if mature code survives the machine

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Bungs banknotes at Bristol boffins to find out if mature code survives the machine
站內正文

待翻譯:Hybrid Search in Qdrant

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:A search result can look plausible and still be wrong. Dense retrieval can return a document on the right topic but miss an exact identifier copied into the query. Sparse retrieval can miss a relevant document when the query describes it with terms the corpus doesn’t use. Either way, your logs record a successful query. Hybrid search runs dense and sparse retrieval over the same query, then merges their result lists. Dense retrieval adds semantic similarity, so paraphrases can rank together. Sparse retrieval adds weighted term matching for exact words and identifiers.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • A search result can look plausible and still be wrong. Dense retrieval can return a document on the right topic but miss an exact identifier copied into the query. Sparse retrieva…
站內正文

待翻譯:Best GPU Neoclouds 2026: CoreWeave, Nebius, Lambda, Crusoe, and Groq Ranked by Published Pricing and Contracted Power

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The five largest GPU neoclouds now run on very different models. CoreWeave and Nebius report to the SEC; Lambda and Crusoe are private and heading toward IPOs; Groq rebuilt itself as an inference cloud after licensing its LPU technology to NVIDIA. This comparison checks each provider's live rate card, Q2 2026 financials, active and contracted gigawatts, anchor contracts, and SemiAnalysis ClusterMAX tier. Nebius posts the lowest H100 rate and the only published B300 price, Lambda has the cheapest B200, Crusoe is the only one with AMD on its card, and CoreWeave commands a 10–15% premium as the sole Platinum-rated provider. Figures verified August 21, 2026. The post Best GPU Neoclouds 2026: CoreWeave, Nebius, Lambda, Crusoe, and Groq Ranked by Published Pricing and Contracted Power appeared first on MarkTechPost.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • The five largest GPU neoclouds now run on very different models. CoreWeave and Nebius report to the SEC; Lambda and Crusoe are private and heading toward IPOs; Groq rebuilt itself…
站內正文

待翻譯:A Dataset-Centric Benchmark of Deep Learning Methods for Grape Leaf Disease Classification and Detection

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.20608v1 Announce Type: new Abstract: Grape leaf disease recognition is important for precision agriculture, enabling early diagnosis, timely intervention, and improved vineyard management. Although deep learning has achieved strong results, many studies rely on few datasets, often acquired under controlled conditions, and may not reflect real vineyard challenges such as complex backgrounds, variable illumination, occlusion, leaf pose, disease severity, and device differences. This paper presents a dataset-centric benchmark of deep learning methods for grape leaf disease classification and detection. We analyze publicly available datasets in terms of disease categories, annotation types, acquisition conditions, image characteristics, class distributions, provenance, and task suitability. Representative models are evaluated in three settings: image-level classification, region-level classification, and object detection. Classification is assessed using accuracy, while detection is evaluated using mAP@50 and mAP@50:95. Cross-dataset experiments further examine transfer between datasets with compatible disease categories but different visual and annotation characteristics. Results show near-saturated classification performance on several controlled or derivative datasets, greater difficulty on heterogeneous datasets, and substantial variation in detection performance across annotation settings. Cross-dataset performance drops sharply, especially for object detection, indicating that shared disease labels do not necessarily define equivalent recognition tasks. The benchmark emphasizes dataset provenance, realistic field evaluation, annotation compatibility, and external validation for reliable vineyard disease recognition.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.20608v1 Announce Type: new Abstract: Grape leaf disease recognition is important for precision agriculture, enabling early diagnosis, timely intervention, and improved…
站內正文

待翻譯:Learning Prostate Anatomy at Test Time for Cancer Detection in Micro-Ultrasound

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.20557v1 Announce Type: new Abstract: Domain shift across clinical centers using different imaging hardware or acquisition protocols remains a fundamental barrier to deploying deep learning models for prostate cancer (PCa) detection. Existing test-time adaptation (TTA) methods address distribution shift through entropy minimization or augmentation-based self-supervision, correcting for statistical differences in image appearance but ignoring the anatomical structure of the target domain. We propose ANT, a segmentation-guided TTA framework that adapts a pretrained cancer detection encoder to the target domain by solving an auxiliary prostate segmentation task at test time, supervised by pseudo-masks from a frozen pretrained segmentation network. By aligning encoder representations to prostate anatomy in the target domain, ANT corrects domain-specific feature drift while preserving cancer-discriminative structure. The model was trained on 693 patients imaged with an earlier-generation micro-ultrasound scanner in a multi-center clinical trial, and evaluated on 118 patients acquired with a newer-generation system across two centers in another clinical trial. Under a leave-one-center-out protocol with identical evaluation conditions across all methods, ANT improves mean AUC by 2.9% and 3.6% at the biopsy-core and patient levels, respectively, over no adaptation, outperforming TTA baselines. Code is available at: https://github.com/ObedDzik/ant.git.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.20557v1 Announce Type: new Abstract: Domain shift across clinical centers using different imaging hardware or acquisition protocols remains a fundamental barrier to dep…
站內正文

待翻譯:Grounded-Exo2Ego: Structured Semantic Grounding for Robust Exocentric-to-Egocentric Video Generation

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.20534v1 Announce Type: new Abstract: Generating egocentric video from a single exocentric video is an emerging and important topic for AR/VR and physical AI. Compared with conventional novel view synthesis, exo-to-ego generation is a significantly harder task because the standard geometric conditioning becomes highly unreliable under extreme view changes and large unobservable regions. We present Grounded-Exo2Ego, a principled framework that addresses these challenges at both the architectural and data levels. Architecturally, Grounded-Exo2Ego is a dual-branch video diffusion model that couples a geometric anchoring branch, which conditions the generation on the rendering of a 3D reconstruction, with a novel semantic grounding branch, which goes beyond the prevailing geometry-based approach and improves quality by synthesizing challenging regions based on object-level context. Additionally, we found that the overlooked issue of camera-reconstruction misalignment severely undermines exo-to-ego learning. We thus introduce a camera re-localization algorithm that resolves this issue and substantially improves quality across all metrics. We further develop a fully automated synthetic data engine that generates and renders rigged 3D characters in procedurally generated environments. Evaluation on the challenging EgoExo4D dataset shows that our method outperforms recent state-of-the-art approaches by large margins across all metrics. Detailed ablations validate improvements from each of our contributions at both the data and architectural level.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.20534v1 Announce Type: new Abstract: Generating egocentric video from a single exocentric video is an emerging and important topic for AR/VR and physical AI. Compared w…
站內正文

待翻譯:Aggregating Visual Information with Optimal Transport for VideoLM Token Compression

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.20473v1 Announce Type: new Abstract: Video language models process videos as dense visual-token sequences with substantial representational redundancy. Compressing these sequences is therefore essential for reducing the visual-token burden on language-model decoding. The central challenge is to preserve visual information dispersed across frames under such compression. To this end, we introduce Aggregating Visual Information with Optimal Transport (AVIOT), which casts video token compression as transporting a dense empirical measure of frame observations onto a compact target measure. The resulting source-to-target coupling induces a distribution over source observations for each target support, directly specifying how the compressed video representation is constructed. We further adapt this construction along task and spatial axes. Question conditioning modulates the transport cost between source frames and target supports, while influencing how many supports are allocated to each temporal segment, thereby directing representation capacity toward question-relevant content. At multiple spatial granularities, AVIOT computes region-specific temporal transport plans and adaptively fuses the representations they yield, allowing different regions within the same compact representation to draw from different moments. Evaluations across varying compression ratios show that AVIOT matches or outperforms the uncompressed baseline on multiple video-understanding benchmarks while retaining strong performance at higher compression ratios.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.20473v1 Announce Type: new Abstract: Video language models process videos as dense visual-token sequences with substantial representational redundancy. Compressing thes…
站內正文

待翻譯:Inhibitory Attention for Clinical Long-Context Reasoning: Characterizing and Mitigating Lost-in-the-Middle Effects in EHR Processing

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.20348v1 Announce Type: new Abstract: Electronic health records now routinely exceed 100,000 tokens per patient. Yet large language models exhibit the lost-in-the-middle (LitM) effect: information near the center of a long context is retrieved less reliably than information near the edges. In clinical use this is not benign: the single most consequential fact in a note can sit at its center. We term this the clinical lost-in-the-middle (CLitM) problem, give its first systematic characterization using MedAlign, and compare context-selection strategies as remedies. Across 2,196 instruction-response pairs and six language models, we observe a 21.9 percentage-point gap between peak accuracy (59.5%, 95% CI [46.3, 71.0], 20-30% decile) and trough accuracy (37.6% [23.2, 52.5] at 70-80%); 67.8% of reference answers fall between the 10th and 90th percentiles of the EHR timeline, inside the CLitM trough. We introduce Query-Conditioned Clinical Suppression (QCCS), a lightweight query-conditioned selection gate, and evaluate it against BM25, BM25 with section-header filtering, dense retrieval, and cross-encoder reranking (N=83 held-out instructions). With Qwen2.5-7B-Instruct (16k context), QCCS outperforms all five comparators under LLM-as-judge scoring: for middle-position instructions QCCS reaches 16.7% versus BM25 3.3%, cross-encoder 0.0%, dense 0.0%, and full context 6.7%; overall QCCS reaches 25.3% versus at most 3.6% for retrieval-only comparators. This advantage is not explained by retrieval recall: at k=20, BM25 retrieves the gold evidence sentence in 98.8% of instructions (QCCS 34.9%), yet retrieval arms stay at most 2.6% accurate even when they retrieve it, whereas QCCS reaches 25.0% even when it does not. In this proof-of-concept evaluation, query-aligned context selection predicts EHR instruction-following accuracy better than gold-sentence retrieval recall.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.20348v1 Announce Type: new Abstract: Electronic health records now routinely exceed 100,000 tokens per patient. Yet large language models exhibit the lost-in-the-middle…
站內正文

待翻譯:Who Do Language Models Think Is Competent? A Mechanistic Analysis of Occupational Bias

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.20347v1 Announce Type: new Abstract: Language models (LMs) often pass behavioral bias evaluations, but it remains unclear whether they no longer represent the underlying associations that give rise to biases, or have merely learned not to express them. In this study, we show that representational biases are often detectable, even when behavioral biases are not visible. We introduce a causal framework that decomposes occupational bias into two measurement points: a model's internal representation of a user's competence, and its observable outputs. We derive steering vectors for representations of user expertise, and verify that they causally mediate model behavior in both a question-answering task and a hiring task. Applying this framework to several open-weight models, we find that demographic attributes, such as gender, race, and socioeconomic status, influence a model's representation of user expertise, even in cases where behavioral metrics detect no disparity between demographics. We show that these model representations can influence downstream behavior under intervention, suggesting failure modes that behavioral metrics alone may not detect.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.20347v1 Announce Type: new Abstract: Language models (LMs) often pass behavioral bias evaluations, but it remains unclear whether they no longer represent the underlyin…
站內正文

主題導航

創業融資 — AI 主題新聞 | AI News Hub