跳到主要内容
AI News HubLIVE
公开文章 107采集文章 117可信度 90刷新频率 30 分钟
健康状态 健康来源类型 研究原文权限 官方原文最近入库 2026-09-29ID apple-ml-research运行状态 已启用

Official research source; confirm reuse terms before enabling full body display.

最新公开文章

待翻译:The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:When language models reason in chain-of-thought or exchange free-text intermediates, they serialize structured information into natural language. How much tree-structured compositional content survives this bottleneck? We propose a round-trip protocol that answers this question empirically for tree-structured expressions. A generator converts a procedurally generated arithmetic expression into a word problem, a separate extractor recovers the expression from the word problem alone, and symbolic equivalence provides an exact oracle. Evaluating all pairwise combinations of sixteen models yields…

Apple Machine Learning Research站内正文待翻译:The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models

待翻译:Faster Rates for Federated Variational Inequalities

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:In this paper, we study federated optimization for solving stochastic variational inequalities (VIs), a problem that has attracted growing attention in recent years. Despite substantial progress, a significant gap remains between existing convergence rates and the state-of-the-art bounds known for federated convex optimization. In this work, we address this limitation by establishing a series of improved convergence rates. First, we show that, for general smooth and monotone variational inequalities, the classical Local Extra SGD algorithm admits tighter guarantees under a refined analysis…

Apple Machine Learning Research站内正文待翻译:Faster Rates for Federated Variational Inequalities

待翻译:A Practical Recipe for Semi-Supervised Federated ASR: Online Pseudo-Labels with Server Update Stabilization

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Semi-supervised federated learning (SSFL) trains models on clients’ unlabeled data using a teacher to generate pseudo-labels, with a small labeled seed dataset on the server. Automatic Speech Recognition (ASR) is particularly fragile here: pseudo-label errors compound across the output sequence and across training rounds into divergence, leaving a large gap to fully-supervised FL. We show that closing this gap turns on two coupled design axes—the teacher (which model generates the pseudo-labels) and the anchor (the server-side updates on labeled data that stabilize training). On the teacher…

Apple Machine Learning Research站内正文待翻译:A Practical Recipe for Semi-Supervised Federated ASR: Online Pseudo-Labels with Server Update Stabilization

待翻译:Compressing Streaming Neural Audio Encoders via Latent-Space Distillation

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:System-wide Dictation on Apple devices runs entirely on-device, and the speech it transcribes reaches the foundation model through a tokenizer: an encoder that maps short windows of waveform onto the representation the language model reads. Because that model is sparsely activated under Instruction-Following Pruning, only a small subset of its experts occupies DRAM at any time, so the always-on tokenizer competes for the same memory, and its parameter count bears directly on power and latency. In this work we study how to compress such a tokenizer by distillation, taking as the supervision…

Apple Machine Learning Research站内正文待翻译:Compressing Streaming Neural Audio Encoders via Latent-Space Distillation

待翻译:How to Guide Your Language Flow

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:We introduce a new method to guide flow matching models. Our approach, which we call probe guidance, uses the frozen internal states of an existing diffusion model to construct a guidance signal. This works using a similar principle as autoguidance, but eliminates the need for an additional forward pass at inference time and provides a reliable path to ensure that the weak and strong model share similar dynamics. We apply and benchmark this method on continuous diffusion language models, where probe guidance sets a new state-of-the-art performance on unconditional generation. When applied to a…

Apple Machine Learning Research站内正文待翻译:How to Guide Your Language Flow

待翻译:Dynamically Scaled Activation Steering

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Activation steering has emerged as a powerful method for guiding the behavior of generative models towards desired outcomes such as toxicity mitigation. However, most existing methods apply interventions uniformly across all inputs, degrading model performance when steering is unnecessary. We introduce Dynamically Scaled Activation Steering (DSAS), a method-agnostic steering framework that decouples when to steer from how to steer. DSAS adaptively modulates the strength of existing steering transformations across layers and inputs, intervening strongly only when undesired behavior is detected…

Apple Machine Learning Research站内正文待翻译:Dynamically Scaled Activation Steering

待翻译:How Value Induction Reshapes LLM Behaviour

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Conversational Large Language Models are post-trained on language that expresses specific behavioural traits, such as curiosity, open-mindedness, and empathy, and values, such as helpfulness, harmlessness, and honesty. This is done to increase utility, ensure safety, and improve the experience of the people interacting with the model. However, values are complex and inter-related – inducing one could modify behaviour on another. Further, inducing certain values can make models more addictive or sycophantic through language used in the generations, with a potential detrimental effect on the…

Apple Machine Learning Research站内正文待翻译:How Value Induction Reshapes LLM Behaviour

待翻译:DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Diffusion large language models are a compelling alternative to autoregressive models, yet existing RL methods for diffusion treat all denoising steps as equally important and rely on biased, high-variance likelihood estimates. We identify two fundamental weaknesses: the absence of temporal credit assignment across the denoising trajectory, and the systematic bias of mean-field likelihood estimates used for policy optimization. To address these, we propose Denoising-Aware Credit Assignment for GRPO (DACA-GRPO), a lightweight, plug-and-play enhancement for any GRPO-style trainer. DACA-GRPO…

Apple Machine Learning Research站内正文待翻译:DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models

待翻译:Glyph: A Multi-Strategy Agentic System for Column Description and Sensitivity-Ontology Tagging of Enterprise Data Catalogs

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Enterprise data lakes accumulate tables faster than human stewards can document or classify them, leaving columns with missing descriptions and unassigned governance labels. This documentation debt undermines data discovery, access control, and regulatory compliance. We present Glyph, a production system that frames two coupled problems, column description generation and column type annotation for data classification, as cooperating LLM agents orchestrated as stateful graphs. The Descriptor grounds generation in the pipeline source code that produces each column, retrieved on demand from an…

Apple Machine Learning Research站内正文待翻译:Glyph: A Multi-Strategy Agentic System for Column Description and Sensitivity-Ontology Tagging of Enterprise Data Catalogs

待翻译:Shared Selective Persistent Memory for Agentic LLM Systems

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Agentic LLM systems that generate code through multi-turn tool use face a fundamental context problem: each session starts from zero, discarding the configuration choices, domain constraints, data schemas, and tool-use patterns that made previous sessions productive. Naively persisting entire conversation histories is both token-inefficient and counterproductive—irrelevant context degrades generation quality. We introduce shared selective persistent memory, a memory architecture for agentic systems that identifies and retains four categories of reusable context—task specifications, data…

Apple Machine Learning Research站内正文待翻译:Shared Selective Persistent Memory for Agentic LLM Systems

待翻译:SimpleDesign: A Joint Model for Protein Sequence and Structure Codesign

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Proteins are fundamental to biological processes, with their function determined by the complex interplay between the amino acid sequence and the three-dimensional structure. Developing generative models capable of understanding this intrinsically multi-modal relationship is crucial for fields like drug discovery and protein engineering. Existing models often rely on a multi-stage training process where autoencoders that tokenize data into latent representations are trained in a first stage. Secondly, a generative model is trained on the latent representation of the autoencoder(s), i.e…

Apple Machine Learning Research站内正文待翻译:SimpleDesign: A Joint Model for Protein Sequence and Structure Codesign

待翻译:Putting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question Answering

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Evaluating video captioning remains a critical challenge for Visual Large Language Models (VLLMs). Existing metrics primarily rely on matching generated text against ground-truth references. This paradigm suffers from the “one-to-many” nature of video description, where high-quality captions are often penalized for lexical mismatches or valid shifts in visual focus. Furthermore, such assessments are typically one-dimensional, failing to provide a fine-grained analysis of caption quality. To address this, we redefine caption quality via information fidelity: A caption must maximize the coverage…

Apple Machine Learning Research站内正文待翻译:Putting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question Answering

待翻译:DiscoSign: Discourse-Aware Text to Sign Language Gloss Translation

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Sign language processing systems have traditionally operated at the sentence level, ignoring critical discourse phenomena fundamental to sign language comprehension. We introduce DiscoSign, a computational approach for discourse-aware text to sign language gloss translation grounded in linguistic research. We address three key phenomena within our modular Large Language Model (LLM)-based translation framework: (i) spatial coreference resolution, where entities maintain consistent spatial locations throughout discourse; (ii) Question-Answer Clauses (QACs), pseudocleft structures serving…

Apple Machine Learning Research站内正文待翻译:DiscoSign: Discourse-Aware Text to Sign Language Gloss Translation

REFACTOR-VLA:无监督的带类型运动程序库学习

苹果机器学习研究团队提出 REFACTOR-VLA,通过“清醒/睡眠”架构让视觉-语言-动作模型无监督学习可复用的带类型运动程序。睡眠阶段用行为等价核在潜在世界模型中聚类动作片段,清醒阶段生成 Hindley–Milner 风格的带类型 lambda 项,并经最小描述长度与回报保持门槛筛选形成技能库。完整 LIBERO 实验显示,将世界模型从 1.88 亿扩到 4.3 亿参数反而使 4 个套件性能全部下降;在阶段 A 加入 InfoNCE 辅助损失后,阶段 C 的技能聚类质量显著提高,NMI 最高达 0.915。

Apple Machine Learning Research站内正文REFACTOR-VLA:无监督的带类型运动程序库学习

待翻译:Agent Seer: Synthesizing Scenarios from Specification Understanding

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Evaluating AI agents that use external tools requires realistic test scenarios that capture how practitioners compose tools and iterate across conversation turns. Constructing such scenarios by hand demands deep domain expertise, does not scale across tool ecosystems, and produces static benchmarks that cannot track evolving APIs. We observe that tool specifications—function names, natural-language descriptions, and typed parameter schemas—already encode sufficient semantic information to synthesize realistic evaluation scenarios without manual curation or live tool execution. Agent Seer…

Apple Machine Learning Research站内正文待翻译:Agent Seer: Synthesizing Scenarios from Specification Understanding

待翻译:LLMs Are Not (Consistently) Bayesian: Quantifying Internal (In)consistencies of LLMs’ Probabilistic Beliefs

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Modern AI systems are being deployed in complex domains such as medicine, science, and law, where there is often not a single correct answer given the observed evidence. Such systems must be able to represent and update uncertain beliefs about the world as new evidence arrives to make rational decisions. We introduce the novel technique of studying LLMs as information processing rules and utilize the information processing gap—the deviation from Bayes updates—to study the internal (in)consistencies of how LLMs update their probabilistic beliefs from evidence. Our extensive experiments evaluate…

Apple Machine Learning Research站内正文待翻译:LLMs Are Not (Consistently) Bayesian: Quantifying Internal (In)consistencies of LLMs’ Probabilistic Beliefs

待翻译:From Preferences to Principles: Rubric-Based Alignment for Grounded Knowledge Answers

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Designing effective reward signals for open-domain question answering is challenging because high-quality responses must simultaneously satisfy multiple aspects of answer quality that are difficult to capture with a holistic scalar objective. We introduce a rubric-based reward framework that generates query-specific rubrics grounded in retrieved evidence and decomposed into multiple quality dimensions, providing fine-grained supervision during post-training. Averaged across three evaluation axes (composition, grounding, and instruction-following), our approach improves over the…

Apple Machine Learning Research站内正文待翻译:From Preferences to Principles: Rubric-Based Alignment for Grounded Knowledge Answers

待翻译:Luce: Relightable Gaussians for 3D Asset Generation

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:High-fidelity image-to-3D generation requires a 3D representation that captures both geometry and appearance. To support relighting and integration into standard rendering pipelines, the representation should include physically based rendering (PBR) modalities such as albedo, metallic-roughness, and surface normals. We propose Luce, a 3D representation that unifies geometry and PBR materials within a voxelized multimodal Gaussian cloud, using dedicated Gaussian primitives for each modality. A variational autoencoder compresses this representation into a unified material-aware latent space. A…

Apple Machine Learning Research站内正文待翻译:Luce: Relightable Gaussians for 3D Asset Generation

待翻译:IDEA Prune: An Integrated Enlarge-and-Prune Pipeline in Generative Language Model Pretraining

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Recent advancements in large language models have intensified the need for efficient and deployable models within limited inference budgets. Structured pruning pipelines have shown promise in token efficiency compared to training target-size models from scratch. In this paper, we advocate incorporating enlarged model pretraining, which is often ignored in previous works, into pruning. We study the enlarge-and-prune pipeline as an integrated system to address two critical questions: whether it is worth pretraining an enlarged model even when the model is never deployed, and how to optimize the…

Apple Machine Learning Research站内正文待翻译:IDEA Prune: An Integrated Enlarge-and-Prune Pipeline in Generative Language Model Pretraining

待翻译:PROOF-Gen: From Optimized Data to Better Distillation

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Supervised fine-tuning on teacher-generated trajectories is the standard first stage for distilling tool-calling capabilities into deployable models. Post-training pipelines that drive shipped tool-calling agents re-run this stage on a daily or weekly cadence, paying the frontier-teacher cost each cycle, yet the mechanism is generate-and-filter (keep the teacher’s passing trajectories, discard the rest) and each cycle leaves behind the same hard scenarios because failures supply no signal. On τ 2-bench, 57% of teacher trials fail, two-thirds of them near-misses (most tool calls correct, undone…

Apple Machine Learning Research站内正文待翻译:PROOF-Gen: From Optimized Data to Better Distillation

待翻译:STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Unified multimodal models that understand, reason over, and generate interleaved text–image sequences remain structurally fragmented: existing approaches either sacrifice visual fidelity through discrete tokenization, impose structural asymmetry by combining causal text generation with iterative diffusion-based denoising, or degrade pretrained understanding when adapting vision-language models for generation. We observe that autoregressive normalizing flows are autoregressive Transformers—sharing the same causal mask, KV-cache mechanism, and left-to-right structure as LLMs—making them the most…

Apple Machine Learning Research站内正文待翻译:STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation

待翻译:Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Multimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, temporal, and embodied environments. By generating intermediate reasoning images, Visual CoT provides an intuitive mechanism for visual foresight but introduces substantial inference overhead, which is particularly problematic for proactive video reasoning. We ask whether models can learn to think visually during training while reasoning directly at inference. We introduce Internalized Visual Thinking (IVT), a post-training framework that jointly optimizes textual prediction and…

Apple Machine Learning Research站内正文待翻译:Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning

待翻译:Progressive Refinement: An Iterative Pseudo-Labeling Approach for Mandarin-English Code-Switching ASR

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Code-switching (CS), alternating languages within the same utterance, poses significant challenges for automatic speech recognition (ASR) due to limited CS training data. This paper applies an iterative pseudo-labeling training approach to CS-ASR for the first time, demonstrating its effectiveness in leveraging unlabeled data to improve CS-ASR performance. The approach comprises three phases: pseudo-label generation, two-stage bilingual model training, and iterative improvements. It begins by generating pseudo-labels from a large unlabeled corpus, creating a semi-supervised dataset. This…

Apple Machine Learning Research站内正文待翻译:Progressive Refinement: An Iterative Pseudo-Labeling Approach for Mandarin-English Code-Switching ASR

待翻译:Scaling Laws for Mixture Pretraining Under Data Constraints

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:As language models scale, the amount of data they require grows – yet many target data sources, such as low-resource languages or specialized domains, are inherently limited in size. A common strategy is to mix this scarce but valuable target data with abundant generic data, which presents a fundamental trade-off: too little target data in the mixture underexposes the model to the target domain, while too much target data repeats the same examples excessively, yielding diminishing returns and eventual overfitting. We study this trade-off across more than 2,000 language-model training runs…

Apple Machine Learning Research站内正文待翻译:Scaling Laws for Mixture Pretraining Under Data Constraints

待翻译:Multilingual Knowledge Transfer under Data Constraints via Lexical Interventions

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Cross-lingual knowledge transfer is critical for building high-performing multilingual language models for languages with insufficient training data. When target language data is scarce, the knowledge required for many downstream tasks involving scientific reasoning, commonsense inference, and world knowledge must be acquired primarily from the high-resource language, making effective knowledge transfer essential. Existing methods for improving such cross-lingual knowledge transfer require large amounts of parallel data, translation systems, auxiliary models, or additional training stages that…

Apple Machine Learning Research站内正文待翻译:Multilingual Knowledge Transfer under Data Constraints via Lexical Interventions

待翻译:The P-Completeness of Inverted Index Traversal: On the Complexity of Evaluating Boolean Query DAGs

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Modern AI agents increasingly rely on search infrastructure to execute complex, neuro-symbolic reasoning workflows. These workflows often compile into deeply nested, non-monotonic Boolean queries over text fields. However, standard query evaluation strategies over inverted indices face severe theoretical limits when handling these structures. Stateful iterator models (Document-at-a-Time) are structurally bounded by NC^1 formula evaluation, suffering a worst-case O(2^|Q|) exponential blowup in query complexity when unrolling re-convergent logic. Conversely, recursive materialization models…

Apple Machine Learning Research站内正文待翻译:The P-Completeness of Inverted Index Traversal: On the Complexity of Evaluating Boolean Query DAGs

待翻译:Examining Human-Like Behaviors in LLMs: A Multi-Dimensional Analysis of Model Behaviors, User Factors, and System Prompts

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Large language models (LLMs) exhibit a wide range of human-like behaviors, from expressing thoughts and emotions, to engaging in relationship-building with users, to refusing requests and maintaining boundaries. Despite their prevalence, researchers and practitioners lack methods and empirical insights to make informed decisions about when and what types of human-like behaviors LLMs should exhibit. To fill this gap, we present a multi-dimensional analysis of the prevalence, potential effects, and controllability of these behaviors using LLM-as-a-judge and human evaluation. Across 21,000…

Apple Machine Learning Research站内正文待翻译:Examining Human-Like Behaviors in LLMs: A Multi-Dimensional Analysis of Model Behaviors, User Factors, and System Prompts

待翻译:A Specialized Semismooth Newton Method for Kernel-Based Optimal Transport

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Kernel-based optimal transport (OT) estimators offer an alternative, functional estimation procedure to address OT problems from samples. Recent works suggest that these estimators are more statistically efficient than plug-in (linear programming-based) OT estimators when comparing probability measures in high-dimensions [Vacher et al., 2021]. Unfortunately, that statistical benefit comes at a very steep computational price: because their computation relies on the short-step interior-point method (SSIPM), which comes with a large iteration count in practice, these estimators quickly become…

Apple Machine Learning Research站内正文待翻译:A Specialized Semismooth Newton Method for Kernel-Based Optimal Transport

待翻译:MVICAD2: Multi-View Independent Component Analysis with Delays and Dilations

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Machine learning techniques in multi-view settings face significant challenges, particularly when integrating heterogeneous data, aligning feature spaces, and managing view-specific biases. These issues are prominent in neuroscience, where data from multiple subjects exposed to the same stimuli are analyzed to uncover brain activity dynamics. In magnetoencephalography (MEG), where signals are captured at the scalp level, estimating the brain’s underlying sources is crucial, especially in group studies where sources are assumed to be similar for all subjects. Common methods, such as Multi-View…

Apple Machine Learning Research站内正文待翻译:MVICAD2: Multi-View Independent Component Analysis with Delays and Dilations

待翻译:GRPO Beyond English: A Large-Scale Study of GRPO in Non-English and Multilingual Settings

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Reinforcement Learning with Verifiable Rewards (RLVR), often optimized with Group Relative Policy Optimization (GRPO), has become a central recipe for improving the reasoning capabilities of pretrained language models but current studies remain heavily English-centric. We conduct a large-scale empirical study of multilingual and non-English GRPO across a wide range of base models, training languages, and different reasoning language rewards. We find that training to reason in the native language often leaves only a small gap to training for English reasoning. We further observe strong…

Apple Machine Learning Research站内正文待翻译:GRPO Beyond English: A Large-Scale Study of GRPO in Non-English and Multilingual Settings

全部来源