Skip to content
AI News HubLIVE
Public articles 888Collected articles 950Trust 75Refresh 360 min
Health HealthySource type ResearchFull-text rights Full text allowedLast ingested 2026-09-28ID arxiv-cs-cvStatus Enabled

Use abstract and metadata; check individual paper license before full text.

Latest public articles

MVAgent: Multi-Agent Video Generation via Consistent Condition Construction and Shot-Level Policy Optimization

arXiv:2609.30609v1 Announce Type: new Abstract: Multi-shot agentic video generation requires consistent character appearance, stable spatial layout across camera angles, and continuous character state between shots. When every shot is a separate request to a frozen generator, repeated text does not determine appearance, layout or state. We therefore recast the problem as condition construction and present MVAgent, a multi-agent pipeline whose agents collaborate through typed conditioning inputs. Because an environment image shows one viewpoint, a Spatial Grounding agent samples views from generated camera-traversal clips and anchors each shot to the view matching its framing. As generated shots drift from the plan, an Observer records how each shot ends in a continuity memory, from which…

arXiv Computer VisionIn-site articleMVAgent: Multi-Agent Video Generation via Consistent Condition Construction and Shot-Level Policy Optimization

Action Forcing: Training World Models on Unsupervised Video by Recovering Underlying Egomotion Bases

arXiv:2609.30595v1 Announce Type: new Abstract: Synchronised action annotations are needed to train controllable world models and these datasets remain elusive. Existing approaches make use of instrumented platforms with calibrated sensors, costly manual annotation, or latent-action models which lack grounding. We instead turn ordinary unlabelled video into action-supervised training data by recovering (without training) a data-derived egomotion basis. We track pixel displacements across frames and exploit the recurring coherent structure induced by egomotion to obtain grounded control signals directly. Using a method as simple as principal components analysis perform this, we find that the leading components provide signed, scalable, and composable throttle--yaw controls, although the me…

arXiv Computer VisionIn-site articleAction Forcing: Training World Models on Unsupervised Video by Recovering Underlying Egomotion Bases

Atlases Are Already Inside: Recovering Population Templates from Pretrained Diffusion Models

arXiv:2609.30566v1 Announce Type: new Abstract: We present a new inference-time sampler for diffusion models that gives a pretrained model a capability it was never trained for: constructing the atlas of the population it synthesizes. The sampler converges from every random seed to the population's central anatomy, which we call the \emph{intrinsic atlas}. The advantage is threefold. (1) It requires no retraining. A diffusion model that has already learned a coherent population, including the released ones, yields its atlas in a single inference pass without involving deformable registration. (2) It applies to multiple domains, such as brain MRI, chest X-ray, faces, and 3D shapes. (3) It extends to subpopulations. One age-conditioned model gives an atlas at any age in its training range,…

arXiv Computer VisionIn-site articleAtlases Are Already Inside: Recovering Population Templates from Pretrained Diffusion Models

The Shape of Events: Edge-Based Inductive Biases via Cross-Domain Distillation

arXiv:2609.30478v1 Announce Type: new Abstract: Convolutional neural networks trained on ImageNet are known to exhibit a strong preference for local high-frequency texture, an inductive bias that translates into fragile robustness against distribution shifts in real-world environments. Event cameras, in contrast, record only changes in scene brightness and are therefore well suited to capturing contour information; however, due to the absence of diagnostic benchmarks in the event domain, the inductive bias that event-camera data instills in vision models has remained underexplored. In this work, we use knowledge distillation from the event domain to the RGB domain so as to exploit the rich evaluation toolkit available in the RGB domain and systematically dissect this inductive bias. Our e…

arXiv Computer VisionIn-site articleThe Shape of Events: Edge-Based Inductive Biases via Cross-Domain Distillation

LensDesigner: A Self-Improving Agent for Optical Lens Design

arXiv:2609.30450v1 Announce Type: new Abstract: Optical lens design is a complex, non-convex optimization challenge that relies heavily on human experience and intuition. Existing optimized-based automatic lens design methods struggle to navigate this vast parameter space without meticulous manual tuning. In this paper, we present LensDesigner, an autonomous agent framework that mirrors the problem-solving workflow of expert opticians. To overcome the initial cold start problem, we construct LensLib100K, an extensive optical lens library, and employ Optics-Aware Retrieval to supply physically valid structural seeds. Within an interactive physical simulation environment, the agent executes macroscopic orchestration while receiving immediate optical feedback. Furthermore, we introduce a con…

arXiv Computer VisionIn-site articleLensDesigner: A Self-Improving Agent for Optical Lens Design

ProCAP: Probabilistic Cross-Attentive Prompt Learning for Vision-Language Models

arXiv:2609.30434v1 Announce Type: new Abstract: Pre-trained vision-language models such as CLIP can recognize new categories via prompting, but they often struggle when labeled data are scarce or the test distribution shifts. Prompt learning adapts only a small set of parameters while keeping the backbone frozen, yet many existing multimodal prompt learners couple the visual and textual branches weakly and can be brittle in low-shot regimes. We propose ProCAP, a probabilistic cross-attentive prompt learning framework that improves cross-modal interaction and training stability without updating any CLIP weights: it learns both visual and textual prompt tokens and links them through stacked bidirectional multi-head cross-attention so the two branches refine each other across prompt depth. T…

arXiv Computer VisionIn-site articleProCAP: Probabilistic Cross-Attentive Prompt Learning for Vision-Language Models

What Improves Multimodal Misinformation Detection? Answers from a Large-Scale Empirical Study

arXiv:2609.30402v1 Announce Type: new Abstract: Multimodal misinformation is increasingly crafted to look convincing by pairing a textual claim with an image that appears to "prove" it. Yet in practice, building effective detectors often hinges on a small set of design choices that are rarely examined in a controlled way. In this paper, we conduct a large-scale study of multimodal design choices for misinformation detection with over 3,375 experiments- spanning three benchmark datasets and a broad range of pre-trained vision and language backbones. Through systematic comparisons and targeted robustness analyses, we distill practical guidance on which design choices help, when do they fail silently, and what aspects of the pipeline most strongly shape model behavior, answering 4 key Resear…

arXiv Computer VisionIn-site articleWhat Improves Multimodal Misinformation Detection? Answers from a Large-Scale Empirical Study

CSCWD: Cross-Scale Channel-wise Knowledge Distillation for Lightweight Tiny Object Detection on Edge Devices

arXiv:2609.30395v1 Announce Type: new Abstract: Real-time tiny object detection in aerial imagery is constrained by the weak spatial evidence of very small objects and the loss of high-resolution detail in lightweight detectors. This study presents Cross-Scale Channel-wise Knowledge Distillation (CSCWD), a training-time framework that transfers high-resolution spatial representations from a YOLO11m-P2 teacher to a compact YOLO11n student without altering the student's inference architecture. Unlike conventional same-scale feature distillation, CSCWD transfers supervision from teacher P2 to student P3 after feature alignment while retaining same-scale distillation at deeper pyramid levels. Under the unified seven-sequence Drone-vs-Bird validation protocol, YOLO11n-CSCWD achieves 50.17% mea…

arXiv Computer VisionIn-site articleCSCWD: Cross-Scale Channel-wise Knowledge Distillation for Lightweight Tiny Object Detection on Edge Devices

LiTe-GS: Oracle-Efficient Next Best View Selection for 3D Gaussian Splatting

arXiv:2609.30393v1 Announce Type: new Abstract: Selecting informative camera views is critical for efficient training and adaptive refinement in 3D Gaussian Splatting, where each observation significantly influences model parameters. However, information-driven view-selection strategies can require repeated evaluations of expensive information-gain oracles as the number of candidate views increases. We propose LiTe-GS, an oracle-efficient method for next best view selection in 3D Gaussian Splatting. LiTe-GS reduces the number of information-oracle evaluations by performing randomized subset evaluation of candidate views rather than exhaustively scoring the full candidate pool. The resulting approach achieves expected $O(M\log(1/\epsilon))$ oracle complexity with respect to the number of c…

arXiv Computer VisionIn-site articleLiTe-GS: Oracle-Efficient Next Best View Selection for 3D Gaussian Splatting

AlphaEarth distinguishes cities but compresses urban variation

arXiv:2609.30356v1 Announce Type: new Abstract: Cities differ in built form, land cover and development history, complicating comparison across places and time. Satellite foundation models map Earth's surface onto common numerical representations. Yet the tasks and targets used to shape them typically do not focus on cities: globally consistent labels for urban function do not exist, and many datasets - especially land cover and land use classifications - collapse the built environment into few classes. Here we audit the representation, focusing on AlphaEarth but with broader applicability to other Earth embeddings, by probing the geometry and geography of embeddings for 1,000 urban areas in 162 countries. We find that cities occupy a shifted but overlapping region on the hypersphere, 62.…

arXiv Computer VisionIn-site articleAlphaEarth distinguishes cities but compresses urban variation

CinematicVQA: Benchmarking Film-Grammar Reasoning in Large Vision-Language Models

arXiv:2609.28813v1 Announce Type: new Abstract: Cinematography, the craft of visual storytelling through framing, lighting, and camera operation, fundamentally shapes how audiences perceive and emotionally engage with video content. While Large Vision Language Models (LVLMs) have made remarkable progress in video question answering, existing benchmarks primarily focus on identifying low-level techniques rather than understanding their storytelling impact. To address this, we introduce CinematicVQA, the first-of-its-kind benchmark for cinematic video understanding that goes beyond technique recognition to evaluate film-grammar reasoning, utilizing our introduced Cinematic Scene Graph (CSG), a structured representation that links filming techniques to their perceptual effects and narrative…

arXiv Computer VisionIn-site articleCinematicVQA: Benchmarking Film-Grammar Reasoning in Large Vision-Language Models

DeltaWAM: Delta World Action Models for Bimanual Manipulation

arXiv:2609.28811v1 Announce Type: new Abstract: World-action models (WAMs) transfer visual and motion priors from pretrained video generators to robot control by jointly modeling visual dynamics and actions. Existing WAMs, however, predict dense future frames during training, repeatedly modeling largely unchanged content and coupling action-conditioned dynamics to nuisance appearance variations. At inference, processing each complete observation with the heavy video expert bottlenecks few-step action generation. Accordingly, we propose DeltaWAM, which jointly predicts visual deltas and actions using dense-anchor, sparse-delta, and action streams, with three architectures that differ in representation and computation sharing. We further develop Streaming Delta Memory (SDM), which updates c…

arXiv Computer VisionIn-site articleDeltaWAM: Delta World Action Models for Bimanual Manipulation

DrGait: Biomechanically Grounded Visual Reasoning for Interpretable Clinical Gait Analysis

arXiv:2609.28796v1 Announce Type: new Abstract: Current automated gait analysis for clinical applications relies on uninterpretable black-box classifiers. Although Vision-Language Models (VLMs) offer strong reasoning capabilities, applying them directly to gait videos often leads to hallucinations, because they struggle to measure subtle geometric deviations from raw visual contexts. To address this, we introduce DrGait, a training-free agentic framework that shifts the VLM's role from a direct visual reasoner to a clinical planner. DrGait decouples semantic reasoning from geometric perception through a structured Triage-Verification-Synthesis (TVS) workflow. Given an input video and a set of basic spatiotemporal metrics, the DrGait agent first performs a heuristic triage to propose diagn…

arXiv Computer VisionIn-site articleDrGait: Biomechanically Grounded Visual Reasoning for Interpretable Clinical Gait Analysis

Small yet Assistive: Spatially-Aware Post-Training for Low Vision

arXiv:2609.28757v1 Announce Type: new Abstract: An estimated 1 billion people worldwide live with vision impairment, yet current vision-language models (VLMs) produce descriptions too vague for safe navigation by blind and low-vision (BLV) users. Large VLMs can generate high-quality audio-description-compliant narrations but cannot run on mobile devices; small VLMs offer competitive latency but lack spatial detail, directional cues, and hazard awareness for navigational assistance. We present Smol-VL-BLV, a compact VLM for blind and low-vision users that closes this gap using a 500M decoder transformer model and two post-training mechanisms: (1) teacher-student distillation and (2) Group Relative Policy Optimization (GRPO) with a composite BLV reward targeting directional language, metric…

arXiv Computer VisionIn-site articleSmall yet Assistive: Spatially-Aware Post-Training for Low Vision

GeoNLI - A Natural Language Interpreter for Satellite Imagery

arXiv:2609.28741v1 Announce Type: new Abstract: Multi-modal multitasking models have shown strong performance on remote sensing datasets. However, because these models are trained on heterogeneous data and vary across tasks, designing a unified model that performs well in captioning, visual question answering (VQA), and visual grounding remains challenging. In this work, we evaluate several models on the VRS Bench and NWPU-VHR-10 datasets. The EarthMind model demonstrates strong results in both captioning and VQA. For grounding, we propose multiple pipelines - RemoteSAM-SAM-v1, RemoteSAM-SAM-v2, and DiffuSAM - and ultimately adopt a majority-voting ensemble across EarthMind, RemoteSAM, SAM3, Falcon, RemoteSAM-SAM3-v1, RemoteSAM-SAM3-v2, and DiffuSAM predictions. Our unified, modular pipel…

arXiv Computer VisionIn-site articleGeoNLI - A Natural Language Interpreter for Satellite Imagery

M-plicits: Neural Implicit Surfaces via Nested Multiscale Residuals

arXiv:2609.28684v1 Announce Type: new Abstract: Encoding input coordinates with sinusoidal functions into multi-layer perceptrons (MLPs) has proven effective for implicit neural representations (INRs) of surfaces defined as zero-level sets. However, existing methods often struggle to balance training efficiency, rendering speed, and noise robustness: single-MLP approaches are expensive at inference, grid-based representations are fast but can limit surface smoothness and overfit input noise, and previous multiscale approaches frequently capture noise and produce artifacts due to hard spectral truncation. To address these limitations, we propose M-plicits, a multiscale framework that models surfaces as a residual sum of MLPs trained via a sequence of nested neighborhoods. Unlike existing r…

arXiv Computer VisionIn-site articleM-plicits: Neural Implicit Surfaces via Nested Multiscale Residuals

PePESeg3D: Perception Prior Enhances Multi-Scale Segmentation for 3D Gaussian Splatting

arXiv:2609.28645v1 Announce Type: new Abstract: Recent advancements in 3D Gaussian Splatting (3DGS) have extended its capabilities to multi-scale segmentation. Existing methods reconstruct a scene with Gaussian primitives and learn multi-scale segmentation features separately, which leaves the geometry unaware of semantic structure and the feature learning dependent on incomplete mask supervision. To address these limitations, we present PePESeg3D, a novel framework that injects perception priors into a multi-scale 3D Gaussian segmentation pipeline. To fully exploit perception priors, we integrate them not only into contrastive feature learning but also into the upstream geometry reconstruction. Specifically, PePE Reconstruction incorporates monocular depth and mask constraints to ensure…

arXiv Computer VisionIn-site articlePePESeg3D: Perception Prior Enhances Multi-Scale Segmentation for 3D Gaussian Splatting

UltraBench 2: Towards Robust Evaluation of Vision Foundation Models on Ultrasound

arXiv:2609.28610v1 Announce Type: new Abstract: Benchmarking is an increasingly critical part of research in machine learning and the domains where it is applied, including healthcare. Yet, despite the steady development of new ultrasound foundation models in recent years, the development of well-designed benchmarks to evaluate them has lagged behind. This deficiency has led to fragmented and inconsistent evaluations of competing models, making it difficult to measure progress. To address this issue, we introduce UltraBench 2, a comprehensive benchmark with wide anatomical and task coverage, and a focus on standardization, reproducibility, and ease-of-use. Using this benchmark, we compare existing vision foundation models for ultrasound image analysis. Our analyses demonstrate that ultras…

arXiv Computer VisionIn-site articleUltraBench 2: Towards Robust Evaluation of Vision Foundation Models on Ultrasound

Token Clustering and Semantic Sequence Mamba for Hyperspectral Image Classification

arXiv:2609.28580v1 Announce Type: new Abstract: Although hyperspectral images (HSIs) provide rich spectral-spatial information, accurate pixel-level classification remains challenging because of spectral-spatial heterogeneity and complex spatial structures. Existing vision state space models (Mamba) typically construct sequences according to predefined spatial neighborhoods, without explicitly accounting for semantic similarity or spatial non-stationarity. To address this limitation, we propose Token Clustering and Semantic Sequence Mamba (STMamba), which organizes sparse tokens into semantically coherent sequences for hyperspectral image classification with the following features. First, at the macro level, a hierarchical encoder decoder progressively selects semantic tokens with the Tok…

arXiv Computer VisionIn-site articleToken Clustering and Semantic Sequence Mamba for Hyperspectral Image Classification

$\unicode{x1F493}$Heartian: Physiology-Aware Relightable Gaussian Head Avatar

arXiv:2609.28539v1 Announce Type: new Abstract: Gaussian head avatars typically model intrinsic facial appearance as temporally static, omitting subtle cardiac-induced skin-color variation. We propose $\unicode{x1F493}$Heartian, a physiology-aware modulation framework that learns cardiac-cycle-dependent per-frame albedo modulation of facial skin-region Gaussians within a relightable head avatar to encode remote photoplethysmography (rPPG) signals. Using synchronized contact PPG supervision, $\unicode{x1F493}$Heartian models the prescribed cardiac waveform as the sum of two Gaussian functions and learns per-frame spatial residuals via a lightweight MLP. Across 152 stationary recordings from UBFC-rPPG, PURE, and MMPD, attribute-space recovery of the supplied signal achieves a pooled recordi…

arXiv Computer VisionIn-site article$\unicode{x1F493}$Heartian: Physiology-Aware Relightable Gaussian Head Avatar

Feed the Panel Dimensions, Not Verdicts: Rubric-Decomposed Fusion of Vision-Language Aesthetic Judges

arXiv:2609.27110v1 Announce Type: new Abstract: Vision-language models (VLMs) are deployed as zero-shot judges of image aesthetics, and panels of several models are recommended, on thin evidence, as the way to make such judges reliable. On two human-rated datasets, EVA and PARA, we find that a panel of holistic judges never significantly beats its best member, whether the verdicts are averaged or fused by a learned combiner. What a panel is worth depends on what it is fed. We therefore have each model score each image on the five dimensions of a frozen, human-written rubric and fuse those scores, alongside each model's verdict, across model families with an out-of-fold combiner. The dimension scores measure what their labels claim: with the overall human score partialled out, a dimension…

arXiv Computer VisionIn-site articleFeed the Panel Dimensions, Not Verdicts: Rubric-Decomposed Fusion of Vision-Language Aesthetic Judges

Pro-Bench: Prompt-Robust Open-Vocabulary Visual Grounding Across Real-World Heterogeneous Environments

arXiv:2609.27076v1 Announce Type: new Abstract: Open-vocabulary visual grounding enables robots to localise task-relevant entities from natural-language queries without dependence on predefined perceptual taxonomies. However, existing benchmarks largely rely on short category labels and web-scraped imagery, leaving it unclear whether open-vocabulary models can robustly ground diverse queries and visual conditions under real deployments. We introduce \textbf{Pro-Bench}, a prompt-conditioned benchmark for open-vocabulary visual grounding in heterogeneous, real-world environments. Pro-Bench includes $13k+$ RGB frames from independent robotic domains (subterranean, industrial, indoor, outdoor, urban), with $74.5k$ manual instance annotations and $515$ target queries covering categorical, attr…

arXiv Computer VisionIn-site articlePro-Bench: Prompt-Robust Open-Vocabulary Visual Grounding Across Real-World Heterogeneous Environments

Adversarial Attacks and Identity Leakage in De-Identification Systems: An Empirical Study

arXiv:2609.27022v1 Announce Type: new Abstract: In this paper, we investigate the impact of adversarial attacks on identity encoders within a realistic de-identification framework. Our experiments show that the transferability of attacks transfers from an external surrogate model to the system model (e.g., CosFace to ArcFace) allows the adversary to cause identity information to leak in a sufficiently sensitive face recognition system. We present experimental evidence and propose strategies to mitigate this vulnerability. Specifically, we show how fine-tuning on adversarial examples helps to mitigate this effect for distortion-based attacks (i.e., snow, fog, etc.), while a simple low-pass filter can attenuate the effect of adversarial noise without affecting the de-identified images. Our…

arXiv Computer VisionIn-site articleAdversarial Attacks and Identity Leakage in De-Identification Systems: An Empirical Study

Anatomy-Aware Synthesis of Post-Contrast Breast MRI from Pre-Contrast Images

arXiv:2609.27015v1 Announce Type: new Abstract: We developed an anatomy-aware deep learning framework to synthesize post-contrast breast MRI from pre-contrast images, emphasizing tumor and background parenchymal enhancement (BPE) regions. This retrospective study included 649 patients with 6,251 paired pre-contrast and post-contrast images. The framework integrates breast mask consistency, lesion-region supervision, and BPE-region supervision into an image-to-image translation model. Evaluation included quantitative image quality metrics, a reader study with two breast radiologists, and downstream Ki-67 classification. The proposed method outperformed Pix2Pix, Pix2PixHD, diffusion-based synthesis, and mask-supervised baselines in whole-image and regional evaluations. Ki-67 classification…

arXiv Computer VisionIn-site articleAnatomy-Aware Synthesis of Post-Contrast Breast MRI from Pre-Contrast Images

HYDRO: Towards Non-Reversible Face De-Identification Using a High-Fidelity Hybrid Diffusion and Target-Oriented Approach

arXiv:2609.27011v1 Announce Type: new Abstract: Target-oriented face de-identification models aim to anonymize the identity of a target individual across different images or video frames, such that the target can no longer be reliably recognized, while maintaining key characteristics of the visual data. Such models commonly leverage generative encoder-decoder architectures to manipulate facial appearances, enabling them to produce realistic high-fidelity de-identification results, while ensuring considerable attribute-retention capabilities. However, target-oriented models also carry the risk of inadvertently preserving subtle identity cues, making them (potentially) reversible and susceptible to reconstruction attacks. To address this problem, we introduce in this paper a novel (robust)…

arXiv Computer VisionIn-site articleHYDRO: Towards Non-Reversible Face De-Identification Using a High-Fidelity Hybrid Diffusion and Target-Oriented Approach

Lessons learned from deploying imaging AI with the open PACS-AI platform

arXiv:2609.26981v1 Announce Type: new Abstract: We describe deploying imaging AI at six hospitals through PACS-AI, an open self-hosted platform. The binding constraint is not model accuracy but infrastructure to route studies, display results, capture feedback, and audit what runs. At one center, angiography models completed 515 of 607 jobs (84.8%); failures reflected absent diagnostic views, and 78.1% of 638 clinician ratings were positive. Publishing honest readiness levels for every model is itself a governance practice.

arXiv Computer VisionIn-site articleLessons learned from deploying imaging AI with the open PACS-AI platform

nnFoundation: 3D Foundation Models for Radiology

arXiv:2609.26924v1 Announce Type: new Abstract: Radiological artificial intelligence has advanced rapidly, yet most systems remain narrowly task-specific, data-intensive, and fragile under domain shift. Foundation models promise more transferable and data-efficient solutions, but existing approaches are limited in scale, evaluated narrowly, and often assume that a single pretrained model can support diverse downstream tasks. Here we present nnFoundation, complementary convolutional and transformer-based 3D radiological foundation models. Developed within the Human Radiome Project (THRP), nnFoundation is trained on 2.1 million CT, MRI, and PET image volumes from 125 institutional and public datasets. We evaluate them across 108 tasks spanning segmentation, detection, classification, report…

arXiv Computer VisionIn-site articlennFoundation: 3D Foundation Models for Radiology

A 3D Pose-Based Ensemble Framework for Cricket Shot Classification and Automated Biomechanical Analysis

arXiv:2609.26923v1 Announce Type: new Abstract: Cricket is one of the most celebrated sports world-wide, and technological advancement has become deeply embedded in how the modern game is analyzed and coached. Cricket shot classification and automated performance analysis add a further dimension to this trend. Traditional approaches rely on RGB video features or static images, which are sensitive to environmental variations such as camera angle, lighting, and background clutter, and often fail to capture the underlying biomechanics of batting actions. In this paper, we propose a system to improve cricket coaching that takes raw video data, extracts batsmen from video frames using YOLO, and extracts 3D pose data from video frames using MeTRAbs. The system produces sequential skeletal pose…

arXiv Computer VisionIn-site articleA 3D Pose-Based Ensemble Framework for Cricket Shot Classification and Automated Biomechanical Analysis

Cross-Modal Contrastive Learning from Histopathology and CT for Automated Renal Cell Carcinoma Grading

arXiv:2609.26920v1 Announce Type: new Abstract: Background: Clear cell renal cell carcinoma (ccRCC) exhibits substantial clinical heterogeneity, and accurate grade assessment is essential for risk stratification and treatment planning. However, conventional grading requires invasive tissue sampling. We developed RCC-Align, a cross-modal contrastive learning framework that leverages paired histopathology and computed tomography (CT) data during training to improve noninvasive CT-based ccRCC grade prediction. Methods: RCC-Align aligns paired whole-slide histopathology images (WSIs) and CT scans through contrastive cross-modal objectives, transferring grade-discriminative information from microscopic tissue morphology to macroscopic radiologic representations. The framework was trained and e…

arXiv Computer VisionIn-site articleCross-Modal Contrastive Learning from Histopathology and CT for Automated Renal Cell Carcinoma Grading

AgroBench: A Reproducible Multimodal Benchmark for Weakly Supervised Crop Yield Learning from County Statistics and Pixel Observations

arXiv:2609.26809v1 Announce Type: new Abstract: Reliable agricultural yield statistics are typically reported at coarse administrative scales, whereas modern geospatial machine learning methods require spatially explicit, pixel level supervision. This mismatch has limited the development of large-scale benchmarks for crop yield learning using multimodal Earth observation data. A reproducible benchmark, AgroBench, is presented for transforming publicly available U.S. county level crop yield statistics into weakly supervised pixel-level crop time series. Each crop pixel time series is paired with a county-level yield value as a weak supervisory signal rather than a directly measured pixel-level yield label. Our geospatial data generation pipeline integrates USDA crop yield statistics with c…

arXiv Computer VisionIn-site articleAgroBench: A Reproducible Multimodal Benchmark for Weakly Supervised Crop Yield Learning from County Statistics and Pixel Observations

All sources