跳到主要內容
AI News HubLIVE
來源內容 · 翻譯待補全6 分鐘閱讀

待翻譯:Biomedical Imaging's Real Bottleneck Is the Data, Not the Model

文章摘要

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Part I — The problemFour sectors, one bottleneckHospitals and health systems generate...

待翻譯:Biomedical Imaging's Real Bottleneck Is the Data, Not the Model
報告錯誤

更正渠道尚未開通,可先複製下方文章資訊留存。

查看更正說明
直接讀正文

AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。

Biomedical Imaging's Real Bottleneck Is the Data, Not the Model | Databricks Blog What it is: A field perspective on why medical imaging AI underdelivers across hospitals, academia, medtech, and pharma; what a stronger data foundation looks like across R&D. The challenge: Imaging is medicine's richest data but its least usable: siloed across PACS, CROs, and central labs, hard to de-identify, and rarely linked to other data modalities that power biomarkers and patient stratification. The hard part is the data, not the model. The takeaway: The frontier isn't a better model, it's the data foundation under it. Govern imaging in one place, make it queryable, de-identify at scale, and link it to EHR, omics, and trial data (via governed lakehouse patterns). Get the data right, and the model becomes the easy part. Part I — The problem Four sectors, one bottleneck Hospitals and health systems generate the raw material. Radiology is the most heavily regulated corner of medical AI anywhere: of the roughly 950 AI/ML-enabled devices the FDA had authorized by mid-2024, about 723 (around 76%) were radiology tools. Almost all are narrow and human-supervised, doing triage, measurement, reconstruction, or worklist prioritization. The real difficulty in a hospital is rarely the model. It's that the data sits locked inside PACS and vendor-neutral archives built to serve images to a viewer, not to answer research questions. Pulling a cohort means extracting images from clinical systems, de-identifying them (PHI hides in DICOM headers and gets burned into the pixels), and finding the compute to use them. The model is the easy 10%. Academic medical centers are where most of the open science happens. The field owes much of its progress to shared datasets like MIMIC-CXR (377,110 chest X-rays), The Cancer Imaging Archive, fastMRI, and the UK Biobank imaging arm, along with open tooling such as MONAI, nnU-Net, and 3D Slicer. Academia also exposes the field's credibility problem. A well-known BMJ review (Nagendran et al., 2020) of deep-learning-versus-clinician imaging studies found that of 81 studies, only a handful were prospective or tested in a real clinical setting, and code and data were unavailable in 93% and 95% of them. That reproducibility debt traces straight back to how hard imaging data is to share and re-run. Medtech and device companies (GE HealthCare, Siemens Healthineers, Philips, Canon) build AI directly into the scanner: deep-learning reconstruction that shortens scans, dose reduction, on-device triage. Their hardest problem is generalization. An algorithm trained on one vendor's scanner, field strength, and protocol can quietly degrade on another's. Curating multi-site training data that represents the real world, then proving it to regulators through the 510(k) and PMA pathways, is both the moat and the cost. Pharma and biotech treat imaging as a measurement instrument. Oncology trials depend on standardized response criteria like RECIST, iRECIST, and RANO, read centrally and blindly to strip out site bias, and quantitative imaging biomarkers can pick up drug response earlier than size-based measures. None of it works unless a measurement means the same thing across sites, scanners, and time, which is why RSNA's QIBA writes profiles that pin down acquisition and analysis to a stated precision. Variability is the enemy of a trial endpoint. Four different jobs, one shared bottleneck. The pixels are everywhere, and queryable nowhere. Fig 1. Hospitals, academia, medtech, and pharma all converge on the same problem: the pixels are everywhere and queryable nowhere. Why collaboration is the whole game, and why it's so hard No single institution holds enough diverse data to build a model that generalizes. The strongest collaborative study to date makes the point. In the EXAM study (Nature Medicine, 2021), 20 institutions trained a shared model to predict COVID-19 oxygen needs from chest X-rays plus EMR data without moving any patient data between them. Only the model weights traveled, and the federated model gained, on average, 16% in AUC and 38% in generalizability over single-site models. The mechanics are less exotic than they sound: each site trains locally, a coordinator averages the model weights rather than the data (the classic federated-averaging recipe), and secure aggregation and differential privacy keep any one site's records from being reconstructed out of those weights. That's the promise. The reality is that collaboration is genuinely difficult, for reasons that are structural rather than technical: Data heterogeneity. Different scanners, protocols, and labeling conventions mean that "the same" exam often isn't. Federated models degrade on this non-IID, acquisition-skewed data unless it's deliberately engineered around with techniques like batch-normalization averaging or site-aware harmonization. Privacy and governance. Every cross-institution study brings per-site IRB approvals, data-use agreements, and de-identification that has to be defensible. No common substrate. Consortia like MIDRC (co-led by the ACR, RSNA, and AAPM) exist precisely because intake, curation, de-identification, labeling, and sharing are so hard that they need a dedicated national effort. Collaboration in imaging isn't a nice-to-have. It's the only path to models that work outside the building they were trained in, and it's gated almost entirely by data plumbing and governance. The complexity most people outside the field never see When people say "medical images," they usually picture a chest X-ray. The reality is a zoo of formats and scales that makes general-purpose data tooling buckle. Imaging typeFormatsDimensions & scaleWhy general-purpose tooling buckles Radiology — X-ray, CT, MRIDICOM, encodable in ~a dozen transfer syntaxes (uncompressed → JPEG 2000 → HTJ2K)2D X-ray → 3D CT/MRI volumes → 4D dynamic cardiacBuilt to move studies between machines, not analyze in bulk; decoders (pydicom, GDCM, pylibjpeg) run single-threaded — a few studies is trivial, a few million is a distributed-systems problem Digital pathology — whole-slide imagesVendor formats: Aperio .svs, Hamamatsu .ndpi, Philips iSyntax (via OpenSlide); almost no DICOMGigapixel — multiple GB, tens of thousands of tiles, read as a pyramid at several magnificationsToo large to open in memory; stain and scanner color vary lab to lab, becoming their own source of model error Ophthalmology, dermatology & othersOCT volumes, clinical photography, each with its own conventions2D and 3D, modality-specificOne more set of formats a folder-of-JPEGs tool was never built to handle "Medical images" isn't one thing — it's a zoo of formats and scales. Research archives reach petabytes, enough that de-identifying them at scale is its own published line of research. There's a second kind of complexity that quietly breaks studies: the numbers themselves aren't comparable across sites. A standardized uptake value or a tumor volume measured on two scanners with two protocols are not the same measurement. This is why quantitative imaging leans on standards like QIBA profiles and the IBSI definitions for radiomics features, and why reproducibility, not raw accuracy, is usually the thing that fails. A tool that handles a folder of JPEGs can't handle any of this, which is a big part of why imaging has lagged genomics and EHR data in becoming analytics-ready. Imaging is the sharpest edge of a bigger problem What's easy to miss when the focus stays only on images: everything here is also true of the rest of R&D. Imaging is just where the pain shows up first, because the data is the biggest and the strangest. Walk into a life-sciences R&D organization today and the real frontier isn't any single data type. It's multi-omics, the genomics, transcriptomics, proteomics, and metabolomics that each arrive with their own formats and silos. It's real-world and clinical-trial data, EHRs and registries and waveforms. It's medical affairs and the literature. The questions that actually move a program, which patients respond and why, what the image plus the molecular signature plus the outcome say together, live in the joins between these domains, not inside any one of them. Every one of those data types carries the same affliction imaging does. It's rich, high dimensional, mostly unstructured, trapped in domain specific systems, hard to govern, and harder to link. A whole-slide image, a genomic variant call, and a trial endpoint are each hard enough on their own. The value comes from putting them in the same sentence. That's a data-foundation problem before it's ever an AI problem, and it's the one teams underestimate most. This isn't only a healthcare story. A recent Bain analysis of why AI budgets keep climbing while returns don't, found that data access and integration is the single biggest barrier to AI, cited by 41% of companies and named even more often by the leaders than the laggards. Their blunt version: most organizations still can't reliably get to their own data. Medicine just makes it harder to look away from. Part II — The foundation This is where it gets concrete and more technical. For readers who aren't building the platform, the short version is simple: govern everything in one place, make the pixels queryable, de-identify at scale, and link imaging to the rest of the patient. The rest of this section is how. What the data foundation actually has to do Strip away the sector specifics and the requirements come out the same. The platform has to: Govern unstructured imaging files and their extracted metadata in one place, with access control, lineage, and auditability that hold up to a regulator's questions. Process petabyte-scale, gigapixel, multi-dimensional data (2D X-ray, 3D CT/MRI volumes, 4D dynamic studies, gigapixel pathology slides) without falling over. De-identify at scale, both in the headers and in the pixels. Link imaging to EHR, genomics, waveforms, and text, so a single research question can touch all of them. Make cross-institution sharing a configuration step, not a six-month negotiation. That's the shape of a lakehouse. Here's how it maps in practice, using patterns that work on Databricks. A medallion architecture for Pixels The mental model that makes imaging tractable is the same medallion pattern teams already use for tabular data, adapted for binary files. Bronze: raw studies first land in a restricted-access cloud object storage or Volume, where a de-identification job runs on headers and pixels before anything else; the de-identified files then form the governed bronze layer in Unity Catalog Volumes, one catalog row per file, with Auto Loader handling incremental arrival so new studies flow in continuously rather than in fragile nightly batches. Silver: the text-valued DICOM tags are extracted into Delta tables (binary blobs like the pixel data stay in the file), so the metadata that was trapped inside the files becomes queryable with plain SQL. This is the step that turns "a viewer can open it" into "an analyst can query it." The open-source Pixels accelerator does exactly this, cataloging files in parallel and extracting tags at scale. Gold: curated, analysis-ready cohorts, joined to clinical and omics tables and ready for training, BI, or a regulator-facing study. The reason this is non-trivial is throughput. DICOM parsing libraries are single-core; the trick is to distribute them. In practice that means pandas UDFs and mapInPandas to fan parsing across a cluster, or a custom Spark data source. In work the Pixels team published, a zipdcm data source reads DICOM metadata straight out of zip archives in memory, leaving the original files compressed in place: it cataloged more than 107,000 DICOMs in about 3.5 minutes on two 8-core workers, roughly 7x faster than prior approaches. The bottleneck shifts from disk and network I/O to pure CPU parsing, whi [truncated for AI cost control]

展開要點與分析

文章情報

工程師進階

要點

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Part I — The problemFour sectors, one bottleneckHospitals and health systems generate...

技術影響

可能影響 GPU、推理集羣、算力成本和供應鏈規劃。

要點與分析由自動化流程生成,可能有誤,請結合原始來源核實。