AI News HubLIVE
Original source2 min read

Data Fusion and Contrastive Alignment for Unconstrained IR Molecular Structure Elucidation

Researchers propose modifications to encoder-decoder transformers for automated molecular structure elucidation from IR spectroscopy without requiring pre-determined chemical formulas. Using a Mixture-of-Experts (MoE) decoder with non-additive aggregation (linear-order statistics and Choquet integral) and an auxiliary contrastive alignment loss, they improve Top-K prediction accuracy by over 10 percentage points over baseline IR-only models. Analysis confirms IR spectra encode most relevant chemical information, broadening AI's utility in analytical chemistry.

SourcearXiv Machine LearningAuthor: Ethan J. Mick, Campbell A. Sweet, Matthias J. Young, Derek T. Anderson

-->

[Submitted on 28 Jul 2026]

Title:Data Fusion and Contrastive Alignment for Unconstrained IR Molecular Structure Elucidation

View a PDF of the paper titled Data Fusion and Contrastive Alignment for Unconstrained IR Molecular Structure Elucidation, by Ethan J. Mick and 3 other authors

View PDF HTML (experimental)

Abstract:Automated molecular structure elucidation from infrared (IR) spectroscopy data has seen significant advancements in recent years, but its broad applicability is limited by a reliance on pre-determined chemical formulas provided as auxiliary model inputs. This limits model predictions to isomer identification rather than full molecular structure prediction. Although transformer models have been shown to identify molecular isomers with high accuracy, their reliability for unconstrained structure elucidation is comparatively low and poorly understood. In this work, we propose and evaluate key modifications to the traditional encoder-decoder transformer. To better address the vast chemical space of the unconstrained problem, we implement a novel Mixture-of-Experts (MoE) decoder module that utilizes non-additive aggregation via linear-order statistics and the Choquet integral. We further modify the transformer to utilize these non-additive operators when aggregating spectral representations as well. Together with an auxiliary contrastive alignment loss term, these enhancements improve Top-K prediction accuracy by over 10 percentage points compared to baseline IR-only models. Through sub-structure fragment analysis of molecular predictions, we further confirm that infrared spectra encode the vast majority of relevant chemical information, implying that the higher performance of isomer-ranking models is largely due to underrepresented or overlapping absorption bands for molecules in the explored chemical space. Ultimately, by demonstrating the efficacy of automated molecular structure elucidation from measured IR spectra, this work serves to significantly broaden the utility of AI in analytical chemistry.

Subjects:

Machine Learning (cs.LG)

Cite as: arXiv:2607.26164 [cs.LG]

(or arXiv:2607.26164v1 [cs.LG] for this version)

https://doi.org/10.48550/arXiv.2607.26164

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ethan Mick [view email] [v1] Tue, 28 Jul 2026 18:15:46 UTC (1,670 KB)

Full-text links:

Access Paper:

View a PDF of the paper titled Data Fusion and Contrastive Alignment for Unconstrained IR Molecular Structure Elucidation, by Ethan J. Mick and 3 other authors

View PDF

HTML (experimental)

TeX Source

view license

Current browse context:

cs.LG

new | recent | 2026-07

Change to browse by:

cs

References & Citations

NASA ADS

Google Scholar

Semantic Scholar

Loading...

Data provided by:

Bibliographic Tools

Bibliographic and Citation Tools

Bibliographic Explorer Toggle

Bibliographic Explorer (What is the Explorer?)

Connected Papers Toggle

Connected Papers (What is Connected Papers?)

Litmaps Toggle

Litmaps (What is Litmaps?)

scite.ai Toggle

scite Smart Citations (What are Smart Citations?)

Code, Data, Media

Code, Data and Media Associated with this Article

alphaXiv Toggle

alphaXiv (What is alphaXiv?)

Links to Code Toggle

CatalyzeX Code Finder for Papers (What is CatalyzeX?)

DagsHub Toggle

DagsHub (What is DagsHub?)

GotitPub Toggle

Gotit.pub (What is GotitPub?)

Huggingface Toggle

Hugging Face (What is Huggingface?)

ScienceCast Toggle

ScienceCast (What is ScienceCast?)

Demos

Demos

Replicate Toggle

Replicate (What is Replicate?)

Spaces Toggle

Hugging Face Spaces (What is Spaces?)

Spaces Toggle

TXYZ.AI (What is TXYZ.AI?)

Related Papers

Recommenders and Search Tools

Link to Influence Flower

Influence Flower (What are Influence Flowers?)

Core recommender toggle

CORE Recommender (What is CORE?)

IArxiv recommender toggle

IArxiv Recommender (What is IArxiv?)

Author

Venue

Institution

Topic

About arXivLabs

arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)