Data Fusion and Contrastive Alignment for Unconstrained IR Molecular Structure Elucidation
Researchers propose modifications to encoder-decoder transformers for automated molecular structure elucidation from IR spectroscopy without requiring pre-determined chemical formulas. Using a Mixture-of-Experts (MoE) decoder with non-additive aggregation (linear-order statistics and Choquet integral) and an auxiliary contrastive alignment loss, they improve Top-K prediction accuracy by over 10 percentage points over baseline IR-only models. Analysis confirms IR spectra encode most relevant chemical information, broadening AI's utility in analytical chemistry.
-->
[Submitted on 28 Jul 2026]
Title:Data Fusion and Contrastive Alignment for Unconstrained IR Molecular Structure Elucidation
View a PDF of the paper titled Data Fusion and Contrastive Alignment for Unconstrained IR Molecular Structure Elucidation, by Ethan J. Mick and 3 other authors
View PDF HTML (experimental)
Abstract:Automated molecular structure elucidation from infrared (IR) spectroscopy data has seen significant advancements in recent years, but its broad applicability is limited by a reliance on pre-determined chemical formulas provided as auxiliary model inputs. This limits model predictions to isomer identification rather than full molecular structure prediction. Although transformer models have been shown to identify molecular isomers with high accuracy, their reliability for unconstrained structure elucidation is comparatively low and poorly understood. In this work, we propose and evaluate key modifications to the traditional encoder-decoder transformer. To better address the vast chemical space of the unconstrained problem, we implement a novel Mixture-of-Experts (MoE) decoder module that utilizes non-additive aggregation via linear-order statistics and the Choquet integral. We further modify the transformer to utilize these non-additive operators when aggregating spectral representations as well. Together with an auxiliary contrastive alignment loss term, these enhancements improve Top-K prediction accuracy by over 10 percentage points compared to baseline IR-only models. Through sub-structure fragment analysis of molecular predictions, we further confirm that infrared spectra encode the vast majority of relevant chemical information, implying that the higher performance of isomer-ranking models is largely due to underrepresented or overlapping absorption bands for molecules in the explored chemical space. Ultimately, by demonstrating the efficacy of automated molecular structure elucidation from measured IR spectra, this work serves to significantly broaden the utility of AI in analytical chemistry.
Subjects:
Machine Learning (cs.LG)
Cite as: arXiv:2607.26164 [cs.LG]
(or arXiv:2607.26164v1 [cs.LG] for this version)
https://doi.org/10.48550/arXiv.2607.26164
arXiv-issued DOI via DataCite (pending registration)
Submission history
From: Ethan Mick [view email] [v1] Tue, 28 Jul 2026 18:15:46 UTC (1,670 KB)
Full-text links:
Access Paper:
View a PDF of the paper titled Data Fusion and Contrastive Alignment for Unconstrained IR Molecular Structure Elucidation, by Ethan J. Mick and 3 other authors
View PDF
HTML (experimental)
TeX Source
view license
Current browse context:
cs.LG
new | recent | 2026-07
Change to browse by:
cs
References & Citations
NASA ADS
Google Scholar
Semantic Scholar
Loading...
Data provided by:
Bibliographic Tools
Bibliographic and Citation Tools
Bibliographic Explorer Toggle
Bibliographic Explorer (What is the Explorer?)
Connected Papers Toggle
Connected Papers (What is Connected Papers?)
Litmaps Toggle
Litmaps (What is Litmaps?)
scite.ai Toggle
scite Smart Citations (What are Smart Citations?)
Code, Data, Media
Code, Data and Media Associated with this Article
alphaXiv Toggle
alphaXiv (What is alphaXiv?)
Links to Code Toggle
CatalyzeX Code Finder for Papers (What is CatalyzeX?)
DagsHub Toggle
DagsHub (What is DagsHub?)
GotitPub Toggle
Gotit.pub (What is GotitPub?)
Huggingface Toggle
Hugging Face (What is Huggingface?)
ScienceCast Toggle
ScienceCast (What is ScienceCast?)
Demos
Demos
Replicate Toggle
Replicate (What is Replicate?)
Spaces Toggle
Hugging Face Spaces (What is Spaces?)
Spaces Toggle
TXYZ.AI (What is TXYZ.AI?)
Related Papers
Recommenders and Search Tools
Link to Influence Flower
Influence Flower (What are Influence Flowers?)
Core recommender toggle
CORE Recommender (What is CORE?)
IArxiv recommender toggle
IArxiv Recommender (What is IArxiv?)
Author
Venue
Institution
Topic
About arXivLabs
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)