PubChemLite plus Collision Cross Section (CCS) values for enhanced interpretation of non-target environmental data

Emma Schymanski,Anjana Elapavalore,Dylan Ross,Valentin Groues,Dagny Aurich,Allison Krinsky,Sunghwan Kim,Paul Thiessen,Jian Zhang,James Dodds,Erin Baker,Evan Bolton,Libin Xu
DOI: https://doi.org/10.26434/chemrxiv-2024-2xcsq
2024-11-22
Abstract:Finding relevant chemicals in the vast (known) chemical space is a major challenge for environmental and exposomics studies leveraging non-target high resolution mass spectrometry (NT-HRMS) methods. Chemical databases now contain hundreds of millions of chemicals, yet many are not relevant. This article details an extensive collaborative, open science effort to provide a dynamic collection of chemicals for environmental, metabolomics and exposomics research, along with supporting information about their relevance to assist researchers in the interpretation of candidate hits. The PubChemLite for Exposomics collection is compiled from ten annotation categories within PubChem, enhanced with patent, literature and annotation counts, predicted partition coefficient (logP) values, as well as predicted collision cross section (CCS) values using CCSbase. Monthly versions are archived on Zenodo under a CC-BY license, supporting reproducible research, and a new interface has been developed, including the chemical stripes on patent and literature data, for researchers to browse the collection. This article further describes how PubChemLite can support researchers in environmental/exposomics studies, describes efforts to increase the availability of experimental CCS values, and explores known limitations and potential for future developments. The data and code behind these efforts are openly available. PubChemLite content can be explored at https://pubchemlite.lcsb.uni.lu.
Chemistry
What problem does this paper attempt to address?