Overview: Systematic Process for the Identification of Unknowns
by GC-MS and LC-MS

Introduction
Our philosophy is simple: many analytical laboratories are not trying to identify completely novel molecules. Instead, they are trying to identify compounds that have already been reported in the scientific literature or included in commercial databases but remain unknown to the analyst. We call these compounds "known unknowns."
This systematic process is described in detail in our February 2013 LCGC article, "Identifying Known Unknowns in Commercial Products by Mass Spectrometry." The accompanying Known Unknown Overview further summarizes the use of large "spectra-less" databases, such as the CAS Registry (SciFinder) and ChemSpider, in combination with accurate mass data.
LCGC Article(opens in new tab)
Overview Spectra-less "Known Unknown"(opens in new tab)
The challenge is not necessarily discovering a new molecule—it is efficiently recognizing an existing one. Our goal is to combine high-quality mass spectrometry, systematic database searching, and expert interpretation to identify these compounds rapidly and confidently.
The identification of unknowns is both a science and an art. Modern software can rapidly generate candidate structures, but reliable identification still requires an understanding of mass spectrometry, chemistry, fragmentation mechanisms, isotope patterns, and complementary databases. Throughout this website, we present a practical workflow developed over nearly five decades of industrial and research experience to help analysts identify unknown compounds with greater confidence.
No single software package identifies unknowns automatically. Successful identification comes from combining high-quality experimental data with multiple complementary tools and a logical decision-making process.
Modern NIST software(opens in new tab)—including Hybrid Search, accurate-mass support, integrated deconvolution and library searching, MS Interpreter, and structure searching—has significantly expanded the ability to identify compounds that are not exact library matches. These library searches are routinely complemented by accurate-mass, molecular formula, and molecular weight searches of large "spectra-less" databases such as the CAS Registry (SciFinder)(opens in new tab) and ChemSpider(opens in new tab), as described in our other publications.
Candidate structures are evaluated using all available evidence, including:
- EI or MS/MS spectra
- NIST Identity and Hybrid Searches
- Accurate mass and molecular formula
- Isotope patterns
- Fragmentation pathways
- Chemical ionization data
- Adduct chemistry
- Exchangeable hydrogens
- Number of associated literature references
- Relevant keywords
- Structural plausibility
- Sample history and analytical context
The correct identification is rarely based on a single measurement. Confidence comes from the agreement of multiple independent pieces of evidence.
Our goal is not simply to teach software—we teach a systematic way of thinking about unknown identification. While instrumentation, software, and databases continue to evolve, the underlying scientific process remains the same: acquire high-quality data, generate reasonable candidate structures, evaluate all available evidence, and arrive at the most chemically defensible identification.
It is also important to recognize that not every unknown is a “known unknown.” Some compounds may not be represented in either commercial mass spectral libraries or large chemical structure databases. These “unknown unknowns” require a more fundamental investigative approach based on chemical knowledge, sample history, accurate-mass and isotope data, NIST hybrid search results, fragmentation interpretation, reaction chemistry, and other available analytical evidence. Their identification may require proposing and evaluating structures rather than simply locating an existing database entry. Although the tools and level of difficulty differ, the same systematic principle applies: combine all available evidence to arrive at the most chemically defensible conclusion.
Fair Winds and Following Seas!

