Detailansicht

Flexible mass spectral networking
making sense of mass spectral data through data mining, statistical prioritization, and interactive visualization
Kevin Amédé Mildau
Art der Arbeit
Dissertation
Universität
Universität Wien
Fakultät
Fakultät für Chemie
Studiumsbezeichnung bzw. Universitätlehrgang (ULG)
Doktoratsstudium Naturwissenschaften: Chemie
Betreuer*in
Jürgen Zanghellini
Volltext in Browser öffnen
Alle Rechte vorbehalten / All rights reserved
DOI
10.25365/thesis.77944
URN
urn:nbn:at:at-ubw:1-23355.16321.770630-3
Link zu u:search
(Print-Exemplar eventuell in Bibliothek verfügbar)

Abstracts

Abstract
(Deutsch)
Hintergrund: Das Gebiet der ungezielten Metabolomik (“untargeted metabolomics”) befasst sich mit der umfassenden Messung und Charakterisierung der Gesamtheit kleiner Moleküle, d. h. Metaboliten, in biologischen Proben. In der Praxis können jedoch durchschnittlich nur etwa 10% der gemessenen Merkmale (“features”) identifiziert werden. Diese geringe Identifizierungsrate erschwert die wissenschaftliche Interpretation experimenteller Ergebnisse erheblich. Denn ohne chemische Annotationen bleibt die erfasste Datenmenge eine Ansammlung abstrakter Unbekannter. Die computergestützte Metabolomik hilft Forschenden, ihre Daten besser zu verstehen, indem sie Werkzeuge zur verbesserten Verarbeitung, Organisation, Exploration und Annotation bereitstellt. Resultate: In dieser Arbeit wurden drei Computerprogramme entwickelt, die Forschende bei der Analyse ihrer gemessenen, holistischen Metabolome unterstützt: homologueDiscoverer, specXplore und msFeaST. homologueDiscoverer adressiert die Probleme der Datenredundanz und der fehlenden Gruppierung von spektralen Peaks, die homologen Serien in LC-MS-Daten zugeordnet sind. Homologe Serien treten häufig in LC-MS-Daten auf und zeigen charakteristische Signaturen. Der homologueDiscoverer-Algorithmus nutzt diese systematischen Muster, um homologe Serien zu identifizieren und zu gruppieren. Dies reduziert die Datenkomplexität und erlaubt entweder a) die Ausklammerung dieser Merkmalsgruppen, wenn homologe Serien als Kontamination betrachtet werden, oder b) die gezielte Analyse interessanter homologer Serien. Interaktive Visualisierungen und benutzerfreundliche Exportformate erleichtern die effiziente Verwaltung dieser Gruppen. specXplore löst die Einschränkungen der Starrheit des so genannten “Molecular Networking” (MN). Diese Methode organisiert Massenspektren anhand ihrer spektralen Ähnlichkeit mittels Netzwerkanalyse und ist eine zentrale Analyse Methode in der ungezielten Metabolomik. Bestehende Implementierungen sind jedoch empfindlich gegenüber Parametereinstellungen und schwer an individuelle Analyse Anforderungen anzupassen. Die Methode verfängt Visualisierungs- und Clusteranalyseaspekte ineinander durch topologische Verarbeitungsschritte, was die Flexibilität einschränkt. Das specXplore-Dashboard bietet eine flexible Plattform zur Datenexploration, in der Einstellungen frei angepasst und Beziehungen zwischen spectralen Merkmalen interaktiv untersucht werden können. msFeaST adressiert das Problem der formalen, statistischen Analyse von Spektralmerkmalgruppen, wie zum Beispiel molekularer Familien, zur verbesserten Priorisierung. Anstatt einzelne Merkmale zu priorisieren, die ohne chemische Interpretation oder Identität oft wenig aussagekräftig sind, hebt msFeaST Gruppen von Merkmalen hervor. Aufbauend auf der Entkopplung von Clusteranalys und Visualisierung in specXplore, verwendet msFeaST interaktive Strategien zu top-K Ego-Netzwerk-Exploration. Dies ermöglicht die Untersuchung spektraler Nachbarschaften unabhängig von vorgegebenen Ähnlichkeitsschwellen. Gemeinsam ist allen drei Projekten der Fokus auf Data-Mining, flexible Priorisierung von Merkmalen und visuelle Datenexploration. Ein besonderer Schwerpunkt liegt dabei auf der Rolle interaktiver Visualisierungen, die es Forschenden erleichtern, die in der Metabolomik anfallenden riesigen Datenmenge effektiv zu analysieren. Ohne eine visuelle und interaktive Darstellungen wäre eine Bewertung dieser Daten kaum möglich. Ausblick: Ungezielte Metabolomik-Analysen erzeugen umfangreiche Datenmengen, deren Auswertung ohne geeignete Programme zur Verarbeitung und Visualisierung kaum möglich wäre Die in dieser Arbeit entwickelten Programme leisten hierzu einen Beitrag, in dem sie 1) automatisiertes Data-Mining von LC-MS-Quantifikationstabellen, 2) Organisation und Integration von LC-MS/MS-Spektraldaten zur Generierung von Übersichten und Exploration, sowie 3) Kombination spektralähnlichkeitsbasierter Gruppierungen mit quantitativer Informationen für kontextualisierte Priorisierung und Analyse ermöglichen.
Abstract
(Englisch)
Background: The field of untargeted metabolomics deals with the comprehensive measurement and characterization of the collection of small molecules, i.e., metabolites, found in biological samples. However, in practice only an average of 10% of features measured can be identified. This low identification rate impedes the scientific interpretation of experimental results. Without chemical annotations data collected are a large collection of abstract unknowns. Computational Metabolomics tools assists practitioners in the field of untargeted metabolomics in making sense of their data by providing tools to better process, organize, explore, and annotate the generated data. Results: In this thesis, we have developed three tools to assist researchers in making sense of their data, namely homologueDiscoverer, specXplore, and msFeaST. In developing homologueDiscoverer, we have tackled data redundancy issues and a lack of grouping of peaks that are part of homologous series in LC-MS data. Homologous series are frequently encountered in LC-MS data and tend to exhibit highly regular LC-MS patterns. The homologueDiscoverer algorithm capitalizes on systematic LC-MS trends exhibited by homologous series to group them. The resulting grouping allows reducing data complexity and either a) provides feature sets for exclusion from further analysis if homologous series are considered a contamination, thereby allowing researchers to focus on biologically relevant information, or b) provides feature sets of interest for further analysis if particular homologous series are of interest. Interactive visualization and convenient export formats enable the straightforward managing of these series. In developing specXplore, we have tackled rigidity limitations of mass spectral molecular networking, that is, the organization of mass spectral similarity data using network visualization approaches. Molecular networking is a core tool in the untargeted metabolomics pipelines. However, current implementations, while sensitive to settings used, are difficult to adjust and tune. They entangle visualization and clustering considerations via topological processing steps, making it difficult to tailor the tool to the needs of the individual researcher. The specXplore dashboard provides a flexible data exploration platform within which settings can be modified liberally and relationships between features can be studied in an interactive and interactive fashion. In the development of msFeaST, we tackled the issue of formalizing statistical inferences on sets of spectral features such as molecular families. Here, the aim is to provide an improved workflow for prioritization. Rather than prioritizing individual features, which often lack chemical identity information and are thus not insightful outside of their group context, groups of features are highlighted jointly. Building on top of the disentangling of clustering from visualization or topology considerations explored in specXplore, msFeaST makes use of interactive top-K ego-network neighborhood exploration visualization strategies, allowing exploration of neighbors irrespective of similarity cutoffs. Common to all three projects are the central themes of data mining, flexible feature prioritization, and visual data exploration. Special emphasis was placed on how interactive visualization can assist researchers in their efforts. The latter is essential as the plethora of information produced would be impossible to evaluate without visual and interactive intermediary representations. Outlook: Untargeted metabolomics analysis pipelines can produce a tremendous amount of information, yet can be difficult to analyze without appropriate processing and visualization tools. The tools developed as part of this thesis contribute to this by developing tools aiming to assist researchers with 1) automatic data mining of LC-MS quantification tables, 2) LCMS/ MS spectral data organization and integration for overview generation and exploration, and 3) merging of spectral similarity-based clustering and feature-specific quantification information for more contextualized prioritization and inference.

Schlagwörter

Schlagwörter
(Deutsch)
Metabolomik Ungezielte Metabolomik Datenvisualisierung Bioinformatik
Schlagwörter
(Englisch)
metabolomics untargeted metabolomics data visualization bioinformatics
Autor*innen
Kevin Amédé Mildau
Haupttitel (Englisch)
Flexible mass spectral networking
Hauptuntertitel (Englisch)
making sense of mass spectral data through data mining, statistical prioritization, and interactive visualization
Publikationsjahr
2024
Umfangsangabe
165 Seiten in verschiedenen Seitenzählungen : Illustrationen
Sprache
Englisch
Beurteiler*innen
Louis-Félix Nothias ,
Wolfram Weckwerth
Klassifikationen
35 Chemie > 35.06 Computeranwendungen ,
35 Chemie > 35.26 Massenspektrometrie
AC Nummer
AC17467343
Utheses ID
74393
Studienkennzahl
UA | 796 | 605 | 419 |
Universität Wien, Universitätsbibliothek, 1010 Wien, Universitätsring 1