Detailansicht
RNA structuredness of viral genomes
Teodora Bucaciuc Mracica
Art der Arbeit
Masterarbeit
Universität
Universität Wien
Fakultät
Fakultät für Informatik
Studiumsbezeichnung bzw. Universitätlehrgang (ULG)
Masterstudium Bioinformatik
Betreuer*in
Hofacker Ivo
DOI
10.25365/thesis.72639
URN
urn:nbn:at:at-ubw:1-28965.79136.475438-2
Link zu u:search
(Print-Exemplar eventuell in Bibliothek verfügbar)
Abstracts
Abstract
(Deutsch)
Das Ziel dieser Masterarbeit ist es, die Strukturiertheit von allen vollständigen RNS-Virengenomen aus der NCBI Datenbank zu bestimmen. Dazu wurde eine Methode entwickelt, den Z-Score für alle Nukleotide eines Genoms zu berechnen, mithilfe dessen die gesamten thermodynamischen Fähigkeiten des Genoms, stabile sekundäre Strukturen zu bilden, ausgedrückt werden können. Das RNAfold-Modul des ViennaRNA-Programms wurde in Verbindung mit einer Sliding-Window-Methode verwendet, um ein ganzes Genom zu scannen; in jedem Sliding Window wurde die Minimum Free Energy (MFE) berechnet sowie ein Z-Score bestimmt. Der Z-Score wird kalkuliert, indem die native MFE mit dem Mittelwert der MFEs mehrerer zufälliger Sequenzen verglichen wird, die dieselbe Länge und Nukleotidzusammensetzung haben wie die native MFE; die Signifikanz einer errechneten MFE wird in Form des Z-Scores als die Anzahl der Standardabweichungen vom Mittelwert dargestellt. Folglich weist ein negativer Z-Score auf eine im Vergleich zu zufälligen RNS-Sequenzen stabilere native RNS-Sequenz hin. Für jede Position im Genom wurde ein durchschnittlicher Z-Score berechnet, indem die MFE-Werte aus all jenen Sliding Windows gemittelt wurden, die ebendiese Position enthalten. Ein automatisiertes Python-Programm wurde entwickelt, um den durchschnittlichen Z-Score in jedem Organismus aus einem Taxon zu bestimmen. Für jede Virusspezies wurden automatisch eigene SLURM-Dateien erstellt und an das Computer-Cluster des Instituts für Theoretische Chemie gesendet. Dadurch werden neben der Z-Score-Berechnung für jedes Nukleotid die Output-Daten gleichzeitig vor- und nachbearbeitet; dazu zählen die Visualisierung der Genom-Annotationen zusammen mit ihren durchschnittlichen Z-Scores oder die Integration anderer ViennaRNA-Module in Form von Bash-Skripten und statistischen Datenanalysen. Insgesamt wurden 4.142 RNS-Genome einzeln analysiert, annotiert, verarbeitet und visualisiert. Zusätzlich zur Strukturanalyse des Genoms anhand der Z-Scores wurde ebenfalls eine ausführliche Vergleichsanalyse zur Verteilung der Z-Scores in den kodierenden und nicht-kodierenden Teilen der RNS durchgeführt. Die RNALfold- und RNAplfold-Module aus dem ViennaRNA-Programmpaket wurden verwendet, um die Stabilität lokaler Sekundärstrukturen zu analysieren sowie den Wert der benötigten freien Energie zu berechnen, um diese Sequenzen zu entfalten; so soll die Ausdrucksqualität eines spezifischen Proteins oder einer nicht-kodierenden RNS bestimmt werden. Diese Masterarbeit zeigt die globale RNS-Strukturiertheit in verschiedenen RNS-Virengruppen und erfasst wichtige Erkenntnisse zu den Strukturen der jeweiligen RNS-Spezies, welche nicht nur für den Lebenszyklus eines Virus sondern potentiell auch für die medizinische sowie biotechnologische Forschung entscheidende Implikationen aufzeigen.
Abstract
(Englisch)
The goal of this work is to assess the structuredness of all complete RNA viral genomes available in the NCBI database. To this end, a method was developed to determine the Z-score of all nucleotides in a given genome, as a way to express the overall thermodynamic ability of the viral genome to form stable secondary structures. A sliding window approach was employed that computes a local measure of structuredness for each position by averaging over all enclosing windows. In order to make the measure of structuredness independent of sequence composition, a Z-score was used, comparing the folding energy of a sequence window to randomized sequences of the same composition. Folding energies were determined using the RNAfold program from the ViennaRNA package. Apart from the genome's overall assessment of structuredness using the Z-score, a more in depth analysis was carried out to compare the distribution of Z-scores between the coding regions and non-coding regions in the genome. Two other different methods were also used to assess structuredness, namely the free opening energies and locally stable structure formation in the whole genome. The usage of different methods to investigate the ability of a genome to form secondary structures were employed as a way to understand these contribute to the stability of the genome, thus protecting it from degradation. A Python automated approach was developed to determine the mean Z-scores per organism for a given taxon, by producing analysis files for each virus species and sending each script on a computer cluster, as well as for pre and post processing the output data, ranging from visualization of the annotated genomes together with their mean Z-scores, to integrating modules from the ViennaRNA package in the form of bash scripts and statistical evaluation of the data. A total of 4142 RNA genomes were individually analyzed, annotated, processed and visualized. The modules RNALfold and RNAplfold from the ViennaRNA package were used to better estimate the stability of local secondary structures and the value of the free energy needed to unfold those structures, as a way to determine the quality of expression of a specific protein or non-coding RNA.This thesis provides a global picture of RNA structuredness across different virus groups, as well as relevant knowledge about each species' RNA structure, which has paramount functional implications not only for the viral life-cycle, but also potentially for medical research and biotechnology. We find that structuredness varies significantly between different groups of viruses, as well as between different genomic regions. In particular the single stranded positive sense RNA viruses form more stable structures than expected by chance, with their non-coding region being more structured than the coding region. In contrast, the single stranded negative sense RNA viruses appear to be more structured in the coding regions, with no significant differences between the coding and non-coding region on both forward and backward strand. Double stranded RNA viruses have on average less structured genomes than the other two groups, regardless of genomic region.
Schlagwörter
Schlagwörter
(Deutsch)
RNA Bioinformatik RNA sekundaer Struktur RNA Virengenome Z-score
Schlagwörter
(Englisch)
RNA Bioinformatics RNA secondary structure RNA viral genomes Z-score
Autor*innen
Teodora Bucaciuc Mracica
Haupttitel (Englisch)
RNA structuredness of viral genomes
Paralleltitel (Deutsch)
RNA Strukturiertheit von viralen Genomen
Publikationsjahr
2022
Umfangsangabe
101 Seiten : Illustrationen
Sprache
Englisch
Beurteiler*in
Hofacker Ivo
Klassifikation
30 Naturwissenschaften allgemein > 30.99 Naturwissenschaften allgemein: Sonstiges
AC Nummer
AC16679534
Utheses ID
65144
Studienkennzahl
UA | 066 | 875 | |
