Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Bioinformatics (56)
- Genomics (31)
- Genetics (21)
- Physical Sciences and Mathematics (19)
- Biology (18)
-
- Molecular Genetics (18)
- Biochemistry, Biophysics, and Structural Biology (13)
- Medicine and Health Sciences (12)
- Computer Sciences (11)
- Molecular Biology (9)
- Biotechnology (8)
- Cell and Developmental Biology (7)
- Ecology and Evolutionary Biology (7)
- Statistics and Probability (7)
- Laboratory and Basic Science Research (6)
- Systems Biology (6)
- Evolution (5)
- Immunology and Infectious Disease (5)
- Microbiology (5)
- Neuroscience and Neurobiology (5)
- Numerical Analysis and Scientific Computing (5)
- Structural Biology (5)
- Applied Mathematics (4)
- Artificial Intelligence and Robotics (4)
- Biostatistics (4)
- Behavioral Neurobiology (3)
- Biochemistry (3)
- Institution
-
- The Texas Medical Center Library (9)
- Augustana College (8)
- Dartmouth College (7)
- University of Nebraska - Lincoln (6)
- COBRA (5)
-
- University of Kentucky (5)
- Old Dominion University (3)
- Virginia Commonwealth University (3)
- Clemson University (2)
- Munster Technological University (2)
- University of Connecticut (2)
- Duke Law (1)
- Jacksonville State University (1)
- LSU New Orleans (1)
- Louisiana State University (1)
- Loyola University Chicago (1)
- Medical University of South Carolina (1)
- Mississippi State University (1)
- Nova Southeastern University (1)
- Rowan University (1)
- Seton Hall University (1)
- The University of Southern Mississippi (1)
- University of Arkansas, Fayetteville (1)
- University of Louisville (1)
- University of Montana (1)
- University of Nebraska Medical Center (1)
- University of New Mexico (1)
- West Virginia University (1)
- Wilfrid Laurier University (1)
- Publication Year
- Publication
-
- Dissertations and Theses (Open Access) (9)
- Meiothermus ruber Genome Analysis Project (8)
- Dartmouth Scholarship (6)
- Computer Science Faculty Publications (3)
- Theses and Dissertations (3)
-
- COBRA Preprint Series (2)
- Honors Scholar Theses (2)
- The University of Michigan Department of Biostatistics Working Paper Series (2)
- Theses and Dissertations--Biology (2)
- All Dissertations (1)
- All Theses (1)
- Bioconductor Project Working Papers (1)
- Biomedical Engineering Undergraduate Honors Theses (1)
- Biomedical Sciences ETDs (1)
- Complex Biosystems Program: Dissertations and Student Research (1)
- Computer Science Senior Theses (1)
- Computer Science: Faculty Publications and Other Works (1)
- Department of Biological Sciences Publications (1)
- Department of Computer Science Publications (1)
- Department of Electrical and Computer Engineering: Dissertations, Theses, and Student Research (1)
- Department of Food Science and Technology: Dissertations, Theses, and Student Research (1)
- Dissertations (1)
- Electronic Theses and Dissertations (1)
- Faculty Scholarship (1)
- Graduate School of Biomedical Sciences Theses and Dissertations (1)
- Graduate Student Theses, Dissertations, & Professional Papers (1)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (1)
- HCNSO Student Theses and Dissertations (1)
- Honors Program: Senior Projects (Public) (1)
- LSU Doctoral Dissertations (1)
- Publication Type
Articles 1 - 30 of 70
Full-Text Articles in Computational Biology
Biologically Informed Negative Samplingfor Antibody Chain Pairing Classification, Ishita Singh
Biologically Informed Negative Samplingfor Antibody Chain Pairing Classification, Ishita Singh
Computer Science Senior Theses
Antibody heavy and light chain (H/L) pairing is fundamental to antigen recognition and stability. While single-cell sequencing preserves native pairing information, widely used bulk repertoire and spatial transcriptomics platforms do not, motivating the need for efficient ML methods to infer H/L pairing. Training a binary classifier for this task faces the methodological challenge of a lack of true biological negatives, since natural selection eliminates B cells with incompatible H/L pairs.
In this thesis, I introduce a biologically informed negative sampling strategy for H/L pairing classification, drawing on known V-gene biases in heavy and light chain pairing. Pseudo-negatives are constructed by …
Identifying Rna Splicing Changes During Alcohol Withdrawal Using An Optimized Rna-Seq Analysis Pipeline, Yasaswi Veera, Luana Martins De Carvalho, Amy Lasek
Identifying Rna Splicing Changes During Alcohol Withdrawal Using An Optimized Rna-Seq Analysis Pipeline, Yasaswi Veera, Luana Martins De Carvalho, Amy Lasek
Undergraduate Research Posters
Alcohol use disorder (AUD) causes long-lasting changes in brain gene expression and RNA splicing, particularly in the ventral hippocampus. This study analyzes RNA-Seq data from rats exposed to chronic alcohol and withdrawal to identify transcript-level and splicing alterations. An optimized RNA-Seq pipeline using STAR, FeatureCounts, edgeR, and WGCNA improved efficiency by 25-40% while maintaining consistency across 25 datasets. Results reveal changes in neural signaling and stress-response pathways associated with withdrawal. This work provides both biological insight into AUD and a reproducible computational framework for transcriptomic analysis.
A Persistent Homology Framework For Scrna-Seq: Assessing Clustering Robustness And Quantifying Preprocessing And Integration Effects On Topological Features., Jonah Daneshmand
A Persistent Homology Framework For Scrna-Seq: Assessing Clustering Robustness And Quantifying Preprocessing And Integration Effects On Topological Features., Jonah Daneshmand
Electronic Theses and Dissertations
As single-cell RNA sequencing (scRNA-seq) data expands, robust methods for integrating diverse datasets are critical. This dissertation applies Persistent Homology (PH), a technique from Topological Data Analysis (TDA), to a collection of scRNA-seq datasets spanning eight tissue types to quantify how data integration affects topological features and biological interpretability. We assessed global topological structure using Betti curves, Euler characteristics, and persistence landscapes across raw, normalized, and integrated data representations. Our analysis revealed a performance inversion: while conventional methods excelled on unintegrated data, high-granularity topological methods, particularly those sensitive to global data structure, became superior after integration. This suggests a synergy …
Bioinformatic Analysis Of Pogz Variants In Relation To White Sutton Syndrome, Hannah Rollins
Bioinformatic Analysis Of Pogz Variants In Relation To White Sutton Syndrome, Hannah Rollins
Theses
White-Sutton syndrome (WHSUS) is a rare neurodevelopmental disorder caused by mutations in the Pogo Transposable Element with ZNF Domain (POGZ) gene, which encodes pogo-transposable element with ZNF domain, a chromatin regulator essential for proper mitotic progression and DNA repair. This study uses a bioinformatic framework to evaluate the structural and functional impact of missense mutations in the conserved amino acid region (positions 500–800) of the POGZ protein. Protein modeling, variant effect prediction, conservation analysis, and molecular dynamics simulations were employed to gain an understanding of the effects of POGZ missense mutations on protein structure and movement with specific emphasis on …
Cazyme Gene Cluster Diversity In Human Gut Microbiome, Yi Xing
Cazyme Gene Cluster Diversity In Human Gut Microbiome, Yi Xing
Department of Food Science and Technology: Dissertations, Theses, and Student Research
In gut microbiome research, carbohydrate-active enzyme gene clusters (CGCs) have emerged as key functional units for understanding microbial glycan degradation. Unlike taxonomic or broad pathway annotations, CGCs offer gene-cluster-level resolution and capture substrate-specific microbial functions. However, their diversity and distribution in relation to host metabolic phenotypes, such as obesity, remain poorly characterized. This study tests the hypothesis that the composition and abundance of fiber-targeting CGCs vary between obese and healthy human gut microbiomes, reflecting distinct microbial carbohydrate utilization strategies. To examine this, we constructed a high-quality reference CGC dataset comprising 94,019 clusters from the Unified Human Gastrointestinal Genome and profiled …
Investigating The Effects Of Transcription Factor Binding And Genetic Variants In The Striatum Of Post-Mortem Cohorts With Opioid Use Disorder, Rajashree Chakraborty
Investigating The Effects Of Transcription Factor Binding And Genetic Variants In The Striatum Of Post-Mortem Cohorts With Opioid Use Disorder, Rajashree Chakraborty
Theses & Dissertations
The opioid crisis has emerged as one of the most pressing public health challenges of our time, with Opioid Use Disorder (OUD) affecting millions of lives across the globe. Studying OUD is not merely an academic pursuit but a critical necessity in addressing this multifaceted epidemic. The urgency of this research is underscored by the staggering prevalence of OUD, with an estimated 3.7% of U.S. adults requiring treatment in 2022 alone. Despite the availability of effective medications for OUD, a significant treatment gap persists, with only a quarter of those in need receiving these life-saving interventions. The far-reaching consequences of …
Benchmarking Batch-Effect Correction Methods Towards The Construction Of A Triple-Negative Breast Cancer Cell Atlas, Peter Scheible, Amy H. Tang, Jing He, Jiangwen Sun
Benchmarking Batch-Effect Correction Methods Towards The Construction Of A Triple-Negative Breast Cancer Cell Atlas, Peter Scheible, Amy H. Tang, Jing He, Jiangwen Sun
Computer Science Faculty Publications
Triple-negative breast cancer (TNBC) requires detailed cellular mapping given its aggressive nature, immense tumor heterogeneity and genetic diversity. We integrated 156,794 cells from six scRNA-seq datasets—including tumors, metastases, and cell lines—to build a TNBC scRNA cell atlas, focusing on batch effect mitigation while maintaining biological and molecular details. Preprocessing f ilters noise, normalizes data, and leverages PCA for integration readiness. We utilized scANVI, a semi-supervised tool, to align datasets, preserving TNBC’s complex tumor heterogeneity via marker annotations [1]. UMAPs demonstrate biological clustering in integrated data, contrasted with datasetdriven unintegrated patterns. Assessments verifying effective batch correction. This method aligns with NASA’s …
Foundation Models In Bioinformatics, Fei Guo, Renchu Guan, Yaohang Li, Qi Liu, Xiaowo Wang, Can Yang, Jianxin Wang
Foundation Models In Bioinformatics, Fei Guo, Renchu Guan, Yaohang Li, Qi Liu, Xiaowo Wang, Can Yang, Jianxin Wang
Computer Science Faculty Publications
With the adoption of foundation models (FMs), artificial intelligence (AI) has become increasingly significant in bioinformatics and has successfully addressed many historical challenges, such as pre-training frameworks, model evaluation and interpretability. FMs demonstrate notable proficiency in managing large-scale, unlabeled datasets, because experimental procedures are costly and labor intensive. In various downstream tasks, FMs have consistently achieved noteworthy results, demonstrating high levels of accuracy in representing biological entities. A new era in computational biology has been ushered in by the application of FMs, focusing on both general and specific biological issues. In this review, we introduce recent advancements in bioinformatics FMs …
Early Onset Alzheimer’S Disease Markers In Mouse Hippocampus Unveiled By Single-Cell Transcriptomic Analysis Following Cranial Radiotherapy, Tuba Aksoy
Dissertations and Theses (Open Access)
Cranial radiation therapy plays an integral role in the treatment of brain tumors but can lead to progressive cognitive deficits in survivors by mechanisms that are poorly understood. To develop preventive or mitigative strategies, it is crucial to better understand the underlying pathogenesis of radiation-induced cognitive impairments. The study investigated single-cell transcriptomics and DNA methylation changes as potential drivers of persistent cellular dysfunction after radiation exposure, specifically concentrating on the CA1-3 regions of the hippocampus and the prefrontal cortex due to their role in cognitive functions. Thirteen-week-old mice underwent whole-brain radiation at clinically relevant doses. Following whole-brain radiation, an assessment …
Transcriptomic Profiling Of Engineered Human Gene Integration In The Mouse Genome Following In Vitro Gene Editing, Ethan Potts, Made Harumi Padmaswari, Christopher Nelson
Transcriptomic Profiling Of Engineered Human Gene Integration In The Mouse Genome Following In Vitro Gene Editing, Ethan Potts, Made Harumi Padmaswari, Christopher Nelson
Biomedical Engineering Undergraduate Honors Theses
Gene replacement is a promising method of therapy for genetic diseases. However, safety and efficacy are areas that need more research. This experiment aims to use RNA sequencing and bioinformatic techniques to provide answers to these questions and provide direction to future studies to develop a gene replacement therapeutic. C2C12 mouse myoblast cells were transfected with a vector containing a CRISPR-Cas9 system and the Human Factor IX (hF9) gene in order to hijack target genes and integrate the hF9 gene. The two target genes, myoglobin (Mb) and creatine kinase (Ckm), were chosen for their high rate of expression and low …
From Code To Crops: Harnessing Bioinformatics And Artificial Intelligence (Ai) In Agricultural Omics, Lakshay Anand
From Code To Crops: Harnessing Bioinformatics And Artificial Intelligence (Ai) In Agricultural Omics, Lakshay Anand
Theses and Dissertations--Plant and Soil Sciences
Global agricultural faces numerous challenges, such as climate change, resource limitations, novel pests and diseases, increasing costs, and the ever-increasing human population. To tackle these challenges, we need innovative strategies that combine new technologies and data analytics approaches to enhance agricultural output, promote sustainable methods, and optimize resource allocation. The key to this innovation lies in understanding the complex molecular web within plants that governs their growth, defense, and adaptability mechanisms. By mastering this molecular network, we can cultivate crops that are more resilient, sustainable, and suitable for different climatic terrains. Moreover, studying the symbiotic relationship between plants and microorganisms …
Machine Learning And Rna Bioinformatics, Jason Rafe Miller
Machine Learning And Rna Bioinformatics, Jason Rafe Miller
Graduate Theses, Dissertations, and Problem Reports (ETD)
The applied science of bioinformatics encompasses computational analysis of molecular biology data. Advances in genomics and DNA sequencing technology have enabled computational analysis of ribonucleic acids (RNAs), which play diverse and critical roles in most cells. To assist the study of human RNA, we trained machine learning models on RNA nucleotide sequences, devoid of domain knowledge. We built models that distinguish long non-coding lncRNA from protein-coding mRNA, and models that predict the cytoplasmic vs. nuclear preferences of lncRNAs. In a review of published lncRNA subcellular localization classifiers, we show that the commonly used validation protocol generates optimistic performance measures, and …
Convolutional Neural Network-Based Gene Prediction Using Buffalograss As A Model System, Michael Morikone
Convolutional Neural Network-Based Gene Prediction Using Buffalograss As A Model System, Michael Morikone
Complex Biosystems Program: Dissertations and Student Research
The task of gene prediction has been largely stagnant in algorithmic improvements compared to when algorithms were first developed for predicting genes thirty years ago. Rather than iteratively improving the underlying algorithms in gene prediction tools by utilizing better performing models, most current approaches update existing tools through incorporating increasing amounts of extrinsic data to improve gene prediction performance. The traditional method of predicting genes is done using Hidden Markov Models (HMMs). These HMMs are constrained by having strict assumptions made about the independence of genes that do not always hold true. To address this, a Convolutional Neural Network (CNN) …
Integrating Omim And Intact Data For The Analysis Of Gene-Phenotype Interactions In Complex Diseases: A Linux-Based Computational Tool For Network Analysis, Devin Keane
All Theses
The field of genetics is constantly evolving. New advances in bioinformatics and computational approaches are leading to exciting new developments in our ability to treat and prevent diseases. Computational genetics provides valuable insights into the complex mechanisms and layers of biological communication that shape an organism's phenotype. Understanding these mechanisms is critical to advancing human health.
The study of diseases in genetics requires a comprehensive understanding of the interactions between various biological processes, including gene expression, protein synthesis, RNA, metabolism, and cell-cell communication. To effectively address the root causes of such diseases, multi-disciplinary approaches that integrate information from different levels …
The Genomics Of Autism-Related Genes Il1rapl1 And Il1rapl2: Insights Into Their Cortical Distribution, Cell-Type Specificity, And Developmental Trajectories, Jacob Weaver
MUSC Theses and Dissertations
Neuropsychiatric disorders have a significant impact on modern society. These disorders affect a large percentage of the population: schizophrenia has a world-wide prevalence of 1% and autism spectrum disorders (ASD) affects 1 in 59 school-aged children in the US. There is substantial evidence that most neuropsychiatric disorders have a genetic component. Thus, with the advent of high throughput sequencing much effort has gone into identifying genetic variants associated with these disorders. The emerging picture from these studies is a complex one where hundreds of genes with small effects interact with a varied landscape of common variants to result in disease. …
Methods And Tools To Improve Performance Of Plant Genome Analysis, Drew Ferrell
Methods And Tools To Improve Performance Of Plant Genome Analysis, Drew Ferrell
Theses and Dissertations
Multi -omics data analysis and integration facilitates hypothesis building toward an understanding of genes and pathway responses driven by environments. Methods designed to estimate and analyze gene expression, with regard to treatments or conditions, can be leveraged to understand gene-level responses in the cell. However, genes often interact and signal within larger structures such as pathways and networks. Complex studies guided toward describing dynamic genetic pathways and networks require algorithms or methods designed for inference based on gene interactions and related topologies. Classes of algorithms and methods may be integrated into generalized workflows for comparative genomics studies, as multi -omics …
Modeling Electrostatics In Molecular Biology And Its Relevance With Molecular Mechanisms Of Diseases, Mahesh Koirala
Modeling Electrostatics In Molecular Biology And Its Relevance With Molecular Mechanisms Of Diseases, Mahesh Koirala
All Dissertations
Electrostatics plays an essential role in molecular biology. Modeling electrostatics in molecular biology is complicated due to the water phase, mobile ions, and irregularly shaped inhomogeneous biological macromolecules. This dissertation presents the popular DelPhi package that solves PBE and delivers the electrostatic potential distribution of biomolecules. We used the newly developed DelPhiForce steered Molecular Dynamics (DFMD) approach to model the binding of barstar to barnase and demonstrated that the first-principles method could also model the binding. This dissertation also reflects the use of existing computational approaches to model the effects of Single Amino Acid Variations (SAVs) to reveal molecular mechanisms …
Characterizing Endogenous Dicer Products To Unravel Novel Rnai Biogenesis Pathways, Jacob Oche Peter
Characterizing Endogenous Dicer Products To Unravel Novel Rnai Biogenesis Pathways, Jacob Oche Peter
Dissertations
ABSTRACT
RNA interference (RNAi) is a pervasive gene regulatory mechanism in eukaryotes based on the action of multiple classes of small RNA (sRNA). Exploiting RNAi pathways in non-model systems have great potential for creating potent RNAi technologies. Here, we accessed RNAi-mediated control of gene expression in the two-spotted spider mite, Tetranychus urticae (T. urticae) using engineered dsRNA designed to modulate the host RNAi pathway and increase RNAi efficacy. Analysis of Dicer (Dcr) generated fragments revealed how exogenous RNAs access the host RNAi pathway in this animal, opening avenues for designing RNAi technology for their control. Further, some organisms …
Comparative Analyses Of De Novo Transcriptome Assembly Pipelines For Diploid Wheat, Natasha Pavlovikj
Comparative Analyses Of De Novo Transcriptome Assembly Pipelines For Diploid Wheat, Natasha Pavlovikj
School of Computing: Dissertations, Theses, and Student Research
Gene expression and transcriptome analysis are currently one of the main focuses of research for a great number of scientists. However, the assembly of raw sequence data to obtain a draft transcriptome of an organism is a complex multi-stage process usually composed of pre-processing, assembling, and post-processing. Each of these stages includes multiple steps such as data cleaning, error correction and assembly validation. Different combinations of steps, as well as different computational methods for the same step, generate transcriptome assemblies with different accuracy. Thus, using a combination that generates more accurate assemblies is crucial for any novel biological discoveries. Implementing …
Alterations Of The Gut Mycobiome In Patients With Ms - A Bioinformatic Approach, Saumya Shah
Alterations Of The Gut Mycobiome In Patients With Ms - A Bioinformatic Approach, Saumya Shah
Honors Scholar Theses
The mycobiome is the fungal component of the gut microbiome and is implicated in several autoimmune diseases. However, its role in multiple sclerosis (MS) has not been studied. We performed descriptive and formal statistical tests using the R language to characterize the gut mycobiome in people with MS (pwMS) and healthy controls. We found that the microbiome composition of multiple sclerosis patients is different from healthy people. The mycobiome had significantly higher alpha diversity and inter-subject variation in pwMS than controls. Additionally, Saccharomyces and Aspergillus were over-represented in pwMS. Different mycobiome profiles, defined as mycotypes, were associated with different bacterial …
An Investigation Of Epigenetic Mechanisms Driving The Biology Of Head And Neck Squamous Cell Carcinoma, Scot Carson Callahan
An Investigation Of Epigenetic Mechanisms Driving The Biology Of Head And Neck Squamous Cell Carcinoma, Scot Carson Callahan
Dissertations and Theses (Open Access)
Head and neck squamous cell carcinoma (HNSCC) is the 6th most common cancer worldwide and is associated with significant morbidity and mortality. To date, the majority of work in the field has focused on genomic alterations such as mutations and copy number alterations. However, the clinical success of targeted therapies that exploit known genomic alterations, such as EGFR mutations, has remained mixed. Over the past decade, the importance of epigenetic regulators has come to the forefront, with the realization that many of these genes are mutated in cancer. Despite this realization, the role of epigenetics in regulating tumorigenesis, progression and …
Unveiling Global Roles Of G-Quadruplexes And G4-22 In Human Genetics, Ruth Barros De Paula
Unveiling Global Roles Of G-Quadruplexes And G4-22 In Human Genetics, Ruth Barros De Paula
Dissertations and Theses (Open Access)
G-quadruplexes are non-B DNA structures formed by four or more runs of repeated guanines that confer unique features to living organism’s genomes. These sequences are enriched in regulatory regions, such as promoters and 5’ UTRs, and have distinct regulatory roles in both health and disease states. Even though previous studies showed the impact of G4 in gene expression, none of them summarized the location-specific effect of G4. Also, there is no broad understanding about the most common G4 repeat in the human genome, named here as G4-22, and how it links to the evolution of mammals and their biology. In …
Comparative Genomics Methods And Applications, Emily N. Alden
Comparative Genomics Methods And Applications, Emily N. Alden
Biomedical Sciences ETDs
Virtually all fields of biology have benefited from the advancements in comparative genomics technologies, specifically in the study of evolution. In this dissertation I develop and use comparative genomic technologies to investigate the novel SARS-CoV-2 virus, assembly the first genome of the black lace domestic angelfish and identify germline genetic variants associated with altered breast cancer-specific survival. Our genome tiling array for the novel coronavirus presents a rapid and cost-effective method to sequence the entire viral genome and can be used to track the rapid evolution of viral variants in the population. The domestic angelfish is a member of the …
Distribution And Diversity Of Heliothine And Other Lepidopteran Nudiviruses, Emrah Ozel
Distribution And Diversity Of Heliothine And Other Lepidopteran Nudiviruses, Emrah Ozel
Theses and Dissertations--Entomology
Helicoverpa zea nudivirus 2 (HzNV-2) is the only known sterilizing and sexually-transmitted insect virus and causes pathological symptoms in H. zea reproductive tissues. HzNV-2 has features that make it a candidate as a H. zea (corn earworm) control agent, such as the ability to cause asymptomatic (latent) and symptomatic (lytic) infections and the ability to influence mating behavior of its host to favor virus spread. HzNV pathology has been studied and its genome sequenced, however, its prevalence in natural populations is largely unknown. In this study, we developed and used a low-cost PCR-based molecular survey to investigate HzNV-2 prevalence and …
Composition And Homology In The Taxonomic Classification Of Escherichia Coli, Tanya Irani
Composition And Homology In The Taxonomic Classification Of Escherichia Coli, Tanya Irani
Theses and Dissertations (Comprehensive)
As new techniques have been introduced, specifically the possibility of complete genome sequencing, better methods of defining bacterial species have also been proposed. One of the most recently proposed methods, using bioinformatic techniques, is to calculate the average nucleotide identity (ANI) between the homologous genome segments of different isolates. Another method for species discrimination that has been tested successfully is the similarity of DNA compositional signatures. However, in a recent update, DNA signatures split the available Escherichia coli complete genomes into three groups. To check if this result was consistent with such genomes belonging to different species, we tested methods …
Analysis Of Subtelomeric Rextal Assemblies Using Quast, Tunazzina Islam, Desh Ranjan, Mohammad Zubair, Eleanor Young, Ming Xiao, Harold Riethman
Analysis Of Subtelomeric Rextal Assemblies Using Quast, Tunazzina Islam, Desh Ranjan, Mohammad Zubair, Eleanor Young, Ming Xiao, Harold Riethman
Computer Science Faculty Publications
Genomic regions of high segmental duplication content and/or structural variation have led to gaps and misassemblies in the human reference sequence, and are refractory to assembly from whole-genome short-read datasets. Human subtelomere regions are highly enriched in both segmental duplication content and structural variations, and as a consequence are both impossible to assemble accurately and highly variable from individual to individual. Recently, we developed a pipeline for improved region-specific assembly called Regional Extension of Assemblies Using Linked-Reads (REXTAL). In this study, we evaluate REXTAL and genome-wide assembly (Supernova) approaches on 10X Genomics linked-reads data sets partitioned and barcoded using the …
Ensemble Protein Inference Evaluation, Kyle Lee Lucke
Ensemble Protein Inference Evaluation, Kyle Lee Lucke
Graduate Student Theses, Dissertations, & Professional Papers
The Protein inference problem is becoming an increasingly important tool that aids in the characterization of complex proteomes and analysis of complex protein samples. In bottom-up shotgun proteomics experiments the metrics for evaluation (like AUC and calibration error) are based on an often imperfect target-decoy database. These metrics make the inherent assumption that all of the proteins in the target set are present in the sample being analyzed. In general, this is not the case, they are typically a mix of present and absent proteins. To objectively evaluate inference methods, protein standard datasets are used. These datasets are special in …
Decoding The Evolutionary Response To Prostate Cancer Therapy Using Plasma Genome Sequencing, Naveen Ramesh
Decoding The Evolutionary Response To Prostate Cancer Therapy Using Plasma Genome Sequencing, Naveen Ramesh
Dissertations and Theses (Open Access)
Investigating genome evolution in response to therapy is difficult in human tissue samples due to the difficulty in accessing metastatic tumor sites and logistical challenges of collecting longitudinal samples. To overcome these issues, we developed an unbiased whole-genome plasma DNA sequencing approach called PEGASUS that concurrently measures genomic copy number and exome mutations from archival cryostored plasma samples. This approach was applied to study longitudinal blood plasma samples from prostate cancer patients. A molecular characterization of archival plasma DNA from 233 patients and genomic profiling of 101 patients identified clinical correlations of aneuploid plasma DNA profiles with poor survival, increased …
Investigation Of Proliferation Suppressors In Genetic Fitness Screens, Walter Frank Lenoir Iv
Investigation Of Proliferation Suppressors In Genetic Fitness Screens, Walter Frank Lenoir Iv
Dissertations and Theses (Open Access)
Innovation of CRISPR gene-editing technology has provided scientists genome manipulation tools that allowed rapid advancement of scientific capabilities and thus improved our ability to systematically study mammalian genetic functional profiles. Genome-wide CRISPR knockout screens conducted in collections of human cell lines can knock out genes at multiple loci, and have provided new insights into functional roles for independent genes. This method has launched massive efforts in looking across genetic backgrounds for context specific genetic vulnerabilities within cancer. Much of the research effort thus far has been spent on optimizing phenotype distinctions between essential, genes required for cell fitness, and non-essential, …
Polerovirus Genomic Variation And Mechanisms Of Silencing Suppression By P0 Protein, Natalie M. Holste
Polerovirus Genomic Variation And Mechanisms Of Silencing Suppression By P0 Protein, Natalie M. Holste
School of Biological Sciences: Dissertations, Theses, and Student Research
The family Luteoviridae consists of three genera: Luteovirus, Enamovirus, and Polerovirus. The genus Polerovirus contains 32 virus species. All are transmitted by aphids and can infect a wide variety of crops from cereals and wheat to cucurbits and peppers. However, little is known about how this wide range of hosts and vectors developed. In poleroviruses, aphid transmission and virion formation is mediated by the coat protein read-through domain (CPRT) while silencing suppression and phloem limitation is mediated by Protein 0 (P0)—a protein unique to poleroviruses. P0 gives poleroviruses a great advantage amongst plant viruses and diversifies polerovirus species, but the …