Open Access. Powered by Scholars. Published by Universities.®

Genomics Commons

Open Access. Powered by Scholars. Published by Universities.®

Physical Sciences and Mathematics

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 31 - 60 of 113

Full-Text Articles in Genomics

Convolutional Neural Network-Based Gene Prediction Using Buffalograss As A Model System, Michael Morikone Nov 2023

Convolutional Neural Network-Based Gene Prediction Using Buffalograss As A Model System, Michael Morikone

Complex Biosystems Program: Dissertations and Student Research

The task of gene prediction has been largely stagnant in algorithmic improvements compared to when algorithms were first developed for predicting genes thirty years ago. Rather than iteratively improving the underlying algorithms in gene prediction tools by utilizing better performing models, most current approaches update existing tools through incorporating increasing amounts of extrinsic data to improve gene prediction performance. The traditional method of predicting genes is done using Hidden Markov Models (HMMs). These HMMs are constrained by having strict assumptions made about the independence of genes that do not always hold true. To address this, a Convolutional Neural Network (CNN) …


Collagene Enables Privacy-Aware Federated And Collaborative Genomic Data Analysis, Wentao Li, Miran Kim, Kai Zhang, Han Chen, Xiaoqian Jiang, Arif Harmanci Sep 2023

Collagene Enables Privacy-Aware Federated And Collaborative Genomic Data Analysis, Wentao Li, Miran Kim, Kai Zhang, Han Chen, Xiaoqian Jiang, Arif Harmanci

Faculty, Staff and Student Publications

Growing regulatory requirements set barriers around genetic data sharing and collaborations. Moreover, existing privacy-aware paradigms are challenging to deploy in collaborative settings. We present COLLAGENE, a tool base for building secure collaborative genomic data analysis methods. COLLAGENE protects data using shared-key homomorphic encryption and combines encryption with multiparty strategies for efficient privacy-aware collaborative method development. COLLAGENE provides ready-to-run tools for encryption/decryption, matrix processing, and network transfers, which can be immediately integrated into existing pipelines. We demonstrate the usage of COLLAGENE by building a practical federated GWAS protocol for binary phenotypes and a secure meta-analysis protocol. COLLAGENE is available at https://zenodo.org/record/8125935 …


Minimal Positional Substring Cover Is A Haplotype Threading Alternative To Li And Stephens Model, Ahsan Sanaullah, Degui Zhi, Shaojie Zhang Jul 2023

Minimal Positional Substring Cover Is A Haplotype Threading Alternative To Li And Stephens Model, Ahsan Sanaullah, Degui Zhi, Shaojie Zhang

Faculty, Staff and Student Publications

The Li and Stephens (LS) hidden Markov model (HMM) models the process of reconstructing a haplotype as a mosaic copy of haplotypes in a reference panel. For small panels, the probabilistic parameterization of LS enables modeling the uncertainties of such mosaics. However, LS becomes inefficient when sample size is large, because of its linear time complexity. Recently the PBWT, an efficient data structure capturing the local haplotype matching among haplotypes, was proposed to offer a fast method for giving some optimal solution (Viterbi) to the LS HMM. Previously, we introduced the minimal positional substring cover (MPSC) problem as an alternative …


The Integrative Studies On The Functional A-To-I Rna Editing Events In Human Cancers, Sijia Wu, Zhiwei Fan, Pora Kim, Liyu Huang, Xiaobo Zhou Jun 2023

The Integrative Studies On The Functional A-To-I Rna Editing Events In Human Cancers, Sijia Wu, Zhiwei Fan, Pora Kim, Liyu Huang, Xiaobo Zhou

Faculty, Staff and Student Publications

Adenosine-to-inosine (A-to-I) RNA editing, constituting nearly 90% of all RNA editing events in humans, has been reported to contribute to the tumorigenesis in diverse cancers. However, the comprehensive map for functional A-to-I RNA editing events in cancers is still insufficient. To fill this gap, we systematically and intensively analyzed multiple tumorigenic mechanisms of A-to-I RNA editing events in samples across 33 cancer types from The Cancer Genome Atlas. For individual candidate among ∼ 1,500,000 quantified RNA editing events, we performed diverse types of downstream functional annotations. Finally, we identified 24,236 potentially functional A-to-I RNA editing events, including the cases …


Deephtlv: A Deep Learning Framework For Detecting Human T-Lymphotrophic Virus 1 Integration Sites, Johnathan Jia, Johnathan Jia May 2023

Deephtlv: A Deep Learning Framework For Detecting Human T-Lymphotrophic Virus 1 Integration Sites, Johnathan Jia, Johnathan Jia

Dissertations and Theses (Open Access)

In the 1980s, researchers found the first human oncogenic retrovirus called human T-lymphotrophic virus type 1 (HTLV-1). Since then, HTLV-1 has been identified as the causative agent behind several diseases such as adult T-cell leukemia/lymphoma (ATL) and a HTLV-1 associated myelopathy or tropical spastic paraparesis (HAM/TSP). As part of its normal replication cycle, the genome is converted into DNA and integrated into the genome. With several hundreds to thousands of unique viral integration sites (VISs) distributed with indeterminate preference throughout the genome, detection of HTLV-1 VISs is a challenging task. Experimental studies typically use molecular biology …


Physio-Psycho-Social Interaction Mechanism In Dyadic Health Of Young And Middle-Aged Stroke Survivors And Their Spousal Caregivers: A Longitudinal Observational Study Protocol, Dandan Xiang, Zhen-Xiang Zhang, Song Ge, Wen Na Wang, Bei-Lei Lin, Su-Yan Chen, Er-Feng Guo, Peng-Bo Zhang, Zhi-Wei Liu, Hui Li, Yong-Xia Mei Apr 2023

Physio-Psycho-Social Interaction Mechanism In Dyadic Health Of Young And Middle-Aged Stroke Survivors And Their Spousal Caregivers: A Longitudinal Observational Study Protocol, Dandan Xiang, Zhen-Xiang Zhang, Song Ge, Wen Na Wang, Bei-Lei Lin, Su-Yan Chen, Er-Feng Guo, Peng-Bo Zhang, Zhi-Wei Liu, Hui Li, Yong-Xia Mei

Faculty, Staff and Student Publications

Introduction: In recent years, stroke has become more common among young people. Stroke not only has a profound impact on patients' health but also incurs stress and health threats to their caregivers, especially spousal caregivers. Moreover, the health of stroke survivors and their caregivers is interdependent. To our knowledge, no study has explored dyadic health of young and middle-aged stroke survivors and their spousal caregivers from physiological, psychological and social perspectives. Therefore, this proposed study aims to explore the mechanism of how physiological, psychological and social factors affect dyadic health of young and middle-aged stroke survivors and their spousal caregivers. …


Physiological And Transcriptomic Responses Of Two Artemisia Californica Populations To Drought: Implications For Restoring Drought-Resilient Native Communities, Hagop S. Atamian Dr., Jennifer L. Funk Apr 2023

Physiological And Transcriptomic Responses Of Two Artemisia Californica Populations To Drought: Implications For Restoring Drought-Resilient Native Communities, Hagop S. Atamian Dr., Jennifer L. Funk

Biology, Chemistry, and Environmental Sciences Faculty Articles and Research

As climate change brings drier and more variable rainfall patterns to many arid and semi-arid regions, land managers must re-assemble appropriate plant communities for these conditions. Transcriptome sequencing can elucidate the molecular mechanisms underlying plant responses to changing environmental conditions, potentially enhancing our ability to screen suitable genotypes and species for restoration. We examined physiological and morphological traits and transcriptome sequences of coastal and inland populations of California sagebrush (Artemisia californica), a critical shrub used to restore coastal sage scrub vegetation communities, grown under low and high rainfall environments. The populations are located approximately 36 km apart but …


Intellectual Disability Related To De Novo Germline Loss Of The Distal End Of The P-Arm Of Chromosome 17: A Case Report, Eden Pope, Matthew Huertas, Amar Paul, Braden Cunningham, Matthew Jennings, Ryan Perry, Stephanie Chavez, John A. Kriak, Kyle B. Bills, David W. Sant Feb 2023

Intellectual Disability Related To De Novo Germline Loss Of The Distal End Of The P-Arm Of Chromosome 17: A Case Report, Eden Pope, Matthew Huertas, Amar Paul, Braden Cunningham, Matthew Jennings, Ryan Perry, Stephanie Chavez, John A. Kriak, Kyle B. Bills, David W. Sant

Annual Research Symposium

Hypothesis/Purpose: In this report we present a case of a 20-year-old female with congenital intellectual disability, stunted growth, and hypothyroidism. Competitive genetic hybridization (CHG) revealed a loss of 17p13.3, and the deletion was not present in either parent. This deletion has not previously been characterized, but mutations on the p-arm of chromosome 17 are responsible for Miller-Dieker Syndrome and Isolated Lissencephaly Sequence, both of which share symptoms in common with the patient.

Methods: Peripheral mononuclear cells (PBMCs) were used for karyotyping and competitive genetic hybridization (CHG). Bioinformatic analysis was carried out using the Genome Data Viewer (ncbi.nlm.nih.gov/genome/gdv).

Results: Karyotype was …


Presentation Of Paired P- And Q-Arm Mosaic Deletions On Chromosome 18 Associated With Neuropsychiatric Symptoms, Jackson Nielsen, Laura Minor, John Dougherty Jr., Paige Moore, Kailee Edwards, Brandon Burrell, Jameson Williams, John A. Kriak, David W. Sant, Kyle B. Bills Feb 2023

Presentation Of Paired P- And Q-Arm Mosaic Deletions On Chromosome 18 Associated With Neuropsychiatric Symptoms, Jackson Nielsen, Laura Minor, John Dougherty Jr., Paige Moore, Kailee Edwards, Brandon Burrell, Jameson Williams, John A. Kriak, David W. Sant, Kyle B. Bills

Annual Research Symposium

No abstract provided.


Extracting High-Molecular Weight Dna From Cyanobacteria Using Promega's Wizard® Hmw Dna Extraction Kit With A Modified Protocol, Metis, Megan A. Hept, Lesley H. Greene Jan 2023

Extracting High-Molecular Weight Dna From Cyanobacteria Using Promega's Wizard® Hmw Dna Extraction Kit With A Modified Protocol, Metis, Megan A. Hept, Lesley H. Greene

Chemistry & Biochemistry Faculty Publications

Extraction of high molecular weight (HMW) DNA for long read sequencing with little to no fragmentation and high purity is difficult to acquire from cyanobacterial species. Here we describe a modified method of extraction using Promega's Wizard® HMW DNA Extraction Kit to acquire high molecular weight DNA from cyanobacterial species. The protocol used in the kit is the “3.D. Isolating HMW DNA from Gram-Positive and Gram-Negative Bacteria” protocol. During a key step in the protocol, the lingering remnants of the mucilage layer of the cyanobacterial species is removed, preventing it from sticking to the DNA pellet produced. This customized modification …


Novel Feature Evaluation In Ultra-High Dimensional Right-Censored Data, With Applications To Head And Neck Cancer, Atika Farzana Urmi Jan 2023

Novel Feature Evaluation In Ultra-High Dimensional Right-Censored Data, With Applications To Head And Neck Cancer, Atika Farzana Urmi

Graduate Research Posters

Background: Head and neck cancer is the 6th most common cancer worldwide with an expected 1.08 million new cases each year. Such cancer data are ultra-high dimensional with thousands of clinical features and gene expressions, making it challenging for the traditional analytical tools to extract the potential biomarker for the cancer survival and control false discoveries. In addition, presence of heavy censoring can affect the screening procedures based on Kaplan-Meier (K-M) survival estimates.

Aim: To propose a model free, ultra-high dimensional feature screening method with two-dimensional survival outcome allowing false discovery rate (FDR) control.

Method: 516 primary tumor patients with …


Privacy-Aware Estimation Of Relatedness In Admixed Populations, Su Wang, Miran Kim, Wentao Li, Xiaoqian Jiang, Han Chen, Arif Harmanci Nov 2022

Privacy-Aware Estimation Of Relatedness In Admixed Populations, Su Wang, Miran Kim, Wentao Li, Xiaoqian Jiang, Han Chen, Arif Harmanci

Faculty, Staff and Student Publications

BACKGROUND: Estimation of genetic relatedness, or kinship, is used occasionally for recreational purposes and in forensic applications. While numerous methods were developed to estimate kinship, they suffer from high computational requirements and often make an untenable assumption of homogeneous population ancestry of the samples. Moreover, genetic privacy is generally overlooked in the usage of kinship estimation methods. There can be ethical concerns about finding unknown familial relationships in third-party databases. Similar ethical concerns may arise while estimating and reporting sensitive population-level statistics such as inbreeding coefficients for the concerns around marginalization and stigmatization.

RESULTS: Here, we present SIGFRIED, which makes …


The Evolving Privacy And Security Concerns For Genomic Data Analysis And Sharing As Observed From The Idash Competition, Tsung-Ting Kuo, Xiaoqian Jiang, Haixu Tang, Xiaofeng Wang, Arif Harmanci, Miran Kim, Kai Post, Diyue Bu, Tyler Bath, Jihoon Kim, Weijie Liu, Hongbo Chen, Lucila Ohno-Machado Nov 2022

The Evolving Privacy And Security Concerns For Genomic Data Analysis And Sharing As Observed From The Idash Competition, Tsung-Ting Kuo, Xiaoqian Jiang, Haixu Tang, Xiaofeng Wang, Arif Harmanci, Miran Kim, Kai Post, Diyue Bu, Tyler Bath, Jihoon Kim, Weijie Liu, Hongbo Chen, Lucila Ohno-Machado

Faculty, Staff and Student Publications

Concerns regarding inappropriate leakage of sensitive personal information as well as unauthorized data use are increasing with the growth of genomic data repositories. Therefore, privacy and security of genomic data have become increasingly important and need to be studied. With many proposed protection techniques, their applicability in support of biomedical research should be well understood. For this purpose, we have organized a community effort in the past 8 years through the integrating data for analysis, anonymization and sharing consortium to address this practical challenge. In this article, we summarize our experience from these competitions, report lessons learned from the events …


Scgwas: Landscape Of Trait-Cell Type Associations By Integrating Single-Cell Transcriptomics-Wide And Genome-Wide Association Studies, Peilin Jia, Ruifeng Hu, Fangfang Yan, Yulin Dai, Zhongming Zhao Oct 2022

Scgwas: Landscape Of Trait-Cell Type Associations By Integrating Single-Cell Transcriptomics-Wide And Genome-Wide Association Studies, Peilin Jia, Ruifeng Hu, Fangfang Yan, Yulin Dai, Zhongming Zhao

Faculty, Staff and Student Publications

BACKGROUND: The rapid accumulation of single-cell RNA sequencing (scRNA-seq) data presents unique opportunities to decode the genetically mediated cell-type specificity in complex diseases. Here, we develop a new method, scGWAS, which effectively leverages scRNA-seq data to achieve two goals: (1) to infer the cell types in which the disease-associated genes manifest and (2) to construct cellular modules which imply disease-specific activation of different processes.

RESULTS: scGWAS only utilizes the average gene expression for each cell type followed by virtual search processes to construct the null distributions of module scores, making it scalable to large scRNA-seq datasets. We demonstrated scGWAS in …


Svat: Secure Outsourcing Of Variant Annotation And Genotype Aggregation, Miran Kim, Su Wang, Xiaoqian Jiang, Arif Harmanci Oct 2022

Svat: Secure Outsourcing Of Variant Annotation And Genotype Aggregation, Miran Kim, Su Wang, Xiaoqian Jiang, Arif Harmanci

Faculty, Staff and Student Publications

BACKGROUND: Sequencing of thousands of samples provides genetic variants with allele frequencies spanning a very large spectrum and gives invaluable insight into genetic determinants of diseases. Protecting the genetic privacy of participants is challenging as only a few rare variants can easily re-identify an individual among millions. In certain cases, there are policy barriers against sharing genetic data from indigenous populations and stigmatizing conditions.

RESULTS: We present SVAT, a method for secure outsourcing of variant annotation and aggregation, which are two basic steps in variant interpretation and detection of causal variants. SVAT uses homomorphic encryption to encrypt the data at …


Evaluation Of Vicinity-Based Hidden Markov Models For Genotype Imputation, Su Wang, Miran Kim, Xiaoqian Jiang, Arif Ozgun Harmanci Aug 2022

Evaluation Of Vicinity-Based Hidden Markov Models For Genotype Imputation, Su Wang, Miran Kim, Xiaoqian Jiang, Arif Ozgun Harmanci

Faculty, Staff and Student Publications

BACKGROUND: The decreasing cost of DNA sequencing has led to a great increase in our knowledge about genetic variation. While population-scale projects bring important insight into genotype-phenotype relationships, the cost of performing whole-genome sequencing on large samples is still prohibitive. In-silico genotype imputation coupled with genotyping-by-arrays is a cost-effective and accurate alternative for genotyping of common and uncommon variants. Imputation methods compare the genotypes of the typed variants with the large population-specific reference panels and estimate the genotypes of untyped variants by making use of the linkage disequilibrium patterns. Most accurate imputation methods are based on the Li-Stephens hidden Markov …


Development Of Graphical Models And Statistical Physics Motivated Approaches To Genomic Investigations, Yashwanth Lagisetty Aug 2022

Development Of Graphical Models And Statistical Physics Motivated Approaches To Genomic Investigations, Yashwanth Lagisetty

Dissertations and Theses (Open Access)

Identifying genes involved in disease pathology has been a goal of genomic research since the early days of the field. However, as technology improves and the body of research grows, we are faced with more questions than answers. Among these is the pressing matter of our incomplete understanding of the genetic underpinnings of complex diseases. Many hypotheses offer explanations as to why direct and independent analyses of variants, as done in genome-wide association studies (GWAS), may not fully elucidate disease genetics. These range from pointing out flaws in statistical testing to invoking the complex dynamics of epigenetic processes. In the …


Halodash: The Deep And Shallow History Of Aquatic Life's Passages Between Marine And Freshwater Habitats, Eric T. Schultz, Lisa Park Boush May 2022

Halodash: The Deep And Shallow History Of Aquatic Life's Passages Between Marine And Freshwater Habitats, Eric T. Schultz, Lisa Park Boush

EEB Articles

This series of papers highlights research into how biological exchanges between salty and freshwater habitats have transformed the biosphere. Life in the ocean and in freshwaters have long been intertwined; multiple major branches of the tree of life originated in the oceans and then adapted to and diversified in freshwaters. Similar exchanges continue to this day, including some species that continually migrate between marine and fresh waters. The series addresses key themes of transitions, transformations, and current threats with a series of questions: When did major colonizations of fresh waters happen? What physiographic changes facilitated transitions? What organismal characteristics facilitate …


Computational Methods To Analyze Next-Generation Sequencing Data In Genomics And Metagenomics, Saidi Wang Jan 2022

Computational Methods To Analyze Next-Generation Sequencing Data In Genomics And Metagenomics, Saidi Wang

Electronic Theses and Dissertations, 2020-2023

This thesis focuses on two important computational problems in genomics and metagenomics with the public available next-generation sequencing data. One is about gene regulation, for which we explore how distal regulatory elements may interact with the proximal regulatory elements. The other is about metagenomics, in which we study how to reconstruct bacterial strain genomes from shotgun reads. Studying gene regulation, especially distal gene regulation, is important because regulatory elements, including those in distal regulatory regions, orchestrate when, where and how much a gene is activated under every experimental condition. Their dysfunction results in various types of diseases. Moreover, the current …


Genomics Of Postprandial Lipidomics In The Genetics Of Lipid-Lowering Drugs And Diet Network Study, Marguerite R. Irvin, May E. Montasser, Tobias Kind, Sili Fan, Dinesh K. Barupal, Amit Patki, Rikki M. Tanner, Nicole D. Armstrong, Kathleen A. Ryan, Steven A. Claas, Jeffrey R. O’Connell, Hemant K. Tiwari, Donna K. Arnett Nov 2021

Genomics Of Postprandial Lipidomics In The Genetics Of Lipid-Lowering Drugs And Diet Network Study, Marguerite R. Irvin, May E. Montasser, Tobias Kind, Sili Fan, Dinesh K. Barupal, Amit Patki, Rikki M. Tanner, Nicole D. Armstrong, Kathleen A. Ryan, Steven A. Claas, Jeffrey R. O’Connell, Hemant K. Tiwari, Donna K. Arnett

Epidemiology and Environmental Health Faculty Publications

Postprandial lipemia (PPL) is an important risk factor for cardiovascular disease. Inter-individual variation in the dietary response to a meal is known to be influenced by genetic factors, yet genes that dictate variation in postprandial lipids are not completely characterized. Genetic studies of the plasma lipidome can help to better understand postprandial metabolism by isolating lipid molecular species which are more closely related to the genome. We measured the plasma lipidome at fasting and 6 h after a standardized high-fat meal in 668 participants from the Genetics of Lipid-Lowering Drugs and Diet Network study (GOLDN) using ultra-performance liquid chromatography coupled …


Mixture Model Approaches To Integrative Analysis Of Multi-Omics Data And Spatially Correlated Genomic Data, Ziqiao Wang May 2021

Mixture Model Approaches To Integrative Analysis Of Multi-Omics Data And Spatially Correlated Genomic Data, Ziqiao Wang

Dissertations and Theses (Open Access)

Integrative genomic data analysis is a powerful tool to study the complex biological processes behind a disease. Statistical methods can model the interrelationships of the involved gene activities through jointly analyzing multiple types of genomic data from different platforms (vertical integration), or improve the power of a study through aggregating the same type of genomic data across studies (horizontal integration). In this dissertation, we propose statistical methods and strategies for integrative multi-omics data in association analysis of disease phenotypes, with an emphasis on cancer applications.

We develop a new strategy based on horizontal integration by leveraging publicly available datasets into …


An Ensemble Of The Icluster Method To Analyze Longitudinal Lncrna Expression Data For Psoriasis Patients, Suyan Tian, Chi Wang Apr 2021

An Ensemble Of The Icluster Method To Analyze Longitudinal Lncrna Expression Data For Psoriasis Patients, Suyan Tian, Chi Wang

Internal Medicine Faculty Publications

BACKGROUND: Psoriasis is an immune-mediated, inflammatory disorder of the skin with chronic inflammation and hyper-proliferation of the epidermis. Since psoriasis has genetic components and the diseased tissue of psoriasis is very easily accessible, it is natural to use high-throughput technologies to characterize psoriasis and thus seek targeted therapies. Transcriptional profiles change correspondingly after an intervention. Unlike cross-sectional gene expression data, longitudinal gene expression data can capture the dynamic changes and thus facilitate causal inference.

METHODS: Using the iCluster method as a building block, an ensemble method was proposed and applied to a longitudinal gene expression dataset for psoriasis, with the …


Nanopore Guided Regional Assembly, Eleni Adam, Desh Ranjan, Harold Riethman Apr 2021

Nanopore Guided Regional Assembly, Eleni Adam, Desh Ranjan, Harold Riethman

College of Sciences Posters

The telomeres are the “caps” of the chromosomes and their vital role is to protect them. Possible telomere dysfunction caused by telomere rearrangements can be fatal for the cell and result in age-related diseases, including cancer. The telomeres and subtelomeres are regions that are hard to investigate. The current technology cannot provide their complete sequence, instead the DNA is given in multiple pieces. Current methods of assembling the pieces of these regions are not accurate enough due to the region’s high variability and complex repeated patterns. We propose a hybrid assembly method, the NPGREAT, which utilizes two of the latest …


Machine Learning Approaches For The Prediction Of Bone Mineral Density By Using Genomic And Phenotypic Data Of 5130 Older Men, Qing Wu, Fatma Nasoz, Jongyun Jung, Bibek Bhattarai, Mira V. Han, Robert A. Greenes, Kenneth G. Saag Feb 2021

Machine Learning Approaches For The Prediction Of Bone Mineral Density By Using Genomic And Phenotypic Data Of 5130 Older Men, Qing Wu, Fatma Nasoz, Jongyun Jung, Bibek Bhattarai, Mira V. Han, Robert A. Greenes, Kenneth G. Saag

School of Medicine Faculty Research

The study aimed to utilize machine learning (ML) approaches and genomic data to develop a prediction model for bone mineral density (BMD) and identify the best modeling approach for BMD prediction. The genomic and phenotypic data of Osteoporotic Fractures in Men Study (n = 5130) was analyzed. Genetic risk score (GRS) was calculated from 1103 associated SNPs for each participant after a comprehensive genotype imputation. Data were normalized and divided into a training set (80%) and a validation set (20%) for analysis. Random forest, gradient boosting, neural network, and linear regression were used to develop BMD prediction models separately. Ten-fold …


Deep Learning For Multi-Tissue Cancer Classification Of Gene Expressions, Tarek Khorshed Jan 2021

Deep Learning For Multi-Tissue Cancer Classification Of Gene Expressions, Tarek Khorshed

Theses and Dissertations

We contribute in saving the lives of cancer patients through early detection and diagnosis, since one of the major challenges in cancer treatment is that patients are diagnosed at very late stages when appropriate medical interventions become less effective and full curative treatment is no longer achievable. Cancer classification using gene expressions is extremely challenging given the complexity and high dimensionality of the data. Current classification methods typically rely on samples collected from a single tissue type and perform a prerequisite of gene feature selection to avoid processing the full set of genes. These methods fall short in taking advantage …


Statistical Methods In Genetic Studies, Cheng Gao Jan 2021

Statistical Methods In Genetic Studies, Cheng Gao

Dissertations, Master's Theses and Master's Reports

This dissertation includes three Chapters. A brief description of each chapter is organized as follows.

In Chapter 1, we proposed a new method, called MF-TOWmuT, for genome-wide association studies with multiple genetic variants and multiple phenotypes using family samples. MF-TOWmuT uses kinship matrix to account for sample relatedness. It is worth mentioning that in simulations, we considered hidden polygenic effects and varied the proportion of variance contributed by it to generate phenotypes. Simulation studies show that MF-TOWmuT can preserve the type I error rates and is more powerful than several existing methods in different simulation scenarios, MFTOWmuT is also quite …


Analysis Of Subtelomeric Rextal Assemblies Using Quast, Tunazzina Islam, Desh Ranjan, Mohammad Zubair, Eleanor Young, Ming Xiao, Harold Riethman Jan 2021

Analysis Of Subtelomeric Rextal Assemblies Using Quast, Tunazzina Islam, Desh Ranjan, Mohammad Zubair, Eleanor Young, Ming Xiao, Harold Riethman

Computer Science Faculty Publications

Genomic regions of high segmental duplication content and/or structural variation have led to gaps and misassemblies in the human reference sequence, and are refractory to assembly from whole-genome short-read datasets. Human subtelomere regions are highly enriched in both segmental duplication content and structural variations, and as a consequence are both impossible to assemble accurately and highly variable from individual to individual. Recently, we developed a pipeline for improved region-specific assembly called Regional Extension of Assemblies Using Linked-Reads (REXTAL). In this study, we evaluate REXTAL and genome-wide assembly (Supernova) approaches on 10X Genomics linked-reads data sets partitioned and barcoded using the …


Statistical Approaches Of Gene Set Analysis With Quantitative Trait Loci For High-Throughput Genomic Studies., Samarendra Das Dec 2020

Statistical Approaches Of Gene Set Analysis With Quantitative Trait Loci For High-Throughput Genomic Studies., Samarendra Das

Electronic Theses and Dissertations

Recently, gene set analysis has become the first choice for gaining insights into the underlying complex biology of diseases through high-throughput genomic studies, such as Microarrays, bulk RNA-Sequencing, single cell RNA-Sequencing, etc. It also reduces the complexity of statistical analysis and enhances the explanatory power of the obtained results. Further, the statistical structure and steps common to these approaches have not yet been comprehensively discussed, which limits their utility. Hence, a comprehensive overview of the available gene set analysis approaches used for different high-throughput genomic studies is provided. The analysis of gene sets is usually carried out based on …


Statistical Methods For Resolving Intratumor Heterogeneity With Single-Cell Dna Sequencing, Alexander Davis Aug 2020

Statistical Methods For Resolving Intratumor Heterogeneity With Single-Cell Dna Sequencing, Alexander Davis

Dissertations and Theses (Open Access)

Tumor cells have heterogeneous genotypes, which drives progression and treatment resistance. Such genetic intratumor heterogeneity plays a role in the process of clonal evolution that underlies tumor progression and treatment resistance. Single-cell DNA sequencing is a promising experimental method for studying intratumor heterogeneity, but brings unique statistical challenges in interpreting the resulting data. Researchers lack methods to determine whether sufficiently many cells have been sampled from a tumor. In addition, there are no proven computational methods for determining the ploidy of a cell, a necessary step in the determination of copy number. In this work, software for calculating probabilities from …


Statistical Inference Of Adaptation At Multiple Genomic Scales Using Supervised Classification And A Hidden Markov Model, Lauren A. Sugden May 2020

Statistical Inference Of Adaptation At Multiple Genomic Scales Using Supervised Classification And A Hidden Markov Model, Lauren A. Sugden

Biology and Medicine Through Mathematics Conference

No abstract provided.