Open Access. Powered by Scholars. Published by Universities.®

Biostatistics Commons™

Open Access. Powered by Scholars. Published by Universities.®

Genetics and Genomics

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 31 - 60 of 105

Full-Text Articles in Biostatistics

Statistical Inference Of Adaptation At Multiple Genomic Scales Using Supervised Classification And A Hidden Markov Model, Lauren A. Sugden May 2020

Statistical Inference Of Adaptation At Multiple Genomic Scales Using Supervised Classification And A Hidden Markov Model, Lauren A. Sugden

Biology and Medicine Through Mathematics Conference

No abstract provided.


Spatial And Temporal Genetic Structure Of Winter-Run Steelhead (Oncorhynchus Mykiss) Returning To The Mad River, California, Steven R. Fong Jan 2020

Spatial And Temporal Genetic Structure Of Winter-Run Steelhead (Oncorhynchus Mykiss) Returning To The Mad River, California, Steven R. Fong

Cal Poly Humboldt theses and projects

Distinct populations of steelhead in the wild are in decline. The propagation of steelhead in hatcheries has been used to boost population numbers for recreational fisheries and for use in conservation. However, hatchery breeding practices of steelhead can result in changes in genetic structure. I investigated the genetic structure of winter-run steelhead (Oncorhynchus mykiss) returning to the Mad River, California, where a hatchery has been used enhance production for recreational fisheries since 1971. Genetic variability in Mad River steelhead was evaluated using 96 single nucleotide polymorphisms (SNPs) among 4203 individuals, including the Mad River and nearby locations, and …


Quantifying Pollen Traits To Build A Mathematical Model Of Pollen Competition - A Biologist's Perspective, Rob Swanson, Alex Capaldi Oct 2019

Quantifying Pollen Traits To Build A Mathematical Model Of Pollen Competition - A Biologist's Perspective, Rob Swanson, Alex Capaldi

Annual Symposium on Biomathematics and Ecology Education and Research

No abstract provided.


Acute Systemic Inflammatory Response To Lipopolysaccharide Stimulation In Pigs Divergently Selected For Residual Feed Intake, Haibo Liu, Kristina M. Feye, Yet T. Nguyen, Anoosh Rakhshandeh, Crystal L. Loving, Jack C. M. Sekkers, Nicholas K. Gabler, Christopher K. Tuggle Oct 2019

Acute Systemic Inflammatory Response To Lipopolysaccharide Stimulation In Pigs Divergently Selected For Residual Feed Intake, Haibo Liu, Kristina M. Feye, Yet T. Nguyen, Anoosh Rakhshandeh, Crystal L. Loving, Jack C. M. Sekkers, Nicholas K. Gabler, Christopher K. Tuggle

Mathematics & Statistics Faculty Publications

Background: It is unclear whether improving feed efficiency by selection for low residual feed intake (RFI) compromises pigs’ immunocompetence. Here, we aimed at investigating whether pig lines divergently selected for RFI had different inflammatory responses to lipopolysaccharide (LPS) exposure, regarding to clinical presentations and transcriptomic changes in peripheral blood cells.

Results: LPS injection induced acute systemic inflammation in both the low-RFI and high-RFI line (n = 8 per line). At 4 h post injection (hpi), the low-RFI line had a significantly lower (p= 0.0075) mean rectal temperature compared to the high-RFI line. However, no significant differences in complete blood count …


Association Of Copy Number Variations With Chronic Hepatitis B In Chinese Population, Fang Niu Aug 2019

Association Of Copy Number Variations With Chronic Hepatitis B In Chinese Population, Fang Niu

Capstone Experience: Master of Public Health

With one third of the Hepatitis B virus (HBV) infection population of the world, chronic Hepatitis B (CHB) has become a top burden in China. CHB is a lifelong infection with HBV which can cause serious health problems, like cirrhosis, liver cancer or even death. HBV infection is known to result in various clinical conditions, including asymptomatic HBV carriers to chronic hepatitis and primary hepatocellular carcinoma. Several studies have shown that host genetic susceptibility could be an important factor that determines these various outcomes of HBV infection. Many Single Nucleotide Polymorphisms (SNPs) and Copy Number Variations (CNVs) have been associated …


Incorporating Pathway Information Into Feature Selection Towards Better Performed Gene Signatures, Suyan Tian, Chi Wang, Bing Wang Apr 2019

Incorporating Pathway Information Into Feature Selection Towards Better Performed Gene Signatures, Suyan Tian, Chi Wang, Bing Wang

Biostatistics Faculty Publications

To analyze gene expression data with sophisticated grouping structures and to extract hidden patterns from such data, feature selection is of critical importance. It is well known that genes do not function in isolation but rather work together within various metabolic, regulatory, and signaling pathways. If the biological knowledge contained within these pathways is taken into account, the resulting method is a pathway-based algorithm. Studies have demonstrated that a pathway-based method usually outperforms its gene-based counterpart in which no biological knowledge is considered. In this article, a pathway-based feature selection is firstly divided into three major categories, namely, pathway-level selection, …


Feature Selection For Longitudinal Data By Using Sign Averages To Summarize Gene Expression Values Over Time, Suyan Tian, Chi Wang Mar 2019

Feature Selection For Longitudinal Data By Using Sign Averages To Summarize Gene Expression Values Over Time, Suyan Tian, Chi Wang

Biostatistics Faculty Publications

With the rapid evolution of high-throughput technologies, time series/longitudinal high-throughput experiments have become possible and affordable. However, the development of statistical methods dealing with gene expression profiles across time points has not kept up with the explosion of such data. The feature selection process is of critical importance for longitudinal microarray data. In this study, we proposed aggregating a gene’s expression values across time into a single value using the sign average method, thereby degrading a longitudinal feature selection process into a classic one. Regularized logistic regression models with pseudogenes (i.e., the sign average of genes across time as predictors) …


Unified Methods For Feature Selection In Large-Scale Genomic Studies With Censored Survival Outcomes, Lauren Spirko-Burns, Karthik Devarajan Mar 2019

Unified Methods For Feature Selection In Large-Scale Genomic Studies With Censored Survival Outcomes, Lauren Spirko-Burns, Karthik Devarajan

COBRA Preprint Series

One of the major goals in large-scale genomic studies is to identify genes with a prognostic impact on time-to-event outcomes which provide insight into the disease's process. With rapid developments in high-throughput genomic technologies in the past two decades, the scientific community is able to monitor the expression levels of tens of thousands of genes and proteins resulting in enormous data sets where the number of genomic features is far greater than the number of subjects. Methods based on univariate Cox regression are often used to select genomic features related to survival outcome; however, the Cox model assumes proportional hazards …


Gene Co-Expression Networks Analysis Reveal Novel Molecular Endotypes In Alpha-1 Antitrypsin Deficiency, Jen-Hwa Chu, Wenlan Zang Jan 2019

Gene Co-Expression Networks Analysis Reveal Novel Molecular Endotypes In Alpha-1 Antitrypsin Deficiency, Jen-Hwa Chu, Wenlan Zang

Yale Day of Data

Rationale:Alpha-1 antitrypsin deficiency (AATD) is a genetic condition that predisposes to early onset pulmonary emphysema and airways obstruction. The exact mechanism through which AATD leads to lung disease is incompletely understood.

Objectives: To investigate the effect of AAT genotype and augmentation therapy on bronchoalveolar lavage (BAL) and peripheral blood mononuclear cells (PBMC) transcriptome, while examining the link between gene expression profiles, and clinical features of AATD.

Methods: We performed RNA-Seq on RNA extracted from BAL and PBMC on samples obtained from 89 AATD patients enrolled in the Genomic Research in Alpha-1 Antitrypsin Deficiency and Sarcoidosis (GRADS) study. Differential …


A Novel Pathway-Based Distance Score Enhances Assessment Of Disease Heterogeneity In Gene Expression, Yunqing Liu, Xiting Yan Jan 2019

A Novel Pathway-Based Distance Score Enhances Assessment Of Disease Heterogeneity In Gene Expression, Yunqing Liu, Xiting Yan

Yale Day of Data

Distance-based unsupervised clustering of gene expression data is commonly used to identify heterogeneity in biologic samples. However, high noise levels in gene expression data and the relatively high correlation between genes are often encountered, so traditional distances such as Euclidean distance may not be effective at discriminating the biological differences between samples. In this study, we developed a novel computational method to assess the biological differences based on pathways by assuming that ontologically defined biological pathways in biologically similar samples have similar behavior. Application of this distance score results in more accurate, robust, and biologically meaningful clustering results in both …


The Effect Of Maternal Dietary Habits During Pregnancy On Neonate Leptin Methylation Patterns And Gestational Age, Sean Fitzpatrick Jan 2019

The Effect Of Maternal Dietary Habits During Pregnancy On Neonate Leptin Methylation Patterns And Gestational Age, Sean Fitzpatrick

Legacy Theses & Dissertations (2009 - 2024)

The health of a newborn baby is inextricably linked to the health status of its mother and in turn the mother’s diet during pregnancy. Leptin (LEP) is an adipokine hormone involved in metabolism regulation and has been linked fetal development through the hypothalamic-pituitary-adrenal axis (HPA). Prior work suggests that gestational epigenetic alterations the LEP gene may be sensitive to adverse exposures during pregnancy, which in turn could explain variation in neonate outcomes. However, no prior work has examined this possibility explicitly. The objective of this study was to investigate the association between dietary patterns of mothers during pregnancy and their …


Genome-Wide Systems Genetics Of Alcohol Consumption And Dependence, Kristin Mignogna Jan 2019

Genome-Wide Systems Genetics Of Alcohol Consumption And Dependence, Kristin Mignogna

Theses and Dissertations

Widely effective treatment for alcohol use disorder is not yet available, because the exact biological mechanisms that underlie this disorder are not completely understood. One way to gain a better understanding of these mechanisms is to examine the genetic frameworks that contribute to the risk for developing this disorder. This dissertation examines genetic association data in combination with gene expression networks in the brain to identify functional groups of genes associated with alcohol consumption and dependence.

The first study took advantage of the behavioral complexity of human samples, and experimental capabilities provided by mouse models, by co-analyzing gene expression networks …


Large-Scale Genome-Wide Meta-Analysis Of Polycystic Ovary Syndrome Suggests Shared Genetic Architecture For Different Diagnosis Criteria, Felix Day, Tugce Karaderi, Michelle R. Jones, Cindy Meun, Chunyan He, Alex Drong, Peter Kraft, Nan Lin, Hongyan Huang, Linda Broer, Reedik Magi, Richa Saxena, Triin Laisk, Margrit Urbanek, M. Geoffrey Hayes, Gudmar Thorleifsson, Juan Fernandez-Tajes, Anubha Mahajan, Benjamin H. Mullin, Bronwyn G. A. Stuckey, Timothy D. Spector, Scott G. Wilson, Mark O. Goodarzi, Lea Davis, Barbara Obermayer-Pietsch, André G. Uitterlinden, Verneri Anttila, Benjamin M. Neale, Marjo-Riitta Jarvelin, Bart Fauser Dec 2018

Large-Scale Genome-Wide Meta-Analysis Of Polycystic Ovary Syndrome Suggests Shared Genetic Architecture For Different Diagnosis Criteria, Felix Day, Tugce Karaderi, Michelle R. Jones, Cindy Meun, Chunyan He, Alex Drong, Peter Kraft, Nan Lin, Hongyan Huang, Linda Broer, Reedik Magi, Richa Saxena, Triin Laisk, Margrit Urbanek, M. Geoffrey Hayes, Gudmar Thorleifsson, Juan Fernandez-Tajes, Anubha Mahajan, Benjamin H. Mullin, Bronwyn G. A. Stuckey, Timothy D. Spector, Scott G. Wilson, Mark O. Goodarzi, Lea Davis, Barbara Obermayer-Pietsch, André G. Uitterlinden, Verneri Anttila, Benjamin M. Neale, Marjo-Riitta Jarvelin, Bart Fauser

Internal Medicine Faculty Publications

Polycystic ovary syndrome (PCOS) is a disorder characterized by hyperandrogenism, ovulatory dysfunction and polycystic ovarian morphology. Affected women frequently have metabolic disturbances including insulin resistance and dysregulation of glucose homeostasis. PCOS is diagnosed with two different sets of diagnostic criteria, resulting in a phenotypic spectrum of PCOS cases. The genetic similarities between cases diagnosed based on the two criteria have been largely unknown. Previous studies in Chinese and European subjects have identified 16 loci associated with risk of PCOS. We report a fixed-effect, inverse-weighted-variance meta-analysis from 10,074 PCOS cases and 103,164 controls of European ancestry and characterisation of PCOS related …


A Logitudinal Feature Selection Method Identifies Relevant Genes To Distinguish Complicated Injury And Uncomplicated Injury Over Time, Suyan Tian, Chi Wang, Howard H. Chang Dec 2018

A Logitudinal Feature Selection Method Identifies Relevant Genes To Distinguish Complicated Injury And Uncomplicated Injury Over Time, Suyan Tian, Chi Wang, Howard H. Chang

Biostatistics Faculty Publications

Background: Feature selection and gene set analysis are of increasing interest in the field of bioinformatics. While these two approaches have been developed for different purposes, we describe how some gene set analysis methods can be utilized to conduct feature selection.

Methods: We adopted a gene set analysis method, the significance analysis of microarray gene set reduction (SAMGSR) algorithm, to carry out feature selection for longitudinal gene expression data.

Results: Using a real-world application and simulated data, it is demonstrated that the proposed SAMGSR extension outperforms other relevant methods. In this study, we illustrate that a gene’s expression profiles over …


Gaw20: Methods And Strategies For The New Frontiers Of Epigenetics And Pharmacogenomics, Nathan L. Tintle, David W. Fardo, Marzia De Andrade, Stella Aslibekyan, Julia N. Bailey, Justo Lorenzo Bermejo, Rita M. Cantor, Saurabh Ghosh, Philip Melton, Xuexua Wang, Jean W. Maccluer, Laura Almasy Sep 2018

Gaw20: Methods And Strategies For The New Frontiers Of Epigenetics And Pharmacogenomics, Nathan L. Tintle, David W. Fardo, Marzia De Andrade, Stella Aslibekyan, Julia N. Bailey, Justo Lorenzo Bermejo, Rita M. Cantor, Saurabh Ghosh, Philip Melton, Xuexua Wang, Jean W. Maccluer, Laura Almasy

Biostatistics Faculty Publications

GAW20 provided a platform for developing and evaluating statistical methods to analyze human lipid-related phenotypes, DNA methylation, and single-nucleotide markers in a study involving a pharmaceutical intervention. In this article, we present an overview of the data sets and the contributions analyzing these data. The data, donated by the Genetics of Lipid Lowering Drugs and Diet Network (GOLDN) investigators, included data from 188 families (N = 1105) which included genome-wide DNA methylation data before and after a 3-week treatment with fenofibrate, single-nucleotide polymorphisms, metabolic syndrome components before and after treatment, and a variety of covariates. The contributions from individual …


Association Analyses Of Repeated Measures On Triglyceride And High-Density Lipoprotein Levels: Insights From Gaw20, Saurabh Ghosh, David W. Fardo Sep 2018

Association Analyses Of Repeated Measures On Triglyceride And High-Density Lipoprotein Levels: Insights From Gaw20, Saurabh Ghosh, David W. Fardo

Biostatistics Faculty Publications

Background: The GAW20 group formed on the theme of methods for association analyses of repeated measures comprised 4sets of investigators. The provided “real” data set included genotypes obtained from a human whole-genome association study based on longitudinal measurements of triglycerides (TGs) and high-density lipoprotein in addition to methylation levels before and after administration of fenofibrate. The simulated data set contained 200 replications of methylation levels and posttreatment TGs, mimicking the real data set.

Results: The different investigators in the group focused on the statistical challenges unique to family-based association analyses of phenotypes measured longitudinally and applied a wide spectrum of …


Bayesian Analytical Approaches For Metabolomics : A Novel Method For Molecular Structure-Informed Metabolite Interaction Modeling, A Novel Diagnostic Model For Differentiating Myocardial Infarction Type, And Approaches For Compound Identification Given Mass Spectrometry Data., Patrick J. Trainor Aug 2018

Bayesian Analytical Approaches For Metabolomics : A Novel Method For Molecular Structure-Informed Metabolite Interaction Modeling, A Novel Diagnostic Model For Differentiating Myocardial Infarction Type, And Approaches For Compound Identification Given Mass Spectrometry Data., Patrick J. Trainor

Electronic Theses and Dissertations

Metabolomics, the study of small molecules in biological systems, has enjoyed great success in enabling researchers to examine disease-associated metabolic dysregulation and has been utilized for the discovery biomarkers of disease and phenotypic states. In spite of recent technological advances in the analytical platforms utilized in metabolomics and the proliferation of tools for the analysis of metabolomics data, significant challenges in metabolomics data analyses remain. In this dissertation, we present three of these challenges and Bayesian methodological solutions for each. In the first part we develop a new methodology to serve a basis for making higher order inferences in metabolomics, …


The Impact Of Truncating Data On The Predictive Ability For Single-Step Genomic Best Linear Unbiased Prediction, Jeremy T. Howard, Thomas A. Rathje, Caitlyn E. Bruns, Danielle F. Wilson-Wells, Stephen D. Kachman, Matthew L. Spangler Jan 2018

The Impact Of Truncating Data On The Predictive Ability For Single-Step Genomic Best Linear Unbiased Prediction, Jeremy T. Howard, Thomas A. Rathje, Caitlyn E. Bruns, Danielle F. Wilson-Wells, Stephen D. Kachman, Matthew L. Spangler

Department of Animal Science: Faculty Publications

Simulated and swine industry data sets were utilized to assess the impact of removing older data on the predictive ability of selection candidate estimated breeding values (EBV) when using single-step genomic best linear unbiased prediction (ssGBLUP). Simulated data included thirty replicates designed to mimic the structure of swine data sets. For the simulated data, varying amounts of data were truncated based on the number of ancestral generations back from the selection candidates. The swine data sets consisted of phenotypic and genotypic records for three traits across two breeds on animals born from 2003 to 2017. Phenotypes and genotypes were iteratively …


An Investigation Of Atomic Structures Derived From X-Ray Crystallography And Cryo-Electron Microscopy Using Distal Blocks Of Side-Chains, Lin Chen, Jing He, Salim Sazzed, Rayshawn Walker Jan 2018

An Investigation Of Atomic Structures Derived From X-Ray Crystallography And Cryo-Electron Microscopy Using Distal Blocks Of Side-Chains, Lin Chen, Jing He, Salim Sazzed, Rayshawn Walker

Computer Science Faculty Publications

Cryo-electron microscopy (cryo-EM) is a structure determination method for large molecular complexes. As more and more atomic structures are determined using this technique, it is becoming possible to perform statistical characterization of side-chain conformations. Two data sets were involved to characterize block lengths for each of the 18 types of amino acids. One set contains 9131 structures resolved using X-ray crystallography from density maps with better than or equal to 1.5 Å resolutions, and the other contains 237 protein structures derived from cryo-EM density maps with 2-4 Å resolutions. The results show that the normalized probability density function of block …


Bayesian Prediction Intervals For Assessing P-Value Variability In Prospective Replication Studies, Olga A. Vsevolozhskaya, Gabriel Ruiz, Dmitri Zaykin Dec 2017

Bayesian Prediction Intervals For Assessing P-Value Variability In Prospective Replication Studies, Olga A. Vsevolozhskaya, Gabriel Ruiz, Dmitri Zaykin

Biostatistics Faculty Publications

Increased availability of data and accessibility of computational tools in recent years have created an unprecedented upsurge of scientific studies driven by statistical analysis. Limitations inherent to statistics impose constraints on the reliability of conclusions drawn from data, so misuse of statistical methods is a growing concern. Hypothesis and significance testing, and the accompanying P-values are being scrutinized as representing the most widely applied and abused practices. One line of critique is that P-values are inherently unfit to fulfill their ostensible role as measures of credibility for scientific hypotheses. It has also been suggested that while P-values …


Enrichment Of Putatively Damaging Rare Variants In The Dyx2 Locus And The Reading-Related Genes Ccdc136 And Flnc, Andrew K. Adams, Shelley D. Smith, Dongnhu T. Truong, Erik G. Willcutt, Richard K. Olson, John C. Defries, Bruce F. Pennington, Jeffrey R. Gruen Nov 2017

Enrichment Of Putatively Damaging Rare Variants In The Dyx2 Locus And The Reading-Related Genes Ccdc136 And Flnc, Andrew K. Adams, Shelley D. Smith, Dongnhu T. Truong, Erik G. Willcutt, Richard K. Olson, John C. Defries, Bruce F. Pennington, Jeffrey R. Gruen

Psychology: Faculty Scholarship

Eleven loci with prior evidence for association with reading and language phenotypes were sequenced in 96 unrelated subjects with significant impairment in reading performance drawn from the Colorado Learning Disability Research Center collection. Out of 148 total individual missense variants identified, the chromosome 7 genes CCDC136 and FLNC contained 19. In addition, a region corresponding to the well-known DYX2 locus for RD contained 74 missense variants. Both allele sets were filtered for a minor allele frequency ≤0.01 and high Polyphen-2 scores. To determine if observations of these alleles are occurring more frequently in our cases than expected by chance in …


Systems Biology Approach To Late-Onset Alzheimer's Disease Genome-Wide Association Study Identifies Novel Candidate Genes Validated Using Brain Expression Data And Caenorhabditis Elegans Experiments, Shubhabrata Mukherjee, Joshua C. Russell, Daniel T. Carr, Jeremy D. Burgess, Mariet Allen, Daniel J. Serie, Kevin L. Boehme, John S. K. Kauwe, Adam C. Naj, David W. Fardo, Dennis W. Dickson, Thomas J. Montine, Nilufer Ertekin-Taner, Matt R. Kaeberlein, Paul K. Crane Oct 2017

Systems Biology Approach To Late-Onset Alzheimer's Disease Genome-Wide Association Study Identifies Novel Candidate Genes Validated Using Brain Expression Data And Caenorhabditis Elegans Experiments, Shubhabrata Mukherjee, Joshua C. Russell, Daniel T. Carr, Jeremy D. Burgess, Mariet Allen, Daniel J. Serie, Kevin L. Boehme, John S. K. Kauwe, Adam C. Naj, David W. Fardo, Dennis W. Dickson, Thomas J. Montine, Nilufer Ertekin-Taner, Matt R. Kaeberlein, Paul K. Crane

Biostatistics Faculty Publications

Introduction—We sought to determine whether a systems biology approach may identify novel late-onset Alzheimer's disease (LOAD) loci.

Methods—We performed gene-wide association analyses and integrated results with human protein-protein interaction data using network analyses. We performed functional validation on novel genes using a transgenic Caenorhabditis elegans Aβ proteotoxicity model and evaluated novel genes using brain expression data from people with LOAD and other neurodegenerative conditions.

Results—We identified 13 novel candidate LOAD genes outside chromosome 19. Of those, RNA interference knockdowns of the C. elegans orthologs of UBC, NDUFS3, EGR1, and ATP5H were associated with Aβ …


Increased Birth Weight Is Associated With Altered Gene Expression In Neonatal Foreskin, Leryn J. Reynolds, Rebecca I. Pollack, Richard J. Charnigo, Cetewayo S. Rashid, Arnold J. Stromberg, Shu Shen, John O'Brien, Kevin J. Pearson Oct 2017

Increased Birth Weight Is Associated With Altered Gene Expression In Neonatal Foreskin, Leryn J. Reynolds, Rebecca I. Pollack, Richard J. Charnigo, Cetewayo S. Rashid, Arnold J. Stromberg, Shu Shen, John O'Brien, Kevin J. Pearson

Pharmacology and Nutritional Sciences Faculty Publications

Elevated birth weight is linked to glucose intolerance and obesity health-related complications later in life. No studies have examined if infant birth weight is associated with gene expression markers of obesity and inflammation in a tissue that comes directly from the infant following birth. We evaluated the association between birth weight and gene expression on fetal programming of obesity. Foreskin samples were collected following circumcision, and gene expression analyzed comparing the 15% greatest birth weight infants (n = 7) v. the remainder of the cohort (n = 40). Multivariate linear regression models were fit to relate expression levels on differentially …


Impact Of Home Visit Capacity On Genetic Association Studies Of Late-Onset Alzheimer's Disease, David W. Fardo, Laura E. Gibbons, Shubhabrata Mukherjee, M. Maria Glymour, Wayne Mccormick, Susan M. Mccurry, James D. Bowen, Eric B. Larson, Paul K. Crane Aug 2017

Impact Of Home Visit Capacity On Genetic Association Studies Of Late-Onset Alzheimer's Disease, David W. Fardo, Laura E. Gibbons, Shubhabrata Mukherjee, M. Maria Glymour, Wayne Mccormick, Susan M. Mccurry, James D. Bowen, Eric B. Larson, Paul K. Crane

Biostatistics Faculty Publications

INTRODUCTION—Findings for genetic correlates of late-onset Alzheimer's disease (LOAD) in studies that rely solely on clinic visits may differ from those with capacity to follow participants unable to attend clinic visits.

METHODS—We evaluated previously identified LOAD-risk single nucleotide variants in the prospective Adult Changes in Thought study, comparing hazard ratios (HRs) estimated using the full data set of both in-home and clinic visits (n = 1697) to HRs estimated using only data that were obtained from clinic visits (n = 1308). Models were adjusted for age, sex, principal components to account for ancestry, and additional health indicators.

RESULTS …


Testing The Independence Hypothesis Of Accepted Mutations For Pairs Of Adjacent Amino Acids In Protein Sequences, Jyotsna Ramanan, Peter Revesz Jul 2017

Testing The Independence Hypothesis Of Accepted Mutations For Pairs Of Adjacent Amino Acids In Protein Sequences, Jyotsna Ramanan, Peter Revesz

School of Computing: Faculty Publications

Evolutionary studies usually assume that the genetic mutations are independent of each other. However, that does not imply that the observed mutations are independent of each other because it is possible that when a nucleotide is mutated, then it may be biologically beneficial if an adjacent nucleotide mutates too. With a number of decoded genes currently available in various genome libraries and online databases, it is now possible to have a large-scale computer-based study to test whether the independence assumption holds for pairs of adjacent amino acids. Hence the independence question also arises for pairs of adjacent amino acids within …


Statistical Methods For Two Problems In Cancer Research: Analysis Of Rna-Seq Data From Archival Samples And Characterization Of Onset Of Multiple Primary Cancers, Jialu Li May 2017

Statistical Methods For Two Problems In Cancer Research: Analysis Of Rna-Seq Data From Archival Samples And Characterization Of Onset Of Multiple Primary Cancers, Jialu Li

Dissertations and Theses (Open Access)

My dissertation is focused on quantitative methodology development and application for two important topics in translational and clinical cancer research.

The first topic was motivated by the challenge of applying transcriptome sequencing (RNA-seq) to formalin-fixation and paraffin-embedding (FFPE) tumor samples for reliable diagnostic development. We designed a biospecimen study to directly compare gene expression results from different protocols to prepare libraries for RNA-seq from human breast cancer tissues, with randomization to fresh-frozen (FF) or FFPE conditions. To comprehensively evaluate the FFPE RNA-seq data quality for expression profiling, we developed multiple computational methods for assessment, such as the uniformity and continuity …


Identification Of Prognostic Genes And Gene Sets For Early-Stage Non-Small Cell Lung Cancer Using Bi-Level Selection Methods, Suyan Tian, Chi Wang, Howard H. Chang, Jianguo Sun Apr 2017

Identification Of Prognostic Genes And Gene Sets For Early-Stage Non-Small Cell Lung Cancer Using Bi-Level Selection Methods, Suyan Tian, Chi Wang, Howard H. Chang, Jianguo Sun

Biostatistics Faculty Publications

In contrast to feature selection and gene set analysis, bi-level selection is a process of selecting not only important gene sets but also important genes within those gene sets. Depending on the order of selections, a bi-level selection method can be classified into three categories – forward selection, which first selects relevant gene sets followed by the selection of relevant individual genes; backward selection which takes the reversed order; and simultaneous selection, which performs the two tasks simultaneously usually with the aids of a penalized regression model. To test the existence of subtype-specific prognostic genes for non-small cell lung cancer …


Estimating The Probability Of Clonal Relatedness Of Pairs Of Tumors In Cancer Patients, Audrey Mauguen, Venkatraman E. Seshan, Irina Ostrovnaya, Colin B. Begg Feb 2017

Estimating The Probability Of Clonal Relatedness Of Pairs Of Tumors In Cancer Patients, Audrey Mauguen, Venkatraman E. Seshan, Irina Ostrovnaya, Colin B. Begg

Memorial Sloan-Kettering Cancer Center, Dept. of Epidemiology & Biostatistics Working Paper Series

Next generation sequencing panels are being used increasingly in cancer research to study tumor evolution. A specific statistical challenge is to compare the mutational profiles in different tumors from a patient to determine the strength of evidence that the tumors are clonally related, i.e. derived from a single, founder clonal cell. The presence of identical mutations in each tumor provides evidence of clonal relatedness, although the strength of evidence from a match is related to how commonly the mutation is seen in the tumor type under investigation. This evidence must be weighed against the evidence in favor of independent tumors …


Detecting Discordance Enrichment Among A Series Of Two-Sample Genome-Wide Expression Data Sets, Yinglei Lai, Fanni Zhang, Tapan Nayak, Reza Modarres, Norman H. Lee, Timothy A. Mccaffrey Jan 2017

Detecting Discordance Enrichment Among A Series Of Two-Sample Genome-Wide Expression Data Sets, Yinglei Lai, Fanni Zhang, Tapan Nayak, Reza Modarres, Norman H. Lee, Timothy A. Mccaffrey

Epidemiology Faculty Publications

Background

With the current microarray and RNA-seq technologies, two-sample genome-wide expression data have been widely collected in biological and medical studies. The related differential expression analysis and gene set enrichment analysis have been frequently conducted. Integrative analysis can be conducted when multiple data sets are available. In practice, discordant molecular behaviors among a series of data sets can be of biological and clinical interest.

Methods

In this study, a statistical method is proposed for detecting discordance gene set enrichment. Our method is based on a two-level multivariate normal mixture model. It is statistically efficient with linearly increased parameter space when …


Statistical Analyses To Detect And Refine Genetic Associations With Neurodegenerative Diseases, Yuriko Katsumata Jan 2017

Statistical Analyses To Detect And Refine Genetic Associations With Neurodegenerative Diseases, Yuriko Katsumata

Theses and Dissertations--Epidemiology and Biostatistics

Dementia is a clinical state caused by neurodegeneration and characterized by a loss of function in cognitive domains and behavior. Alzheimer’s disease (AD) is the most common form of dementia. Although the amyloid β (Aβ) protein and hyperphosphorylated tau aggregates in the brain are considered to be the key pathological hallmarks of AD, the exact cause of AD is yet to be identified. In addition, clinical diagnoses of AD can be error prone. Many previous studies have compared the clinical diagnosis of AD against the gold standard of autopsy confirmation and shown substantial AD misdiagnosis Hippocampal sclerosis of aging (HS-Aging) …