Feature Selection For Longitudinal Data By Using Sign Averages To Summarize Gene Expression Values Over Time,
2019
The First Hospital of Jilin University, China
Feature Selection For Longitudinal Data By Using Sign Averages To Summarize Gene Expression Values Over Time, Suyan Tian, Chi Wang
Biostatistics Faculty Publications
With the rapid evolution of high-throughput technologies, time series/longitudinal high-throughput experiments have become possible and affordable. However, the development of statistical methods dealing with gene expression profiles across time points has not kept up with the explosion of such data. The feature selection process is of critical importance for longitudinal microarray data. In this study, we proposed aggregating a gene’s expression values across time into a single value using the sign average method, thereby degrading a longitudinal feature selection process into a classic one. Regularized logistic regression models with pseudogenes (i.e., the sign average of genes across time as predictors) …
Variable Selection In Accelerated Failure Time (Aft) Frailty Models: An Application Of Penalized Quasi-Likelihood,
2019
Georgia Southern University
Variable Selection In Accelerated Failure Time (Aft) Frailty Models: An Application Of Penalized Quasi-Likelihood, Sarbesh R. Pandeya
College of Graduate Studies: Theses & Dissertations
Variable selection is one of the standard ways of selecting models in large scale datasets. It has applications in many fields of research study, especially in large multi-center clinical trials. One of the prominent methods in variable selection is the penalized likelihood, which is both consistent and efficient. However, the penalized selection is significantly challenging under the influence of random (frailty) covariates. It is even more complicated when there is involvement of censoring as it may not have a closed-form solution for the marginal log-likelihood. Therefore, we applied the penalized quasi-likelihood (PQL) approach that approximates the solution for such a …
Spatiotemporal Dynamics Of Nitrogen And Carbon Biogeochemistry In A Wetland-Stream Sequence,
2019
University of Montana
Spatiotemporal Dynamics Of Nitrogen And Carbon Biogeochemistry In A Wetland-Stream Sequence, Patrick E. Hurley
Graduate Student Theses, Dissertations, & Professional Papers
Studies of aquatic ecosystems often segregate streams from the influential ponds, lakes, and wetland zones that act as important transitions between terrestrial and fluvial systems. Across the aquatic landscape, these zones interact to form linked ecosystems that function as discrete nutrient processing domains, shifting biogeochemical signals due to spatial and temporal variability in hydrologic and biologic controls. Using a mass-balance approach, we profiled nutrient dynamics along a 23-km wetland-stream sequence over three seasons. Hydrologic, morphologic, and biologic conditions, as well as landscape attributes, were quantified to determine potential controls on biogeochemical cycling in a tributary of the Upper Clark Fork …
Bayesian Nonparametric Analysis Of Longitudinal Data With Non-Ignorable Non-Monotone Missingness,
2019
Virginia Commonwealth University
Bayesian Nonparametric Analysis Of Longitudinal Data With Non-Ignorable Non-Monotone Missingness, Yu Cao
Theses and Dissertations
In longitudinal studies, outcomes are measured repeatedly over time, but in reality clinical studies are full of missing data points of monotone and non-monotone nature. Often this missingness is related to the unobserved data so that it is non-ignorable. In such context, pattern-mixture model (PMM) is one popular tool to analyze the joint distribution of outcome and missingness patterns. Then the unobserved outcomes are imputed using the distribution of observed outcomes, conditioned on missing patterns. However, the existing methods suffer from model identification issues if data is sparse in specific missing patterns, which is very likely to happen with a …
On Cluster Robust Models,
2019
Claremont Graduate University
On Cluster Robust Models, José Bayoán Santiago Calderón
CGU Theses & Dissertations
Cluster robust models are a kind of statistical models that attempt to estimate parameters considering potential heterogeneity in treatment effects. Absent heterogeneity in treatment effects, the partial and average treatment effect are the same. When heterogeneity in treatment effects occurs, the average treatment effect is a function of the various partial treatment effects and the composition of the population of interest. The first chapter explores the performance of common estimators as a function of the presence of heterogeneity in treatment effects and other characteristics that may influence their performance for estimating average treatment effects. The second chapter examines various approaches …
Global Warming Statistical Analysis,
2019
The University of Akron
Global Warming Statistical Analysis, Jared Skinner
Williams Honors College, Honors Research Projects
This paper will investigate global warming and its effects on natural disasters. I will review the historic movements of climate change and activism, as well as the current discussions surrounding global warming. Secondly, I will examine various datasets, paying attention to the severity and frequency of specific natural disasters. I will then touch briefly on the topic of catastrophe modeling as it relates to the increased risk and losses associated with the discussed natural disasters and how those put the problem of global warming in a framework which financial and government institutions can grasp. I will also be analyzing economic …
Bayesian Analysis For The Intraclass Model And For The Quantile Semiparametric Mixed-Effects Double Regression Models,
2019
Michigan Technological University
Bayesian Analysis For The Intraclass Model And For The Quantile Semiparametric Mixed-Effects Double Regression Models, Duo Zhang
Dissertations, Master's Theses and Master's Reports
This dissertation consists of three distinct but related research projects. The first two projects focus on objective Bayesian hypothesis testing and estimation for the intraclass correlation coefficient in linear models. The third project deals with Bayesian quantile inference for the semiparametric mixed-effects double regression models. In the first project, we derive the Bayes factors based on the divergence-based priors for testing the intraclass correlation coefficient (ICC). The hypothesis testing of the ICC is used to test the uncorrelatedness in multilevel modeling, and it has not well been studied from an objective Bayesian perspective. Simulation results show that the two sorts …
Latent Growth Model Approach To Characterize Maternal Prenatal Dna Methylation Trajectories,
2019
Virginia Commonwealth University
Latent Growth Model Approach To Characterize Maternal Prenatal Dna Methylation Trajectories, Dana Lapato
Theses and Dissertations
Background. DNA methylation (DNAm) is a removable chemical modification to the DNA sequence intimately associated with genomic stability, cellular identity, and gene expression. DNAm patterning reflects joint contributions from genetic, environmental, and behavioral factors. As such, differences in DNAm patterns may explain interindividual variability in risk liability for complex traits like major depression (MD). Hundreds of significant DNAm loci have been identified using cross-sectional association studies. This dissertation builds on that foundational work to explore novel statistical approaches for longitudinal DNAm analyses. Methods. Repeated measures of genome-wide DNAm and social and environmental determinants of health were collected up to six …
Toward Using High-Frequency Coastal Radars For Calibration Of S-Ais Based Ocean Vessel Tracking Models,
2019
Wilfrid Laurier University
Toward Using High-Frequency Coastal Radars For Calibration Of S-Ais Based Ocean Vessel Tracking Models, Ben Freidrich
Theses and Dissertations (Comprehensive)
Most of the world relies on ships for transportation, shipping, and tourism. Automatic Identification System messages are transmitted from ships and provide a wealth of positional data on these open ocean vessels. This data is being utilized to determine the optimal path for ships, as well as predicting where a ship may be going in the near future. It has only been in the past decade that Automatic Identification Systems (AIS) signals have been easily received with satellites (S-AIS) so there have been few studies that look at using available information and pairing it with the new abundance of ship …
Statistical Methods For Mixed Frequency Data Sampling Models,
2019
Michigan Technological University
Statistical Methods For Mixed Frequency Data Sampling Models, Yun Liu
Dissertations, Master's Theses and Master's Reports
The MIDAS models are developed to handle different sampling frequencies in one regression model, preserving information in the higher sampling frequency. Time averaging has been the traditional parametric approach to handle mixed sampling frequencies. However, it ignores information potentially embedded in high frequency. MIDAS regression models provide a concise way to utilize additional information in HF variables. While a parametric MIDAS model provides a parsimonious way to summarize information in HF data, nonparametric models would maintain more flexibility at the expense of the computational complexity. Moreover, one parametric form may not necessarily be appropriate for all cross-sectional subjects. This thesis …
A Logitudinal Feature Selection Method Identifies Relevant Genes To Distinguish Complicated Injury And Uncomplicated Injury Over Time,
2018
The First Hospital of Jilin University, China
A Logitudinal Feature Selection Method Identifies Relevant Genes To Distinguish Complicated Injury And Uncomplicated Injury Over Time, Suyan Tian, Chi Wang, Howard H. Chang
Biostatistics Faculty Publications
Background: Feature selection and gene set analysis are of increasing interest in the field of bioinformatics. While these two approaches have been developed for different purposes, we describe how some gene set analysis methods can be utilized to conduct feature selection.
Methods: We adopted a gene set analysis method, the significance analysis of microarray gene set reduction (SAMGSR) algorithm, to carry out feature selection for longitudinal gene expression data.
Results: Using a real-world application and simulated data, it is demonstrated that the proposed SAMGSR extension outperforms other relevant methods. In this study, we illustrate that a gene’s expression profiles over …
Instances Of Influenza In The United States Visualized,
2018
CUNY New York City College of Technology
Instances Of Influenza In The United States Visualized, Parth Patel
Publications and Research
The Tycho Project collects large data sets related to healthcare and in particular, instances and geographical information of diseases. We look at the instance counts and locations of Influenza from 1919-1951 across the United States. We hope to find seasonal and geographical insight to the spread of the disease.
Role Of Misclassification Estimates In Estimating Disease Prevalence And A Non-Linear Approach To Study Synchrony Using Heart Rate Variability In Chickens,
2018
University of Nebraska-Lincoln
Role Of Misclassification Estimates In Estimating Disease Prevalence And A Non-Linear Approach To Study Synchrony Using Heart Rate Variability In Chickens, Dola Pathak
Department of Statistics: Dissertations, Theses, and Student Research
Infectious disease assays can be imperfect. When estimating disease prevalence, these imperfections are accounted for by incorporating assay sensitivity and specificity into point and variance estimates. Unfortunately, these accuracy measures are often treated as fixed constants, rather than acknowledging that they are estimates from an assay validation process. The purpose of this study is to show the detrimental effect of not taking into account this sampling variability when samples are obtained through group testing (aka, pooled testing). We show that confidence interval coverage can dramatically decline as the sample size increases for the main sample of interest. As a remedy …
Snakebite Dynamics Of Colombia: Effects Of Precipitation Seasonality Of Incidence,
2018
Illinois State University
Snakebite Dynamics Of Colombia: Effects Of Precipitation Seasonality Of Incidence, Carlos Cruz
Annual Symposium on Biomathematics and Ecology Education and Research
No abstract provided.
Monitoring And Evaluating The Influences Of Class V Injection Wells On Urban Karst Hydrology,
2018
Western Kentucky University
Monitoring And Evaluating The Influences Of Class V Injection Wells On Urban Karst Hydrology, James Adam Shelley
Masters Theses & Specialist Projects
The response of a karst aquifer to storm events is often faster and more severe than that of a non-karst aquifer. This distinction is often problematic for planners and municipalities, because karst flooding does not typically occur along perennial water courses; thus, traditional flood management strategies are usually ineffective. The City of Bowling Green (CoBG), Kentucky is a representative example of an area plagued by karst flooding. The CoBG, is an urban karst area (UKA), that uses Class V Injection Wells to lessen the severity of flooding. The overall effectiveness, siting, and flooding impact of Injection Wells in UKA’s is …
Association Analyses Of Repeated Measures On Triglyceride And High-Density Lipoprotein Levels: Insights From Gaw20,
2018
Indian Statistical Institute, India
Association Analyses Of Repeated Measures On Triglyceride And High-Density Lipoprotein Levels: Insights From Gaw20, Saurabh Ghosh, David W. Fardo
Biostatistics Faculty Publications
Background: The GAW20 group formed on the theme of methods for association analyses of repeated measures comprised 4sets of investigators. The provided “real” data set included genotypes obtained from a human whole-genome association study based on longitudinal measurements of triglycerides (TGs) and high-density lipoprotein in addition to methylation levels before and after administration of fenofibrate. The simulated data set contained 200 replications of methylation levels and posttreatment TGs, mimicking the real data set.
Results: The different investigators in the group focused on the statistical challenges unique to family-based association analyses of phenotypes measured longitudinally and applied a wide spectrum of …
Longitudinal Data Methods For Evaluating Genome-By-Epigenome Interactions In Families,
2018
University of Kentucky
Longitudinal Data Methods For Evaluating Genome-By-Epigenome Interactions In Families, Justin C. Strickland, I-Chen Chen, Chanung Wang, David W. Fardo
Psychology Faculty Publications
Background: Longitudinal measurement is commonly employed in health research and provides numerous benefits for understanding disease and trait progression over time. More broadly, it allows for proper treatment of correlated responses within clusters. We evaluated 3 methods for analyzing genome-by-epigenome interactions with longitudinal outcomes from family data.
Results: Linear mixed-effect models, generalized estimating equations, and quadratic inference functions were used to test a pharmacoepigenetic effect in 200 simulated posttreatment replicates. Adjustment for baseline outcome provided greater power and more accurate control of Type I error rates than computation of a pre-to-post change score.
Conclusions: Comparison of all modeling approaches indicated …
The Influence Of Parental Control And Parent-Child Relational Qualities On Adolescent Internet Addiction: A 3-Year Longitudinal Study In Hong Kong,
2018
University of Kentucky
The Influence Of Parental Control And Parent-Child Relational Qualities On Adolescent Internet Addiction: A 3-Year Longitudinal Study In Hong Kong, Daniel T. L. Shek, Xiaoqin Zhu, Cecilia M. S. Ma
Pediatrics Faculty Publications
This study investigated how parental behavioral control, parental psychological control, and parent-child relational qualities predicted the initial level and rate of change in adolescent internet addiction (IA) across the junior high school years. The study also investigated the concurrent and longitudinal effects of different parenting factors on adolescent IA. Starting from the 2009/2010 academic year, 3,328 Grade 7 students (Mage = 12.59 ± 0.74 years) from 28 randomly selected secondary schools in Hong Kong responded on a yearly basis to a questionnaire measuring multiple constructs including socio-demographic characteristics, perceived parenting characteristics, and IA. Individual growth curve (IGC) analyses …
Longitudinal Tracking Of Physiological State With Electromyographic Signals.,
2018
University of Louisville
Longitudinal Tracking Of Physiological State With Electromyographic Signals., Robert Warren Stallard
Electronic Theses and Dissertations
Electrophysiological measurements have been used in recent history to classify instantaneous physiological configurations, e.g., hand gestures. This work investigates the feasibility of working with changes in physiological configurations over time (i.e., longitudinally) using a variety of algorithms from the machine learning domain. We demonstrate a high degree of classification accuracy for a binary classification problem derived from electromyography measurements before and after a 35-day bedrest. The problem difficulty is increased with a more dynamic experiment testing for changes in astronaut sensorimotor performance by taking electromyography and force plate measurements before, during, and after a jump from a small platform. A …
Self-Reported Risk And Delinquent Behavior And Problem Behavioral Intention In Hong Kong Adolescents: The Role Of Moral Competence And Spirituality,
2018
University of Kentucky
Self-Reported Risk And Delinquent Behavior And Problem Behavioral Intention In Hong Kong Adolescents: The Role Of Moral Competence And Spirituality, Daniel T. L. Shek, Xiaoqin Zhu
Pediatrics Faculty Publications
Based on the six-wave data collected from Grade 7 to Grade 12 students (N = 3,328 at Wave 1), this pioneer study examined the development of problem behaviors (risk and delinquent behavior and problem behavioral intention) and the predictors (moral competence and spirituality) among adolescents in Hong Kong. Individual growth curve models revealed that while risk and delinquent behavior accelerated and then slowed down in the high school years, adolescent problem behavioral intention slightly accelerated over time. After controlling the background socio-demographic factors, moral competence and spirituality were negatively associated with risk and delinquent behavior as well as problem …
