Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Longitudinal data

Discipline
Institution
Publication Year
Publication
Publication Type

Articles 1 - 30 of 52

Full-Text Articles in Statistics and Probability

Flexible Modeling Of Non-Gaussian Longitudinal Data: Some Approaches Using Copula, Subhajit Chattopadhyay Mar 2026

Flexible Modeling Of Non-Gaussian Longitudinal Data: Some Approaches Using Copula, Subhajit Chattopadhyay

Doctoral Theses

Longitudinal data are common in medical and biological sciences, where measurements are gathered from subjects over time to explore relationships with explanatory variables (covariates) and to uncover the underlying mechanisms of dependence among these measurements. The responses observed at each instance can be either discrete or continuous. One of the primary challenges in longitudinal data analysis lies in the non-Gaussian nature of the response variables. As a result, there are relatively few multivariate models in the literature that effectively address the specific characteristics observed in such datasets. In this dissertation, we address four problems concerning longitudinal data analysis by developing …


Propensity Score Matching Accounting For Longitudinal Trends Before Baseline With Group-Based Trajectory Modeling, Dustin R. Bastaich Jan 2026

Propensity Score Matching Accounting For Longitudinal Trends Before Baseline With Group-Based Trajectory Modeling, Dustin R. Bastaich

Theses and Dissertations

Propensity score matching is used in observational studies to balance baseline attributes between a treatment of interest and a control group. Propensity score matching typically relies on baseline variables, but longitudinal trends in patient characteristics can also influence treatment decisions and subsequent health outcomes. This dissertation extends standard approaches by explicitly incorporating longitudinal trajectories of key variables into the propensity score estimation process.

Trends in a longitudinal variable prior to baseline were characterized using group-based trajectory modeling. A two-step modeling approach was implemented where trajectory groups of a key variable were first estimated and then included as covariates in the …


Functional Data Analysis On Life Expectancy And Healthcare Expenditure, Hagen Sanchez Jul 2025

Functional Data Analysis On Life Expectancy And Healthcare Expenditure, Hagen Sanchez

Theses and Dissertations

This study analyzes the life expectancy for 237 countries from 1950-2023 and the healthcare expenditure for 50 countries from 1970-2022, and how life expectancy and healthcare expenditure relate to each other. Functional Principal Components Analysis was used to analyze the life expectancy and healthcare expenditure for each of the countries. Additionally, the regions of the countries were analyzed to identify any regional trends for the life expectancy data. Due to missing data and the structure of the healthcare data, multiple imputation methods and Principal Component Analysis techniques were explored for the healthcare data. Furthermore, a simulation study was conducted to …


Bayesian Joint Modeling Of Longitudinal Data And Interval-Censored Failure Time Data, Yuchen Mao Apr 2025

Bayesian Joint Modeling Of Longitudinal Data And Interval-Censored Failure Time Data, Yuchen Mao

Theses and Dissertations

Longitudinal data are a collection of repeated observations of the same subjects at different points in time. Interval-censored data arise when the time to the event of interest for each subject is never exactly observed but known to fall between two consecutive points in time. Joint analysis of longitudinal data and interval-censored failure time data can lead to more accurate estimates compared to separate modeling when the correlation among events of interest or intracluster correlation is present. The aim of this dissertation is to develop efficient and reliable joint analyses of longitudinal and interval-censored failure time data using Bayesian methods. …


A New Robust Imputation Method For Longitudinal Data With Non-Normal Continuous Outcomes, Ahmed M. Gad Prof, Yasmie A. Mohamed, Nesma M. Daewish Dr, Abdelnaser S. Abdrabou Prof, Wafaa M. Ibrahim Dr Jan 2025

A New Robust Imputation Method For Longitudinal Data With Non-Normal Continuous Outcomes, Ahmed M. Gad Prof, Yasmie A. Mohamed, Nesma M. Daewish Dr, Abdelnaser S. Abdrabou Prof, Wafaa M. Ibrahim Dr

Business Administration

Missing values is very common in longitudinal data and it is the main challenge in analysis of longitudinal data. Missing values have a significant effect on longitudinal data analysis because they lead to loss of information, biased estimates, and misleading results. In practice there is a need for an imputation method to deal with missing values.

Aim: In this study a new robust regression-based imputation method to deal with missing values in longitudinal data is proposed. This method utilizes the modified adaptive linear regression model and does not require the normality of the responses. It is a novel robust imputation …


Model-Free Organization Of Patient Reported Outcomes Data: Geometrical Rep-Resentation Of The Modified Compartmen-Talization Method, Manasi Sheth, N. Rao Chaganty Jan 2025

Model-Free Organization Of Patient Reported Outcomes Data: Geometrical Rep-Resentation Of The Modified Compartmen-Talization Method, Manasi Sheth, N. Rao Chaganty

Mathematics & Statistics Faculty Publications

There is a recent advancement in the field of mathematics and statistics to understand the geometry or connectedness of the data due to the massive amounts of data being generated. The data provided for analyses are usually very large and need to be organized and minimized in order to make it more useful and meaningful. In biostatistics or medical field, it is important for patients to have access to high-quality, safe and effective and/ or efficacious medical products. It is quite necessary to ascertain that the patients and their care-partners stay at the center of the regulatory decision-making process. In …


Statistical Approaches To Estimate Bidirectional And Time-Varying Causal Effects Using Mendelian Randomization, Jinhao Zou Aug 2023

Statistical Approaches To Estimate Bidirectional And Time-Varying Causal Effects Using Mendelian Randomization, Jinhao Zou

Dissertations and Theses (Open Access)

Mendelian Randomization (MR) is an epidemiological framework using genetic variants as instrumental variables (IVs) to examine the causal effect of an exposure on an outcome. It is widely used to detect causal factors of diseases and provide insight into the biological pathway of diseases. Current methods under the MR framework are built to estimate the unidirectional causal effects of exposures on outcomes and neglect the potential bidirectional causal effects. However, a bidirectional causal effect creates a feedback loop that biases the casual inference in MR studies. Furthermore, current MR methods estimate the causal effect as a single value using cross-sectional …


Nonlinear Mixed-Effects Models For Hiv Viral Load Trajectories Before And After Antiretroviral Therapy Interruption, Incorporating Left Censoring, Sihaoyu Gao, Lang Wu, Tingting Yu, Roger Kouyos, Huldrych F. Gunthard, Rui Wang Jan 2022

Nonlinear Mixed-Effects Models For Hiv Viral Load Trajectories Before And After Antiretroviral Therapy Interruption, Incorporating Left Censoring, Sihaoyu Gao, Lang Wu, Tingting Yu, Roger Kouyos, Huldrych F. Gunthard, Rui Wang

Harvard University Biostatistics Working Paper Series

Characterizing features of the viral rebound trajectories and identifying host, virological, and immunological factors that are predictive of the viral rebound trajectories are central to HIV cure research. In this paper, we investigate if key features of HIV viral decay and CD4 trajectories during antiretroviral therapy (ART) are associated with characteristics of HIV viral rebound following ART interruption. Nonlinear mixed effect (NLME) models are used to model viral load trajectories before and following ART interruption, incorporating left censoring due to lower detection limits of viral load assays. A stochastic approximation EM (SAEM) algorithm is used for parameter estimation and inference. …


An Ensemble Of The Icluster Method To Analyze Longitudinal Lncrna Expression Data For Psoriasis Patients, Suyan Tian, Chi Wang Apr 2021

An Ensemble Of The Icluster Method To Analyze Longitudinal Lncrna Expression Data For Psoriasis Patients, Suyan Tian, Chi Wang

Internal Medicine Faculty Publications

BACKGROUND: Psoriasis is an immune-mediated, inflammatory disorder of the skin with chronic inflammation and hyper-proliferation of the epidermis. Since psoriasis has genetic components and the diseased tissue of psoriasis is very easily accessible, it is natural to use high-throughput technologies to characterize psoriasis and thus seek targeted therapies. Transcriptional profiles change correspondingly after an intervention. Unlike cross-sectional gene expression data, longitudinal gene expression data can capture the dynamic changes and thus facilitate causal inference.

METHODS: Using the iCluster method as a building block, an ensemble method was proposed and applied to a longitudinal gene expression dataset for psoriasis, with the …


Incorporation And Measurement Of Uncertainty In Clustered And Spatial Data, Yuan Hong Oct 2020

Incorporation And Measurement Of Uncertainty In Clustered And Spatial Data, Yuan Hong

Theses and Dissertations

Analyzing population representative datasets for local estimation and predictions over time is important for monitoring related public health issues, however, there are many statistical challenges associated with such analyses. Mixed effect models are one of the common options which can incorporate time and spatial effect in the model and related inference is well established.

In the first part of this dissertation, to estimate area-level prevalence using individuallevel data, small area estimation (SAE) with post-stratified mixed effect models were used where sampling weights were also incorporated into it. However, if poststratification which requires more computation effort can improve estimation accuracy is …


Joint Models Of Longitudinal Outcomes And Informative Time, Jangdong Seo Jun 2020

Joint Models Of Longitudinal Outcomes And Informative Time, Jangdong Seo

Journal of Modern Applied Statistical Methods

Longitudinal data analyses commonly assume that time intervals are predetermined and have no information regarding the outcomes. However, there might be irregular time intervals and informative time. Presented are joint models and asymptotic behaviors of the parameter estimates. Also, the models are applied for real data sets.


Beta Regression Models For Repeated-Measures Data Analysis, Nicholas A. Hein Aug 2019

Beta Regression Models For Repeated-Measures Data Analysis, Nicholas A. Hein

Theses & Dissertations

Bounded data often give rise to uncorrectable skew and heteroscedasticity. Bounded data are a relatively frequent occurrence in clinical and research settings. For example, in neuropsychology, most neurocognitive tests are bounded, and subjects are repeatedly measured over time. The statistician needs to choose a model that accounts for the correlated nature of the repeated measures. The Beta distribution is a natural choice for modeling bounded data. Currently, generalized linear mixed models (GLMM) and generalized estimating equations (GEE) are two methods that can be used to model Beta distributed data with repeated measures. However, GLMMs and GEEs have limitations, i.e., GLMMs …


Feature Selection For Longitudinal Data By Using Sign Averages To Summarize Gene Expression Values Over Time, Suyan Tian, Chi Wang Mar 2019

Feature Selection For Longitudinal Data By Using Sign Averages To Summarize Gene Expression Values Over Time, Suyan Tian, Chi Wang

Biostatistics Faculty Publications

With the rapid evolution of high-throughput technologies, time series/longitudinal high-throughput experiments have become possible and affordable. However, the development of statistical methods dealing with gene expression profiles across time points has not kept up with the explosion of such data. The feature selection process is of critical importance for longitudinal microarray data. In this study, we proposed aggregating a gene’s expression values across time into a single value using the sign average method, thereby degrading a longitudinal feature selection process into a classic one. Regularized logistic regression models with pseudogenes (i.e., the sign average of genes across time as predictors) …


Bayesian Nonparametric Analysis Of Longitudinal Data With Non-Ignorable Non-Monotone Missingness, Yu Cao Jan 2019

Bayesian Nonparametric Analysis Of Longitudinal Data With Non-Ignorable Non-Monotone Missingness, Yu Cao

Theses and Dissertations

In longitudinal studies, outcomes are measured repeatedly over time, but in reality clinical studies are full of missing data points of monotone and non-monotone nature. Often this missingness is related to the unobserved data so that it is non-ignorable. In such context, pattern-mixture model (PMM) is one popular tool to analyze the joint distribution of outcome and missingness patterns. Then the unobserved outcomes are imputed using the distribution of observed outcomes, conditioned on missing patterns. However, the existing methods suffer from model identification issues if data is sparse in specific missing patterns, which is very likely to happen with a …


The Nested Joint Clustering Via Dirichlet Process Mixture Model, Shengtong Han, Hongmei Zhang, Wenhui Sheng, Hasan Arshad Jan 2019

The Nested Joint Clustering Via Dirichlet Process Mixture Model, Shengtong Han, Hongmei Zhang, Wenhui Sheng, Hasan Arshad

Mathematical and Statistical Science Faculty Research and Publications

This article focuses on the clustering problem based on Dirichlet process (DP) mixtures. To model both time invariant and temporal patterns, different from other existing clustering methods, the proposed semi-parametric model is flexible in that both the common and unique patterns are taken into account simultaneously. Furthermore, by jointly clustering subjects and the associated variables, the intrinsic complex shared patterns among subjects and among variables are expected to be captured. The number of clusters and cluster assignments are directly inferred with the use of DP. Simulation studies illustrate the effectiveness of the proposed method. An application to wheal size data …


Association Analyses Of Repeated Measures On Triglyceride And High-Density Lipoprotein Levels: Insights From Gaw20, Saurabh Ghosh, David W. Fardo Sep 2018

Association Analyses Of Repeated Measures On Triglyceride And High-Density Lipoprotein Levels: Insights From Gaw20, Saurabh Ghosh, David W. Fardo

Biostatistics Faculty Publications

Background: The GAW20 group formed on the theme of methods for association analyses of repeated measures comprised 4sets of investigators. The provided “real” data set included genotypes obtained from a human whole-genome association study based on longitudinal measurements of triglycerides (TGs) and high-density lipoprotein in addition to methylation levels before and after administration of fenofibrate. The simulated data set contained 200 replications of methylation levels and posttreatment TGs, mimicking the real data set.

Results: The different investigators in the group focused on the statistical challenges unique to family-based association analyses of phenotypes measured longitudinally and applied a wide spectrum of …


Missing Data In Longitudinal Surveys: A Comparison Of Performance Of Modern Techniques, Paola Zaninotto, Amanda Sacker Dec 2017

Missing Data In Longitudinal Surveys: A Comparison Of Performance Of Modern Techniques, Paola Zaninotto, Amanda Sacker

Journal of Modern Applied Statistical Methods

Using a simulation study, the performance of complete case analysis, full information maximum likelihood, multivariate normal imputation, multiple imputation by chained equations and two-fold fully conditional specification to handle missing data were compared in longitudinal surveys with continuous and binary outcomes, missing covariates, and an interaction term.


Joint Modelling Of Longitudinal Measurements And Time-To-Event Data : Application To Hiv Study, Mirna Walid Halawani May 2017

Joint Modelling Of Longitudinal Measurements And Time-To-Event Data : Application To Hiv Study, Mirna Walid Halawani

Theses, Dissertations and Culminating Projects

Longitudinal and survival data are frequently collected in biomedical studies. The research questions of interest in these studies often require separate analysis of the outcomes. But in many occasions interest also lies in studying their association structures, such as in biomarker research, where the clinical studies are designed to identify biomarkers with strong prognostic capabilities for event time outcomes. In the separate analyses, a linear mixed-effects model is used for modeling the longitudinal data to study the changing trend of the response overtime when controlling some covariates and a survival model is used to model the time-to-event data. A common …


A Semi-Parametric Approach For Analyzing Longitudinal Measurements With Non-Ignorable Missingness Using Regression Spline, Taban Baghfalaki, Saeide Sefidi, Mojtaba Ganjali Jun 2015

A Semi-Parametric Approach For Analyzing Longitudinal Measurements With Non-Ignorable Missingness Using Regression Spline, Taban Baghfalaki, Saeide Sefidi, Mojtaba Ganjali

Applications and Applied Mathematics: An International Journal (AAM)

In longitudinal studies with missingness, shared parameter models (SPM) provide appropriate framework for the joint modeling of the measurements and missingness process. These models use a set of random effects to account for the interdependence between two processes. Sometimes the longitudinal responses may not be fitted well by using a linear model and some non-parametric methods have to be used. Also, parametric assumptions are typically made for the random effects distribution, and violation of those may affect the parameter estimates and standard errors. To overcome these problems, we propose a semi-parametric model for the joint modelling of longitudinal markers and …


Statistical Modeling And Prediction Of Hiv/Aids Prognosis: Bayesian Analyses Of Nonlinear Dynamic Mixtures, Xiaosun Lu Jul 2014

Statistical Modeling And Prediction Of Hiv/Aids Prognosis: Bayesian Analyses Of Nonlinear Dynamic Mixtures, Xiaosun Lu

USF Tampa Graduate Theses and Dissertations

Statistical analyses and modeling have contributed greatly to our understanding of the pathogenesis of HIV-1 infection; they also provide guidance for the treatment of AIDS patients and evaluation of antiretroviral (ARV) therapies. Various statistical methods, nonlinear mixed-effects models in particular, have been applied to model the CD4 and viral load trajectories. A common assumption in these methods is all patients come from a homogeneous population following one mean trajectories. This assumption unfortunately obscures important characteristic difference between subgroups of patients whose response to treatment and whose disease trajectories are biologically different. It also may lack the robustness against population heterogeneity …


Variable-Domain Functional Regression For Modeling Icu Data, Jonathan E. Gellar, Elizabeth Colantuoni, Dale M. Needham, Ciprian M. Crainiceanu May 2014

Variable-Domain Functional Regression For Modeling Icu Data, Jonathan E. Gellar, Elizabeth Colantuoni, Dale M. Needham, Ciprian M. Crainiceanu

Johns Hopkins University, Dept. of Biostatistics Working Papers

We introduce a class of scalar-on-function regression models with subject-specific functional predictor domains. The fundamental idea is to consider a bivariate functional parameter that depends both on the functional argument and on the width of the functional predictor domain. Both parametric and nonparametric models are introduced to fit the functional coefficient. The nonparametric model is theoretically and practically invariant to functional support transformation, or support registration. Methods were motivated by and applied to a study of association between daily measures of the Intensive Care Unit (ICU) Sequential Organ Failure Assessment (SOFA) score and two outcomes: in-hospital mortality, and physical impairment …


Mediation Analysis With Time-Varying Exposures And Mediators, Tyler J. Vanderweele, Eric Tchetgen Tchetgen Mar 2014

Mediation Analysis With Time-Varying Exposures And Mediators, Tyler J. Vanderweele, Eric Tchetgen Tchetgen

Harvard University Biostatistics Working Paper Series

In this paper we consider mediation analysis when exposures and mediators vary over time. We give non-parametric identification results, discuss parametric implementation, and also provide a weighting approach to direct and indirect effects based on combining the results of two marginal structural models. We also discuss how our results give rise to a causal interpretation of the effect estimates produced from longitudinal structural equation models. When there are no time-varying confounders affected by prior exposure and mediator values, identification of direct and indirect effects is achieved by a longitudinal version of Pearl's mediation formula. When there are time-varying confounders affected …


Regularization Methods For Predicting An Ordinal Response Using Longitudinal High-Dimensional Genomic Data, Jiayi Hou Nov 2013

Regularization Methods For Predicting An Ordinal Response Using Longitudinal High-Dimensional Genomic Data, Jiayi Hou

Theses and Dissertations

Ordinal scales are commonly used to measure health status and disease related outcomes in hospital settings as well as in translational medical research. Notable examples include cancer staging, which is a five-category ordinal scale indicating tumor size, node involvement, and likelihood of metastasizing. Glasgow Coma Scale (GCS), which gives a reliable and objective assessment of conscious status of a patient, is an ordinal scaled measure. In addition, repeated measurements are common in clinical practice for tracking and monitoring the progression of complex diseases. Classical ordinal modeling methods based on the likelihood approach have contributed to the analysis of data in …


Regression Trees For Longitudinal Data, Madan Gopal Kundu, Jaroslaw Harezlak Sep 2013

Regression Trees For Longitudinal Data, Madan Gopal Kundu, Jaroslaw Harezlak

COBRA Preprint Series

Often when a longitudinal change is studied in a population of interest we find that changes over time are heterogeneous (in terms of time and/or covariates' effect) and a traditional linear mixed effect model [Laird and Ware, 1982] on the entire population assuming common parametric form for covariates and time may not be applicable to the entire population. This is usually the case in studies when there are many possible predictors influencing the response trajectory. For example, Raudenbush [2001] used depression as an example to argue that it is incorrect to assume that all the people in a given population …


Jmasm 32: Multiple Imputation Of Missing Multilevel, Longitudinal Data: A Case When Practical Considerations Trump Best Practices?, Jennifer E. V. Lloyd, Jelena Obradović, Richard M. Carpiano, Frosso Motti-Stefanidi May 2013

Jmasm 32: Multiple Imputation Of Missing Multilevel, Longitudinal Data: A Case When Practical Considerations Trump Best Practices?, Jennifer E. V. Lloyd, Jelena Obradović, Richard M. Carpiano, Frosso Motti-Stefanidi

Journal of Modern Applied Statistical Methods

A pedagogical tool is presented for applied researchers dealing with incomplete multilevel, longitudinal data. It explains why such data pose special challenges regarding missingness. Syntax created to perform a multiply-imputed growth modeling procedure in Stata Version 11 (StataCorp, 2009) is also described.


Analysis Of Continuous Longitudinal Data With Arma(1, 1) And Antedependence Correlation Structures, Sirisha Mushti Apr 2013

Analysis Of Continuous Longitudinal Data With Arma(1, 1) And Antedependence Correlation Structures, Sirisha Mushti

Mathematics & Statistics Theses & Dissertations

Longitudinal or repeated measure data are common in biomedical and clinical trials. These data are often collected on individuals at scheduled times resulting in dependent responses. Inference methods for studying the behavior of responses over time as well as methods to study the association with certain risk factors or covariates taking into account the dependencies are of great importance. In this research we focus our study on the analysis of continuous longitudinal data. To model the dependencies of the responses over time, we consider appropriate correlation structures generated by the stationary and non-stationary time-series models. We develop new estimation procedures …


Vertically Shifted Mixture Models For Clustering Longitudinal Data By Shape, Brianna C. Heggeseth, Nicholas P. Jewell Mar 2013

Vertically Shifted Mixture Models For Clustering Longitudinal Data By Shape, Brianna C. Heggeseth, Nicholas P. Jewell

U.C. Berkeley Division of Biostatistics Working Paper Series

Longitudinal studies play a prominent role in health, social and behavioral sciences as well as in the biological sciences, economics, and marketing. By following subjects over time, temporal changes in an outcome of interest can be directly observed and studied. An important question concerns the existence of distinct trajectory patterns. One way to determine these distinct patterns is through cluster analysis, which seeks to separate objects (subjects, patients, observational units) into homogeneous groups. Many methods have been adapted for longitudinal data, but almost all of them fail to explicitly group trajectories according to distinct pattern shapes. To fulfill the need …


Likelihood Ratio Tests For The Mean Structure Of Correlated Functional Processes, Ana-Maria Staicu, Yingxing Li, Ciprian Crainiceanu, David M. Ruppert Nov 2012

Likelihood Ratio Tests For The Mean Structure Of Correlated Functional Processes, Ana-Maria Staicu, Yingxing Li, Ciprian Crainiceanu, David M. Ruppert

Johns Hopkins University, Dept. of Biostatistics Working Papers

The paper introduces a general framework for testing hypotheses about the structure of the mean function of complex functional processes. Important particular cases of the proposed framework are: 1) testing the null hypotheses that the mean of a functional process is parametric against a nonparametric alternative; and 2) testing the null hypothesis that the means of two possibly correlated functional processes are equal or differ by only a simple parametric function. A global pseudo likelihood ratio test is proposed and its asymptotic distribution is derived. The size and power properties of the test are confirmed in realistic simulation scenarios. Finite …


Longitudinal Functional Models With Structured Penalties, Madan G. Kundu, Jaroslaw Harezlak, Timothy W. Randolph Nov 2012

Longitudinal Functional Models With Structured Penalties, Madan G. Kundu, Jaroslaw Harezlak, Timothy W. Randolph

Johns Hopkins University, Dept. of Biostatistics Working Papers

Collection of functional data is becoming increasingly common including longitudinal observations in many studies. For example, we use magnetic resonance (MR) spectra collected over a period of time from late stage HIV patients. MR spectroscopy (MRS) produces a spectrum which is a mixture of metabolite spectra, instrument noise and baseline profile. Analysis of such data typically proceeds in two separate steps: feature extraction and regression modeling. In contrast, a recently-proposed approach, called partially empirical eigenvectors for regression (PEER) (Randolph, Harezlak and Feng, 2012), for functional linear models incorporates a priori knowledge via a scientifically-informed penalty operator in the regression function …


Prediction In Several Conventional Contexts, Bertrand Clarke, Jennifer Clarke Jan 2012

Prediction In Several Conventional Contexts, Bertrand Clarke, Jennifer Clarke

Department of Statistics: Faculty Publications

We review predictive techniques from several traditional branches of statistics. Starting with prediction based on the normal model and on the empirical distribution function, we proceed to techniques for various forms of regression and classification. Then, we turn to time series, longitudinal data, and survival analysis. Our focus throughout is on the mechanics of prediction more than on the properties of predictors.