Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Statistical Methodology (9)
- Statistical Models (9)
- Biostatistics (8)
- Data Science (3)
- Statistical Theory (3)
-
- Survival Analysis (3)
- Computer Sciences (2)
- Longitudinal Data Analysis and Time Series (2)
- Medicine and Health Sciences (2)
- Multivariate Analysis (2)
- Artificial Intelligence and Robotics (1)
- Categorical Data Analysis (1)
- Discrete Mathematics and Combinatorics (1)
- Economics (1)
- Mathematics (1)
- Medical Biomathematics and Biometrics (1)
- Medical Sciences (1)
- Numerical Analysis and Scientific Computing (1)
- Other Economics (1)
- Other Public Health (1)
- Probability (1)
- Public Health (1)
- Regional Economics (1)
- Social Statistics (1)
- Social and Behavioral Sciences (1)
- Theory and Algorithms (1)
- Vital and Health Statistics (1)
- Keyword
-
- Statistics (14)
- Biostatistics (4)
- Bayesian Statistics (2)
- ACS (1)
- American Community Survey (1)
-
- Bayesian (1)
- CAR (1)
- Computable Phenotypes (1)
- Computational Immunology (1)
- Computer Science (1)
- Conditional Autoregressive Model (1)
- Counterfactual (1)
- Deep Generative Model (1)
- Economics (1)
- Electronic Health Records (1)
- Functional principal component analysis (1)
- GPT (1)
- Genetics (1)
- ICAR (1)
- Immunology (1)
- Induced dependent censoring (1)
- Joint modeling (1)
- Keywords Identification (1)
- LLMs (1)
- Large Language Models (1)
- Multiple Instance Classification (1)
- Multivariate recurrent event modeling (1)
- Policy decisions (1)
- Predictive Model (1)
- Spatial (1)
Articles 1 - 15 of 15
Full-Text Articles in Applied Statistics
Statistical Methodologies For Count Time Series Analysis And Topological Data Analysis Of Medical Images, Yuhyeong Jang
Statistical Methodologies For Count Time Series Analysis And Topological Data Analysis Of Medical Images, Yuhyeong Jang
Statistical Science Theses and Dissertations
This dissertation addresses two distinct topics related to count time series analysis and topological medical image analysis, respectively. The first part of the dissertation comprises an application of a count time series model to analysis of US monthly sex trafficking data and development of a new model for multivariate count data that exhibits serial dependence and overdispersion. By imposing a family of multivariate mixed Poisson distributions on the count random vector, the proposed model can accommodate a broad range of overdispersion as well as positive contemporaneous correlations. For maximum likelihood estimation, a computationally feasible EM-type algorithm is derived based on …
Bayesian Spatiotemporal Model For Counterfactual Estimation In Socioeconomic Studies, Duwani W. Gonzalez
Bayesian Spatiotemporal Model For Counterfactual Estimation In Socioeconomic Studies, Duwani W. Gonzalez
Statistical Science Theses and Dissertations
Impact evaluations of regional development programs often require estimating counterfactual outcomes for a small number of treated regions using survey-based areal data. In practice, evaluators typically rely on two-group quasi-experimental methods such as propensity score matching (PSM) and Difference-in-Differences (DiD). These approaches perform poorly when only a few regions receive treatment, and when the set of observed covariates is limited or only partially relevant. Moreover, they typically do not explicitly exploit the spatial and temporal dependence present in survey-based areal data such as in ACS (American Community Survey). This dissertation develops a family of Bayesian spatial predictive models for directly …
Smarter Disease Detection From Electronic Health Record Data: An End-To-End Ai-Augmented Pipeline For Computable Phenotyping, Dylan Owens
Statistical Science Theses and Dissertations
Electronic Health Records (EHR) contain a wealth of structured and unstructured patient data that can be leveraged for computable phenotyping, the process of algorithmically identifying patient cohorts with specific diseases or conditions. Traditional rule-based phenotyping approaches, while interpretable, often struggle with scalability, portability across institutions, and effective use of unstructured clinical narratives. Recent advances in large language models (LLMs) present new opportunities for synthesizing complex free-text information into concise, clinically meaningful representations. However, integrating LLMs into phenotyping workflows requires careful design to maintain transparency, interpretability, and measurable uncertainty—features essential for clinical adoption and downstream applications such as decision support.
We …
Statistical Methods For Joint Outcome Modeling And Dynamic Assessment Of Recurrent Events, Zifang Kong
Statistical Methods For Joint Outcome Modeling And Dynamic Assessment Of Recurrent Events, Zifang Kong
Statistical Science Theses and Dissertations
Recurrent event data frequently arise in clinical studies where individuals experience repeated, possibly related, events over time. These data are often accompanied by sparse and irregular longitudinal measurements, creating challenges for traditional joint modeling approaches that struggle to account for time-dependent associations and within-subject correlations. We propose FRAILTY (Functional Regression with AutoRegressIve fraiLTY), a novel two-step framework that integrates functional principal component analysis (PACE) with a dynamic frailty model featuring autoregressive structure. FRAILTY accommodates both scalar and functional predictors and captures within-subject dependence across recurrent events. To further extend its utility, we develop a multivariate joint modeling framework that simultaneously …
Bayesian And Deep Generative Modeling In Immunology, Yuqiu Yang
Bayesian And Deep Generative Modeling In Immunology, Yuqiu Yang
Statistical Science Theses and Dissertations
Due to the accumulation of a large volume of data of different natures such as sequencing data, proteomics data, and clinical data, statistical methods and deep learning algorithms have become increasingly important in the field of immunology. By leveraging the diverse datasets as well as interdisciplinary knowledge from areas like biology and public health, these quantitative methods have revolutionized this field by providing powerful tools for data analysis, modeling, and prediction. This has led to a deeper understanding of the immune system, accelerated the development of novel therapies, and paved the way for personalized and precision medicine approaches in immunology. …
Bayesian Variational Inference In Keyword Identification And Multiple Instance Classification, Yaofang Hu
Bayesian Variational Inference In Keyword Identification And Multiple Instance Classification, Yaofang Hu
Statistical Science Theses and Dissertations
This dissertation investigates (1) Variational Bayesian Semi-supervised Keyword Extraction and (2) Variational Bayesian Multimodal Multiple Instance Classification.
The expansion of textual data, stemming from various sources such as online product reviews and scholarly publications on scientific discoveries, has created a demand for the extraction of succinct yet comprehensive information. As a result, in recent years, efforts have been spent in developing novel methodologies for keyword extraction. Although many methods have been proposed to automatically extract keywords in the contexts of both unsupervised and fully supervised learning, how to effectively use partially observed keywords, such as author-specified keywords, remains an under-explored …
Bayesian Statistical Modeling Of Spatially Resolved Transcriptomics Data, Xi Jiang
Bayesian Statistical Modeling Of Spatially Resolved Transcriptomics Data, Xi Jiang
Statistical Science Theses and Dissertations
Spatially resolved transcriptomics (SRT) quantifies expression levels at different spatial locations, providing a new and powerful tool to investigate novel biological insights. As experimental technologies enhance both in capacity and efficiency, there arises a growing demand for the development of analytical methodologies.
One question in SRT data analysis is to identify genes whose expressions exhibit spatially correlated patterns, called spatially variable (SV) genes. Most current methods to identify SV genes are built upon the geostatistical model with Gaussian process, which could limit the models' ability to identify complex spatial patterns. In order to overcome this challenge and capture more types …
A Comparison Of Confidence Intervals In State Space Models, Jinyu Du
A Comparison Of Confidence Intervals In State Space Models, Jinyu Du
Statistical Science Theses and Dissertations
This thesis develops general procedures for constructing confidence intervals (CIs) of the error disturbance parameters (standard deviations) and transformations of the error disturbance parameters in time-invariant state space models (ssm). With only a set of observations, estimating individual error disturbance parameters accurately in the presence of other unknown parameters in ssm is a very challenging problem. We attempted to construct four different types of confidence intervals, Wald, likelihood ratio, score, and higher-order asymptotic intervals for both the simple local level model and the general time-invariant state space models (ssm). We show that for a simple local level model, both the …
Optimizing Tumor Xenograft Experiments Using Bayesian Linear And Nonlinear Mixed Modelling And Reinforcement Learning, Mary Lena Bleile
Optimizing Tumor Xenograft Experiments Using Bayesian Linear And Nonlinear Mixed Modelling And Reinforcement Learning, Mary Lena Bleile
Statistical Science Theses and Dissertations
Tumor xenograft experiments are a popular tool of cancer biology research. In a typical such experiment, one implants a set of animals with an aliquot of the human tumor of interest, applies various treatments of interest, and observes the subsequent response. Efficient analysis of the data from these experiments is therefore of utmost importance. This dissertation proposes three methods for optimizing cancer treatment and data analysis in the tumor xenograft context. The first of these is applicable to tumor xenograft experiments in general, and the second two seek to optimize the combination of radiotherapy with immunotherapy in the tumor xenograft …
Dynamic Prediction For Alternating Recurrent Events Using A Semiparametric Joint Frailty Model, Jaehyeon Yun
Dynamic Prediction For Alternating Recurrent Events Using A Semiparametric Joint Frailty Model, Jaehyeon Yun
Statistical Science Theses and Dissertations
Alternating recurrent events data arise commonly in health research; examples include hospital admissions and discharges of diabetes patients; exacerbations and remissions of chronic bronchitis; and quitting and restarting smoking. Recent work has involved formulating and estimating joint models for the recurrent event times considering non-negligible event durations. However, prediction models for transition between recurrent events are lacking. We consider the development and evaluation of methods for predicting future events within these models. Specifically, we propose a tool for dynamically predicting transition between alternating recurrent events in real time. Under a flexible joint frailty model, we derive the predictive probability of …
Bayesian Semi-Supervised Keyphrase Extraction And Jackknife Empirical Likelihood For Assessing Heterogeneity In Meta-Analysis, Guanshen Wang
Bayesian Semi-Supervised Keyphrase Extraction And Jackknife Empirical Likelihood For Assessing Heterogeneity In Meta-Analysis, Guanshen Wang
Statistical Science Theses and Dissertations
This dissertation investigates: (1) A Bayesian Semi-supervised Approach to Keyphrase Extraction with Only Positive and Unlabeled Data, (2) Jackknife Empirical Likelihood Confidence Intervals for Assessing Heterogeneity in Meta-analysis of Rare Binary Events.
In the big data era, people are blessed with a huge amount of information. However, the availability of information may also pose great challenges. One big challenge is how to extract useful yet succinct information in an automated fashion. As one of the first few efforts, keyphrase extraction methods summarize an article by identifying a list of keyphrases. Many existing keyphrase extraction methods focus on the unsupervised setting, …
Improved Statistical Methods For Time-Series And Lifetime Data, Xiaojie Zhu
Improved Statistical Methods For Time-Series And Lifetime Data, Xiaojie Zhu
Statistical Science Theses and Dissertations
In this dissertation, improved statistical methods for time-series and lifetime data are developed. First, an improved trend test for time series data is presented. Then, robust parametric estimation methods based on system lifetime data with known system signatures are developed.
In the first part of this dissertation, we consider a test for the monotonic trend in time series data proposed by Brillinger (1989). It has been shown that when there are highly correlated residuals or short record lengths, Brillinger’s test procedure tends to have significance level much higher than the nominal level. This could be related to the discrepancy between …
Samples, Unite! Understanding The Effects Of Matching Errors On Estimation Of Total When Combining Data Sources, Benjamin Williams
Samples, Unite! Understanding The Effects Of Matching Errors On Estimation Of Total When Combining Data Sources, Benjamin Williams
Statistical Science Theses and Dissertations
Much recent research has focused on methods for combining a probability sample with a non-probability sample to improve estimation by making use of information from both sources. If units exist in both samples, it becomes necessary to link the information from the two samples for these units. Record linkage is a technique to link records from two lists that refer to the same unit but lack a unique identifier across both lists. Record linkage assigns a probability to each potential pair of records from the lists so that principled matching decisions can be made. Because record linkage is a probabilistic …
Modeling Stochastically Intransitive Relationships In Paired Comparison Data, Ryan Patrick Alexander Mcshane
Modeling Stochastically Intransitive Relationships In Paired Comparison Data, Ryan Patrick Alexander Mcshane
Statistical Science Theses and Dissertations
If the Warriors beat the Rockets and the Rockets beat the Spurs, does that mean that the Warriors are better than the Spurs? Sophisticated fans would argue that the Warriors are better by the transitive property, but could Spurs fans make a legitimate argument that their team is better despite this chain of evidence?
We first explore the nature of intransitive (rock-scissors-paper) relationships with a graph theoretic approach to the method of paired comparisons framework popularized by Kendall and Smith (1940). Then, we focus on the setting where all pairs of items, teams, players, or objects have been compared to …
Developing Statistical Methods For Data From Platforms Measuring Gene Expression, Gaoxiang Jia
Developing Statistical Methods For Data From Platforms Measuring Gene Expression, Gaoxiang Jia
Statistical Science Theses and Dissertations
This research contains two topics: (1) PBNPA: a permutation-based non-parametric analysis of CRISPR screen data; (2) RCRnorm: an integrated system of random-coefficient hierarchical regression models for normalizing NanoString nCounter data from FFPE samples.
Clustered regularly-interspaced short palindromic repeats (CRISPR) screens are usually implemented in cultured cells to identify genes with critical functions. Although several methods have been developed or adapted to analyze CRISPR screening data, no single spe- cific algorithm has gained popularity. Thus, rigorous procedures are needed to overcome the shortcomings of existing algorithms. We developed a Permutation-Based Non-Parametric Analysis (PBNPA) algorithm, which computes p-values at the gene level …