Open Access. Powered by Scholars. Published by Universities.®

Biostatistics Commons™

Open Access. Powered by Scholars. Published by Universities.®

Southern Methodist University

Discipline
Keyword
Publication Year
Publication
Publication Type

Articles 1 - 20 of 20

Full-Text Articles in Biostatistics

Statistical Methodologies For Count Time Series Analysis And Topological Data Analysis Of Medical Images, Yuhyeong Jang Aug 2026

Statistical Methodologies For Count Time Series Analysis And Topological Data Analysis Of Medical Images, Yuhyeong Jang

Statistical Science Theses and Dissertations

This dissertation addresses two distinct topics related to count time series analysis and topological medical image analysis, respectively. The first part of the dissertation comprises an application of a count time series model to analysis of US monthly sex trafficking data and development of a new model for multivariate count data that exhibits serial dependence and overdispersion. By imposing a family of multivariate mixed Poisson distributions on the count random vector, the proposed model can accommodate a broad range of overdispersion as well as positive contemporaneous correlations. For maximum likelihood estimation, a computationally feasible EM-type algorithm is derived based on …


Deep Learning Frameworks For Biological Data Integration And Generation, Alexa Beachum Apr 2026

Deep Learning Frameworks For Biological Data Integration And Generation, Alexa Beachum

Statistical Science Theses and Dissertations

Data integration represents a key area of research for analyzing the rapidly growing volume of high-dimensional biological data across sources, stages, and modalities. To model and understand these complex, often non-linear relationships, deep learning has become an increasingly powerful tool. Here, we present two novel deep learning frameworks that address distinct but complementary integration challenges. The first framework aligns single-cell omics data across temporal stages, and the second bridges imaging and omics modalities to generate patient-level molecular profiles.

In Chapter 1, we briefly summarize existing approaches---both statistical and deep learning-based---for single-cell omics data integration and discuss their limitations for handling …


Smarter Disease Detection From Electronic Health Record Data: An End-To-End Ai-Augmented Pipeline For Computable Phenotyping, Dylan Owens Oct 2025

Smarter Disease Detection From Electronic Health Record Data: An End-To-End Ai-Augmented Pipeline For Computable Phenotyping, Dylan Owens

Statistical Science Theses and Dissertations

Electronic Health Records (EHR) contain a wealth of structured and unstructured patient data that can be leveraged for computable phenotyping, the process of algorithmically identifying patient cohorts with specific diseases or conditions. Traditional rule-based phenotyping approaches, while interpretable, often struggle with scalability, portability across institutions, and effective use of unstructured clinical narratives. Recent advances in large language models (LLMs) present new opportunities for synthesizing complex free-text information into concise, clinically meaningful representations. However, integrating LLMs into phenotyping workflows requires careful design to maintain transparency, interpretability, and measurable uncertainty—features essential for clinical adoption and downstream applications such as decision support.

We …


Statistical Methods For Joint Outcome Modeling And Dynamic Assessment Of Recurrent Events, Zifang Kong Aug 2025

Statistical Methods For Joint Outcome Modeling And Dynamic Assessment Of Recurrent Events, Zifang Kong

Statistical Science Theses and Dissertations

Recurrent event data frequently arise in clinical studies where individuals experience repeated, possibly related, events over time. These data are often accompanied by sparse and irregular longitudinal measurements, creating challenges for traditional joint modeling approaches that struggle to account for time-dependent associations and within-subject correlations. We propose FRAILTY (Functional Regression with AutoRegressIve fraiLTY), a novel two-step framework that integrates functional principal component analysis (PACE) with a dynamic frailty model featuring autoregressive structure. FRAILTY accommodates both scalar and functional predictors and captures within-subject dependence across recurrent events. To further extend its utility, we develop a multivariate joint modeling framework that simultaneously …


Towards Reliable Clinical Applications Of Ai Models In Radiotherapy, Biling Wang Aug 2025

Towards Reliable Clinical Applications Of Ai Models In Radiotherapy, Biling Wang

Statistical Science Theses and Dissertations

Over the past decade, artificial intelligence (AI), particularly through deep learning (DL) techniques, has made significant strides in fields like computer vision (CV) and natural language processing (NLP), leading to transformative advancements across numerous applications. This progress has sparked considerable enthusiasm within the medical field, where DL-related research has grown exponentially since 2015. However, despite these promising developments, the real-world deployment of DL models in healthcare remains limited, especially in safety-critical domains such as radiotherapy (RT), where reliability, safety, and sustained performance are critical. This thesis addresses three core challenges associated with the clinical application of DL models: (1) post-deployment …


Bayesian And Deep Generative Modeling In Immunology, Yuqiu Yang Aug 2024

Bayesian And Deep Generative Modeling In Immunology, Yuqiu Yang

Statistical Science Theses and Dissertations

Due to the accumulation of a large volume of data of different natures such as sequencing data, proteomics data, and clinical data, statistical methods and deep learning algorithms have become increasingly important in the field of immunology. By leveraging the diverse datasets as well as interdisciplinary knowledge from areas like biology and public health, these quantitative methods have revolutionized this field by providing powerful tools for data analysis, modeling, and prediction. This has led to a deeper understanding of the immune system, accelerated the development of novel therapies, and paved the way for personalized and precision medicine approaches in immunology. …


Statistical Approaches For The Early Detection Of Colorectal Cancer Using Longitudinal Biomarkers, Emily Berry May 2024

Statistical Approaches For The Early Detection Of Colorectal Cancer Using Longitudinal Biomarkers, Emily Berry

Statistical Science Theses and Dissertations

Colorectal cancer (CRC) is the third leading cause of cancer-related death in the United States [45]. CRC is believed to advance from adenomatous polyps creating a unique opportunity for both early detection and cancer prevention [4, 23]. Like other diseases, CRC screening reduces mortality by detecting cancer at earlier, more treatable stages; however, it can also reduce incidence through the removal of precancerous lesions [4]. As a result, screening is recommended for average-risk adults ≥ 45 years of age and includes a variety of tests [4, 12]. Despite alternate screening options, colonoscopy capacity is often cited as a barrier to …


Deep Learning For Microbiome-Based Integrative Modeling And Microbial Biomarkers Identification, Sen Yang Dec 2023

Deep Learning For Microbiome-Based Integrative Modeling And Microbial Biomarkers Identification, Sen Yang

Statistical Science Theses and Dissertations

The human microbiome, comprising trillions of microorganisms, plays a pivotal role in modulating host physiology via molecular and metabolite exchanges. One of the major challenges in this field lies in the effective integration of microbiome and metabolomics data, an achievement that holds the promise of substantially enhancing the precision of disease prediction. However, many datasets prioritize microbiome data while neglecting paired metabolome information. Additionally, the prevalent analytical tools face challenges in effectively merging these intricate datasets, leading to possible misinterpretations and reduced prediction accuracies.

To address these challenges, the first part of this research introduces the Microbiome-based Supervised Contrastive Learning …


Bayesian Statistical Modeling Of Spatially Resolved Transcriptomics Data, Xi Jiang Oct 2023

Bayesian Statistical Modeling Of Spatially Resolved Transcriptomics Data, Xi Jiang

Statistical Science Theses and Dissertations

Spatially resolved transcriptomics (SRT) quantifies expression levels at different spatial locations, providing a new and powerful tool to investigate novel biological insights. As experimental technologies enhance both in capacity and efficiency, there arises a growing demand for the development of analytical methodologies.

One question in SRT data analysis is to identify genes whose expressions exhibit spatially correlated patterns, called spatially variable (SV) genes. Most current methods to identify SV genes are built upon the geostatistical model with Gaussian process, which could limit the models' ability to identify complex spatial patterns. In order to overcome this challenge and capture more types …


Optimizing Tumor Xenograft Experiments Using Bayesian Linear And Nonlinear Mixed Modelling And Reinforcement Learning, Mary Lena Bleile May 2023

Optimizing Tumor Xenograft Experiments Using Bayesian Linear And Nonlinear Mixed Modelling And Reinforcement Learning, Mary Lena Bleile

Statistical Science Theses and Dissertations

Tumor xenograft experiments are a popular tool of cancer biology research. In a typical such experiment, one implants a set of animals with an aliquot of the human tumor of interest, applies various treatments of interest, and observes the subsequent response. Efficient analysis of the data from these experiments is therefore of utmost importance. This dissertation proposes three methods for optimizing cancer treatment and data analysis in the tumor xenograft context. The first of these is applicable to tumor xenograft experiments in general, and the second two seek to optimize the combination of radiotherapy with immunotherapy in the tumor xenograft …


Regression Modeling Of Complex Survival Data Based On Pseudo-Observations, Rong Rong Dec 2022

Regression Modeling Of Complex Survival Data Based On Pseudo-Observations, Rong Rong

Statistical Science Theses and Dissertations

The restricted mean survival time (RMST) is a clinically meaningful summary measure in studies with survival outcomes. Statistical methods have been developed for regression analysis of RMST to investigate impacts of covariates on RMST, which is a useful alternative to the Cox regression analysis. However, existing methods for regression modeling of RMST are not applicable to left-truncated right-censored data that arise frequently in prevalent cohort studies, for which the sampling bias due to left truncation and informative censoring induced by the prevalent sampling scheme must be properly addressed. Meanwhile, statistical methods have been developed for regression modeling of the cumulative …


Sars-Cov-2 Pandemic Analytical Overview With Machine Learning Predictability, Anthony Tanaydin, Jingchen Liang, Daniel W. Engels Jan 2021

Sars-Cov-2 Pandemic Analytical Overview With Machine Learning Predictability, Anthony Tanaydin, Jingchen Liang, Daniel W. Engels

SMU Data Science Review

Understanding diagnostic tests and examining important features of novel coronavirus (COVID-19) infection are essential steps for controlling the current pandemic of 2020. In this paper, we study the relationship between clinical diagnosis and analytical features of patient blood panels from the US, Mexico, and Brazil. Our analysis confirms that among adults, the risk of severe illness from COVID-19 increases with pre-existing conditions such as diabetes and immunosuppression. Although more than eight months into pandemic, more data have become available to indicate that more young adults were getting infected. In addition, we expand on the definition of COVID-19 test and discuss …


Bayesian Semi-Supervised Keyphrase Extraction And Jackknife Empirical Likelihood For Assessing Heterogeneity In Meta-Analysis, Guanshen Wang Dec 2020

Bayesian Semi-Supervised Keyphrase Extraction And Jackknife Empirical Likelihood For Assessing Heterogeneity In Meta-Analysis, Guanshen Wang

Statistical Science Theses and Dissertations

This dissertation investigates: (1) A Bayesian Semi-supervised Approach to Keyphrase Extraction with Only Positive and Unlabeled Data, (2) Jackknife Empirical Likelihood Confidence Intervals for Assessing Heterogeneity in Meta-analysis of Rare Binary Events.

In the big data era, people are blessed with a huge amount of information. However, the availability of information may also pose great challenges. One big challenge is how to extract useful yet succinct information in an automated fashion. As one of the first few efforts, keyphrase extraction methods summarize an article by identifying a list of keyphrases. Many existing keyphrase extraction methods focus on the unsupervised setting, …


Compressed Dna Representation For Efficient Amr Classification, John Partee, Robert Hazell, Anjli Solsi, John Santerre Aug 2020

Compressed Dna Representation For Efficient Amr Classification, John Partee, Robert Hazell, Anjli Solsi, John Santerre

SMU Data Science Review

In this paper, we explore a representation methodology for the compression of DNA isolates. Using lossless string compression via tokenization of frequently repeated segments of DNA, we reduce the length of the isolates to be counted as k-mers for classification. With this new representation, we apply a previously established feature sampling method to dramatically reduce the feature space. In understanding the genetic diversity, we also look at conserving biological function across these spaces. Using a random forest model we were able to predict the resistance or susceptibility of bacteria with 85-90\% accuracy, with a 30-50\% reduction in overall isolate length, …


Sensitivity Analysis For Incomplete Data And Causal Inference, Heng Chen May 2020

Sensitivity Analysis For Incomplete Data And Causal Inference, Heng Chen

Statistical Science Theses and Dissertations

In this dissertation, we explore sensitivity analyses under three different types of incomplete data problems, including missing outcomes, missing outcomes and missing predictors, potential outcomes in \emph{Rubin causal model (RCM)}. The first sensitivity analysis is conducted for the \emph{missing completely at random (MCAR)} assumption in frequentist inference; the second one is conducted for the \emph{missing at random (MAR)} assumption in likelihood inference; the third one is conducted for one novel assumption, the ``sixth assumption'' proposed for the robustness of instrumental variable estimand in causal inference.


Inference Of Heterogeneity In Meta-Analysis Of Rare Binary Events And Rss-Structured Cluster Randomized Studies, Chiyu Zhang Dec 2019

Inference Of Heterogeneity In Meta-Analysis Of Rare Binary Events And Rss-Structured Cluster Randomized Studies, Chiyu Zhang

Statistical Science Theses and Dissertations

This dissertation contains two topics: (1) A Comparative Study of Statistical Methods for Quantifying and Testing Between-study Heterogeneity in Meta-analysis with Focus on Rare Binary Events; (2) Estimation of Variances in Cluster Randomized Designs Using Ranked Set Sampling.

Meta-analysis, the statistical procedure for combining results from multiple studies, has been widely used in medical research to evaluate intervention efficacy and safety. In many practical situations, the variation of treatment effects among the collected studies, often measured by the heterogeneity parameter, may exist and can greatly affect the inference about effect sizes. Comparative studies have been done for only one or …


Sample Size Calculation Of Clinical Trials With Correlated Outcomes, Dateng Li Aug 2019

Sample Size Calculation Of Clinical Trials With Correlated Outcomes, Dateng Li

Statistical Science Theses and Dissertations

In this thesis, we investigate sample size calculation for three kinds of clinical trials: (1). Randomized controlled trials (RCTs) with longitudinal count outcomes; (2). Cluster randomized trials (CRTs) with count outcomes; (3). CRTs with multiple binary co-primary endpoints.


Robust And Adaptive Design Approaches For Stepped Wedge Cluster Randomized Trials, Jijia Wang Jan 2019

Robust And Adaptive Design Approaches For Stepped Wedge Cluster Randomized Trials, Jijia Wang

Statistical Science Theses and Dissertations

The stepped wedge (SW) cluster randomized design has been increasingly employed by pragmatic trials in health services research. In this study, based on the GEE approach, I present a closed-form sample size that is applicable to both closed-cohort and cross-sectional SW trials with outcomes from the exponential family. On the other hand, I proposed a Bayesian adaptive design for cross-sectional SW cluster randomized trials. It is more adaptable than traditional designs because it allows early termination of the trial when interim data indicate that the intervention is sufficient efficacious or inefficacious. A decision to terminate or continue the trial will …


Association Tests For Genetic Effect And Its Interaction With Environmental Factors, Zhengyang Zhou Jul 2018

Association Tests For Genetic Effect And Its Interaction With Environmental Factors, Zhengyang Zhou

Statistical Science Theses and Dissertations

My research is in the area of statistical genetics, and it contains three projects: (1) Differentiating the Cochran-Armitage (CA) trend test and Pearson’s chi-square test: location and dispersion; (2) Decomposing Pearson’s chi-square test: a linear regression and its departure from linearity; (3) Testing nonlinear gene-environment (GxE) interaction through varying coefficient and linear mixed models.

(1) In genetic case-control association studies, a standard practice is to perform the CA trend test with 1 degree-of-freedom (df) under the assumption of an additive model. However, when the true genetic model is recessive or near recessive, it is outperformed by Pearson’s chi-square test with …


Developing Statistical Methods For Data From Platforms Measuring Gene Expression, Gaoxiang Jia Apr 2018

Developing Statistical Methods For Data From Platforms Measuring Gene Expression, Gaoxiang Jia

Statistical Science Theses and Dissertations

This research contains two topics: (1) PBNPA: a permutation-based non-parametric analysis of CRISPR screen data; (2) RCRnorm: an integrated system of random-coefficient hierarchical regression models for normalizing NanoString nCounter data from FFPE samples.

Clustered regularly-interspaced short palindromic repeats (CRISPR) screens are usually implemented in cultured cells to identify genes with critical functions. Although several methods have been developed or adapted to analyze CRISPR screening data, no single spe- cific algorithm has gained popularity. Thus, rigorous procedures are needed to overcome the shortcomings of existing algorithms. We developed a Permutation-Based Non-Parametric Analysis (PBNPA) algorithm, which computes p-values at the gene level …