Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Biostatistics (108)
- Applied Statistics (39)
- Life Sciences (38)
- Medicine and Health Sciences (33)
- Statistical Models (29)
-
- Statistical Methodology (23)
- Applied Mathematics (19)
- Social and Behavioral Sciences (14)
- Longitudinal Data Analysis and Time Series (13)
- Genetics and Genomics (12)
- Data Science (9)
- Design of Experiments and Sample Surveys (9)
- Multivariate Analysis (9)
- Public Health (9)
- Mathematics (8)
- Bioinformatics (7)
- Microarrays (7)
- Probability (7)
- Psychology (7)
- Computer Sciences (6)
- Genomics (6)
- Other Applied Mathematics (6)
- Categorical Data Analysis (5)
- Ecology and Evolutionary Biology (5)
- Genetics (5)
- Ordinary Differential Equations and Applied Dynamics (5)
- Dynamic Systems (4)
- Engineering (4)
- Keyword
-
- Epidemiology (8)
- Biostatistics (6)
- Neuroscience (6)
- Bayesian (5)
- Medicine (5)
-
- Clinical trial (4)
- Environment (4)
- Genomics (4)
- Bayesian Hierarchical Models (3)
- Gene expression (3)
- Longitudinal data (3)
- Machine learning (3)
- Mixed integer programming (3)
- Other (3)
- Power (3)
- Simulation (3)
- Support vector machine (3)
- Adaptive design (2)
- Algorithm (2)
- Chemicals (2)
- Classification (2)
- Cluster analysis (2)
- Covariance (2)
- DNA Methylation (2)
- Data analysis (2)
- Dimensionality reduction (2)
- Dropout (2)
- Evolution (2)
- Forensic Science (2)
- Genetics (2)
- Publication Year
- Publication
- Publication Type
Articles 1 - 30 of 176
Full-Text Articles in Statistics and Probability
Delineating Differences In Firing Rate Estimates Of Healthy And Parkinsonian Single-Unit Basal Ganglia Recordings, Richard R. Foster, Cheng Ly
Delineating Differences In Firing Rate Estimates Of Healthy And Parkinsonian Single-Unit Basal Ganglia Recordings, Richard R. Foster, Cheng Ly
Biology and Medicine Through Mathematics Conference
No abstract provided.
Mitigating Parameter Identifiability Issues Through Model Calibration On The Data-Informed Active Subspace: An Example In Tumor Growth, Allison L. Lewis, Rebecca A. Everett
Mitigating Parameter Identifiability Issues Through Model Calibration On The Data-Informed Active Subspace: An Example In Tumor Growth, Allison L. Lewis, Rebecca A. Everett
Biology and Medicine Through Mathematics Conference
No abstract provided.
Propensity Score Matching Accounting For Longitudinal Trends Before Baseline With Group-Based Trajectory Modeling, Dustin R. Bastaich
Propensity Score Matching Accounting For Longitudinal Trends Before Baseline With Group-Based Trajectory Modeling, Dustin R. Bastaich
Theses and Dissertations
Propensity score matching is used in observational studies to balance baseline attributes between a treatment of interest and a control group. Propensity score matching typically relies on baseline variables, but longitudinal trends in patient characteristics can also influence treatment decisions and subsequent health outcomes. This dissertation extends standard approaches by explicitly incorporating longitudinal trajectories of key variables into the propensity score estimation process.
Trends in a longitudinal variable prior to baseline were characterized using group-based trajectory modeling. A two-step modeling approach was implemented where trajectory groups of a key variable were first estimated and then included as covariates in the …
Identifying Mobility Hub Suitability In The Mid-Sized United States Urban Area Using Weighted Overlay And Percentile Based Local Peak Analysis, Eric Asplund
Theses and Dissertations
IDENTIFYING MOBILITY HUB SUITABILITY IN THE MID-SIZED UNITED STATES URBAN AREA USING WEIGHTED OVERLAY AND PERCENTILE‑BASED LOCAL PEAK ANALYSIS
By Eric Asplund
A thesis submitted in partial fulfillment of the requirements for the degree of Master of Urban and Regional Planning at Virginia Commonwealth University.
Virginia Commonwealth University, 2026.
Major Director: Dr. Ivan Suen, Ph. D., Associate Professor, Faculty of Urban and Regional Studies and Planning
Shared mobility hubs are increasingly posited in transportation planning as interventions supporting multimodality, sustainability, and equitable access, but guidance on evaluation of potential hub locations remains uneven, particularly in mid-sized United States cities where …
Machine Learning-Based Spatio-Temporal Modeling Of Climate Dynamics And Desertification In The Sahara–Sahel Region, Stephen M. Tivenan
Machine Learning-Based Spatio-Temporal Modeling Of Climate Dynamics And Desertification In The Sahara–Sahel Region, Stephen M. Tivenan
Theses and Dissertations
Arid climate classifications are threshold-dependent and easily interpretable mappings that are widely used in ecological, agricultural, and climate-related studies. These classifications inform scientific understanding, support policy and land management decisions, and provide an intuitive summary of environmental conditions. Despite their usefulness, traditional arid climate classifications often fail to quantify uncertainty, incorporate spatial context, or account for complex relationships among relevant environmental variables. Existing approaches to uncertainty assessment have largely relied on comparing classifications across multiple datasets or alternative formulas, but these methods generally overlook important spatial dependence and latent structure in the data.
This dissertation develops three machine learning-based statistical …
Dimension Reduction Involving Exogenous Variables With Applications In Manufacturing And Healthcare, Linxi Li
Dimension Reduction Involving Exogenous Variables With Applications In Manufacturing And Healthcare, Linxi Li
Theses and Dissertations
High-dimensional data analysis presents diverse challenges, including the curse of dimensionality, the complexities of working with datasets that combine large feature spaces with limited sample sizes, and difficulties in identifying meaningful relationships among variables. As datasets grow in size and complexity across different fields, it is increasingly important to develop practical approaches for extracting essential information from such data. Dimension reduction methods address these challenges by alleviating the effects of high dimensionality, enhancing the ability to reveal hidden patterns, and uncovering latent structures within the data to support further analysis. Some methods reduce dimensionality while preserving all relevant information, offering …
An Integrated Data-Driven Framework For Arctic Shipping: Analyzing Vessel Speed, Environmental And Ecological Factors Through Innovative Statistical Spatio-Temporal Methods, Inverse Optimization And Machine Learning, Mauli Pant
Theses and Dissertations
This dissertation develops an integrated data-driven framework to analyze vessel navigation and ecological risk in the United States Arctic from 2010 to 2019. As environmental change and maritime activity increase in the region, understanding how vessels respond to dynamic conditions and how those responses interact with marine ecosystems has become increasingly important. A central theme of this dissertation is the treatment of vessel speed as both an observed outcome and a decision variable reflecting trade- offs among operational, environmental, and ecological factors. The first chapter develops a predictive framework for vessel speed over ground (SOG) using Gaussian Process Boosting (GPBoost), …
Beyond Homogeneity: Exploring Causal Heterogeneity In Psychopathology, Philip B. Vinh
Beyond Homogeneity: Exploring Causal Heterogeneity In Psychopathology, Philip B. Vinh
Theses and Dissertations
Traditional models in psychiatric research often impose assumptions of causal homogeneity, treating population-level associations as reflective of uniform underlying mechanisms. This dissertation challenges that assumption by introducing statistical and machine learning frameworks designed to detect and model causal heterogeneity in the development of psychopathology. Central to this approach is the advancement of finite mixture structural equation modeling (FM-SEM) to identify latent subgroups characterized by distinct, and sometimes opposing, causal pathways.
The dissertation comprises three integrated empirical studies. The first introduces mixDoC, a finite mixture extension of the classical Direction of Causation (DoC) model applied to twin data, enabling the detection …
Predictive Inference For Ion Concentration With Machine Learning And Bayesian Methods, Alexandra B. Ulbing
Predictive Inference For Ion Concentration With Machine Learning And Bayesian Methods, Alexandra B. Ulbing
Theses and Dissertations
Ultraviolet--visible (UV--Vis) spectroscopy produces high-dimensional signals that are strongly collinear, shift with concentration, and exhibit heteroskedastic, non-Gaussian noise. These features make supervised regression from spectra to ionic concentrations statistically challenging and limit the reliability of methods that assume linear structure or homoscedastic errors.
This dissertation develops two complementary frameworks for prediction and uncertainty quantification in UV--Vis spectroscopic regression: (1) frequentist stacked ensembles combined with distribution-free conformal prediction, and (2) Bayesian hierarchical modeling and Bayesian stacking. Together, they provide a unified view of model-based and distribution-free uncertainty across nickel and nickel--cobalt datasets.
The frequentist component builds ensembles of Functional Data Analysis …
Sealing The Deal: A Case Study Of A Private, Southeastern, Regional College’S Student Onboarding Practices, Alicia R. Gaston, Kristen Meyer, Tyler Ogden, Jennifer Perkinson, Casey Yocum Michaels
Sealing The Deal: A Case Study Of A Private, Southeastern, Regional College’S Student Onboarding Practices, Alicia R. Gaston, Kristen Meyer, Tyler Ogden, Jennifer Perkinson, Casey Yocum Michaels
Doctor of Education Capstones
Each year, higher education institutions offer support to new, incoming students through an onboarding process that requires dedication and collaboration across multiple departments. Offering streamlined guidance and clear communication throughout the onboarding process is essential to ensure incoming students understand action steps without becoming overwhelmed with new terminology, processes, and environments. This explanatory case study is set to understand the current communication practices and technology use across onboarding departments at Brightside College – also referred to as Brightside or BC (pseudonym). The research team aimed to understand Brightside's onboarding staff's perspective of current processes and practices. With emphasis on the …
Optimal Data Splitting Methods, Sujay Mudalgi
Optimal Data Splitting Methods, Sujay Mudalgi
Theses and Dissertations
In predictive modeling, effective data splitting is crucial for creating statistically representative training and validation sets. The state-of-the-art data splitting methods are based on minimizing the energy distance between the split subsets. However, there are a number of limitations in the existing methods, which this dissertation aims to address. First, the existing methods were computationally inefficient. Thus, Chapter 2 proposes a method to scale up these approaches for big data. Here, we introduce scalable Twinning (s-Twinning), which significantly improves the execution speed of data splitting without sacrificing accuracy. Second, the existing methods did not consider the predictive relationship in the …
Machine Learning Models Leveraging Patient-Similarity And Clinical Temporality For Disease Prognoses, Ahmad F. Al Musawi
Machine Learning Models Leveraging Patient-Similarity And Clinical Temporality For Disease Prognoses, Ahmad F. Al Musawi
Theses and Dissertations
Electronic Health Records (EHRs) constitute a comprehensive and high-dimensional repository of clinical data, encompassing a wide array of patient-level information such as diagnoses, procedures, medications, laboratory results, and unstructured clinical narratives. These data hold immense potential for advancing predictive modeling in healthcare, including tasks such as disease progression modeling, hospital readmission prediction, and length of stay (LoS) estimation. However, the intrinsic complexity of EHR data—manifested in its heterogeneity, sparsity, and temporal dynamics—poses significant analytical challenges that limit the generalizability and interpretability of conventional machine learning models. Recent methodological advancements in deep learning and graph-based learning, particularly Graph Neural Networks (GNNs), …
An Adaptive Method For Covariate Balancing In Block Randomized Clinical Trials, Ren Rasnick
An Adaptive Method For Covariate Balancing In Block Randomized Clinical Trials, Ren Rasnick
Theses and Dissertations
Clinical trials are randomized in part to limit allocation bias, but also to ensure comparability between treatment arms for a baseline variable of concern. Comparable with regard to a baseline variable of concern is necessary for the validity of statistical methods. However, comparability is not guaranteed for trials of any size and is even more likely in trials with < 200 total participants. We propose a new method for adapting the allocation of participants in a sequentially allocated two-armed study with a small sample size to better ensure comparability.
The proposed method calculates the expected final imbalance (lack of comparability) based on the current participant values. Unlike several other methods, our method ensures the final desired sample size for each treatment arm, utilizes the expected final imbalance, increases comparability between …
Time Scale Separation In Life-Long Ovarian Follicles Population Dynamics Model, Romain Yvinec, Frédérique Clément, Guillaume Ballif
Time Scale Separation In Life-Long Ovarian Follicles Population Dynamics Model, Romain Yvinec, Frédérique Clément, Guillaume Ballif
Biology and Medicine Through Mathematics Conference
No abstract provided.
Multi-Type Branching Processes In Time-Varying Environments, Arash Jamshidpey
Multi-Type Branching Processes In Time-Varying Environments, Arash Jamshidpey
Biology and Medicine Through Mathematics Conference
No abstract provided.
Modeling Human Temporal Eeg Responses To Vr Visual Stimuli, Richard R. Foster, Connor Delaney, Dean J. Krusienski, Cheng Ly
Modeling Human Temporal Eeg Responses To Vr Visual Stimuli, Richard R. Foster, Connor Delaney, Dean J. Krusienski, Cheng Ly
Biology and Medicine Through Mathematics Conference
No abstract provided.
Developing Machine Learning And Time-Series Analysis Methods With Applications In Diverse Fields, Muhammed Aljifri
Developing Machine Learning And Time-Series Analysis Methods With Applications In Diverse Fields, Muhammed Aljifri
Theses and Dissertations
This dissertation introduces methodologies that combine machine learning models with time-series analysis to tackle data analysis challenges in varied fields. The first study enhances the traditional cumulative sum control charts with machine learning models to leverage their predictive power for better detection of process shifts, applying this advanced control chart to monitor hospital readmission rates. The second project develops multi-layer models for predicting chemical concentrations from ultraviolet-visible spectroscopy data, specifically addressing the challenge of analyzing chemicals with a wide range of concentrations. The third study presents a new method for detecting multiple changepoints in autocorrelated ordinal time series, using the …
Bayesian Estimation Of Hierarchical Linear Models From Incomplete Data: Cluster-Level Non-Linear Effects And Small Sample Sizes, Dongho Shin
Theses and Dissertations
We consider Bayesian estimation of a hierarchical linear model (HLM) from small sample sizes. The continuous response Y and covariates C are partially observed and assumed missing at random. With C having linear effects, the HLM may be efficiently estimated by available methods. When C includes cluster-level covariates having interactive or other nonlinear effects given small sample sizes, however, maximum likelihood estimation is suboptimal, and existing Gibbs samplers are based on a Bayesian joint distribution compatible with the HLM, but impute missing values of C by a Metropolis algorithm via a proposal density having a constant variance while the target …
Title: I: L1-Norm Matrix Completion For Recommender Systems Ii: Conjecturing-Based Classification, Fatemeh Valizadeh Gamchi
Title: I: L1-Norm Matrix Completion For Recommender Systems Ii: Conjecturing-Based Classification, Fatemeh Valizadeh Gamchi
Theses and Dissertations
Recommendation systems are essential for providing personalized user experiences, but their performance can be affected by outliers especially in traditional collaborative filtering methods that use the L2-norm. To address this challenge, we developed two new algorithms, SharpEl1rs and SharpEl1rs-Impute, based on the L1-norm to improve resistance against extreme values and effectively handle missing data. Our experimental setting was designed to compare these proposed methods with existing techniques. Then our algorithms are applied to real datasets to assess their performance, with findings indicating that our proposed models offer improved accuracy in some cases and solid performance in others for industrial-scale recommendation …
Redesign, Evaluation, And Validation Of A Commercially Viable High-Resolution Melt Based Mixture Screening Tool, Chastyn Smith
Redesign, Evaluation, And Validation Of A Commercially Viable High-Resolution Melt Based Mixture Screening Tool, Chastyn Smith
Theses and Dissertations
Analysis of evidentiary samples containing DNA from multiple contributors (“mixtures”) is a time intensive process for a forensic analyst and one where the contributor nature of a sample is not revealed until the end of the traditional forensic workflow. Often, at this stage, retesting or additional testing of mixture samples may not be possible, particularly if the DNA collection device did not preserve the DNA well enough; consequently leaving only trace amounts of a contributor’s DNA present. Thus, a new collection device that would allow for the increased preservation/integrity of evidentiary samples as well as a method that would allow …
The Genetic Architecture Of Cervical Change During Pregnancy: From Modeling To Mechanism — Does The Cervix Mediate Maternal Risk For Spontaneous Preterm Birth?, Hope M. Wolf
Theses and Dissertations
This project leverages clinical data and biospecimens from a prospective longitudinal cohort of pregnant women to study the genetic and phenotypic relationships between cervical shortening and the duration of pregnancy. Sonographic cervical length (CL) was measured throughout pregnancy in a cohort of 5,160 Black/African American women in Detroit, Michigan. Maternal DNA samples were sequenced with a next-generation low-pass whole genome platform. The heritability of cervical change during pregnancy and its genetic correlation with gestational age at delivery (GAD) were estimated using Genome-Wide Complex Trait Analysis. These estimates suggest that cervical change is heritable (h²CL = 51%) and highly polygenic trait. …
Analytical Approach For Monitoring The Behavior Of Patients With Pancreatic Adenocarcinoma At Different Stages As A Function Of Time, Aditya Chakaborty Dr, Chris P. Tsokos Dr
Analytical Approach For Monitoring The Behavior Of Patients With Pancreatic Adenocarcinoma At Different Stages As A Function Of Time, Aditya Chakaborty Dr, Chris P. Tsokos Dr
Biology and Medicine Through Mathematics Conference
No abstract provided.
A Novel Family Of Chain Binomial Models To Investigate Correlated Vaccination And Infection Rates In Sveirs Epidemic Dynamics, Divine Wanduku
A Novel Family Of Chain Binomial Models To Investigate Correlated Vaccination And Infection Rates In Sveirs Epidemic Dynamics, Divine Wanduku
Biology and Medicine Through Mathematics Conference
No abstract provided.
Predicting Dengue Incidence In Central Argentina Using Google Trends Data, Sahil Chindal, Elizabet Estallo, Yanjun Qian, Michael Robert
Predicting Dengue Incidence In Central Argentina Using Google Trends Data, Sahil Chindal, Elizabet Estallo, Yanjun Qian, Michael Robert
Biology and Medicine Through Mathematics Conference
No abstract provided.
Novel Feature Evaluation In Ultra-High Dimensional Right-Censored Data, With Applications To Head And Neck Cancer, Atika Farzana Urmi
Novel Feature Evaluation In Ultra-High Dimensional Right-Censored Data, With Applications To Head And Neck Cancer, Atika Farzana Urmi
Graduate Research Posters
Background: Head and neck cancer is the 6th most common cancer worldwide with an expected 1.08 million new cases each year. Such cancer data are ultra-high dimensional with thousands of clinical features and gene expressions, making it challenging for the traditional analytical tools to extract the potential biomarker for the cancer survival and control false discoveries. In addition, presence of heavy censoring can affect the screening procedures based on Kaplan-Meier (K-M) survival estimates.
Aim: To propose a model free, ultra-high dimensional feature screening method with two-dimensional survival outcome allowing false discovery rate (FDR) control.
Method: 516 primary tumor patients with …
Variability In Causal Effects On A Binary Outcome And Noncompliance In A Multisite Randomized Trial, Xinxin Sun
Variability In Causal Effects On A Binary Outcome And Noncompliance In A Multisite Randomized Trial, Xinxin Sun
Theses and Dissertations
Noncompliance to treatment assignment is widespread in randomized trials and presents challenges in causal inference. In the presence of noncompliance, the most commonly estimated effect of treatment assignment, also known as intent-to-treat (ITT) effect, is biased. Of interest in this setting is the complier average causal effect (CACE), the ITT effect among compliers. Further complication arises when the outcome variable is partially observed.
My research focuses on estimating the distribution of a site-specific CACE in a multisite randomized controlled trial (MRCT) by maximum likelihood (ML). Assuming compliance missing at random (MAR). We express the likelihood as an integral with respect …
Reassessing Replication: Addressing The Replication Crisis From A Statistical Perspective, Alicia Richards Phd
Reassessing Replication: Addressing The Replication Crisis From A Statistical Perspective, Alicia Richards Phd
Theses and Dissertations
In 2015, Open Science Framework directly replicated 100 psychology studies and found astonishingly low replication rates. Since, researchers have suggested factors that may have influenced the low rates, including the metrics used to assess replications. The definitions used to decide whether a replication study was successful all suffer from flaws. Therefore, we propose a new metric for assessing replication that can estimate the likelihood a study successfully replicated rather than forcing a binary choice and accounts for study design limitations.
Using equivalence study techniques, we first propose a new metric to assess replication, defining a successful replication as one where …
Early Termination In Phase Ii Clinical Trials: Admissible Designs Using Decreasingly Informative Priors, Chen Wang
Theses and Dissertations
In Phase II clinical trials, Thall and Simon’s Bayesian posterior probability design is commonly implemented to allow for an early termination to determine whether a new treatment warrants further investigation in a larger-scale Phase III trial; this in turn requires a pre-selected prior distribution based on known clinical opinion or historical information. Moreover, this Bayesian approach can result in an issue of inflating type I error rate by monitoring interim data to inform early termination decisions. Alternatively, a Bayesian approach with the decreasingly informative prior (DIP), which is an informative yet skeptical prior, can be implemented to overcome the contentious …
Model-Based Imputation Of Below Detection Limit Missing Data And Group Selection In Bayesian Group Index Regression, Matthew Carli
Model-Based Imputation Of Below Detection Limit Missing Data And Group Selection In Bayesian Group Index Regression, Matthew Carli
Theses and Dissertations
Investigations into the association between chemical exposure and health outcomes are increasingly focused on the role of chemical mixtures, as opposed to individual chemicals. The analysis of chemical mixture data required the development of novel statistical methods, one of these being Bayesian group index regression. A statistical challenge common to all chemical mixture analyses is the ubiquitous presence of below detection limit (BDL) data. We propose an extension of Bayesian group index regression that treats both regression effects and missing BDL observations as parameters in a model estimated through a Markov Chain Monte Carlo algorithm that we refer to as …
Integrative Post-Gwas Analyses Of Psychiatric Disorders: Identifying Putative Risk Genes And Gene Sets Using Transcriptome, Proteome And Methylome Information, Huseyin Gedik
Theses and Dissertations
Genome-wide association studies (GWAS) of psychiatric disorders (PD) yield numerous loci with significant signals, but often they do not implicate specific protein coding genes. Because GWAS risk loci are enriched in expression/protein/methylation quantitative loci (e/p/mQTL, hereafter xQTL), transcriptome/proteome/methylome-wide association studies (T/P/MWAS, hereafter XWAS), which integrate information from GWAS and x-level (mRNA, protein or DNA methylation levels) coming from largest xQTL studies, can link GWAS signals to effects on specific genes. For gene level analyses, researchers use mendelian randomization (MR) methods to fine-map the association between x-levels and trait. However, none of the previous studies ever jointly analyzed XWAS of multiple …