Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Medicine and Health Sciences (197)
- Life Sciences (156)
- Public Health (118)
- Applied Statistics (105)
- Statistical Models (103)
-
- Statistical Methodology (77)
- Epidemiology (74)
- Bioinformatics (60)
- Computer Sciences (44)
- Genetics and Genomics (38)
- Longitudinal Data Analysis and Time Series (37)
- Clinical Trials (35)
- Multivariate Analysis (35)
- Data Science (33)
- Survival Analysis (32)
- Social and Behavioral Sciences (27)
- Applied Mathematics (25)
- Genetics (24)
- Diseases (23)
- Medical Specialties (23)
- Mathematics (22)
- Environmental Sciences (21)
- Categorical Data Analysis (20)
- Ecology and Evolutionary Biology (19)
- Statistical Theory (19)
- Biochemistry, Biophysics, and Structural Biology (18)
- Engineering (16)
- Institution
-
- Virginia Commonwealth University (96)
- University of Kentucky (52)
- University of South Carolina (52)
- University of Louisville (45)
- The Texas Medical Center Library (36)
-
- University at Albany, State University of New York (36)
- Loma Linda University (34)
- University of South Florida (32)
- University of Nevada, Las Vegas (28)
- New Jersey Institute of Technology (23)
- Southern Methodist University (18)
- Old Dominion University (16)
- Walden University (14)
- University of Nebraska Medical Center (13)
- University of Arkansas, Fayetteville (12)
- Michigan Technological University (11)
- University of Texas at El Paso (10)
- California Polytechnic State University, San Luis Obispo (8)
- East Tennessee State University (8)
- Georgia Southern University (7)
- University of Alabama at Birmingham (7)
- Dartmouth College (6)
- University of New Mexico (6)
- Clemson University (5)
- Illinois State University (5)
- James Madison University (5)
- Purdue University (5)
- University of Arkansas Little Rock (5)
- University of North Florida (5)
- Missouri University of Science and Technology (4)
- Keyword
-
- Statistics (28)
- Biostatistics (26)
- Machine learning (18)
- Bioinformatics (15)
- Bayesian (13)
-
- Survival analysis (12)
- Causal inference (11)
- Survival Analysis (10)
- Longitudinal data (8)
- Missing data (8)
- Bayesian inference (7)
- Gene expression (7)
- HIV (7)
- Simulation (7)
- Survival (7)
- Variable selection (7)
- Biological sciences (6)
- Clinical trial (6)
- Diabetes (6)
- Epidemiology (6)
- GWAS (6)
- Machine Learning (6)
- Maximum likelihood (6)
- Meta-analysis (6)
- Physical Sciences and Mathematics, Statistics and Probability, Biostatistics (6)
- Propensity score (6)
- Breast cancer (5)
- Cancer (5)
- Classification (5)
- Cluster analysis (5)
- Publication Year
- Publication
-
- Theses and Dissertations (162)
- Electronic Theses and Dissertations (61)
- Dissertations and Theses (Open Access) (36)
- Loma Linda University Electronic Theses, Dissertations & Projects (34)
- USF Tampa Graduate Theses and Dissertations (32)
-
- Legacy Theses & Dissertations (2009 - 2024) (28)
- UNLV Theses, Dissertations, Professional Papers, and Capstones (28)
- Theses and Dissertations--Epidemiology and Biostatistics (24)
- Theses (22)
- Theses and Dissertations--Statistics (20)
- Statistical Science Theses and Dissertations (18)
- Walden Dissertations and Doctoral Studies (14)
- Graduate Theses and Dissertations (12)
- Dissertations, Master's Theses and Master's Reports (11)
- Open Access Theses & Dissertations (10)
- Electronic Theses & Dissertations (2024 - present) (8)
- Mathematics & Statistics Theses & Dissertations (8)
- Theses & Dissertations (8)
- College of Graduate Studies: Theses & Dissertations (7)
- Dartmouth College Ph.D Dissertations (6)
- Mathematics & Statistics ETDs (6)
- Statistics (6)
- All Dissertations (5)
- All ETDs from UAB (5)
- Capstone Experience: Master of Public Health (5)
- Master's Theses (5)
- UNF Graduate Theses and Dissertations (5)
- Dissertations (4)
- Doctoral Dissertations (4)
- Open Access Dissertations (4)
Articles 1 - 30 of 679
Full-Text Articles in Biostatistics
Robust Statistical Methods For Microbiome Abundance Data, Yiming Shi
Robust Statistical Methods For Microbiome Abundance Data, Yiming Shi
WUSM Theses and Dissertations – All Programs
Differential abundance analysis in microbiome studies aims to identify taxa whose abundance differs across biological or clinical conditions. The observed data are typically taxon-specific sequencing read counts, representing reads assigned to different taxa within each sample. These counts are indirect measurements of the underlying microbial abundance profile and are constrained by sample-specific library sizes. Microbiome count data are also typically sparse, overdispersed, and heteroscedastic. Together, these characteristics create substantial challenges for differential abundance analysis and make the results highly sensitive to normalization procedures, model specification, and the statistical methods used for inference.
Normalization defines the scale on which samples are …
Statistical Methodologies For Count Time Series Analysis And Topological Data Analysis Of Medical Images, Yuhyeong Jang
Statistical Methodologies For Count Time Series Analysis And Topological Data Analysis Of Medical Images, Yuhyeong Jang
Statistical Science Theses and Dissertations
This dissertation addresses two distinct topics related to count time series analysis and topological medical image analysis, respectively. The first part of the dissertation comprises an application of a count time series model to analysis of US monthly sex trafficking data and development of a new model for multivariate count data that exhibits serial dependence and overdispersion. By imposing a family of multivariate mixed Poisson distributions on the count random vector, the proposed model can accommodate a broad range of overdispersion as well as positive contemporaneous correlations. For maximum likelihood estimation, a computationally feasible EM-type algorithm is derived based on …
Statistical Methods For Mendelian Randomization Under Nonlinearity And Non-Normality, Mary Appah
Statistical Methods For Mendelian Randomization Under Nonlinearity And Non-Normality, Mary Appah
ETDs from 2020-2029
Instrumental variable (IV) methods are widely used for estimating causal effects in observational studies where unmeasured confounding may bias traditional regression estimates. The core idea is to use a variable referred to as an instrument that is associated with the exposure of interest, independent of unmeasured confounders, and influences the outcome only through the exposure. While originally developed in econometrics, IV methods have been increasingly adopted in epidemiology and genetic research under the framework of Mendelian Randomization (MR), where genetic variants, most commonly single nucleotide polymorphisms (SNPs), serve as instruments. MR provides a powerful tool for investigating causal relationships between …
Developing A Humpback Whale Vocalization Detector Using Machine Learning Models, Lucas Kantorowski
Developing A Humpback Whale Vocalization Detector Using Machine Learning Models, Lucas Kantorowski
Master's Theses
Humpback whale songs are notoriously complex. Identification of humpback whale song units requires bioacousticians to tediously listen, analyze, and annotate collected sound data. Even sparse data requires listening to the entirety of the collected acoustic data. In this study, three hours of audio containing over one-thousand humpback whale song units was collected in Monterey Bay, California.
Prior studies have seen success using convolutional neural networks by performing image classification on hundreds of hours worth of spectrograms. Our study uses traditional machine learning models, as they are less computationally demanding, and require less data.
We use time splitting and Mel-frequency cepstrum …
Effect Of The Sava Syndemic On Hiv Viral Suppression Among People Living With Hiv In The United States: A Scoping Review And Meta-Analysis, Jacquelyn Rodriguez
Effect Of The Sava Syndemic On Hiv Viral Suppression Among People Living With Hiv In The United States: A Scoping Review And Meta-Analysis, Jacquelyn Rodriguez
UNLV Theses, Dissertations, Professional Papers, and Capstones
The SAVA syndemic highlights the interconnected and mutually reinforcing nature of substance use, violence victimization, and HIV/AIDS. The synergistic interaction of these conditions creates structural and behavioral barriers that disrupt the HIV care continuum and contribute to significant health inequities. Achieving a suppressed viral load is a critical clinical outcome for people living with HIV, while viral non-suppression can be an indicator of poor health, elevated transmission risk, and disease progression. Still, the association between SAVA factors and viral suppression/non-suppression outcomes within U.S. populations remains inconsistently characterized due to heterogeneous methodologies and variable inclusion of high-risk groups. This mixed-methods study …
The Legacy Of Lead: Lead Exposure's Harmful Effects And Its Concentration In Poc Communities, Layla Sophronia Barber
The Legacy Of Lead: Lead Exposure's Harmful Effects And Its Concentration In Poc Communities, Layla Sophronia Barber
Student Theses 2015-Present
This thesis examines the disproportionate burden of lead exposure carried by low income, POC communities. The systemic nature of this problem is a symptom of a longstanding legacy of environmental injustice in the United States. Decades of federal neglect are reflected in the higher statistics of lead exposure and poisoning in predominantly black communities. While it is understood that lead exposure poses a serious threat to physical health and early cognitive development, there is a discouraging lack of urgency to remove the toxin from non-wealthy communities. The material covered by this thesis aims to identify and correct the discriminatory social …
Deep Learning Frameworks For Biological Data Integration And Generation, Alexa Beachum
Deep Learning Frameworks For Biological Data Integration And Generation, Alexa Beachum
Statistical Science Theses and Dissertations
Data integration represents a key area of research for analyzing the rapidly growing volume of high-dimensional biological data across sources, stages, and modalities. To model and understand these complex, often non-linear relationships, deep learning has become an increasingly powerful tool. Here, we present two novel deep learning frameworks that address distinct but complementary integration challenges. The first framework aligns single-cell omics data across temporal stages, and the second bridges imaging and omics modalities to generate patient-level molecular profiles.
In Chapter 1, we briefly summarize existing approaches---both statistical and deep learning-based---for single-cell omics data integration and discuss their limitations for handling …
Database-Driven Revelations In Epilepsy: Patient Reporting Accuracy And The Circadian Timing Of Seizures, Ithay Biton
Database-Driven Revelations In Epilepsy: Patient Reporting Accuracy And The Circadian Timing Of Seizures, Ithay Biton
Theses and Dissertations
For over thirty years, patients have been visiting the Arkansas Epilepsy Program for diagnosis and treatment for seizures and seizure-like episodes. As part of their clinical evaluation, patients often undergo ambulatory EEG (electroencephalogram) monitoring. This routine process produces valuable data for treating the patient. In this study, over 2000 ambulatory EEG reports from 1998 to 2016 were reviewed. A large database of seizures was created from the reports, with information on 407 patients, 1611 EEG-confirmed seizures, and 1726 patient-reported seizures (IRB Protocol #17-093). The database was used to address two important issues in epilepsy. The first topic of the study …
Handling Missing Data In Copd Research, Chia-Ying Chiu
Handling Missing Data In Copd Research, Chia-Ying Chiu
ETDs from 2020-2029
Chronic Obstructive Pulmonary Disease (COPD) remains a major global health concern and one of the leading causes of death in the United States, affecting approximately 4.6% of adults and reaching a prevalence of 9.4% in Alabama according to 2024 National Health Interview Survey. The disease imposes a substantial burden on quality of life and healthcare costs and contributes to increased disability-adjusted life years and years of life lost, as highlighted by the 2022 Lancet Commission report. Despite its im-pact, COPD is frequently diagnosed only after irreversible lung damage has occurred, largely due to under-recognized symptoms and limited diagnostic tools. Early …
Type Ii Diabetes Treatment Comparison Via Compartment Modeling, Abigail M. Collins
Type Ii Diabetes Treatment Comparison Via Compartment Modeling, Abigail M. Collins
Theses and Dissertations
Type II diabetes mellitus affects one in ten adults worldwide, yet the effects of treatment type and adherence level on developing complications and quality of life have not been well characterized at the population level, and mathematical modeling offers a structured way to examine these dynamics. This thesis adapts the Boutayeb et al. (2004) model to incorporate dynamic treatment types and levels of adherence, producing nine scenarios in which complication development rate and complication recovery rate differed, to compare peak complications and quality of life across treatment and adherence conditions. Using a system of ordinary differential equations and compartment modeling, …
Using Camera-Based Unmarked Spatial Capture-Recapture Modeling To Estimate Reintroduced Elk (Cervus Canadensis) Population Parameters And Distribution In Southeastern Kentucky, Claire Marie Muia
Theses and Dissertations--Forestry and Natural Resources
Estimation of population parameters is important for wildlife management decisions. Elk reintroduced to southeastern Kentucky experienced early irruptive population growth and are currently monitored using a statewide harvest-based statistical population reconstruction model (SPR) across the Kentucky Elk Restoration Zone (KERZ). Because the SPR model is spatially coarse and difficult to scale to the smaller management units comprising the KERZ, we conducted a spatially explicit capture-recapture study using a clustered camera-trapping array deployed for 10 weeks from June–August 2024 to estimate elk population parameters within Management Unit 4. Due to a lack of resights of GPS-marked elk, population parameters were estimated …
Beyond The Lace Index: Benchmarking Machine Learning Architectures And Explaining 30-Day Hospital Readmission Risk With Shap Analysis, Carl E. Hughes Iii
Beyond The Lace Index: Benchmarking Machine Learning Architectures And Explaining 30-Day Hospital Readmission Risk With Shap Analysis, Carl E. Hughes Iii
Williams Honors College, Honors Research Projects
Unplanned 30-day hospital readmission remains a fundamental challenge in US healthcare, associated with increased risk to patient recovery and representing an estimated $52.4 billion in annual expenses (Beauvais et al., 2022). While the rigorously validated LACE index serves as the clinical standard for readmission modeling, its linear structure and four explanatory variables lack the complexity to capture the high-dimensional and interactive nature of patient risk. This study utilizes an admission granularity level cohort of the MIMIC-IV database to develop and compare machine learning architectures against the baseline LACE index. Due to the imbalanced prevalence of readmission, the penalized logistic regression, …
Propensity Score Matching Accounting For Longitudinal Trends Before Baseline With Group-Based Trajectory Modeling, Dustin R. Bastaich
Propensity Score Matching Accounting For Longitudinal Trends Before Baseline With Group-Based Trajectory Modeling, Dustin R. Bastaich
Theses and Dissertations
Propensity score matching is used in observational studies to balance baseline attributes between a treatment of interest and a control group. Propensity score matching typically relies on baseline variables, but longitudinal trends in patient characteristics can also influence treatment decisions and subsequent health outcomes. This dissertation extends standard approaches by explicitly incorporating longitudinal trajectories of key variables into the propensity score estimation process.
Trends in a longitudinal variable prior to baseline were characterized using group-based trajectory modeling. A two-step modeling approach was implemented where trajectory groups of a key variable were first estimated and then included as covariates in the …
Advancing Periodic Time Series Analysis: Application, Bias Assessment, And Optimal Window Selection In The Variable Bandpass Periodic Block Bootstrap Method, Yanan Sun
Electronic Theses & Dissertations (2024 - present)
Time series analysis is essential for understanding long-term patterns, periodic behavior, and underlying correlations in complex datasets. The periodically correlated (PC) time series is a type of time series where the correlation structure repeats over fixed intervals. The Variable Bandpass Periodic Block Bootstrap (VBPBB) has recently been proposed as a resampling method that preserves PC structures through the use of periodogram, bandpass filters, and block bootstrap resampling. Although promising, VBPBB remains underutilized, and its limitations have not been fully examined. This dissertation advances both the application and methodological development of the VBPBB.
The first project applies the VBPBB to a …
Bayesian Modelling On Periodically And Multiple Periodically Correlated Time Series Data, Jie Yao
Bayesian Modelling On Periodically And Multiple Periodically Correlated Time Series Data, Jie Yao
Electronic Theses & Dissertations (2024 - present)
Time series with multiple periodically correlated (MPC) components present a complex challenge, with relatively limited prior research. Most existing models are designed for simpler periodically correlated (PC) components and often struggle with over-parameterization, optimization issues, and capturing complex PC patterns within a time series. Frequency separation techniques can help preserve the correlation structure of individual PC components, while Bayesian methods can integrate new and prior information to refine beliefs about these components. This study proposes a two-stage approach that combines frequency separation and Bayesian techniques to forecast PC and MPC time series data. This method aims to demonstrate improved effectiveness …
Two-Stage Response-Adaptive Randomization Designs For Multi-Arm Trials With Normal Outcome, Tanjin Tamanna Happy
Two-Stage Response-Adaptive Randomization Designs For Multi-Arm Trials With Normal Outcome, Tanjin Tamanna Happy
UNF Graduate Theses and Dissertations
This study focuses on improving how clinical trials compare new treatments with a standard treatment, when the response is quantitative (normally distributed). In a common two-stage design, several new treatments are first evaluated, and the best-performing one is selected if it appears better than the standard. In the second stage, this selected treatment is compared again with the standard using additional data to confirm its effectiveness. This approach is known to be efficient in terms of accuracy and sample size savings. We extend this design by introducing an adaptive method for assigning patients to treatments in the second stage. Instead …
The Maxima Method For Identification Of Principal Periodic Components In Time Series Analysis, Megan Di Maio
The Maxima Method For Identification Of Principal Periodic Components In Time Series Analysis, Megan Di Maio
Electronic Theses & Dissertations (2024 - present)
This dissertation investigates methods for mean estimation in periodically correlated time series, focusing on the Variable Bandpass Periodic Block Bootstrap (VBPBB) and a novel data-driven maxima method. Time series require specific methods because of the temporal correlation in the data. Traditional methods like the General Seasonal Block Bootstrap (GSBB) account for this correlation but often produce wide confidence intervals because they cannot isolate multiple periodicities, allowing noise and other frequencies to interfere with analysis. The VBPBB method addresses this by applying a Kolmogorov-Zurbenko Fourier Transform (KZFT) filter to the data before bootstrapping, which suppresses interfering frequencies and results in narrower, …
Time Series Decomposition And Forecasting Of Alzheimer’S Disease Mortality Using A Kolmogorov-Zurbenko Filter, Jack D. Farrell
Time Series Decomposition And Forecasting Of Alzheimer’S Disease Mortality Using A Kolmogorov-Zurbenko Filter, Jack D. Farrell
Electronic Theses & Dissertations (2024 - present)
Alzheimer’s disease mortality has substantially risen in recent history, placing a significant burden on public health infrastructure and highlighting the need for improved analytical methods to better understand mortality data patterns and offer reliable predictions. Time series methods often struggle to balance both accuracy and interpretability, which hinders the ability to gather meaningful insights from time series data. To address these limitations, this study applies the Kolmogorov-Zurbenko (KZ) filter to monthly U.S. Alzheimer’s mortality data spanning from 1999 to 2023 and decomposes the series into long-term trend, seasonal, and noise components on a logarithmic scale. Long-term trend accounts for 86.49% …
Spectral Analysis Of Traffic Accidents In New York's Capital Region, Michael Barr
Spectral Analysis Of Traffic Accidents In New York's Capital Region, Michael Barr
Electronic Theses & Dissertations (2024 - present)
In this paper we estimate the spectral density of traffic accident events in the Capital Distict, NY area using a band-pass filter known as the Kolmogorov-Zurbenko Fourier Transform (KZFT). The source data is provided by Moosavi, et al. (2019) and originally captured from various public entities and sensors in the road network. Spectral density estimation with KZFT suppresses noise to reveal the constituent frequencies embedded in the noisy signal. Signal reconstruction based on KZFT produces an approximate weekly accident arrivals for this noisy signal, or in other words a pattern which is proportionate to the event expectation viewed over a …
Cross-Temporal Statistical Approaches For Evaluating Predictor-Outcome Relationships, Jeff Joseph
Cross-Temporal Statistical Approaches For Evaluating Predictor-Outcome Relationships, Jeff Joseph
Dartmouth College Ph.D Dissertations
Central auditory function is linked with cognitive deficits, but few research projects use existing statistical approaches or develop new ones to forecast cognitive deficits using the results of central auditory tests. To address this limitation, we use a series of statistical learning frameworks for predicting a child’s cognitive abilities based on his/her central auditory performances and demographic factors. Two key challenges exist. First, children may start the study at a time when they are unable to perform the central auditory tests or cognitive tests. Second, cognitive performance is age-dependent, particularly in the early formative years of childhood and adolescence. To …
A Comparative Evaluation Of Data Imbalance Handling Techniques In Machine Learning Models For One-Year Mortality Prediction In Liver Cirrhosis, Sumiya Hasan Trisha
A Comparative Evaluation Of Data Imbalance Handling Techniques In Machine Learning Models For One-Year Mortality Prediction In Liver Cirrhosis, Sumiya Hasan Trisha
UNF Graduate Theses and Dissertations
Liver cirrhosis is associated with substantial morbidity and mortality, making one-year mortality prediction a clinically relevant problem. Using a liver cirrhosis dataset as the motivating application, this thesis evaluates five machine learning classifiers—Logistic Regression, Random Forest, XGBoost, LightGBM, and CatBoost—under five class-imbalance handling strategies: Baseline learning, Random Oversampling, SMOTE-NC, ADASYN, and Cost-Sensitive Learning. Hyperparameter tuning was conducted using randomized search, and predictive performance was assessed over 200 iterations of Monte Carlo Cross-Validation using Accuracy, Precision, Recall, Fl-score, and ROC-AUC.
The results suggest that imbalance-handling strategies can materially affect predictive performance, particularly recall. Because the outcome of interest is death within …
Ridit-Based Adaptive Allocation In Two-Stage Clinical Trials With Binary Outcomes, Dewan Fahim
Ridit-Based Adaptive Allocation In Two-Stage Clinical Trials With Binary Outcomes, Dewan Fahim
UNF Graduate Theses and Dissertations
This study explores better ways to assign patients to treatments in clinical trials with binary outcomes, such as success or failure. Adaptive methods are used to learn from early results and adjust treatment assignments during the trial, helping more patients receive better performing treatments while maintaining reliable conclusions. We focus on trials comparing multiple treatments using a two-stage design. In the first stage, several treatments are tested to identify the most promising one; in the second stage, that treatment is compared with a control. Unlike traditional equal assignment, we use adaptive allocation in the second stage to make better use …
A Comparative Study Of Classification Methods For Healthcare Analytics, Xueting Zhao
A Comparative Study Of Classification Methods For Healthcare Analytics, Xueting Zhao
UNF Graduate Theses and Dissertations
This thesis presents a comparative study of logistic regression, Linear Discriminant Analy- sis (LDA), and Quadratic Discriminant Analysis (QDA) for binary classification in healthcare analytics, integrating theoretical derivation, simulation, and real-data application. A facto- rial simulation study crosses the covariance structure (equal vs. unequal), predictor correla- tion (ρ ∈ {0, 0.5, 0.9}), dimensionality (p ∈ {2, 5, 10}) and sample size (n ∈ {50, 100, 200}) across 54 scenarios with 1,000 Monte Carlo replicates each. Three main findings emerge. Logistic regression and LDA are nearly interchangeable when the assumption of equal-covariance holds. QDA achieves substantially better discrimi- nation when class-specific …
Generating Predictive Gene Expression Signatures For Alzheimer's Disease Using Postmortem Brain Tissue, Ashley Duche
Generating Predictive Gene Expression Signatures For Alzheimer's Disease Using Postmortem Brain Tissue, Ashley Duche
Pharmaceutical Sciences (PhD) Dissertations
Background: Alzheimer’s Disease (AD) is a progressive neurodegenerative disorder characterized by the accumulation of amyloid-beta (Aβ) plaques and tau protein aggregates. These pathological features develop in specific brain regions, but why some areas are more vulnerable to early AD-related changes remains unclear. To address this, predictive gene expression signatures were developed to explore the molecular mechanisms underlying regional susceptibility to AD pathology.
Methods: This was performed using postmortem brain (PMB) tissue from participants in the Religious Orders Study and Memory and Aging Project (ROSMAP), Mayo Clinic, and Mount Sinai Brain Bank (MSBB) to generate gene expression signatures from six brain …
Flexible Spatial Priors In Bayesian Neuroimaging: Gmrf, Nngp, And Deep Gmrf, Boyoung Hur
Flexible Spatial Priors In Bayesian Neuroimaging: Gmrf, Nngp, And Deep Gmrf, Boyoung Hur
All Dissertations
Structural neuroimaging is essential for understanding neurological disorders such as Alzheimer’s disease, enabling accurate delineation of brain regions through image segmentation. Among various segmentation methods, multi-atlas-based approaches like label fusion have become leading techniques. In statistics, Bayesian hierarchical models for label fusion are increasingly favored for their ability to incorporate uncertainty and prior knowledge. Also, a key challenge in modeling neuroimaging data is spatial dependence among image voxels, making the choice of spatial prior critical—particularly in high-resolution settings where segmentation accuracy and computational efficiency are both essential.
This dissertation proposes fully Bayesian spatial hierarchical models that explore two flex- ible …
A Persistent Homology Framework For Scrna-Seq: Assessing Clustering Robustness And Quantifying Preprocessing And Integration Effects On Topological Features., Jonah Daneshmand
A Persistent Homology Framework For Scrna-Seq: Assessing Clustering Robustness And Quantifying Preprocessing And Integration Effects On Topological Features., Jonah Daneshmand
Electronic Theses and Dissertations
As single-cell RNA sequencing (scRNA-seq) data expands, robust methods for integrating diverse datasets are critical. This dissertation applies Persistent Homology (PH), a technique from Topological Data Analysis (TDA), to a collection of scRNA-seq datasets spanning eight tissue types to quantify how data integration affects topological features and biological interpretability. We assessed global topological structure using Betti curves, Euler characteristics, and persistence landscapes across raw, normalized, and integrated data representations. Our analysis revealed a performance inversion: while conventional methods excelled on unintegrated data, high-granularity topological methods, particularly those sensitive to global data structure, became superior after integration. This suggests a synergy …
Taxonomy Of Endophytic Fungi Associated With Vallisneria Neotropicalis, Md. Arafat Rashid
Taxonomy Of Endophytic Fungi Associated With Vallisneria Neotropicalis, Md. Arafat Rashid
Graduate Theses and Dissertations (2019 - present)
This study presents the first comprehensive taxonomic and ecological investigation of endophytic fungi (EF) associated with the submerged aquatic macrophyte Vallisneria neotropicalis in the southern United States. Over a 12 month period, from April 2023 to March 2024, leaf samples were collected from two distinct sites in Mobile Bay, Alabama a disturbed, brackish Causeway location and a cleaner, less impacted site at Meaher State Park. Using culture dependent methods, a total of 257 fungal endophytes were isolated from 1,200 leaf segments. All isolates belonged to the phylum Ascomycota, distributed across 3 classes, 6 orders, 10 families, and 19 taxa. The …
Statistical Challenges And Simulation Results For Pilot Clinical Trials, Weiliang Cen
Statistical Challenges And Simulation Results For Pilot Clinical Trials, Weiliang Cen
USF Tampa Graduate Theses and Dissertations
Background: The effect size estimated from a pilot trial is often an inaccurate reflection of the true effect size observed in a large trial, leading to either underestimation or overestimation. Published data suggest that effect sizes from large trials are typically smaller than those reported in their corresponding pilot trials. To address this discrepancy, conservative or discount adjustment methods are widely recommended to modify pilot trial effect sizes when calculating sample sizes, thereby maintaining adequate statistical power. This study aims to assess effect sizes from both pilot and large trials and to evaluate the performance of existing adjustment methods.
Methods: …
Smarter Disease Detection From Electronic Health Record Data: An End-To-End Ai-Augmented Pipeline For Computable Phenotyping, Dylan Owens
Statistical Science Theses and Dissertations
Electronic Health Records (EHR) contain a wealth of structured and unstructured patient data that can be leveraged for computable phenotyping, the process of algorithmically identifying patient cohorts with specific diseases or conditions. Traditional rule-based phenotyping approaches, while interpretable, often struggle with scalability, portability across institutions, and effective use of unstructured clinical narratives. Recent advances in large language models (LLMs) present new opportunities for synthesizing complex free-text information into concise, clinically meaningful representations. However, integrating LLMs into phenotyping workflows requires careful design to maintain transparency, interpretability, and measurable uncertainty—features essential for clinical adoption and downstream applications such as decision support.
We …
Spatiotemporal Modeling Of Maternal Mortality In South Carolina 2018-2023, Leah Wood, Ray Bai, Emily Mann
Spatiotemporal Modeling Of Maternal Mortality In South Carolina 2018-2023, Leah Wood, Ray Bai, Emily Mann
Senior Theses
Maternal death serves as a public health indicator due to fact that it is considered preventable with the availability of modern biomedicine, however, it persists broadly throughout the United States. Current literature outlines national trends in maternal mortality with complicating, preexisting conditions, and structural upstream factors often cited as being the largest contributors to increased risk. This study utilizes publicly available, county-level data for maternal death in addition to demographic and descriptive data in order to estimate maternal mortality rates in each of South Carolina’s 46 counties from 2018 to 2023. In order to address sparsity in the outcome variable …