Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Southern Methodist University

Discipline
Keyword
Publication Year
Publication
Publication Type

Articles 1 - 30 of 119

Full-Text Articles in Statistics and Probability

Statistical Methodologies For Count Time Series Analysis And Topological Data Analysis Of Medical Images, Yuhyeong Jang Aug 2026

Statistical Methodologies For Count Time Series Analysis And Topological Data Analysis Of Medical Images, Yuhyeong Jang

Statistical Science Theses and Dissertations

This dissertation addresses two distinct topics related to count time series analysis and topological medical image analysis, respectively. The first part of the dissertation comprises an application of a count time series model to analysis of US monthly sex trafficking data and development of a new model for multivariate count data that exhibits serial dependence and overdispersion. By imposing a family of multivariate mixed Poisson distributions on the count random vector, the proposed model can accommodate a broad range of overdispersion as well as positive contemporaneous correlations. For maximum likelihood estimation, a computationally feasible EM-type algorithm is derived based on …


An Analytical Framework For Quantifying Urban And Community Resilience To Natural Hazards From Cell-Phone Gps-Location And Traffic-Flow Data, Georgios Chatzikyriakidis May 2026

An Analytical Framework For Quantifying Urban And Community Resilience To Natural Hazards From Cell-Phone Gps-Location And Traffic-Flow Data, Georgios Chatzikyriakidis

Civil and Environmental Engineering Theses and Dissertations

Urban areas are increasingly exposed to natural hazards while accommodating a growing share of the global population, yet a consistent science-based framework for quantifying urban and community resilience remains lacking. This dissertation develops a physics-based analytical framework grounded in statistical mechanics and the quantitative theory of Brownian motion. A city is conceptualized as a complex medium in which citizens move analogously to Brownian particles within a viscoelastic environment, influenced by socioeconomic interactions and infrastructure functionality.

A central premise is that urban resilience, interpreted as engineering resilience (an outcome), can be quantified through a single metric: the mean-square displacement MSD=⟨r²(t)⟩, of …


Bayesian Spatiotemporal Model For Counterfactual Estimation In Socioeconomic Studies, Duwani W. Gonzalez May 2026

Bayesian Spatiotemporal Model For Counterfactual Estimation In Socioeconomic Studies, Duwani W. Gonzalez

Statistical Science Theses and Dissertations

Impact evaluations of regional development programs often require estimating counterfactual outcomes for a small number of treated regions using survey-based areal data. In practice, evaluators typically rely on two-group quasi-experimental methods such as propensity score matching (PSM) and Difference-in-Differences (DiD). These approaches perform poorly when only a few regions receive treatment, and when the set of observed covariates is limited or only partially relevant. Moreover, they typically do not explicitly exploit the spatial and temporal dependence present in survey-based areal data such as in ACS (American Community Survey). This dissertation develops a family of Bayesian spatial predictive models for directly …


Deep Learning Frameworks For Biological Data Integration And Generation, Alexa Beachum Apr 2026

Deep Learning Frameworks For Biological Data Integration And Generation, Alexa Beachum

Statistical Science Theses and Dissertations

Data integration represents a key area of research for analyzing the rapidly growing volume of high-dimensional biological data across sources, stages, and modalities. To model and understand these complex, often non-linear relationships, deep learning has become an increasingly powerful tool. Here, we present two novel deep learning frameworks that address distinct but complementary integration challenges. The first framework aligns single-cell omics data across temporal stages, and the second bridges imaging and omics modalities to generate patient-level molecular profiles.

In Chapter 1, we briefly summarize existing approaches---both statistical and deep learning-based---for single-cell omics data integration and discuss their limitations for handling …


Automating Cardiff Model Data Capture In Emergency Departments: Ambient Nlp Integration With Oracle-Cerner Fhir Systems, Simi Augustine, Marco A. Lopez, Jacquelyn Cheun, Chris Papesh Mar 2026

Automating Cardiff Model Data Capture In Emergency Departments: Ambient Nlp Integration With Oracle-Cerner Fhir Systems, Simi Augustine, Marco A. Lopez, Jacquelyn Cheun, Chris Papesh

SMU Data Science Review

Violence and overdose events in Las Vegas occur at rates above the national average, with fewer than half of violent injuries reported to law enforcement [2,7]. The Cardiff Model offers a proven framework for standardized data collection and sharing between hospitals and public safety partners, yet many implementations still rely on manual entry. We propose an ambient triage pipeline integrated with Oracle-Cerner electronic health record systems to listen to nurse–patient dialogue, convert speech to text, extract Cardiff fields, and write standards-based FHIR Bundles for analytics. Using SMART on FHIR standards and Cerner Millennium APIs, the study evaluates whether ambient capture …


Availability Model To Evaluate Ai Data Centers’ Role In Grid Stability, Troy Mcsimov, Trevor S. Kunz, Jeffrey Billo Mar 2026

Availability Model To Evaluate Ai Data Centers’ Role In Grid Stability, Troy Mcsimov, Trevor S. Kunz, Jeffrey Billo

SMU Data Science Review

The United States has made it clear; it is imperative that the US wins the global AI race. This paper focuses on one of the most challenging puzzle pieces surfaced at the POWER Data Center conference (San Antonio, Sept. 30.); for Electric Reliability Council of Texas (ERCOT) the limiting factor is not generation alone but the need to balance generation and load to preserve grid reliability.

The regulatory landscape fundamentally changed with the passage of Texas Senate Bill 6 in June 2025, which mandates new large loads must "contribute to the recovery of the interconnecting electric utility’s costs" (Texas Legislature, …


Smarter Disease Detection From Electronic Health Record Data: An End-To-End Ai-Augmented Pipeline For Computable Phenotyping, Dylan Owens Oct 2025

Smarter Disease Detection From Electronic Health Record Data: An End-To-End Ai-Augmented Pipeline For Computable Phenotyping, Dylan Owens

Statistical Science Theses and Dissertations

Electronic Health Records (EHR) contain a wealth of structured and unstructured patient data that can be leveraged for computable phenotyping, the process of algorithmically identifying patient cohorts with specific diseases or conditions. Traditional rule-based phenotyping approaches, while interpretable, often struggle with scalability, portability across institutions, and effective use of unstructured clinical narratives. Recent advances in large language models (LLMs) present new opportunities for synthesizing complex free-text information into concise, clinically meaningful representations. However, integrating LLMs into phenotyping workflows requires careful design to maintain transparency, interpretability, and measurable uncertainty—features essential for clinical adoption and downstream applications such as decision support.

We …


Statistical Methods For Joint Outcome Modeling And Dynamic Assessment Of Recurrent Events, Zifang Kong Aug 2025

Statistical Methods For Joint Outcome Modeling And Dynamic Assessment Of Recurrent Events, Zifang Kong

Statistical Science Theses and Dissertations

Recurrent event data frequently arise in clinical studies where individuals experience repeated, possibly related, events over time. These data are often accompanied by sparse and irregular longitudinal measurements, creating challenges for traditional joint modeling approaches that struggle to account for time-dependent associations and within-subject correlations. We propose FRAILTY (Functional Regression with AutoRegressIve fraiLTY), a novel two-step framework that integrates functional principal component analysis (PACE) with a dynamic frailty model featuring autoregressive structure. FRAILTY accommodates both scalar and functional predictors and captures within-subject dependence across recurrent events. To further extend its utility, we develop a multivariate joint modeling framework that simultaneously …


Towards Reliable Clinical Applications Of Ai Models In Radiotherapy, Biling Wang Aug 2025

Towards Reliable Clinical Applications Of Ai Models In Radiotherapy, Biling Wang

Statistical Science Theses and Dissertations

Over the past decade, artificial intelligence (AI), particularly through deep learning (DL) techniques, has made significant strides in fields like computer vision (CV) and natural language processing (NLP), leading to transformative advancements across numerous applications. This progress has sparked considerable enthusiasm within the medical field, where DL-related research has grown exponentially since 2015. However, despite these promising developments, the real-world deployment of DL models in healthcare remains limited, especially in safety-critical domains such as radiotherapy (RT), where reliability, safety, and sustained performance are critical. This thesis addresses three core challenges associated with the clinical application of DL models: (1) post-deployment …


Advancing Statistical Methods For Multivariate And Network Meta-Analysis, Yifei Wang Jul 2025

Advancing Statistical Methods For Multivariate And Network Meta-Analysis, Yifei Wang

Statistical Science Theses and Dissertations

Multivariate meta-analysis (MMA) and network meta-analysis (NMA) are essential tools for synthesizing evidence across multiple correlated outcomes and treatments. However, these tools face practical challenges, including outcome reporting bias (ORB), unreported within-study correlations, and computational burden. ORB can distort effect estimates in MMA, while missing within-study correlations in multivariate NMA may lead to biased conclusions. To address these challenges, this dissertation introduces two novel statistical methods. For MMA, we propose SemiMMA, a semiparametric and scalable approach that treats ORB as a missing-not-at-random problem and combines inverse propensity weighting (IPW) with the generalized method of moments (GMM). For multivariate NMA, we …


Keynote - Data? We Don't Have Time For Data: A Realistic Look At Law Enforcement Use Of And Need For Human Trafficking Data, Doug Gilmer Phd Jun 2025

Keynote - Data? We Don't Have Time For Data: A Realistic Look At Law Enforcement Use Of And Need For Human Trafficking Data, Doug Gilmer Phd

SMU Human Trafficking Data Conference

Drawing on over 35 years of law enforcement experience (25 years with the Department of Homeland Security), Dr. Gilmer will speak from a government and law enforcement perspective on the need and use for human trafficking data. Some agencies and components of the U.S. government, and individual states, are heavily invested in collecting data to satisfy their reporting requirements. From a law enforcement perspective, however, big human trafficking data sets are rarely examined. Data science in law enforcement is a relatively new phenomenon, and most law enforcement officers do not have the time, resources, or background to collect or analyze …


Hybrid Graph-Recurrent Architecture For Citation Recommendation Via Future Embedding Forecasting, Mohammad Ausaf Ali Haqqani May 2025

Hybrid Graph-Recurrent Architecture For Citation Recommendation Via Future Embedding Forecasting, Mohammad Ausaf Ali Haqqani

Computer Science and Engineering Theses and Dissertations

The rapid expansion of scientific literature has intensified the challenge of identifying relevant citations, particularly for newly published or under-cited papers. Traditional citation recommendation systems typically model static relationships or respond to past citation activity, offering limited predictive power for emerging works. In response, this thesis presents a temporal modeling framework for citation recommendation that anticipates future scholarly relevance by forecasting the latent representations of academic papers.

Building on prior work that utilized Temporal Graph Networks (TGNs) to model dynamic citation flows, we propose Graph-Time, a hybrid architecture that integrates a Graph Transformer with a GRU-based time series predictor. The …


Enhancing Animal Shelter Operations With Time Series And Machine Learning, Sakava L. Kiv, Donald L. Anderson, Shivam Negi, Jacquelyn Cheun Apr 2025

Enhancing Animal Shelter Operations With Time Series And Machine Learning, Sakava L. Kiv, Donald L. Anderson, Shivam Negi, Jacquelyn Cheun

SMU Data Science Review

Enhancing animal shelter operations through machine learning involves employing a variety of advanced techniques aimed at increasing efficiency, promoting animal welfare, and optimizing resource allocation. This paper explores predictive analytics for adoption rates using regression models to estimate the likelihood of adoption based on historical data, encompassing variables such as breed, health status, and previous adoption trends. Additionally, classification algorithms are utilized to categorize animals by adoption probability, facilitating better resources and marketing prioritization. Clustering algorithms are employed to group animals according to behavior patterns and/or physical health, enabling tailored medical care and enrichment activities that improve their mental and …


Bayesian And Deep Generative Modeling In Immunology, Yuqiu Yang Aug 2024

Bayesian And Deep Generative Modeling In Immunology, Yuqiu Yang

Statistical Science Theses and Dissertations

Due to the accumulation of a large volume of data of different natures such as sequencing data, proteomics data, and clinical data, statistical methods and deep learning algorithms have become increasingly important in the field of immunology. By leveraging the diverse datasets as well as interdisciplinary knowledge from areas like biology and public health, these quantitative methods have revolutionized this field by providing powerful tools for data analysis, modeling, and prediction. This has led to a deeper understanding of the immune system, accelerated the development of novel therapies, and paved the way for personalized and precision medicine approaches in immunology. …


Bayesian Variational Inference In Keyword Identification And Multiple Instance Classification, Yaofang Hu Aug 2024

Bayesian Variational Inference In Keyword Identification And Multiple Instance Classification, Yaofang Hu

Statistical Science Theses and Dissertations

This dissertation investigates (1) Variational Bayesian Semi-supervised Keyword Extraction and (2) Variational Bayesian Multimodal Multiple Instance Classification.

The expansion of textual data, stemming from various sources such as online product reviews and scholarly publications on scientific discoveries, has created a demand for the extraction of succinct yet comprehensive information. As a result, in recent years, efforts have been spent in developing novel methodologies for keyword extraction. Although many methods have been proposed to automatically extract keywords in the contexts of both unsupervised and fully supervised learning, how to effectively use partially observed keywords, such as author-specified keywords, remains an under-explored …


A Symbolic Approach To Nonlinear Time Series Analysis, Ranjan Karki, Nibhrat Lohia, Michael B. Schulte May 2024

A Symbolic Approach To Nonlinear Time Series Analysis, Ranjan Karki, Nibhrat Lohia, Michael B. Schulte

SMU Data Science Review

Current nonlinear time series methods such as neural networks forecast well. However, they act as a black box and are difficult to interpret, leaving the researchers and the audience with little insight into why the forecasts are the way they are. There is a need for a method that forecasts accurately while also being easy to interpret. This paper aims to develop a method to build an interpretable model for univariate and multivariate nonlinear time series data using wavelets and symbolic regression. The final method relies on multilayer perceptron (MLP) neural networks as a form of dimensionality reduction and the …


Reevaluating Texas Energy Market Forecasts In The Wake Of Recent Extreme Weather Events, Robert A. Derner, Richard W. Butler Ii, Alexandria Neff, Adam R. Ruthford May 2024

Reevaluating Texas Energy Market Forecasts In The Wake Of Recent Extreme Weather Events, Robert A. Derner, Richard W. Butler Ii, Alexandria Neff, Adam R. Ruthford

SMU Data Science Review

This paper provides updated forecasts of energy demand in Texas and recognizes the impact of sustainable energy. It is important that the forecasts of the adoption of sustainable energy are reexamined after Winter Storm Uri crippled the Texas power grid and left many without power. This storm highlighted the issues the Texas power grid had and has continued to struggle with in supplying the state with energy. This paper will offer an overview of the relevant literature on the adoption of sustainable energy and relevant events that have occurred in the state of Texas that will give the reader the …


Leveraging Transformer Models For Genre Classification, Andreea C. Craus, Ben Berger, Yves Hughes, Hayley Horn May 2024

Leveraging Transformer Models For Genre Classification, Andreea C. Craus, Ben Berger, Yves Hughes, Hayley Horn

SMU Data Science Review

As the digital music landscape continues to expand, the need for effective methods to understand and contextualize the diverse genres of lyrical content becomes increasingly critical. This research focuses on the application of transformer models in the domain of music analysis, specifically in the task of lyric genre classification. By leveraging the advanced capabilities of transformer architectures, this project aims to capture intricate linguistic nuances within song lyrics, thereby enhancing the accuracy and efficiency of genre classification. The relevance of this project lies in its potential to contribute to the development of automated systems for music recommendation and genre-based playlist …


Context Aware Music Recommendation And Playlist Generation, Elias Mann May 2024

Context Aware Music Recommendation And Playlist Generation, Elias Mann

SMU Journal of Undergraduate Research

There are many reasons people listen to music, and the type of music is largely determined by what the listener may be doing while they listen. For example, one may listen to one type of music while commuting, another while exercising, and yet another while relaxing. Without access to the physiological state of the user, current music recommendation methods rely on collaborative filtering - recommending music based on what other similar users listen to - and content based filtering - recommending songs based on their similarities to songs the user already prefers. With the rise in popularity of smart devices …


Statistical Approaches For The Early Detection Of Colorectal Cancer Using Longitudinal Biomarkers, Emily Berry May 2024

Statistical Approaches For The Early Detection Of Colorectal Cancer Using Longitudinal Biomarkers, Emily Berry

Statistical Science Theses and Dissertations

Colorectal cancer (CRC) is the third leading cause of cancer-related death in the United States [45]. CRC is believed to advance from adenomatous polyps creating a unique opportunity for both early detection and cancer prevention [4, 23]. Like other diseases, CRC screening reduces mortality by detecting cancer at earlier, more treatable stages; however, it can also reduce incidence through the removal of precancerous lesions [4]. As a result, screening is recommended for average-risk adults ≥ 45 years of age and includes a variety of tests [4, 12]. Despite alternate screening options, colonoscopy capacity is often cited as a barrier to …


Interpretable Word-Level Sentiment Analysis With Attention-Based Multiple Instance Classification Models, Chenyu Yang Dec 2023

Interpretable Word-Level Sentiment Analysis With Attention-Based Multiple Instance Classification Models, Chenyu Yang

Statistical Science Theses and Dissertations

In this study, our main objective is to tackle the black-box nature of popular machine learning models in sentiment analysis and enhance model interpretability. We aim to gain more insight into the decision-making process of sentiment analysis models, which is often obscure in those complex models. To achieve this goal, we introduce two word-level sentiment analysis models.

The first model is called the attention-based multiple instance classification (AMIC) model. It combines the transparent model structure of multiple instance classification and the self-attention mechanism in deep learning to incorporate the contextual information from documents. As demonstrated by a wine review dataset …


Deep Learning For Microbiome-Based Integrative Modeling And Microbial Biomarkers Identification, Sen Yang Dec 2023

Deep Learning For Microbiome-Based Integrative Modeling And Microbial Biomarkers Identification, Sen Yang

Statistical Science Theses and Dissertations

The human microbiome, comprising trillions of microorganisms, plays a pivotal role in modulating host physiology via molecular and metabolite exchanges. One of the major challenges in this field lies in the effective integration of microbiome and metabolomics data, an achievement that holds the promise of substantially enhancing the precision of disease prediction. However, many datasets prioritize microbiome data while neglecting paired metabolome information. Additionally, the prevalent analytical tools face challenges in effectively merging these intricate datasets, leading to possible misinterpretations and reduced prediction accuracies.

To address these challenges, the first part of this research introduces the Microbiome-based Supervised Contrastive Learning …


Ohio Recovery Housing: Resident Risk And Outcomes Assessment, Elyjiah Potter, Bivin Sadler Dec 2023

Ohio Recovery Housing: Resident Risk And Outcomes Assessment, Elyjiah Potter, Bivin Sadler

SMU Data Science Review

Addiction and substance abuse disorder is a significant problem in the United States. Over the past two decades, the United States has faced a boom in substance abuse, which has resulted in an increase in death and disruption of families across the nation. The State of Ohio has been particularly hard hit by the crisis, with overdose rates nearly doubling the national average. Established in the mid 1970’s Sober Living Housing is an alcohol and substance use recovery model emphasizing personal responsibility, sober living, and community support. This model has been adopted by the Ohio Recovery Housing organization, which seeks …


Differentiation Of Human, Dog, And Cat Hair Fibers Using Dart Tofms And Machine Learning, Laura Ahumada, Erin R. Mcclure-Price, Chad Kwong, Edgard O. Espinoza, John Santerre Dec 2023

Differentiation Of Human, Dog, And Cat Hair Fibers Using Dart Tofms And Machine Learning, Laura Ahumada, Erin R. Mcclure-Price, Chad Kwong, Edgard O. Espinoza, John Santerre

SMU Data Science Review

Hair is found in over 90% of crime scenes and has long been analyzed as trace evidence. However, recent reviews of traditional hair fiber analysis techniques, primarily morphological examination, have cast doubt on its reliability. To address these concerns, this study employed machine learning algorithms, specifically Linear Discriminant Analysis (LDA) and Random Forest, on Direct Analysis in Real Time time-of-flight mass spectra collected from human, cat, and dog hair samples. The objective was to develop a chemistry- and statistics-based classification method for unbiased taxonomic identification of hair. The results of the study showed that LDA and Random Forest were highly …


Bayesian Statistical Modeling Of Spatially Resolved Transcriptomics Data, Xi Jiang Oct 2023

Bayesian Statistical Modeling Of Spatially Resolved Transcriptomics Data, Xi Jiang

Statistical Science Theses and Dissertations

Spatially resolved transcriptomics (SRT) quantifies expression levels at different spatial locations, providing a new and powerful tool to investigate novel biological insights. As experimental technologies enhance both in capacity and efficiency, there arises a growing demand for the development of analytical methodologies.

One question in SRT data analysis is to identify genes whose expressions exhibit spatially correlated patterns, called spatially variable (SV) genes. Most current methods to identify SV genes are built upon the geostatistical model with Gaussian process, which could limit the models' ability to identify complex spatial patterns. In order to overcome this challenge and capture more types …


Using Geographic Information To Explore Player-Specific Movement And Its Effects On Play Success In The Nfl, Hayley Horn, Eric Laigaie, Alexander Lopez, Shravan Reddy Aug 2023

Using Geographic Information To Explore Player-Specific Movement And Its Effects On Play Success In The Nfl, Hayley Horn, Eric Laigaie, Alexander Lopez, Shravan Reddy

SMU Data Science Review

American Football is a billion-dollar industry in the United States. The analytical aspect of the sport is an ever-growing domain, with open-source competitions like the NFL Big Data Bowl accelerating this growth. With the amount of player movement during each play, tracking data can prove valuable in many areas of football analytics. While concussion detection, catch recognition, and completion percentage prediction are all existing use cases for this data, player-specific movement attributes, such as speed and agility, may be helpful in predicting play success. This research calculates player-specific speed and agility attributes from tracking data and supplements them with descriptive …


Traditional Vs Machine Learning Approaches: A Comparison Of Time Series Modeling Methods, Miguel E. Bonilla Jr., Jason Mcdonald, Tamas Toth, Bivin Sadler Aug 2023

Traditional Vs Machine Learning Approaches: A Comparison Of Time Series Modeling Methods, Miguel E. Bonilla Jr., Jason Mcdonald, Tamas Toth, Bivin Sadler

SMU Data Science Review

In recent years, various new Machine Learning and Deep Learning algorithms have been introduced, claiming to offer better performance than traditional statistical approaches when forecasting time series. Studies seeking evidence to support the usage of ML/DL over statistical approaches have been limited to comparing the forecasting performance of univariate, linear time series data. This research compares the performance of traditional statistical-based and ML/DL methods for forecasting multivariate and nonlinear time series.


A Hybrid Ensemble Of Learning Models, Bivin Sadler, Dhruba Dey, Duy Nguyen, Tavin Weeda Aug 2023

A Hybrid Ensemble Of Learning Models, Bivin Sadler, Dhruba Dey, Duy Nguyen, Tavin Weeda

SMU Data Science Review

Statistical models in time series forecasting have long been challenged to be superseded by the advent of deep learning models. This research proposes a new hybrid ensemble of forecasting models that combines the strengths of several strong candidates from these two model types. The proposed ensemble aims to improve the accuracy of forecasts and reduce computational complexity by leveraging the strengths of each candidate model.


Optimal Experimental Planning Of Reliability Experiments Based On Coherent Systems, Yang Yu Jul 2023

Optimal Experimental Planning Of Reliability Experiments Based On Coherent Systems, Yang Yu

Statistical Science Theses and Dissertations

In industrial engineering and manufacturing, assessing the reliability of a product or system is an important topic. Life-testing and reliability experiments are commonly used reliability assessment methods to gain sound knowledge about product or system lifetime distributions. Usually, a sample of items of interest is subjected to stresses and environmental conditions that characterize the normal operating conditions. During the life-test, successive times to failure are recorded and lifetime data are collected. Life-testing is useful in many industrial environments, including the automobile, materials, telecommunications, and electronics industries.

There are different kinds of life-testing experiments that can be applied for different purposes. …


A Comparison Of Confidence Intervals In State Space Models, Jinyu Du Jul 2023

A Comparison Of Confidence Intervals In State Space Models, Jinyu Du

Statistical Science Theses and Dissertations

This thesis develops general procedures for constructing confidence intervals (CIs) of the error disturbance parameters (standard deviations) and transformations of the error disturbance parameters in time-invariant state space models (ssm). With only a set of observations, estimating individual error disturbance parameters accurately in the presence of other unknown parameters in ssm is a very challenging problem. We attempted to construct four different types of confidence intervals, Wald, likelihood ratio, score, and higher-order asymptotic intervals for both the simple local level model and the general time-invariant state space models (ssm). We show that for a simple local level model, both the …