Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- Central Bank of Nigeria (27)
- Southern Methodist University (25)
- University of Kentucky (25)
- Virginia Commonwealth University (17)
- Georgia Southern University (14)
-
- Claremont Colleges (13)
- Illinois State University (11)
- University of Nebraska - Lincoln (11)
- California Polytechnic State University, San Luis Obispo (10)
- East Tennessee State University (10)
- The University of Akron (9)
- University of Arkansas, Fayetteville (9)
- City University of New York (CUNY) (8)
- Kennesaw State University (8)
- University of Central Florida (8)
- University of Louisville (7)
- Clemson University (6)
- Murray State University (6)
- Old Dominion University (6)
- COBRA (5)
- Michigan Technological University (5)
- Technological University Dublin (5)
- University of Connecticut (5)
- University of Texas Rio Grande Valley (5)
- Wilfrid Laurier University (5)
- Florida Institute of Technology (4)
- Louisiana State University (4)
- Purdue University (4)
- Rochester Institute of Technology (4)
- South Dakota State University (4)
- Keyword
-
- Statistics (41)
- Machine learning (13)
- Regression (13)
- Machine Learning (11)
- Bayesian (9)
-
- Logistic regression (9)
- Simulation (9)
- R (6)
- Time series (6)
- Classification (5)
- Clustering (5)
- Deep Learning (5)
- Artificial Intelligence (4)
- Data science (4)
- EM algorithm (4)
- Epidemiology (4)
- Forecasting (4)
- Generalized linear models (4)
- Linear regression (4)
- Logistic Regression (4)
- Nigeria (4)
- Poisson (4)
- Prediction (4)
- Variable selection (4)
- Algorithm (3)
- Analytics (3)
- Baseball (3)
- Big Data (3)
- Biostatistics (3)
- COVID-19 (3)
- Publication Year
- Publication
-
- CBN Journal of Applied Statistics (JAS) (27)
- Electronic Theses and Dissertations (23)
- Theses and Dissertations (23)
- Theses and Dissertations--Statistics (20)
- SMU Data Science Review (16)
-
- College of Graduate Studies: Theses & Dissertations (13)
- CMC Senior Theses (10)
- Statistical Science Theses and Dissertations (9)
- Williams Honors College, Honors Research Projects (9)
- Annual Symposium on Biomathematics and Ecology Education and Research (7)
- Graduate Theses and Dissertations (7)
- Articles (6)
- Data Science and Data Mining (6)
- All Dissertations (5)
- Dissertations, Master's Theses and Master's Reports (5)
- Published and Grey Literature from PhD Candidates (5)
- Statistics (5)
- Theses and Dissertations (Comprehensive) (5)
- Basic Science Engineering (4)
- Department of Statistics: Dissertations, Theses, and Student Research (4)
- Master's Theses (4)
- Mathematics & Statistics Theses & Dissertations (4)
- SDSU Data Science Symposium (4)
- Al-Bahir (3)
- Biology and Medicine Through Mathematics Conference (3)
- CHIP Documents (3)
- COBRA Preprint Series (3)
- Dissertations, Theses, and Capstone Projects (3)
- Graduate Student Theses, Dissertations, & Professional Papers (3)
- Graduate Theses/Dissertations (3)
- Publication Type
- File Type
Articles 31 - 60 of 374
Full-Text Articles in Statistical Models
Testing For Dice Control At Craps, Stewart N. Ethier
Testing For Dice Control At Craps, Stewart N. Ethier
UNLV Gaming Research & Review Journal
Dice control involves “setting” the dice and then throwing them carefully, in the hope of influencing the outcomes and gaining an advantage at craps. How does one test for this ability? To specify the alternative hypothesis, we need a statistical model of dice control. Two have been suggested in the gambling literature, namely the Smith–Scott model and the Wong–Shackleford model. Both models are parameterized by θ ∈ [0, 1], which measures the shooter’s level of control. We propose and compare four test statistics: (a) the sample proportion of 7s; (b) the sample proportion of pass-line wins; (c) the sample mean …
Modeling Private Debt Using U.S. Consumer Expenditure Data, Stsiapan Dziamentsyeu
Modeling Private Debt Using U.S. Consumer Expenditure Data, Stsiapan Dziamentsyeu
Honors Capstones
This project models private household debt among U.S. consumers using data from the Consumer Expenditure Survey (CES) between 2013 and 2023. The analysis focuses on identifying how demographic and economic characteristics, such as income, housing expenditures, education, and occupation, relate to non-mortgage “other” loan balances. After initial model development produced poor residual behavior due to zero-inflation from imputed debt values, the analysis was refined to include only households reporting verifiable debt. Multiple modeling techniques, including AIC-based variable selection and Lasso regularization, were compared under a five-fold cross-validation framework. The Lasso model achieved superior predictive accuracy (RMSE = 1.55, MAE = …
Coupled Machine Learning Models: Combining Observations And Numerical Analysis In A Physics-Regularized Approach, Austin B. Schmidt
Coupled Machine Learning Models: Combining Observations And Numerical Analysis In A Physics-Regularized Approach, Austin B. Schmidt
LSU New Orleans Theses and Dissertations
This dissertation investigates surrogate modeling for fixed-location environmental forecasting using novel data-combination techniques. The work surveys the landscape of observational measurements and numerically generated data, identifying similar research and gaps in current methodologies. The ratio-coupled training framework is introduced to combine two data sources per predicted feature through a tunable parameter that weights training signal strength. An optimization scheme is developed to simultaneously tune surrogate weights and the coupled signal ratio, allowing relative influence between signals to act as an explicit regularizer. Three case studies demonstrate the methodology and approach in a variety of contexts. The first study is based …
The Moderating Effect Of Income Inequality On The Income–Emissions Relationship In G20 Countries, Zahra Rizky Fadilah, Budiasih Budiasih
The Moderating Effect Of Income Inequality On The Income–Emissions Relationship In G20 Countries, Zahra Rizky Fadilah, Budiasih Budiasih
Economics and Finance in Indonesia
This study analyzes the moderating effect of income inequality on the income–emissions relationship in the environmental Kuznets curve (EKC) framework. Findings indicate that the relationship is inverted U-shaped in middle-income G20 countries, but monotonically increasing in high-income G20 countries. Interestingly, income inequality moderates this relationship only in the latter group. These findings suggest that middle-income G20 countries should focus on raising income per capita to mitigate environmental degradation, while their high-income counterparts need to prioritize reducing income inequality to effectively decouple income from emissions.
On Bayesian Empirical Likelihood-Based Method For Complex Survey Data With Application To Non-Probability Sampling, Md Hasibur Rahman
On Bayesian Empirical Likelihood-Based Method For Complex Survey Data With Application To Non-Probability Sampling, Md Hasibur Rahman
Department of Statistics: Dissertations, Theses, and Student Research
This thesis develops a Bayesian empirical likelihood (BEL) framework for inference under complex survey designs and extends it to non-probability sampling. Parametric likelihood based methods are difficult to apply to complex survey data because the likelihood is rarely available in closed form. EL provides a flexible alternative by replacing the parametric likelihood with an empirical likelihood constructed from moment conditions. The proposed method first integrates empirical likelihood constraints with survey design features then extends BEL to non-probability sampling through selection models and design consistent restrictions. Posterior inference is carried out using a Metropolis–Hastings MCMC algorithm. A real-data analysis further illustrates …
A Leslie System For A Demographic Simulation: From An Actuarial Point Of View, David Kings
A Leslie System For A Demographic Simulation: From An Actuarial Point Of View, David Kings
Electronic Theses and Dissertations
This thesis develops a discrete stochastic linear systems interpretation of age–stage demographic evolution grounded in Leslie operators and realized in a discrete-event simulation implemented with salabim. The central claim is that one annual cycle of the simulation constitutes a cone-preserving, stochastic affine transformation on a high- dimensional population state vector indexed by age, sex, marital status, household type, employment, and education, and that the composition of yearly operators yields a random matrix product whose top Lyapunov exponent is the stochastic counterpart of the Perron–Frobenius growth rate (Caswell, 2001; Tuljapurkar, 1997)[1, 2]. The actuarial bridge is constructed by mapping simulated survival …
Smarter Disease Detection From Electronic Health Record Data: An End-To-End Ai-Augmented Pipeline For Computable Phenotyping, Dylan Owens
Statistical Science Theses and Dissertations
Electronic Health Records (EHR) contain a wealth of structured and unstructured patient data that can be leveraged for computable phenotyping, the process of algorithmically identifying patient cohorts with specific diseases or conditions. Traditional rule-based phenotyping approaches, while interpretable, often struggle with scalability, portability across institutions, and effective use of unstructured clinical narratives. Recent advances in large language models (LLMs) present new opportunities for synthesizing complex free-text information into concise, clinically meaningful representations. However, integrating LLMs into phenotyping workflows requires careful design to maintain transparency, interpretability, and measurable uncertainty—features essential for clinical adoption and downstream applications such as decision support.
We …
Comparative Evaluation Of Estimation Techniques For Purchasing Power Parity In African Countries Using The Country-Product-Dummy Regression Framework, Rokibat Adeola Tijani, Taiwo Abideen Lasisi, Dahud Kehinde Shangodoyin, Olasunkanmi James Oladapo
Comparative Evaluation Of Estimation Techniques For Purchasing Power Parity In African Countries Using The Country-Product-Dummy Regression Framework, Rokibat Adeola Tijani, Taiwo Abideen Lasisi, Dahud Kehinde Shangodoyin, Olasunkanmi James Oladapo
Al-Bahir
Purchasing Power Parity (PPP) is a popular macroeconomic analysis metric used to compare economic productivity and standards of living between countries. This study examines the estimation of PPP within the International Comparison Program (ICP) at Basic Heading (BH) level stage and leverages on the data from the 2011 ICP round. Focusing on five BHs out of 12 BHs across 50 Africa countries, to empirically evaluate the validity of the classical Ordinary Least Square (OLS) assumptions in the estimation of Country Product Dummy (CPD) regressions. Given the widespread use of OLS for BH level PPP computation, a rigorous examination of these …
Nba Player Types And Salaries: Assessing The Disparities In Pay, Nick Riccardi, Rodney J. Paul
Nba Player Types And Salaries: Assessing The Disparities In Pay, Nick Riccardi, Rodney J. Paul
Sport Management - All Scholarship
The purpose of this study was to identify player types that exist in the modern National Basketball Association (NBA), test whether player types are paid differently controlling for performance and other factors and construct successful rosters with cheaper payrolls.
We collected performance statistics and salary data for players and teams across five seasons (2018-19 to 2022-23). Cluster analysis is leveraged to group together player-seasons to identify the player types that exist in the NBA. Linear regression models are run to test for differences in pay by cluster membership while controlling for performance, age, and contractual details. Linear programming simulation models …
Statistical Methods For Joint Outcome Modeling And Dynamic Assessment Of Recurrent Events, Zifang Kong
Statistical Methods For Joint Outcome Modeling And Dynamic Assessment Of Recurrent Events, Zifang Kong
Statistical Science Theses and Dissertations
Recurrent event data frequently arise in clinical studies where individuals experience repeated, possibly related, events over time. These data are often accompanied by sparse and irregular longitudinal measurements, creating challenges for traditional joint modeling approaches that struggle to account for time-dependent associations and within-subject correlations. We propose FRAILTY (Functional Regression with AutoRegressIve fraiLTY), a novel two-step framework that integrates functional principal component analysis (PACE) with a dynamic frailty model featuring autoregressive structure. FRAILTY accommodates both scalar and functional predictors and captures within-subject dependence across recurrent events. To further extend its utility, we develop a multivariate joint modeling framework that simultaneously …
Performance Of The Two Sample Likelihood Ratio Test Under A Nested Dirichlet: A Simulation Study, Edwina Agyeman
Performance Of The Two Sample Likelihood Ratio Test Under A Nested Dirichlet: A Simulation Study, Edwina Agyeman
Electronic Theses and Dissertations
Compositional data analysis (CoDA) addresses multivariate data constrained to a constant sum, such as proportions or percentages. Originating from early warnings regarding misinterpretation by Pearson (1897), the field was formalized by John Aitchison in 1986, whose foundational work remains highly influential. Over time, new modeling techniques and visualization tools have advanced the field, as noted by Greenacre et al. More recently, Turner et al. proposed an approach based on the Nested Dirichlet Distribution (NDD), which accommodates more flexible dependence structures than the standard Dirichlet model. This thesis builds on the methodology of Turner et al. Chapter 1 introduces the nature …
Experimental Design And Analysis For Decision Making: Methodology And Applications, Yezhuo Li
Experimental Design And Analysis For Decision Making: Methodology And Applications, Yezhuo Li
All Dissertations
This dissertation develops and applies advanced statistical and optimization frameworks to enhance decision-making under uncertainty, particularly in engineering and manufacturing contexts. First, we introduce an approach for the optimal design of controlled experiments that accounts for observational covariates, enabling more precise and personalized decisions. Second, we explore the application of constrained Bayesian optimization, using Gaussian process surrogate models, to optimize composite cure processes, significantly reducing computational effort while maintaining high predictive accuracy. Building on this foundation, we extend Bayesian optimization to bivariate Gaussian process models that capture correlations between objective and constraint functions, offering new insights into multidimensional decision landscapes. …
Estimation Methods For Bayesian Exponential Random Graph Models Under The Horseshoe Prior., Pamela Linares
Estimation Methods For Bayesian Exponential Random Graph Models Under The Horseshoe Prior., Pamela Linares
Electronic Theses and Dissertations
Networks are powerful tools for modeling the complexity of social interactions, biological systems, and information spread. A leading statistical frameworks for analyzing network data are Exponential Random Graph Models (ERGMs), which provide a principled approach to capturing structural dependencies. However, ERGMs remain challenging to estimate, especially in sparse or high-dimensional settings where models suffer from degeneracy and unstable parameter inference. This paper proposes a penalized Bayesian approach to ERGMs that utilizes the horseshoe prior, a sparsity-inducing global-local shrinkage prior. This prior offers robust regularization while preserving important signals, improving estimation by shrinking irrelevant parameters and reducing the impact of extreme …
Advancing Real-World Implementation Of The Well Optimized Linear Finder (Wolf) High-Speed Atmospheric Turbulence Compensation Method, Timothy Evan Coon
Advancing Real-World Implementation Of The Well Optimized Linear Finder (Wolf) High-Speed Atmospheric Turbulence Compensation Method, Timothy Evan Coon
Theses and Dissertations
This dissertation advances the real-world implementation of the Well Optimized Linear Finder (WOLF) method for high-speed Atmospheric Turbulence Compensation (ATC). Atmospheric turbulence introduces phase aberrations into optical wavefronts and degrades image quality in terrestrial imaging systems. Traditional phase diversity methods are computationally intensive and poorly suited to real-time operation. The WOLF method addresses these limitations through a novel, point-wise formulation of the optical transfer function (OTF) as a structured autocorrelation of the generalized pupil function (GPF). This formulation enables the estimation of phase aberrations at individual spatial coordinates with distributed computational complexity.
The research begins by developing a MATLAB-based simulation …
Unified Hybrid Censoring Samples From Power Pratibha Distribution And Its Applications, Mahmoud Mansour, Hebatalla H. Mohammad Dr, Khalaf S. Sultan Prof.
Unified Hybrid Censoring Samples From Power Pratibha Distribution And Its Applications, Mahmoud Mansour, Hebatalla H. Mohammad Dr, Khalaf S. Sultan Prof.
Basic Science Engineering
This paper suggests an extensive inferential method for the Power Pratibha Distribution (PPD) under Unified Hybrid Censoring Schemes (UHCSs), since there is a growing interest in flexible models in both reliability and service operations. This work studies the PPD model using standard Maximum Likelihood Estimation methods and modern Bayesian approaches too. Using a complex architecture, UHCS simulates tests more closely to what is done in practice than by using more basic censoring schemes. Using analysis, the probability and statistical ranges are carefully calculated for the parameters. Tests demonstrate that Bayesian estimation gives better results than many other methods for estimation, …
Bandwagon Behavior In Major League Baseball, Daniel E. Erro
Bandwagon Behavior In Major League Baseball, Daniel E. Erro
Master's Theses
This study investigates “bandwagon” behavior among Major League Baseball (MLB) fans by analyzing Google search interest data from 2004 to 2019. Drawing on publicly available information from Google Trends, the analysis explores how fluctuations in search activity align with team performance during both the regular season and postseason. Hierarchical linear models are used to estimate expected levels of fan interest based on team performance and market characteristics. Deviations from these expectations during the regular season are interpreted as evidence of bandwagon or anti-bandwagon behavior. A drop-off in interest following playoff elimination is also examined to capture shifts in fan attention …
Welfare Implication Of Alternative Tax Rates Adjustment Policy In Nigeria: A Dsge Analysis, Umar B. Ibrahim, Isah F. Abubakar
Welfare Implication Of Alternative Tax Rates Adjustment Policy In Nigeria: A Dsge Analysis, Umar B. Ibrahim, Isah F. Abubakar
CBN Journal of Applied Statistics (JAS)
This study sets out to determine the desirable policy adjustment in the tax rate for Nigeria that ensures the least welfare cost. A calibrated small open-economy New Keynesian Dynamic Stochastic General Equilibrium (NKDSGE) model of the Nigerian economy is applied to achieve this objective. Within this framework, we examined the impact of an increase in value-added tax (VAT) rate from 7.5 to 15 percent on key macroeconomic variables relative to the impact of an increase in company income tax (CIT) rate from 30 to 35 percent on macroeconomic variables. Furthermore, we examined the welfare costs of the increases in the …
Expressive And Interpretable User Engagement Prediction Using Multivariate Survival Processes, Akshay Aravamudan
Expressive And Interpretable User Engagement Prediction Using Multivariate Survival Processes, Akshay Aravamudan
Theses and Dissertations
The ability to characterize how information diffuses online is of paramount importance to stakeholders that are interested in tasks such as proposing solutions for mitigating and countering dis/misinformation, predicting user engagement of content in social media, planning marketing campaigns to roll-out products and planning dissemination of political campaign messaging among others. One such facet of learning the dynamics of information diffusion is the ability to predict user engagement or the popularity of a single piece of information as it spreads through an online medium. Existing works in this regard mainly either obfuscate user level information or utilize frameworks that are …
Mortgage Default Classification Modeling For Variable Analysis, Brendan R. Goggins
Mortgage Default Classification Modeling For Variable Analysis, Brendan R. Goggins
Honors College Theses
The financial crisis of the early 2000’s is a prime example of the severe consequences that mortgage default and borrower insolvency can have on economies at large. Mortgage default specifically is a prime case with the popularization of mortgage backed securities and the commonality of this loan structure. Multiple hypotheses and models have been formed to understand the reasons, causes, and consequences of mortgage default. This paper uses both machine learning and statistical classification models to inform an understanding of the variables most significant and impactful to the default outcome of mortgages. Consideration is given to both loan-level microeconomic variables …
Enhancing Animal Shelter Operations With Time Series And Machine Learning, Sakava L. Kiv, Donald L. Anderson, Shivam Negi, Jacquelyn Cheun
Enhancing Animal Shelter Operations With Time Series And Machine Learning, Sakava L. Kiv, Donald L. Anderson, Shivam Negi, Jacquelyn Cheun
SMU Data Science Review
Enhancing animal shelter operations through machine learning involves employing a variety of advanced techniques aimed at increasing efficiency, promoting animal welfare, and optimizing resource allocation. This paper explores predictive analytics for adoption rates using regression models to estimate the likelihood of adoption based on historical data, encompassing variables such as breed, health status, and previous adoption trends. Additionally, classification algorithms are utilized to categorize animals by adoption probability, facilitating better resources and marketing prioritization. Clustering algorithms are employed to group animals according to behavior patterns and/or physical health, enabling tailored medical care and enrichment activities that improve their mental and …
Profiting On The Kentucky Derby, Bailey Korfhage
Profiting On The Kentucky Derby, Bailey Korfhage
Undergraduate Theses
This paper analyzes the quantitative data of horses that ran in the Kentucky Derby to recognize statistically significant variables to predict the horse that comes in first or in-the-money. This analysis is specific to the post-implementation of the points system that began for the 2013 Kentucky Derby. Churchill Downs, the host of the Kentucky Derby, changed the methodology of qualification for a horse to enter the race; instead of qualifying with highest earnings in lifetime starts, the institution implemented a points system that awarded different proportions of points depending on the value of various prep races leading up to the …
Linking Water Quality And Climate Change To Long-Term Trends In Species Abundance In Norwalk Harbor, Viktoria Savatorova, Aidan Kieft, Nicole C. Spiller, Kasey Burns
Linking Water Quality And Climate Change To Long-Term Trends In Species Abundance In Norwalk Harbor, Viktoria Savatorova, Aidan Kieft, Nicole C. Spiller, Kasey Burns
Spora: A Journal of Biomathematics
This study examines the effects of environmental changes on fish populations in Norwalk Harbor, focusing on winter flounder (Pseudopleuronectes americanus), cunner (Tautogolabrus adspersus), northern pipefish (Syngnathus fuscus), and naked goby (Gobiosoma bosci) as examples of species responding to climate-related shifts. We analyze how water temperature, salinity, and dissolved oxygen correlate with fish abundance. To assess statistically significant differences in catch per unit effort (CPUE) across harbor regions, we applied the Kruskal-Wallis test followed by Dunn's post-hoc test. Seasonal variations in CPUE were examined by comparing monthly catch data for each species. K-means …
Discounting Effect Size When Borrowing External Data In Clinical Studies, Zhuanzhuan Ma, Chul Ahn, Bin Wang, Xuefeng Li
Discounting Effect Size When Borrowing External Data In Clinical Studies, Zhuanzhuan Ma, Chul Ahn, Bin Wang, Xuefeng Li
Research Symposium
Background: When borrowing information from external data to augment a current trial, many available methods discount the sample size but retain the effect size from previous studies. Discounting the sample size is just one way to discount the prior information. It may not be appropriate if the underlying assumption of unbiased treatment effect does not hold, for example, when the treatment effect in the historical study is likely higher than the one expected in the current trial.
Methods: To tackle this potential issue, we study some methods to shrink the effect size from previous studies assuming that the prior effect …
Sparse Bayesian Variable Selection Using Global-Local Shrinkage Priors For The Analysis Of Cancer Datasets, Zhuanzhuan Ma
Sparse Bayesian Variable Selection Using Global-Local Shrinkage Priors For The Analysis Of Cancer Datasets, Zhuanzhuan Ma
Research Symposium
Background: With a rapid development of data collection technology, high dimensional data, whose model dimension k may be growing or much larger than the sample size n, is becoming increasingly prevalent in different fields of study, such as ecology, genetics, among others. This data deluge is introducing new challenges to traditional statistical procedures and theories and is thus generating a renewed interest in the problems of variable selection and classification in high dimensional regression models. In large k, small n settings, variable selection is usually the first step for dimension reduction to uncover significant covariates, which contribute to …
Kroger Post-Pandemic Customer Segmentation, Mario Mata, Joey Truitt, Renn Spigelmyer, Dhanuja Kasturiratna, Lisa Holden, Nitish Baidya, Hanna Tafari
Kroger Post-Pandemic Customer Segmentation, Mario Mata, Joey Truitt, Renn Spigelmyer, Dhanuja Kasturiratna, Lisa Holden, Nitish Baidya, Hanna Tafari
Posters-at-the-Capitol
The grocery retail industry landscape has changed greatly in the wake of the pandemic. Specifically, delivery and pickup services have become more popular and customer buying habits have evolved. At the same time, improvements in data collection and analysis have allowed grocery marketing strategies to become highly individualized.
We worked with 84.51, an analytics firm, to identify customer segments for the Kroger Company based on data from 2023. Using clustering techniques, we organized customers into groups, or segments, based on similar characteristics. We identified and profiled four distinct groups of customers. Three segments were characterized by high frequency and spending …
Predicting Superconducting Critical Temperature From Composition-Derived Features: A Transparent Linear And Regularized Regression Study, Md Ahiduzzaman
Predicting Superconducting Critical Temperature From Composition-Derived Features: A Transparent Linear And Regularized Regression Study, Md Ahiduzzaman
Data Science and Data Mining
We study prediction of superconducting critical temperature (Tc) from 81 composition-derived descriptors across 21,263 materials. To keep the analysis transparent and repro- ducible, we focus on linear models: Ordinary Least Squares (OLS), Ridge, Lasso, and Elastic Net (ENet). All models share a single evaluation protocol (5-fold cross-validation with standardized inputs) and are compared on RMSE, MAE, and R2. On this feature set, OLS attains the best cross-validated performance (RMSE = 17.6 K, MAE = 13.3 K , R2 = 0.735), with Lasso/ENet essentially tied next (RMSE ≈ 17.7 K , R2 ≈ 0.734); Ridge underperforms (RMSE = 18.9 K , …
Comparative Analysis Of Lasso, Ridge, And Elastic Net For Variable Selection In High-Dimensional Maize Data, Md Ahiduzzaman
Comparative Analysis Of Lasso, Ridge, And Elastic Net For Variable Selection In High-Dimensional Maize Data, Md Ahiduzzaman
Data Science and Data Mining
In high-dimensional genomic data analysis, traditional linear regression techniques often struggle due to the presence of a large number of predictor variables relative to observations. Penalized regression methods such as LASSO, Ridge, and Elastic Net have emerged as effective solutions by imposing regularization, which helps in managing multicollinearity and enhancing prediction accuracy. This study applies these techniques to the Maize dataset to model the time to male flowering, selecting relevant genetic markers as predictors. Our findings suggest that Elastic Net is particularly effective for high-dimensional data with correlated variables, achieving a balance between prediction accuracy and variable selection. The results …
Analyzing Factors Influencing Employee Turnover In Tech Companies: A Predictive Modeling Approach, Shinjon Ghosh
Analyzing Factors Influencing Employee Turnover In Tech Companies: A Predictive Modeling Approach, Shinjon Ghosh
Theses and Dissertations
Employee turnover poses substantial challenges for technology firms, and understanding its key drivers through predictive modeling is essential for developing effective retention strategies. This study investigates factors influencing employee turnover in technology companies by implementing a predictive modeling approach on the IBM HR Analytics Employee Attrition dataset. The research aims were identifying key factors contributing to employee attrition, developing predictive models to forecast turnover risk, and analyzing interactions among significant predictors. By examining a range of features, the results highlight significant variables (Over Time, Monthly Income, Marital Status, etc.) of attrition and offer actionable insights for developing targeted employee retention …
Hybrid Mixtures Of Factor Analyzers For High Dimensional Data, Kazeem Abiodun Kareem
Hybrid Mixtures Of Factor Analyzers For High Dimensional Data, Kazeem Abiodun Kareem
Dissertations, Master's Theses and Master's Reports
Factor analysis is a powerful tool for modeling latent structures in high-dimensional data, traditional approaches assume a single global structure, limiting their ability to capture heterogeneity. The Mixture of Factor Analyzers (MFA) extends classical factor analysis by modeling data as a mixture of Gaussian-distributed local subspaces, effectively uncovering cluster-specific latent structures. However, MFA relies on Gaussian mixtures, making it sensitive to outliers and ill-suited for heavy-tailed data. The Mixture of $t$-Factor Analyzers (M$t$FA) addresses these limitations by incorporating multivariate $t$-distributions, improving robustness. Despite their advantages, both MFA and M$t$FA face significant computational challenges in high-dimensional settings, particularly due to costly …
Methods In Statistics, Machine Learning, And Deep Learning For Combining Multi-Omics Dataset, Md Mutasim Billah
Methods In Statistics, Machine Learning, And Deep Learning For Combining Multi-Omics Dataset, Md Mutasim Billah
Dissertations, Master's Theses and Master's Reports
Transcriptome-wide association studies (TWAS) have emerged as a powerful strategy to bridge genome-wide association studies (GWAS) with gene regulatory mechanisms by integrating genotypic data with gene expression data. While early TWAS methods typically rely on linear models and single-tissue expression references, recent advances underscore the need for flexible, multi-tissue approaches that can capture heterogeneous regulatory architectures and tissue-specific expression patterns. This dissertation introduces a three‑part research project that advances multi‑tissue transcriptome‑wide association studies (TWAS) along complementary axes of methodology, statistical power, and modelling flexibility.
In chapter One, TWAS‑CTL introduces a two‑stage cross‑tissue learner that trains any user‑chosen single‑tissue imputers (STLs) …