Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- COBRA (107)
- Central Bank of Nigeria (26)
- University of Kentucky (22)
- Southern Methodist University (15)
- Virginia Commonwealth University (10)
-
- Georgia Southern University (9)
- University of Louisville (8)
- Clemson University (7)
- Michigan Technological University (7)
- University of Central Florida (7)
- Kennesaw State University (6)
- University of Arkansas, Fayetteville (6)
- City University of New York (CUNY) (5)
- Technological University Dublin (5)
- The University of Akron (5)
- University of Nevada, Las Vegas (5)
- The Texas Medical Center Library (4)
- University of Nebraska - Lincoln (4)
- California Polytechnic State University, San Luis Obispo (3)
- East Tennessee State University (3)
- James Madison University (3)
- Marshall University (3)
- University of New Hampshire (3)
- University of New Mexico (3)
- University of Texas Rio Grande Valley (3)
- Western Michigan University (3)
- Bucknell University (2)
- Chapman University (2)
- Claremont Colleges (2)
- Dartmouth College (2)
- Keyword
-
- Statistics (21)
- Regression (10)
- Machine Learning (8)
- Simulation (7)
- Bayesian inference (6)
-
- COVID-19 (6)
- Causal inference (6)
- Counting process (6)
- Model selection (6)
- Estimating equation (5)
- Machine learning (5)
- Prediction (5)
- Survival analysis (5)
- Censored data (4)
- Counterfactual (4)
- Cross-validation (4)
- Genetics (4)
- Linear regression (4)
- Logistic regression (4)
- Missing Data (4)
- Quantile regression (4)
- Semiparametric model (4)
- Variable Selection (4)
- Bayesian Statistics (3)
- Bayesian analysis (3)
- Bias (3)
- Censoring (3)
- Classification (3)
- Confidence intervals (3)
- Confounding (3)
- Publication Year
- Publication
-
- U.C. Berkeley Division of Biostatistics Working Paper Series (38)
- Harvard University Biostatistics Working Paper Series (27)
- CBN Journal of Applied Statistics (JAS) (26)
- Theses and Dissertations--Statistics (21)
- The University of Michigan Department of Biostatistics Working Paper Series (15)
-
- Electronic Theses and Dissertations (12)
- Theses and Dissertations (12)
- Johns Hopkins University, Dept. of Biostatistics Working Papers (10)
- COBRA Preprint Series (9)
- College of Graduate Studies: Theses & Dissertations (8)
- Statistical Science Theses and Dissertations (8)
- All Dissertations (7)
- Dissertations, Master's Theses and Master's Reports (7)
- SMU Data Science Review (7)
- UW Biostatistics Working Paper Series (7)
- Data Science and Data Mining (6)
- Articles (5)
- Graduate Theses and Dissertations (5)
- Williams Honors College, Honors Research Projects (5)
- Reactor Campaign (TRP) (4)
- Dissertations (3)
- Dissertations and Theses (Open Access) (3)
- Published and Grey Literature from PhD Candidates (3)
- Theses, Dissertations and Capstones (3)
- Al-Bahir (2)
- CHIP Documents (2)
- Conference papers (2)
- Dartmouth College Ph.D Dissertations (2)
- Dissertations, 2014-2019 (2)
- Dissertations, Theses, and Capstone Projects (2)
- Publication Type
Articles 31 - 60 of 350
Full-Text Articles in Statistical Models
Opening The Black Box With Regal: A Novel Explainable Ai Approach To Uncover Key Predictors In Search And Rescue Success, Brandon Hyunjun Kim
Opening The Black Box With Regal: A Novel Explainable Ai Approach To Uncover Key Predictors In Search And Rescue Success, Brandon Hyunjun Kim
Master's Theses
The outcome of a search and rescue (SAR) operation is influenced by a complex, non-linear interplay among numerous factors, including geographic context, subject-specific characteristics, and environmental conditions. The high dimensionality and intricate dependencies among these variables pose significant challenges to traditional exploratory modeling approaches, limiting their ability to uncover meaningful patterns and relationships associated with mission success. This study introduces Rules Based Explanations for Generated neighborhoods Around Localized cases (REGAL), a novel adaptation of the Local Interpretable Model-agnostic Explanations (LIME) framework to explain deep multimodal neural networks and what key features it assesses to determine search and rescue success. REGAL …
Innovative Methods For The Design And Analysis Of Phase Ii Clinical Trials, Feng Tian
Innovative Methods For The Design And Analysis Of Phase Ii Clinical Trials, Feng Tian
Dissertations and Theses (Open Access)
Drug development has become increasingly time-consuming, costly, and risky in recent years. There is significant potential for improving clinical trial designs, particularly for phase II trials, which play a critical role in the drug development process. Innovative methods are especially necessary for addressing key challenges in phase II trials in terms of dose-ranging study, patient population selection, and decentralized clinical trials (DCTs). This dissertation presents a comprehensive set of methodologies that address these critical issues with three projects. The first project introduces a Bayesian adaptive dose-ranging design that integrates both efficacy and toxicity data to evaluate each dose comprehensively. The …
Discounting Effect Size When Borrowing External Data In Clinical Studies, Zhuanzhuan Ma, Chul Ahn, Bin Wang, Xuefeng Li
Discounting Effect Size When Borrowing External Data In Clinical Studies, Zhuanzhuan Ma, Chul Ahn, Bin Wang, Xuefeng Li
Research Symposium
Background: When borrowing information from external data to augment a current trial, many available methods discount the sample size but retain the effect size from previous studies. Discounting the sample size is just one way to discount the prior information. It may not be appropriate if the underlying assumption of unbiased treatment effect does not hold, for example, when the treatment effect in the historical study is likely higher than the one expected in the current trial.
Methods: To tackle this potential issue, we study some methods to shrink the effect size from previous studies assuming that the prior effect …
Sparse Bayesian Variable Selection Using Global-Local Shrinkage Priors For The Analysis Of Cancer Datasets, Zhuanzhuan Ma
Sparse Bayesian Variable Selection Using Global-Local Shrinkage Priors For The Analysis Of Cancer Datasets, Zhuanzhuan Ma
Research Symposium
Background: With a rapid development of data collection technology, high dimensional data, whose model dimension k may be growing or much larger than the sample size n, is becoming increasingly prevalent in different fields of study, such as ecology, genetics, among others. This data deluge is introducing new challenges to traditional statistical procedures and theories and is thus generating a renewed interest in the problems of variable selection and classification in high dimensional regression models. In large k, small n settings, variable selection is usually the first step for dimension reduction to uncover significant covariates, which contribute to …
Predicting Superconducting Critical Temperature From Composition-Derived Features: A Transparent Linear And Regularized Regression Study, Md Ahiduzzaman
Predicting Superconducting Critical Temperature From Composition-Derived Features: A Transparent Linear And Regularized Regression Study, Md Ahiduzzaman
Data Science and Data Mining
We study prediction of superconducting critical temperature (Tc) from 81 composition-derived descriptors across 21,263 materials. To keep the analysis transparent and repro- ducible, we focus on linear models: Ordinary Least Squares (OLS), Ridge, Lasso, and Elastic Net (ENet). All models share a single evaluation protocol (5-fold cross-validation with standardized inputs) and are compared on RMSE, MAE, and R2. On this feature set, OLS attains the best cross-validated performance (RMSE = 17.6 K, MAE = 13.3 K , R2 = 0.735), with Lasso/ENet essentially tied next (RMSE ≈ 17.7 K , R2 ≈ 0.734); Ridge underperforms (RMSE = 18.9 K , …
Comparative Analysis Of Lasso, Ridge, And Elastic Net For Variable Selection In High-Dimensional Maize Data, Md Ahiduzzaman
Comparative Analysis Of Lasso, Ridge, And Elastic Net For Variable Selection In High-Dimensional Maize Data, Md Ahiduzzaman
Data Science and Data Mining
In high-dimensional genomic data analysis, traditional linear regression techniques often struggle due to the presence of a large number of predictor variables relative to observations. Penalized regression methods such as LASSO, Ridge, and Elastic Net have emerged as effective solutions by imposing regularization, which helps in managing multicollinearity and enhancing prediction accuracy. This study applies these techniques to the Maize dataset to model the time to male flowering, selecting relevant genetic markers as predictors. Our findings suggest that Elastic Net is particularly effective for high-dimensional data with correlated variables, achieving a balance between prediction accuracy and variable selection. The results …
Beyond Homogeneity: Exploring Causal Heterogeneity In Psychopathology, Philip B. Vinh
Beyond Homogeneity: Exploring Causal Heterogeneity In Psychopathology, Philip B. Vinh
Theses and Dissertations
Traditional models in psychiatric research often impose assumptions of causal homogeneity, treating population-level associations as reflective of uniform underlying mechanisms. This dissertation challenges that assumption by introducing statistical and machine learning frameworks designed to detect and model causal heterogeneity in the development of psychopathology. Central to this approach is the advancement of finite mixture structural equation modeling (FM-SEM) to identify latent subgroups characterized by distinct, and sometimes opposing, causal pathways.
The dissertation comprises three integrated empirical studies. The first introduces mixDoC, a finite mixture extension of the classical Direction of Causation (DoC) model applied to twin data, enabling the detection …
Hybrid Mixtures Of Factor Analyzers For High Dimensional Data, Kazeem Abiodun Kareem
Hybrid Mixtures Of Factor Analyzers For High Dimensional Data, Kazeem Abiodun Kareem
Dissertations, Master's Theses and Master's Reports
Factor analysis is a powerful tool for modeling latent structures in high-dimensional data, traditional approaches assume a single global structure, limiting their ability to capture heterogeneity. The Mixture of Factor Analyzers (MFA) extends classical factor analysis by modeling data as a mixture of Gaussian-distributed local subspaces, effectively uncovering cluster-specific latent structures. However, MFA relies on Gaussian mixtures, making it sensitive to outliers and ill-suited for heavy-tailed data. The Mixture of $t$-Factor Analyzers (M$t$FA) addresses these limitations by incorporating multivariate $t$-distributions, improving robustness. Despite their advantages, both MFA and M$t$FA face significant computational challenges in high-dimensional settings, particularly due to costly …
Methods In Statistics, Machine Learning, And Deep Learning For Combining Multi-Omics Dataset, Md Mutasim Billah
Methods In Statistics, Machine Learning, And Deep Learning For Combining Multi-Omics Dataset, Md Mutasim Billah
Dissertations, Master's Theses and Master's Reports
Transcriptome-wide association studies (TWAS) have emerged as a powerful strategy to bridge genome-wide association studies (GWAS) with gene regulatory mechanisms by integrating genotypic data with gene expression data. While early TWAS methods typically rely on linear models and single-tissue expression references, recent advances underscore the need for flexible, multi-tissue approaches that can capture heterogeneous regulatory architectures and tissue-specific expression patterns. This dissertation introduces a three‑part research project that advances multi‑tissue transcriptome‑wide association studies (TWAS) along complementary axes of methodology, statistical power, and modelling flexibility.
In chapter One, TWAS‑CTL introduces a two‑stage cross‑tissue learner that trains any user‑chosen single‑tissue imputers (STLs) …
“Regression To The Mean”: The Confluence Of Eugenics And Statistics In The 19th And 20th Centuries, Emrys G. King
“Regression To The Mean”: The Confluence Of Eugenics And Statistics In The 19th And 20th Centuries, Emrys G. King
Pomona Senior Theses
The work of this thesis is twofold — first, qualitatively characterizing the confluence between the British eugenics and statistics movements in the late 19th and early 20th centuries, and second, quantitatively analyzing the effect of this foundation on pedagogical materials in the growing field of statistics between 1880 and 1970. Towards the first goal, the history of the method of least squares, state statistics, and positive and negative eugenics are outlined, followed by a close reading of the foundational texts authored by Francis Galton and Karl Pearson that introduced linear regression. Towards the latter goal, English-language statistics textbooks published between …
Theoretical Foundations And Applied Performance Of Periodicity-Aware Imputation: Variable Bandpass Block Bootstrap Methods For Incomplete Time Series, Asmaa Ahmad
Electronic Theses & Dissertations (2024 - present)
Time series data are prevalent across a wide range of disciplines, including health surveillance, public policy, and environmental monitoring. In the presence of underlying cyclical patterns, the integrity of time series analysis depends critically on the ability to detect, model, and impute structured missing data without compromising the temporal structure. This dissertation introduces and validates a novel imputation framework that integrates the Variable Bandpass Periodic Block Bootstrap (VBPBB) into multiple imputation procedures, improving the accuracy, robustness, and interpretability of time series models under high rates of missingness and noise. The overarching goal of this dissertation was to develop and evaluate …
Safeguard Cyberspace In Ransomware Era: Risk Analysis & Cyber Insurance, Li Huang
Safeguard Cyberspace In Ransomware Era: Risk Analysis & Cyber Insurance, Li Huang
Electronic Theses & Dissertations (2024 - present)
The increasing frequency and severity of ransomware attacks pose significant challenges for organizational cybersecurity. Fragmentation across disciplines in cyber defense has created practical gaps in the development of the necessary capabilities needed to address rapidly evolving cyber threats. This study explores the impact of ransomware attacks and the evolving role of cyber insurance as a proactive cybersecurity partner. Bridging the gap between actuarial science and cyber risk management, it proposes an interdisciplinary framework that quantifies the impact of ransomware and integrates cyber insurance into cybersecurity strategies.
The primary contribution of this study is methodology. We present a framework that remains …
The Energy Efficiency Price Premium Of Residential Buildings In Three Italian Regions, Elena Giarda, Demetrio Panarello
The Energy Efficiency Price Premium Of Residential Buildings In Three Italian Regions, Elena Giarda, Demetrio Panarello
CBER Conference
The aim of this paper is to investigate whether a higher energy efficiency of residential buildings translates into higher house prices in Italy. We employ novel, and almost unexploited, data on Energy Performance Certificates of three Italian regions (EmiliaRomagna, Lombardy and Piedmont) and merge them with house prices and socioeconomic variables at various aggregation levels. The relationship between house prices and energy efficiency is estimated by means of hedonic regression models, quantile regressions and fixed effects panel data models. Our results reveal the existence of an energy-efficiency price premium in the three regions, with significant differences among them. Heterogeneity is …
Bayesian Approaches In Multi-State Markov Models And High Dimensional Time-To-Event Data., Yuchen Han
Bayesian Approaches In Multi-State Markov Models And High Dimensional Time-To-Event Data., Yuchen Han
Electronic Theses and Dissertations
This dissertation consists of two projects. The first one involves nonparametric methods on Continuous Time Markov Chains (CTMCs). The second one is centered around Bayesian shrinkage models for detecting prognostic and predictive biomarkers in high-dimensional clinical data. Both these projects build on methods from across the frequentist and Bayesian paradigm to offer novel solutions. In the first project, we aim to model the nonlinear effects of continuous variables within multistate framework in a non-parametrically by appealing to the rich mathematical framework of Reproducing Kernel Hilbert Spaces (RKHS). Then we adapted the classical Representer Theorem to penalized (squared norm) log-likelihood which …
Forecasting Commercial Vehicle Miles Traveled (Vmt) In Urban California Areas, Steve Chung, Jaymin Kwon, Yushin Ahn
Forecasting Commercial Vehicle Miles Traveled (Vmt) In Urban California Areas, Steve Chung, Jaymin Kwon, Yushin Ahn
Mineta Transportation Institute
This study investigates commercial truck vehicle miles traveled (VMT) across six diverse California counties from 2000 to 2020. The counties—Imperial, Los Angeles, Riverside, San Bernardino, San Diego, and San Francisco—represent a broad spectrum of California’s demographics, economies, and landscapes. Using a rich dataset spanning demographics, economics, and pollution variables, we aim to understand the factors influencing commercial VMT. We first visually represent the geographic distribution of the counties, highlighting their unique characteristics. Linear regression models, particularly the least absolute shrinkage and selection operator (LASSO) and elastic net regressions are employed to identify key predictors of total commercial VMT. LASSO regression …
Value Added Tax Rate Variation, Import Demand And Sectoral Output In Nigeria, Joshua K. Nomkuha, Aondoawase Asooso, Philip T. Abachi
Value Added Tax Rate Variation, Import Demand And Sectoral Output In Nigeria, Joshua K. Nomkuha, Aondoawase Asooso, Philip T. Abachi
CBN Journal of Applied Statistics (JAS)
This study employs computable general equilibrium (CGE) model to estimate the effect of increase in value added tax (VAT), from 5 per cent to 7.5 per cent, on import demand and sectoral output in Nigeria. The study uses 2020 as the base year for the data analysis. The results show that increase in VAT affects import demand negatively, based on import penetration ratios, with mixed effect across six sectors. The implication of the result is that the VAT policy discourage consumption of foreign products, and constitute excess burden to consumers of such products in Nigeria. The results further reveal that …
Trade Liberalization, Non-Oil Export And Economic Growth In Nigeria, Jerome T. Andohol, Terhemen Tarzoor, Dennis T. Nomor
Trade Liberalization, Non-Oil Export And Economic Growth In Nigeria, Jerome T. Andohol, Terhemen Tarzoor, Dennis T. Nomor
CBN Journal of Applied Statistics (JAS)
The study examines the impact of trade liberalization and non-oil exports on economic growth in Nigeria from 1986 to 2021. The study utilizes an autoregressive distributed lag model and found the combined effect of trade liberalization and non-oil exports to be positive and statistical significant. While trade liberalization alone may have negative consequences, its synergy with a robust non-oil export can drive sustainable economic growth. The study recommends that strategies to enhance non-oil exports should be encouraged to support the effectiveness of trade liberalization in promoting growth.
Stock Market Volatility In The United Kingdom: Simulating Post-Covid-19 Recovery, Bala A. Dahiru, Mohammed Shuaibu, Najibullah Hassanov
Stock Market Volatility In The United Kingdom: Simulating Post-Covid-19 Recovery, Bala A. Dahiru, Mohammed Shuaibu, Najibullah Hassanov
CBN Journal of Applied Statistics (JAS)
This paper investigates the time it would take for the FTSE-100 index to reach its post-COVID-19 peak. The paper utilises an exponential generalised autoregressive conditional heteroscedasticity (EGARCH) model that accounts for leverage effect and asymmetries. The preferred models amongst competing variants was the Autoregressive Moving Average (ARMA)-EGARCH(2,1) specification and was used to predict daily FTSE-100 data from 5th January 2000 to 21st June 2024. The empirical exercise showed that the COVID-19-induced financial crisis negatively affected the United Kingdom’s stock market performance. The results show that the FTSE100 index could reach its post-pandemic peak around 27th August, 2024 (two months after …
A Novel Correction For The Multivariate Ljung-Box Test, Minhao Huang
A Novel Correction For The Multivariate Ljung-Box Test, Minhao Huang
Computational and Data Sciences (PhD) Dissertations
This research introduces an analytical improvement to the Multivariate Ljung-Box test that addresses significant deviations of the original test from the nominal Type I error rates under almost all scenarios. Prior attempts to mitigate this issue have been directed at modification of the test statistics or correction of the test distribution to achieve precise results in finite samples. In previous studies, focused on designing corrections to the univariate Ljung-Box, a method that specifically adjusts the test rejection region has been the most successful of attaining the best Type I error rates. We adopt the same approach for the more complex, …
Representation Learning For Generative Models With Applications To Healthcare, Astronautics, And Aviation, Van Minh Nguyen
Representation Learning For Generative Models With Applications To Healthcare, Astronautics, And Aviation, Van Minh Nguyen
Theses and Dissertations
This dissertation explores applications of representation learning and generative models to challenges in healthcare, astronautics, and aviation.
The first part investigates the use of Generative Adversarial Networks (GANs) to synthesize realistic electronic health record (EHR) data. An initial attempt at training a GAN on the MIMIC-IV dataset encountered stability and convergence issues, motivating a deeper study of 1-Lipschitz regularization techniques for Auxiliary Classifier GANs (AC-GANs). An extensive ablation study on the CIFAR-10 dataset found that Spectral Normalization is key for AC-GAN stability and performance, while Weight Clipping fails to converge without Spectral Normalization. Analysis of the training dynamics provided further …
Code For Care: Hypertension Prediction In Women Aged 18-39 Years, Kruti Sheth
Code For Care: Hypertension Prediction In Women Aged 18-39 Years, Kruti Sheth
Electronic Theses, Projects, and Dissertations
The longstanding prevalence of hypertension, often undiagnosed, poses significant risks of severe chronic and cardiovascular complications if left untreated. This study investigated the causes and underlying risks of hypertension in females aged between 18-39 years. The research questions were: (Q1.) What factors affect the occurrence of hypertension in females aged 18-39 years? (Q2.) What machine learning algorithms are suited for effectively predicting hypertension? (Q3.) How can SHAP values be leveraged to analyze the factors from model outputs? The findings are: (Q1.) Performing Feature selection using binary classification Logistic regression algorithm reveals an array of 30 most influential factors at an …
Assessment Of Method Effects Of Keying And Wording In Instruments: A Mixed-Methods Explanatory Sequential Study, Lin Ma
Electronic Theses and Dissertations
This dissertation presents an innovative approach to examining the keying method, wording method, and construct validity on psychometric instruments. By employing a mixed methods explanatory sequential design, the effects of keying and wording in two psychometric assessments were examined and validated. Those two self-report psychometric assessments were the Effortful Control assessment (Ellis & Rothbart, 2001) and the Grit assessment (Duckworth & Quinn, 2009). Moreover, the quantitative phase utilized structural equation modeling to analyze 2,104 students’ responses and assess the construct of keying and wording. Various hypothetical models were investigated and evaluated. The reliability of each construct in each method was …
Session 6: The Size-Biased Lognormal Mixture With The Entropy Regularized Algorithm, Tatjana Miljkovic, Taehan Bae
Session 6: The Size-Biased Lognormal Mixture With The Entropy Regularized Algorithm, Tatjana Miljkovic, Taehan Bae
SDSU Data Science Symposium
A size-biased left-truncated Lognormal (SB-ltLN) mixture is proposed as a robust alternative to the Erlang mixture for modeling left-truncated insurance losses with a heavy tail. The weak denseness property of the weighted Lognormal mixture is studied along with the tail behavior. Explicit analytical solutions are derived for moments and Tail Value at Risk based on the proposed model. An extension of the regularized expectation–maximization (REM) algorithm with Shannon's entropy weights (ewREM) is introduced for parameter estimation and variability assessment. The left-truncated internal fraud data set from the Operational Riskdata eXchange is used to illustrate applications of the proposed model. Finally, …
Sparse Bayesian Variable Selection In High‐Dimensional Logistic Regression Models With Correlated Priors, Zhuanzhuan Ma, Zifei Han, Souparno Ghosh, Liucang Wu, Min Wang
Sparse Bayesian Variable Selection In High‐Dimensional Logistic Regression Models With Correlated Priors, Zhuanzhuan Ma, Zifei Han, Souparno Ghosh, Liucang Wu, Min Wang
School of Mathematical & Statistical Sciences Faculty Publications
In this paper, we propose a sparse Bayesian procedure with global and local(GL) shrinkage priors for the problems of variable selection and classification in high-dimensional logistic regression models. In particular, we consider two types of GL shrinkage priors for the regression coefficients, the horseshoe (HS)prior and the normal-gamma (NG) prior, and then specify a correlated prior for the binary vector to distinguish models with the same size. The GL priors are then combined with mixture representations of logistic distribution to construct a hierarchical Bayes model that allows efficient implementation of a Markov chain Monte Carlo (MCMC) to generate samples from …
Predicting Superconducting Critical Temperature Using Regression Analysis, Roland Fiagbe
Predicting Superconducting Critical Temperature Using Regression Analysis, Roland Fiagbe
Data Science and Data Mining
This project estimates a regression model to predict the superconducting critical temperature based on variables extracted from the superconductor’s chemical formula. The regression model along with the stepwise variable selection gives a reasonable and good predictive model with a lower prediction error (MSE). Variables extracted based on atomic radius, valence, atomic mass and thermal conductivity appeared to have the most contribution to the predictive model.
Defensive Impact Wins: Developing A New Method To Rate Individual Defense In Nba Games, Dylan J. Stiles
Defensive Impact Wins: Developing A New Method To Rate Individual Defense In Nba Games, Dylan J. Stiles
Honors Theses and Capstones
With the analytics revolution in sports in the past 20 years, it seems that everything that can be quantified is. In basketball though, trying to break the game down into a set of numbers comes with a unique problem. While we've come up with a good set of advanced numbers to measure offensive efficiency, defense is fundamentally harder to quantify. The game is played five on five, but it has often been popular or convenient to model defense as a set of five one on one games. As defenses became more complex into the 2010s, this methodology became more insignificant. …
Machine Learning Approaches For Cyberbullying Detection, Roland Fiagbe
Machine Learning Approaches For Cyberbullying Detection, Roland Fiagbe
Data Science and Data Mining
Cyberbullying refers to the act of bullying using electronic means and the internet. In recent years, this act has been identifed to be a major problem among young people and even adults. It can negatively impact one’s emotions and lead to adverse outcomes like depression, anxiety, harassment, and suicide, among others. This has led to the need to employ machine learning techniques to automatically detect cyberbullying and prevent them on various social media platforms. In this study, we want to analyze the combination of some Natural Language Processing (NLP) algorithms (such as Bag-of-Words and TFIDF) with some popular machine learning …
Variable Selection For High-Dimensional Data With Interaction Effects: Methods, Applications, And Inferences, Leiyue Li
Theses and Dissertations--Statistics
For high-dimensional data where the number of variables greatly exceeds the number of observations, selecting important variables while maintaining the required heredity conditions can be challenging. This dissertation is structured into three interconnected parts. In the first part, we propose a variable selection method by implementing a well-known optimization technique, the Genetic Algorithm. An R package was developed to simplify the implementation and usage of the proposed method. We then propose another variable selection method by extending the study from the Genetic Algorithm to a different but related optimization technique, Simulated Annealing. We consider three different hierarchical structures in both …
Imputation Strategies For Different Categories Of Missing Data, Karthik Chalumuri
Imputation Strategies For Different Categories Of Missing Data, Karthik Chalumuri
Honors Theses and Capstones
Addressing missing data in research is crucial for ensuring the reliability and validity of study findings, yet it remains a significant challenge. This study investigates the impact of missing data on research outcomes and explores the underutilization of existing tools for managing missingness, potentially leading to gaps in critical information with tangible implications for decision-making processes (Dziura et al.).
Focusing on the different categories of missing data—Missing Completely At Random (MCAR), Missing At Random (MAR), and Missing Not At Random (MNAR)—this research examines various imputation strategies tailored to each category. Specifically, we compare the efficacy of several model-based imputation methods, …
Ensemble Classification: An Analysis Of The Random Forest Model, Jarod Korn
Ensemble Classification: An Analysis Of The Random Forest Model, Jarod Korn
Williams Honors College, Honors Research Projects
The random forest model proposed by Dr. Leo Breiman in 2001 is an ensemble machine learning method for classification prediction and regression. In the following paper, we will conduct an analysis on the random forest model with a focus on how the model works, how it is applied in software, and how it performs on a set of data. To fully understand the model, we will introduce the concept of decision trees, give a summary of the CART model, explain in detail how the random forest model operates, discuss how the model is implemented in software, demonstrate the model by …