Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Computer Sciences (13)
- Data Science (11)
- Statistical Models (11)
- Artificial Intelligence and Robotics (9)
- Engineering (7)
-
- Probability (7)
- Numerical Analysis and Scientific Computing (6)
- Statistical Methodology (6)
- Applied Mathematics (5)
- Categorical Data Analysis (4)
- Longitudinal Data Analysis and Time Series (4)
- Other Statistics and Probability (4)
- Business (3)
- Business Analytics (3)
- Databases and Information Systems (3)
- Mathematics (3)
- Multivariate Analysis (3)
- Other Applied Mathematics (3)
- Social and Behavioral Sciences (3)
- Theory and Algorithms (3)
- Astrophysics and Astronomy (2)
- Biostatistics (2)
- Computer Engineering (2)
- Econometrics (2)
- Economics (2)
- Electrical and Computer Engineering (2)
- Environmental Monitoring (2)
- Institution
-
- Southern Methodist University (4)
- COBRA (2)
- Georgia Southern University (2)
- The University of Akron (2)
- West Virginia University (2)
-
- Belmont University (1)
- California Polytechnic State University, San Luis Obispo (1)
- City University of New York (CUNY) (1)
- Claremont Colleges (1)
- Embry-Riddle Aeronautical University (1)
- Florida Institute of Technology (1)
- Louisiana State University (1)
- Michigan Technological University (1)
- Murray State University (1)
- Rose-Hulman Institute of Technology (1)
- Universitas Indonesia (1)
- University at Albany, State University of New York (1)
- University of Kentucky (1)
- University of Nebraska at Omaha (1)
- University of North Florida (1)
- Publication
-
- SMU Data Science Review (4)
- College of Graduate Studies: Theses & Dissertations (2)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (2)
- U.C. Berkeley Division of Biostatistics Working Paper Series (2)
- Williams Honors College, Honors Research Projects (2)
-
- Beyond: Undergraduate Research Journal (1)
- CMC Senior Theses (1)
- Dissertations, Master's Theses and Master's Reports (1)
- Dissertations, Theses, and Capstone Projects (1)
- Electronic Theses & Dissertations (2024 - present) (1)
- Honors College Theses (1)
- Journal of Environmental Science and Sustainable Development (1)
- LSU Doctoral Dissertations (1)
- Master's Theses (1)
- Rose-Hulman Undergraduate Mathematics Journal (1)
- SPARK Symposium Presentations (1)
- Theses and Dissertations (1)
- Theses and Dissertations--Mathematics (1)
- Theses/Capstones/Creative Projects (1)
- UNF Graduate Theses and Dissertations (1)
- Publication Type
Articles 1 - 27 of 27
Full-Text Articles in Applied Statistics
Mapping The Water Quality Of Jamaica Bay, New York (1996-2024): Principal Component Analysis And K-Means Clustering, Sneha Srivastava
Mapping The Water Quality Of Jamaica Bay, New York (1996-2024): Principal Component Analysis And K-Means Clustering, Sneha Srivastava
Dissertations, Theses, and Capstone Projects
Jamaica Bay, located along the southeastern coast of New York City, acts as a biodiverse estuary of wetlands, meadows, and salt marsh islands. The purpose of this study is to analyze the water quality conditions of the region over time, comparing locations around the bay to identify hyperlocal features that influence larger trends in the hydrological system. Ten variables were used as water quality indicators, including total Kjeldahl nitrogen, salinity, pH, Secchi disk depth, and total phosphorus, among others, across five stations in the bay, between 1994 and 2024. After data cleaning and standardization methods were applied, principal component analysis …
Bayesball : A Comprehensive Framework For Predicting Ucl Injury, Brady M. Pinter, Will Best Ph.D.
Bayesball : A Comprehensive Framework For Predicting Ucl Injury, Brady M. Pinter, Will Best Ph.D.
SPARK Symposium Presentations
Ulnar Collateral Ligament (UCL) reconstruction, commonly referred to as Tommy John Surgery, has seen a significant rise among Major League Baseball (MLB) pitchers, prompting growing interest in identifying the mechanical and performance-based factors that contribute to injury risk. While previous studies have examined these relationships using traditional frequentist approaches separately, this study combines multiple different model techniques to present a broad framework for finding significant predictors of UCL Surgery. These models include Lasso and Ridge Regression, Principal Component Regression (PCR) , Partial Least Squares Regression (PLS) , Random Forest, Multiple Linear Regression, and a Bayesian Statistical Model. Using these models, …
Beyond The Lace Index: Benchmarking Machine Learning Architectures And Explaining 30-Day Hospital Readmission Risk With Shap Analysis, Carl E. Hughes Iii
Beyond The Lace Index: Benchmarking Machine Learning Architectures And Explaining 30-Day Hospital Readmission Risk With Shap Analysis, Carl E. Hughes Iii
Williams Honors College, Honors Research Projects
Unplanned 30-day hospital readmission remains a fundamental challenge in US healthcare, associated with increased risk to patient recovery and representing an estimated $52.4 billion in annual expenses (Beauvais et al., 2022). While the rigorously validated LACE index serves as the clinical standard for readmission modeling, its linear structure and four explanatory variables lack the complexity to capture the high-dimensional and interactive nature of patient risk. This study utilizes an admission granularity level cohort of the MIMIC-IV database to develop and compare machine learning architectures against the baseline LACE index. Due to the imbalanced prevalence of readmission, the penalized logistic regression, …
Performance Of Numerical Methods Applied To The Black–Scholes Model, Scott Cameron Williams
Performance Of Numerical Methods Applied To The Black–Scholes Model, Scott Cameron Williams
UNF Graduate Theses and Dissertations
We compare five numerical approaches for approximating solutions to the Black–Scholes partial differential equation for pricing European call options: FTCS, BTCS, Crank– Nicolson, Monte Carlo simulation, and a physics–informed neural network (PINN). These methods span finite difference techniques, probabilistic simulation, and machine learning. Performance is evaluated based on computational efficiency and accuracy relative to the analytical Black–Scholes solution.
Among the methods, Crank–Nicolson and the PINN demonstrated the strongest overall performance. Crank–Nicolson achieved the highest accuracy but exhibited increased runtime as the number of underlying stock price grid points grew. In contrast, the PINN produced slightly less accurate results but with …
A Comparative Study Of Neural Networks And Xgboost Models For Flight Time Prediction, Ioannis Paraschos, Taryn E. Trimble, Eshna Bhargava, Jake Klingler, Benjamin R. Nicolai
A Comparative Study Of Neural Networks And Xgboost Models For Flight Time Prediction, Ioannis Paraschos, Taryn E. Trimble, Eshna Bhargava, Jake Klingler, Benjamin R. Nicolai
Beyond: Undergraduate Research Journal
Flight time prediction plays a crucial role in modern air travel, benefiting airlines and passengers alike. Accurate predictions enable airlines to optimize schedules, allocate resources effectively, and ensure passenger safety and satisfaction. In recent years, machine learning models, such as neural networks and XGBoost, have gained popularity for predicting flight times. This study aims to compare the performance of neural network and XGBoost models in predicting flight times, considering factors such as weather conditions, air traffic control, and aircraft performance. The results indicate that both models are effective, with XGBoost achieving slightly higher accuracy. However, neural networks offer advantages in …
Expressive And Interpretable User Engagement Prediction Using Multivariate Survival Processes, Akshay Aravamudan
Expressive And Interpretable User Engagement Prediction Using Multivariate Survival Processes, Akshay Aravamudan
Theses and Dissertations
The ability to characterize how information diffuses online is of paramount importance to stakeholders that are interested in tasks such as proposing solutions for mitigating and countering dis/misinformation, predicting user engagement of content in social media, planning marketing campaigns to roll-out products and planning dissemination of political campaign messaging among others. One such facet of learning the dynamics of information diffusion is the ability to predict user engagement or the popularity of a single piece of information as it spreads through an online medium. Existing works in this regard mainly either obfuscate user level information or utilize frameworks that are …
Mortgage Default Classification Modeling For Variable Analysis, Brendan R. Goggins
Mortgage Default Classification Modeling For Variable Analysis, Brendan R. Goggins
Honors College Theses
The financial crisis of the early 2000’s is a prime example of the severe consequences that mortgage default and borrower insolvency can have on economies at large. Mortgage default specifically is a prime case with the popularization of mortgage backed securities and the commonality of this loan structure. Multiple hypotheses and models have been formed to understand the reasons, causes, and consequences of mortgage default. This paper uses both machine learning and statistical classification models to inform an understanding of the variables most significant and impactful to the default outcome of mortgages. Consideration is given to both loan-level microeconomic variables …
Machine Learning Methods For Intrusion Detection And Response In Network Security, Ayomide Oyemaja
Machine Learning Methods For Intrusion Detection And Response In Network Security, Ayomide Oyemaja
College of Graduate Studies: Theses & Dissertations
Intrusion Detection Systems (IDS) play a crucial role in computer network security by identifying malicious activities and potential cyberattacks. This thesis combines machine learning and cybersecurity by applying Reinforcement Learning (RL) in intrusion detection and response using the NSL-KDD dataset.
We designed and implemented a Q-learning framework where an agent learns to classify network traffic over time by interacting with the environment and receiving rewards based on detection accuracy. We also look at the importance of feature selection and classification techniques and how effective they are in improving model performance, reducing the complexity of computation, and producing more desirable results. …
Impact Of Urban Development On Uv Exposure: A Clustering And Machine Learning Assessment, Taufik Roni Sahroni Mr., Verdi Yasin, Lulut Alfaris, Reza Ariefka, Ruben Cornelius Siagian, Mohammad Alfin Karim, Nana Rahdiana, Ade Suhara
Impact Of Urban Development On Uv Exposure: A Clustering And Machine Learning Assessment, Taufik Roni Sahroni Mr., Verdi Yasin, Lulut Alfaris, Reza Ariefka, Ruben Cornelius Siagian, Mohammad Alfin Karim, Nana Rahdiana, Ade Suhara
Journal of Environmental Science and Sustainable Development
The relocation of Indonesia's capital city is anticipated to promote inclusive economic growth while embracing cultural diversity. However, this transition may affect ultraviolet (UV) radiation exposure patterns. The study investigated variations in UV exposure in the IKN region, focusing on urban development factors such as land use and population density that affect public health, sun protection, and skin cancer prevention. The research hypothesized that UV radiation is significantly correlated with these factors. UV Index data from 2010-2023, a hierarchical clustering method, identifies complex data patterns without determining the number of clusters. XGBoost, a machine learning model, was used for handling …
A Machine Learning Based Approach For The Identification Of Fake Bills, Tianyang Lu, Hongyang Pang
A Machine Learning Based Approach For The Identification Of Fake Bills, Tianyang Lu, Hongyang Pang
Rose-Hulman Undergraduate Mathematics Journal
Fake or counterfeiting currency, which has been around as long as money has existed, is a major economic problem. Since the US dollar is the most popular form of currency globally, it is the most popular currency to counterfeit. The United States Department of Treasury estimates that between $70 million and $200 million in fake bills are in circulation. The Federal Reserve Bank uses special banknote processing systems to count each bill deposited by the bank and examine them for the possibility of counterfeits. These machines have sensors designed to detect general quality of the bills, including paper type, quality …
Ensemble Classification: An Analysis Of The Random Forest Model, Jarod Korn
Ensemble Classification: An Analysis Of The Random Forest Model, Jarod Korn
Williams Honors College, Honors Research Projects
The random forest model proposed by Dr. Leo Breiman in 2001 is an ensemble machine learning method for classification prediction and regression. In the following paper, we will conduct an analysis on the random forest model with a focus on how the model works, how it is applied in software, and how it performs on a set of data. To fully understand the model, we will introduce the concept of decision trees, give a summary of the CART model, explain in detail how the random forest model operates, discuss how the model is implemented in software, demonstrate the model by …
Sparse Representation Learning For Temporal Networks, Maxwell Mcneil
Sparse Representation Learning For Temporal Networks, Maxwell Mcneil
Electronic Theses & Dissertations (2024 - present)
Temporal networks arise in many domains including activity of social network users, sensor network readings over time, and time course gene expression within the interaction network of a model organism. Data of this type contains a wealth of prior information such as the connectivity among nodes (e.g., a friendship graph), and prior knowledge of expected temporal patterns (e.g., periodicity). Modeling these temporal and network patterns jointly is essential for state-of-the-art performance in temporal network data analysis and mining. Sparse dictionary encoding is one modeling approach for such underlying patterns. However, most classical approaches consider only one dimension of the data …
Application Of Distributed Fiber-Optic Sensing For Pressure Predictions And Multiphase Flow Characterization, Gerald Kelechi Ekechukwu
Application Of Distributed Fiber-Optic Sensing For Pressure Predictions And Multiphase Flow Characterization, Gerald Kelechi Ekechukwu
LSU Doctoral Dissertations
In the oil and gas industry, distributed fiber optics sensing (DFOS) has the potential to revolutionize well and reservoir surveillance applications. Using fiber optic sensors is becoming increasingly common because of its chemically passive and non-magnetic interference properties, the possibility of flexible installations that could be behind the casing, on the tubing, or run on wireline, as well as the potential for densely distributed measurements along the entire length of the fiber. The main objectives of my research are to develop and demonstrate novel signal processing and machine learning computational techniques and workflows on DFOS data for a variety of …
Attempting To Predict The Unpredictable: March Madness, Coleton Kanzmeier
Attempting To Predict The Unpredictable: March Madness, Coleton Kanzmeier
Theses/Capstones/Creative Projects
Each year, millions upon millions of individuals fill out at least one if not hundreds of March Madness brackets. People test their luck every year, whether for fun, with friends or family, or to even win some money. Some people rely on their basketball knowledge whereas others know it is called March Madness for a reason and take a shot in the dark. Others have even tried using statistics to give them an edge. I intend to follow a similar approach, using statistics to my advantage. The end goal is to predict this year’s, 2022, March Madness bracket. To achieve …
Reinforcement Learning: Low Discrepancy Action Selection For Continuous States And Actions, Jedidiah Lindborg
Reinforcement Learning: Low Discrepancy Action Selection For Continuous States And Actions, Jedidiah Lindborg
College of Graduate Studies: Theses & Dissertations
In reinforcement learning the process of selecting an action during the exploration or exploitation stage is difficult to optimize. The purpose of this thesis is to create an action selection process for an agent by employing a low discrepancy action selection (LDAS) method. This should allow the agent to quickly determine the utility of its actions by prioritizing actions that are dissimilar to ones that it has already picked. In this way the learning process should be faster for the agent and result in more optimal policies.
Searching For Anomalous Extensive Air Showers Using The Pierre Auger Observatory Fluorescence Detector, Andrew Puyleart
Searching For Anomalous Extensive Air Showers Using The Pierre Auger Observatory Fluorescence Detector, Andrew Puyleart
Dissertations, Master's Theses and Master's Reports
Anomalous extensive air showers have yet to be detected by cosmic ray observatories. Fluorescence detectors provide a way to view the air showers created by cosmic rays with primary energies reaching up to hundreds of EeV . The resulting air showers produced by these highly energetic collisions can contain features that deviate from average air showers. Detection of these anomalous events may provide information into unknown regions of particle physics, and place constraints on cross-sectional interaction lengths of protons. In this dissertation, I propose measurements of extensive air shower profiles that are used in a machine learning pipeline to distinguish …
Combining Machine Learning And Empirical Engineering Methods Towards Improving Oil Production Forecasting, Andrew J. Allen
Combining Machine Learning And Empirical Engineering Methods Towards Improving Oil Production Forecasting, Andrew J. Allen
Master's Theses
Current methods of production forecasting such as decline curve analysis (DCA) or numerical simulation require years of historical production data, and their accuracy is limited by the choice of model parameters. Unconventional resources have proven challenging to apply traditional methods of production forecasting because they lack long production histories and have extremely variable model parameters. This research proposes a data-driven alternative to reservoir simulation and production forecasting techniques. We create a proxy-well model for predicting cumulative oil production by selecting statistically significant well completion parameters and reservoir information as independent predictor variables in regression-based models. Then, principal component analysis (PCA) …
Orthogonal Recurrent Neural Networks And Batch Normalization In Deep Neural Networks, Kyle Eric Helfrich
Orthogonal Recurrent Neural Networks And Batch Normalization In Deep Neural Networks, Kyle Eric Helfrich
Theses and Dissertations--Mathematics
Despite the recent success of various machine learning techniques, there are still numerous obstacles that must be overcome. One obstacle is known as the vanishing/exploding gradient problem. This problem refers to gradients that either become zero or unbounded. This is a well known problem that commonly occurs in Recurrent Neural Networks (RNNs). In this work we describe how this problem can be mitigated, establish three different architectures that are designed to avoid this issue, and derive update schemes for each architecture. Another portion of this work focuses on the often used technique of batch normalization. Although found to be successful …
How Machine Learning And Probability Concepts Can Improve Nba Player Evaluation, Harrison Miller
How Machine Learning And Probability Concepts Can Improve Nba Player Evaluation, Harrison Miller
CMC Senior Theses
In this paper I will be breaking down a scholarly article, written by Sameer K. Deshpande and Shane T. Jensen, that proposed a new method to evaluate NBA players. The NBA is the highest level professional basketball league in America and stands for the National Basketball Association. They proposed to build a model that would result in how NBA players impact their teams chances of winning a game, using machine learning and probability concepts. I preface that by diving into these concepts and their mathematical backgrounds. These concepts include building a linear model using ordinary least squares method, the bias …
Three Essays On Health Economics And Policy Evaluation, Shishir Shakya
Three Essays On Health Economics And Policy Evaluation, Shishir Shakya
Graduate Theses, Dissertations, and Problem Reports (ETD)
This dissertation consists of three essays on the U.S. Health care policy. Each paragraph below refers to the three abstracts for the three chapters in this dissertation, respectively. I provide quantitative evidence on how much Prescription Drug Monitoring Programs (PDMPs) affects the retail opioid prescribing behaviors. Using the American Community Survey (ACS), I retrieve county-level high dimensional panel data set from 2010 to 2017. I employ three separate identification strategies: difference-in-difference, double selection post-LASSO, and spatial difference-in-difference. I compare how the retail opioid prescribing behaviors of counties, that are mandatory for prescribers to check the PDMP before prescribing controlled substances …
Predicting Wind Turbine Blade Erosion Using Machine Learning, Casey Martinez, Festus Asare Yeboah, Scott Herford, Matt Brzezinski, Viswanath Puttagunta
Predicting Wind Turbine Blade Erosion Using Machine Learning, Casey Martinez, Festus Asare Yeboah, Scott Herford, Matt Brzezinski, Viswanath Puttagunta
SMU Data Science Review
Using time-series data and turbine blade inspection assessments, we present a classification model in order to predict remaining turbine blade life in wind turbines. Capturing the kinetic energy of wind requires complex mechanical systems, which require sophisticated maintenance and planning strategies. There are many traditional approaches to monitoring the internal gearbox and generator, but the condition of turbine blades can be difficult to measure and access. Accurate and cost- effective estimates of turbine blade life cycles will drive optimal investments in repairs and improve overall performance. These measures will drive down costs as well as provide cheap and clean electricity …
Visualization And Machine Learning Techniques For Nasa’S Em-1 Big Data Problem, Antonio P. Garza Iii, Jose Quinonez, Misael Santana, Nibhrat Lohia
Visualization And Machine Learning Techniques For Nasa’S Em-1 Big Data Problem, Antonio P. Garza Iii, Jose Quinonez, Misael Santana, Nibhrat Lohia
SMU Data Science Review
In this paper, we help NASA solve three Exploration Mission-1 (EM-1) challenges: data storage, computation time, and visualization of complex data. NASA is studying one year of trajectory data to determine available launch opportunities (about 90TBs of data). We improve data storage by introducing a cloud-based solution that provides elasticity and server upgrades. This migration will save $120k in infrastructure costs every four years, and potentially avoid schedule slips. Additionally, it increases computational efficiency by 125%. We further enhance computation via machine learning techniques that use the classic orbital elements to predict valid trajectories. Our machine learning model decreases trajectory …
Machine Learning Pipeline For Exoplanet Classification, George Clayton Sturrock, Brychan Manry, Sohail Rafiqi
Machine Learning Pipeline For Exoplanet Classification, George Clayton Sturrock, Brychan Manry, Sohail Rafiqi
SMU Data Science Review
Planet identification has typically been a tasked performed exclusively by teams of astronomers and astrophysicists using methods and tools accessible only to those with years of academic education and training. NASA’s Exoplanet Exploration program has introduced modern satellites capable of capturing a vast array of data regarding celestial objects of interest to assist with researching these objects. The availability of satellite data has opened up the task of planet identification to individuals capable of writing and interpreting machine learning models. In this study, several classification models and datasets are utilized to assign a probability of an observation being an exoplanet. …
Improving Vix Futures Forecasts Using Machine Learning Methods, James Hosker, Slobodan Djurdjevic, Hieu Nguyen, Robert Slater
Improving Vix Futures Forecasts Using Machine Learning Methods, James Hosker, Slobodan Djurdjevic, Hieu Nguyen, Robert Slater
SMU Data Science Review
The problem of forecasting market volatility is a difficult task for most fund managers. Volatility forecasts are used for risk management, alpha (risk) trading, and the reduction of trading friction. Improving the forecasts of future market volatility assists fund managers in adding or reducing risk in their portfolios as well as in increasing hedges to protect their portfolios in anticipation of a market sell-off event. Our analysis compares three existing financial models that forecast future market volatility using the Chicago Board Options Exchange Volatility Index (VIX) to six machine/deep learning supervised regression methods. This analysis determines which models provide best …
Quantifying Human Biological Age: A Machine Learning Approach, Syed Ashiqur Rahman
Quantifying Human Biological Age: A Machine Learning Approach, Syed Ashiqur Rahman
Graduate Theses, Dissertations, and Problem Reports (ETD)
Quantifying human biological age is an important and difficult challenge. Different biomarkers and numerous approaches have been studied for biological age prediction, each with its advantages and limitations. In this work, we first introduce a new anthropometric measure (called Surface-based Body Shape Index, SBSI) that accounts for both body shape and body size, and evaluate its performance as a predictor of all-cause mortality. We analyzed data from the National Health and Human Nutrition Examination Survey (NHANES). Based on the analysis, we introduce a new body shape index constructed from four important anthropometric determinants of body shape and body size: body …
A Scalable Supervised Subsemble Prediction Algorithm, Stephanie Sapp, Mark J. Van Der Laan
A Scalable Supervised Subsemble Prediction Algorithm, Stephanie Sapp, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Subsemble is a flexible ensemble method that partitions a full data set into subsets of observations, fits the same algorithm on each subset, and uses a tailored form of V-fold cross-validation to construct a prediction function that combines the subset-specific fits with a second metalearner algorithm. Previous work studied the performance of Subsemble with subsets created randomly, and showed that these types of Subsembles often result in better prediction performance than the underlying algorithm fit just once on the full dataset. Since the final Subsemble estimator varies depending on the data used to create the subset-specific fits, different strategies for …
Subsemble: An Ensemble Method For Combining Subset-Specific Algorithm Fits, Stephanie Sapp, Mark J. Van Der Laan, John Canny
Subsemble: An Ensemble Method For Combining Subset-Specific Algorithm Fits, Stephanie Sapp, Mark J. Van Der Laan, John Canny
U.C. Berkeley Division of Biostatistics Working Paper Series
Ensemble methods using the same underlying algorithm trained on different subsets of observations have recently received increased attention as practical prediction tools for massive datasets. We propose Subsemble: a general subset ensemble prediction method, which can be used for small, moderate, or large datasets. Subsemble partitions the full dataset into subsets of observations, fits a specified underlying algorithm on each subset, and uses a clever form of V-fold cross-validation to output a prediction function that combines the subset-specific fits. We give an oracle result that provides a theoretical performance guarantee for Subsemble. Through simulations, we demonstrate that Subsemble can be …