Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Computer Sciences (17)
- Data Science (14)
- Artificial Intelligence and Robotics (12)
- Applied Statistics (11)
- Statistical Methodology (8)
-
- Categorical Data Analysis (6)
- Engineering (6)
- Mathematics (6)
- Business (5)
- Databases and Information Systems (5)
- Numerical Analysis and Scientific Computing (5)
- Probability (5)
- Computer Engineering (4)
- Medicine and Health Sciences (4)
- Other Statistics and Probability (4)
- Social and Behavioral Sciences (4)
- Applied Mathematics (3)
- Finance and Financial Management (3)
- Longitudinal Data Analysis and Time Series (3)
- Programming Languages and Compilers (3)
- Survival Analysis (3)
- Theory and Algorithms (3)
- Analysis (2)
- Biostatistics (2)
- Business Analytics (2)
- Cybersecurity (2)
- Data Storage Systems (2)
- Institution
-
- Southern Methodist University (5)
- Claremont Colleges (2)
- Purdue University (2)
- The University of Akron (2)
- Utah State University (2)
-
- Belmont University (1)
- California Polytechnic State University, San Luis Obispo (1)
- California State University, San Bernardino (1)
- Chapman University (1)
- Dartmouth College (1)
- East Tennessee State University (1)
- Florida Institute of Technology (1)
- Institute of Business Administration (1)
- Kennesaw State University (1)
- LSU New Orleans (1)
- Minnesota State University Moorhead (1)
- Mississippi State University (1)
- Murray State University (1)
- Portland State University (1)
- Rose-Hulman Institute of Technology (1)
- University at Albany, State University of New York (1)
- University of Arkansas, Fayetteville (1)
- University of North Florida (1)
- West Virginia University (1)
- Publication
-
- SMU Data Science Review (5)
- CMC Senior Theses (2)
- Williams Honors College, Honors Research Projects (2)
- All Graduate Plan B and other Reports, Spring 1920 to Spring 2023 (1)
- All Graduate Theses and Dissertations, Spring 1920 to Summer 2023 (1)
-
- CBER Conference (1)
- Capstone Projects (1)
- Computational and Data Sciences (PhD) Dissertations (1)
- Dartmouth College Ph.D Dissertations (1)
- Data Science Undergraduate Honors Theses (1)
- Discovery Undergraduate Interdisciplinary Research Internship (1)
- Dissertations, Theses, and Projects (1)
- Doctor of Data Science and Analytics Dissertations (1)
- Electronic Theses & Dissertations (2024 - present) (1)
- Electronic Theses and Dissertations (1)
- Electronic Theses, Projects, and Dissertations (1)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (1)
- Honors College Theses (1)
- LSU New Orleans Theses and Dissertations (1)
- Master's Theses (1)
- Rose-Hulman Undergraduate Mathematics Journal (1)
- SPARK Symposium Presentations (1)
- The Summer Undergraduate Research Fellowship (SURF) Symposium (1)
- Theses and Dissertations (1)
- UNF Graduate Theses and Dissertations (1)
- University Honors Theses (1)
- Publication Type
Articles 1 - 30 of 32
Full-Text Articles in Statistical Models
Evaluating Machine Learning Models On Classification Of Novel Cyber Attacks In The Healthcare Domain, Promise Ehimen
Evaluating Machine Learning Models On Classification Of Novel Cyber Attacks In The Healthcare Domain, Promise Ehimen
Dissertations, Theses, and Projects
The increasing adoption of the Internet of Medical Things (IoMT) has improved healthcare delivery through connected medical devices while simultaneously expanding the cybersecurity risks facing healthcare organizations. Although machine learning based intrusion detection systems have demonstrated high detection accuracy, their ability to respond reliably to previously unseen cyberattacks remains uncertain. This study investigated how a Neural Network model and a Logistic Regression model classified novel cyberattacks within the IoMT environment. The Neural Network and Logistic Regression models were both trained and tested using a subset of the CICIoMT2024 benchmark dataset. The Neural Network achieved 99.82% test accuracy and a 0.94 …
Bayesball : A Comprehensive Framework For Predicting Ucl Injury, Brady M. Pinter, Will Best Ph.D.
Bayesball : A Comprehensive Framework For Predicting Ucl Injury, Brady M. Pinter, Will Best Ph.D.
SPARK Symposium Presentations
Ulnar Collateral Ligament (UCL) reconstruction, commonly referred to as Tommy John Surgery, has seen a significant rise among Major League Baseball (MLB) pitchers, prompting growing interest in identifying the mechanical and performance-based factors that contribute to injury risk. While previous studies have examined these relationships using traditional frequentist approaches separately, this study combines multiple different model techniques to present a broad framework for finding significant predictors of UCL Surgery. These models include Lasso and Ridge Regression, Principal Component Regression (PCR) , Partial Least Squares Regression (PLS) , Random Forest, Multiple Linear Regression, and a Bayesian Statistical Model. Using these models, …
Beyond The Lace Index: Benchmarking Machine Learning Architectures And Explaining 30-Day Hospital Readmission Risk With Shap Analysis, Carl E. Hughes Iii
Beyond The Lace Index: Benchmarking Machine Learning Architectures And Explaining 30-Day Hospital Readmission Risk With Shap Analysis, Carl E. Hughes Iii
Williams Honors College, Honors Research Projects
Unplanned 30-day hospital readmission remains a fundamental challenge in US healthcare, associated with increased risk to patient recovery and representing an estimated $52.4 billion in annual expenses (Beauvais et al., 2022). While the rigorously validated LACE index serves as the clinical standard for readmission modeling, its linear structure and four explanatory variables lack the complexity to capture the high-dimensional and interactive nature of patient risk. This study utilizes an admission granularity level cohort of the MIMIC-IV database to develop and compare machine learning architectures against the baseline LACE index. Due to the imbalanced prevalence of readmission, the penalized logistic regression, …
A Comparative Evaluation Of Data Imbalance Handling Techniques In Machine Learning Models For One-Year Mortality Prediction In Liver Cirrhosis, Sumiya Hasan Trisha
A Comparative Evaluation Of Data Imbalance Handling Techniques In Machine Learning Models For One-Year Mortality Prediction In Liver Cirrhosis, Sumiya Hasan Trisha
UNF Graduate Theses and Dissertations
Liver cirrhosis is associated with substantial morbidity and mortality, making one-year mortality prediction a clinically relevant problem. Using a liver cirrhosis dataset as the motivating application, this thesis evaluates five machine learning classifiers—Logistic Regression, Random Forest, XGBoost, LightGBM, and CatBoost—under five class-imbalance handling strategies: Baseline learning, Random Oversampling, SMOTE-NC, ADASYN, and Cost-Sensitive Learning. Hyperparameter tuning was conducted using randomized search, and predictive performance was assessed over 200 iterations of Monte Carlo Cross-Validation using Accuracy, Precision, Recall, Fl-score, and ROC-AUC.
The results suggest that imbalance-handling strategies can materially affect predictive performance, particularly recall. Because the outcome of interest is death within …
Crop Yield Prediction At Multiple Spatial Scales With Statistical Machine Learning, Vaibhav Charan, Pratishtha Poudel
Crop Yield Prediction At Multiple Spatial Scales With Statistical Machine Learning, Vaibhav Charan, Pratishtha Poudel
Discovery Undergraduate Interdisciplinary Research Internship
Understanding and accurately predicting crop yield is becoming increasingly important today in the face of global food security challenges, and thus, the availability of standardized data and scalable models is the need of the hour. To support this, researchers have developed CY-Bench (Crop Yield Benchmark), a comprehensive dataset that helps forecast maize and wheat yields on a global scale. This research project primarily involved working with the CY-Bench dataset aiming to improve crop yield prediction through machine learning. Initially, papers explaining the CY-Bench dataset and other papers for agriculture modeling were studied and analyzed in detail. The research then progressed …
Forecasting Influenza Rates Using Machine Learning: A Study Of Chatgpt's Predictive Accuracy, Sara Saleh
Forecasting Influenza Rates Using Machine Learning: A Study Of Chatgpt's Predictive Accuracy, Sara Saleh
University Honors Theses
This study evaluates ChatGPT's ability to forecast influenza rates, such as the number of flu cases, hospitalizations, and death during peak season periods using CDC data, and comparing forecasts against actual results to calculate statistical accuracy and consistency. Influenza forecasting is essential for public health planning, but traditional methods may not always provide timely or accurate predictions. In this research study, ChatGPT was utilized to predict the influenza rates for the following week based on the previous week's data obtained from the FluView surveillance system. The predicted rates were compared to the actual influenza rates to assess the model's overall …
Property Testing Ai: An Efficient Frontier, Paul Sopher Lintilhac
Property Testing Ai: An Efficient Frontier, Paul Sopher Lintilhac
Dartmouth College Ph.D Dissertations
In this dissertation, we take a step towards addressing the major problem of a lack of standardized and rigorous approaches to testing and evaluation of AI systems. Taking inspiration from both the fields of Property Testing and Property Based Testing (for programs), we develop a novel taxonomy of partially overlapping classes of properties of AI systems, including simple properties, compound properties, higher order properties, data relation properties, and architecture-utility properties. We argue that this taxonomy categorizes a diverse set of AI traits -- including accuracy, fairness, robustness, monotonicity, point-wise and global privacy properties, sensitivity, and more -- according to the …
Opening The Black Box With Regal: A Novel Explainable Ai Approach To Uncover Key Predictors In Search And Rescue Success, Brandon Hyunjun Kim
Opening The Black Box With Regal: A Novel Explainable Ai Approach To Uncover Key Predictors In Search And Rescue Success, Brandon Hyunjun Kim
Master's Theses
The outcome of a search and rescue (SAR) operation is influenced by a complex, non-linear interplay among numerous factors, including geographic context, subject-specific characteristics, and environmental conditions. The high dimensionality and intricate dependencies among these variables pose significant challenges to traditional exploratory modeling approaches, limiting their ability to uncover meaningful patterns and relationships associated with mission success. This study introduces Rules Based Explanations for Generated neighborhoods Around Localized cases (REGAL), a novel adaptation of the Local Interpretable Model-agnostic Explanations (LIME) framework to explain deep multimodal neural networks and what key features it assesses to determine search and rescue success. REGAL …
Resale Revolution: Trend Implications From Media Presence Transcended To Luxury Retail Markets, Penelope Prochnow
Resale Revolution: Trend Implications From Media Presence Transcended To Luxury Retail Markets, Penelope Prochnow
Capstone Projects
This study aims to deepen understanding of fashion trend decline from peak popularity to obsolescence, with implications for sustainability and producer profit margins. It investigates how the attributes and media presence of fashion items influence their journey from high-end editorial coverage to resale platforms. Using survival analysis to model trend lifetimes and cosine similarity metrics to compare resale and magazine keyword frequencies, alongside machine learning for price prediction, the study uncovers critical temporal patterns. Results show that resale trends reflect magazine content with a lag of approximately 18 to 30 months and draw from long-wave revivals spanning 6 to 14 …
Computer Vision In Soccer: Yolov11 Analytics Engine For Quantifying Game Strategy, Connor S. Maurer
Computer Vision In Soccer: Yolov11 Analytics Engine For Quantifying Game Strategy, Connor S. Maurer
Data Science Undergraduate Honors Theses
Single-shot object detection capabilities significantly reduce computational overhead for real-time computer vision in sports analytics at 60 FPS. YOLO11’s lightweight CNN gives promising accuracy while meeting the low-latency demand of dynamic soccer matches. As data-driven approaches take over the sport of soccer, efficient player tracking systems become critical for informing coach’s strategies. I prototype the ETL (Extract, Transform, Load) process of data collected from a single- shot detection program and evaluate its viability for estimating player fatigue. YOLO11 detects players, the ball, and other characteristics, with the output transformed by homography to estimate the positions in the real world. These …
Expressive And Interpretable User Engagement Prediction Using Multivariate Survival Processes, Akshay Aravamudan
Expressive And Interpretable User Engagement Prediction Using Multivariate Survival Processes, Akshay Aravamudan
Theses and Dissertations
The ability to characterize how information diffuses online is of paramount importance to stakeholders that are interested in tasks such as proposing solutions for mitigating and countering dis/misinformation, predicting user engagement of content in social media, planning marketing campaigns to roll-out products and planning dissemination of political campaign messaging among others. One such facet of learning the dynamics of information diffusion is the ability to predict user engagement or the popularity of a single piece of information as it spreads through an online medium. Existing works in this regard mainly either obfuscate user level information or utilize frameworks that are …
Mortgage Default Classification Modeling For Variable Analysis, Brendan R. Goggins
Mortgage Default Classification Modeling For Variable Analysis, Brendan R. Goggins
Honors College Theses
The financial crisis of the early 2000’s is a prime example of the severe consequences that mortgage default and borrower insolvency can have on economies at large. Mortgage default specifically is a prime case with the popularization of mortgage backed securities and the commonality of this loan structure. Multiple hypotheses and models have been formed to understand the reasons, causes, and consequences of mortgage default. This paper uses both machine learning and statistical classification models to inform an understanding of the variables most significant and impactful to the default outcome of mortgages. Consideration is given to both loan-level microeconomic variables …
Unlocking The Power Of Data: Enhancing Public Policy Through Advanced Data Infrastructure And Language Model Analysis, Zahid Asghar
Unlocking The Power Of Data: Enhancing Public Policy Through Advanced Data Infrastructure And Language Model Analysis, Zahid Asghar
CBER Conference
Data is the fundamental building block for advancements in artificial intelligence (AI), general AI (GAI), machine learning (ML), and large language models (LLMs). This study emphasizes the critical need for robust data infrastructure, arguing that without it, countries cannot fully benefit from technological advancements in various economic sectors. Governments possess vast repositories of both structured and unstructured data across multiple domains such as the judiciary, parliaments, and civil bureaucracy. However, these potential goldmines remain untapped due to inadequate data management capabilities and a lack of appreciation for the necessity of high-quality data. The research identifies key issues in public data …
A Machine Learning Based Approach For The Identification Of Fake Bills, Tianyang Lu, Hongyang Pang
A Machine Learning Based Approach For The Identification Of Fake Bills, Tianyang Lu, Hongyang Pang
Rose-Hulman Undergraduate Mathematics Journal
Fake or counterfeiting currency, which has been around as long as money has existed, is a major economic problem. Since the US dollar is the most popular form of currency globally, it is the most popular currency to counterfeit. The United States Department of Treasury estimates that between $70 million and $200 million in fake bills are in circulation. The Federal Reserve Bank uses special banknote processing systems to count each bill deposited by the bank and examine them for the possibility of counterfeits. These machines have sensors designed to detect general quality of the bills, including paper type, quality …
Code For Care: Hypertension Prediction In Women Aged 18-39 Years, Kruti Sheth
Code For Care: Hypertension Prediction In Women Aged 18-39 Years, Kruti Sheth
Electronic Theses, Projects, and Dissertations
The longstanding prevalence of hypertension, often undiagnosed, poses significant risks of severe chronic and cardiovascular complications if left untreated. This study investigated the causes and underlying risks of hypertension in females aged between 18-39 years. The research questions were: (Q1.) What factors affect the occurrence of hypertension in females aged 18-39 years? (Q2.) What machine learning algorithms are suited for effectively predicting hypertension? (Q3.) How can SHAP values be leveraged to analyze the factors from model outputs? The findings are: (Q1.) Performing Feature selection using binary classification Logistic regression algorithm reveals an array of 30 most influential factors at an …
Ensemble Classification: An Analysis Of The Random Forest Model, Jarod Korn
Ensemble Classification: An Analysis Of The Random Forest Model, Jarod Korn
Williams Honors College, Honors Research Projects
The random forest model proposed by Dr. Leo Breiman in 2001 is an ensemble machine learning method for classification prediction and regression. In the following paper, we will conduct an analysis on the random forest model with a focus on how the model works, how it is applied in software, and how it performs on a set of data. To fully understand the model, we will introduce the concept of decision trees, give a summary of the CART model, explain in detail how the random forest model operates, discuss how the model is implemented in software, demonstrate the model by …
Sparse Representation Learning For Temporal Networks, Maxwell Mcneil
Sparse Representation Learning For Temporal Networks, Maxwell Mcneil
Electronic Theses & Dissertations (2024 - present)
Temporal networks arise in many domains including activity of social network users, sensor network readings over time, and time course gene expression within the interaction network of a model organism. Data of this type contains a wealth of prior information such as the connectivity among nodes (e.g., a friendship graph), and prior knowledge of expected temporal patterns (e.g., periodicity). Modeling these temporal and network patterns jointly is essential for state-of-the-art performance in temporal network data analysis and mining. Sparse dictionary encoding is one modeling approach for such underlying patterns. However, most classical approaches consider only one dimension of the data …
Stressor: An R Package For Benchmarking Machine Learning Models, Samuel A. Haycock
Stressor: An R Package For Benchmarking Machine Learning Models, Samuel A. Haycock
All Graduate Theses and Dissertations, Spring 1920 to Summer 2023
Many discipline specific researchers need a way to quickly compare the accuracy of their predictive models to other alternatives. However, many of these researchers are not experienced with multiple programming languages. Python has recently been the leader in machine learning functionality, which includes the PyCaret library that allows users to develop high-performing machine learning models with only a few lines of code. The goal of the stressor package is to help users of the R programming language access the advantages of PyCaret without having to learn Python. This allows the user to leverage R’s powerful data analysis workflows, while simultaneously …
Application Of Sentiment Analysis And Machine Learning Techniques To Predict Daily Cryptocurrency Price Returns, Edward Wu
CMC Senior Theses
This paper examines the effects of social media sentiment relating to Bitcoin on the daily price returns of Bitcoin and other popular cryptocurrencies by utilizing sentiment analysis and machine learning techniques to predict daily price returns. Many investors think that social media sentiment affects cryptocurrency prices. However, the results of this paper find that social media sentiment relating to Bitcoin does not add significant predictive value to forecasting daily price returns for each of the six cryptocurrencies used for analysis and that machine learning models that do not assume linearity between the current day price return and previous daily price …
Machine Learning Based Restaurant Sales Forecasting, Austin B. Schmidt
Machine Learning Based Restaurant Sales Forecasting, Austin B. Schmidt
LSU New Orleans Theses and Dissertations
To encourage proper employee scheduling for managing crew load, restaurants have a need for accurate sales forecasting. We predict partitions of sales days, so each day is broken up into three sales periods: 10:00 AM-1:59 PM, 2:00 PM-5:59 PM, and 6:00 PM-10:00 PM. This study focuses on the middle timeslot, where sales forecasts should extend for one week. We gather three years of sales between 2016-2019 from a local restaurant, to generate a new dataset for researching sales forecasting methods.
Outlined are methodologies used when going from raw data to a workable dataset. We test many machine learning models on …
How Machine Learning And Probability Concepts Can Improve Nba Player Evaluation, Harrison Miller
How Machine Learning And Probability Concepts Can Improve Nba Player Evaluation, Harrison Miller
CMC Senior Theses
In this paper I will be breaking down a scholarly article, written by Sameer K. Deshpande and Shane T. Jensen, that proposed a new method to evaluate NBA players. The NBA is the highest level professional basketball league in America and stands for the National Basketball Association. They proposed to build a model that would result in how NBA players impact their teams chances of winning a game, using machine learning and probability concepts. I preface that by diving into these concepts and their mathematical backgrounds. These concepts include building a linear model using ordinary least squares method, the bias …
Ordinal Hyperplane Loss, Bob Vanderheyden
Ordinal Hyperplane Loss, Bob Vanderheyden
Doctor of Data Science and Analytics Dissertations
This research presents the development of a new framework for analyzing ordered class data, commonly called “ordinal class” data. The focus of the work is the development of classifiers (predictive models) that predict classes from available data. Ratings scales, medical classification scales, socio-economic scales, meaningful groupings of continuous data, facial emotional intensity and facial age estimation are examples of ordinal data for which data scientists may be asked to develop predictive classifiers. It is possible to treat ordinal classification like any other classification problem that has more than two classes. Specifying a model with this strategy does not fully utilize …
Bias Reduction In Machine Learning Classifiers For Spatiotemporal Analysis Of Coral Reefs Using Remote Sensing Images, Justin J. Gapper
Bias Reduction In Machine Learning Classifiers For Spatiotemporal Analysis Of Coral Reefs Using Remote Sensing Images, Justin J. Gapper
Computational and Data Sciences (PhD) Dissertations
This dissertation is an evaluation of the generalization characteristics of machine learning classifiers as applied to the detection of coral reefs using remote sensing images. Three scientific studies have been conducted as part of this research: 1) Evaluation of Spatial Generalization Characteristics of a Robust Classifier as Applied to Coral Reef Habitats in Remote Islands of the Pacific Ocean 2) Coral Reef Change Detection in Remote Pacific Islands using Support Vector Machine Classifiers 3) A Generalized Machine Learning Classifier for Spatiotemporal Analysis of Coral Reefs in the Red Sea. The aim of this dissertation is to propose and evaluate a …
Repairing Landsat Satellite Imagery Using Deep Machine Learning Techniques, Griffin J. Lane, Patricia Goresen, Robert Slater
Repairing Landsat Satellite Imagery Using Deep Machine Learning Techniques, Griffin J. Lane, Patricia Goresen, Robert Slater
SMU Data Science Review
Satellite Imagery is one of the most widely used sources to analyze geographic features and environments in the world. The data gathered from satellites are used to quantify many vital problems facing our society, such as the impact of natural disasters, shore erosion, rising water levels, and urban growth rates. In this paper, we construct machine learning and deep learning algorithms for repairing anomalies in the Landsat satellite imagery data which arise for various reasons ranging from cloud obstruction to satellite malfunctions. The accuracy of GIS data is crucial to ensuring the models produced from such data are as close …
Visualization And Machine Learning Techniques For Nasa’S Em-1 Big Data Problem, Antonio P. Garza Iii, Jose Quinonez, Misael Santana, Nibhrat Lohia
Visualization And Machine Learning Techniques For Nasa’S Em-1 Big Data Problem, Antonio P. Garza Iii, Jose Quinonez, Misael Santana, Nibhrat Lohia
SMU Data Science Review
In this paper, we help NASA solve three Exploration Mission-1 (EM-1) challenges: data storage, computation time, and visualization of complex data. NASA is studying one year of trajectory data to determine available launch opportunities (about 90TBs of data). We improve data storage by introducing a cloud-based solution that provides elasticity and server upgrades. This migration will save $120k in infrastructure costs every four years, and potentially avoid schedule slips. Additionally, it increases computational efficiency by 125%. We further enhance computation via machine learning techniques that use the classic orbital elements to predict valid trajectories. Our machine learning model decreases trajectory …
An Evaluation Of Training Size Impact On Validation Accuracy For Optimized Convolutional Neural Networks, Jostein Barry-Straume, Adam Tschannen, Daniel W. Engels, Edward Fine
An Evaluation Of Training Size Impact On Validation Accuracy For Optimized Convolutional Neural Networks, Jostein Barry-Straume, Adam Tschannen, Daniel W. Engels, Edward Fine
SMU Data Science Review
In this paper, we present an evaluation of training size impact on validation accuracy for an optimized Convolutional Neural Network (CNN). CNNs are currently the state-of-the-art architecture for object classification tasks. We used Amazon’s machine learning ecosystem to train and test 648 models to find the optimal hyperparameters with which to apply a CNN towards the Fashion-MNIST (Mixed National Institute of Standards and Technology) dataset. We were able to realize a validation accuracy of 90% by using only 40% of the original data. We found that hidden layers appear to have had zero impact on validation accuracy, whereas the neural …
Improving Vix Futures Forecasts Using Machine Learning Methods, James Hosker, Slobodan Djurdjevic, Hieu Nguyen, Robert Slater
Improving Vix Futures Forecasts Using Machine Learning Methods, James Hosker, Slobodan Djurdjevic, Hieu Nguyen, Robert Slater
SMU Data Science Review
The problem of forecasting market volatility is a difficult task for most fund managers. Volatility forecasts are used for risk management, alpha (risk) trading, and the reduction of trading friction. Improving the forecasts of future market volatility assists fund managers in adding or reducing risk in their portfolios as well as in increasing hedges to protect their portfolios in anticipation of a market sell-off event. Our analysis compares three existing financial models that forecast future market volatility using the Chicago Board Options Exchange Volatility Index (VIX) to six machine/deep learning supervised regression methods. This analysis determines which models provide best …
Quantifying Human Biological Age: A Machine Learning Approach, Syed Ashiqur Rahman
Quantifying Human Biological Age: A Machine Learning Approach, Syed Ashiqur Rahman
Graduate Theses, Dissertations, and Problem Reports (ETD)
Quantifying human biological age is an important and difficult challenge. Different biomarkers and numerous approaches have been studied for biological age prediction, each with its advantages and limitations. In this work, we first introduce a new anthropometric measure (called Surface-based Body Shape Index, SBSI) that accounts for both body shape and body size, and evaluate its performance as a predictor of all-cause mortality. We analyzed data from the National Health and Human Nutrition Examination Survey (NHANES). Based on the analysis, we introduce a new body shape index constructed from four important anthropometric determinants of body shape and body size: body …
Rfviz: An Interactive Visualization Package For Random Forests In R, Christopher Beckett
Rfviz: An Interactive Visualization Package For Random Forests In R, Christopher Beckett
All Graduate Plan B and other Reports, Spring 1920 to Spring 2023
Random forests are very popular tools for predictive analysis and data science. They work for both classification (where there is a categorical response variable) and regression (where the response is continuous). Random forests provide proximities, and both local and global measures of variable importance. However, these quantities require special tools to be effectively used to interpret the forest. Rfviz is a sophisticated interactive visualization package and toolkit in R, specially designed for interpreting the results of a random forest in a user-friendly way. Rfviz uses a recently developed R package (loon) from the Comprehensive R Archive Network (CRAN) to create …
Overcoming Small Data Limitations In Heart Disease Prediction By Using Surrogate Data, Alfeo Sabay, Laurie Harris, Vivek Bejugama, Karen Jaceldo-Siegl
Overcoming Small Data Limitations In Heart Disease Prediction By Using Surrogate Data, Alfeo Sabay, Laurie Harris, Vivek Bejugama, Karen Jaceldo-Siegl
SMU Data Science Review
In this paper, we present a heart disease prediction use case showing how synthetic data can be used to address privacy concerns and overcome constraints inherent in small medical research data sets. While advanced machine learning algorithms, such as neural networks models, can be implemented to improve prediction accuracy, these require very large data sets which are often not available in medical or clinical research. We examine the use of surrogate data sets comprised of synthetic observations for modeling heart disease prediction. We generate surrogate data, based on the characteristics of original observations, and compare prediction accuracy results achieved from …