Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

Random Forest

Discipline
Institution
Publication Year
Publication
Publication Type

Articles 1 - 30 of 30

Full-Text Articles in Data Science

Machine Learning-Based Regression For Magnetic Field Prediction From Odmr Spectral Data, Jesse B. Hernandez Dec 2026

Machine Learning-Based Regression For Magnetic Field Prediction From Odmr Spectral Data, Jesse B. Hernandez

Electronic Theses, Projects, and Dissertations

Optically Detected Magnetic Resonance (ODMR) using nitrogen-vacancy (NV) centers in diamond enables sensitive, room-temperature magnetic field sensing, but real ODMR spectra are often noisy and difficult to analyze with traditional peak-fitting methods. This thesis investigates whether machine learning can reliably predict magnetic field strength directly from ODMR spectra, and compares four model families under a single regression task: a random forest, an artificial neural network (ANN), a one-dimensional convolutional neural network (1D-CNN), and a Transformer.

Training data were generated from an NV-ensemble simulation calibrated to real measurements provided by the Ulsan National Institute of Science and Technology (UNIST), spanning 0 …


Identifying Textual Predictors Of Early Termination In Clinical Trials In Medicine: An Explainable Machine-Learning Study, Rohan Ramnarain Jun 2026

Identifying Textual Predictors Of Early Termination In Clinical Trials In Medicine: An Explainable Machine-Learning Study, Rohan Ramnarain

Dissertations, Theses, and Capstone Projects

About one in five clinical trials in medicine ends early, wasting valuable resources and reducing the evidence available for developing life-saving medical treatments. This project uses a method called Trial2Vec, which is a self-supervised machine-learning method that converts clinical trial documents into dense numerical representations that capture their key design and clinical characteristics, to turn each proposed clinical trial’s written protocol into a compact numerical profile (a process referred to as embedding). These profiles are then paired with a predictive machine learning models to identify the words and phrases in the trial documents that can signal a higher risk of …


Understanding Delays In Emergency Department Care: A National Analysis Of Wait Times, Gregory Forsberg May 2026

Understanding Delays In Emergency Department Care: A National Analysis Of Wait Times, Gregory Forsberg

Mathematics, Statistics, and Computer Science Honors Projects

Emergency department (ED) wait times remain a persistent bottleneck in the United States healthcare system, impacting patient outcomes, hospital efficiency, and equitable access to care. This study analyzes nationally representative data from the National Hospital Ambulatory Medical Care Survey (NHAMCS), a complex, multi-stage probability sample. Using survey-weighted analyses and predictive modeling, we examine the effects of patient characteristics, triage acuity, and visit timing. Results indicate that operational and system-level factors, including hospital capacity, geographic region, and temporal variation, are among the most influential predictors of ED wait times


Enhancing Network Security Through Dual-Layer Log Analysis: Integrating Machine Learning Classifiers With Large Language Models For Intelligent Anomaly Detection, Anthony Burton-Cordova, O'Neil Gray, Mohammad Al Rousan Apr 2025

Enhancing Network Security Through Dual-Layer Log Analysis: Integrating Machine Learning Classifiers With Large Language Models For Intelligent Anomaly Detection, Anthony Burton-Cordova, O'Neil Gray, Mohammad Al Rousan

SMU Data Science Review

This paper presents an innovative approach to enhancing network security by integrating machine learning algorithms with fine-tuned large language models (LLMs) to provide an expert assistant querying. The proposed method utilizes machine learning for efficient preprocessing and feature extraction from log data, followed by the application of a fine-tuned LLM to analyze and interpret anomalies with greater accuracy. This dual-layer detection system is designed to improve the identification of subtle and sophisticated security threats. The research team’s extensive evaluation using real-world log datasets indicates that the combined approach increases detection rates and communicates results in an understandable manner, demonstrating its …


Predictive Analytics For Customer Churns In Financial Services., Thant Thiha Jan 2025

Predictive Analytics For Customer Churns In Financial Services., Thant Thiha

ICT

This project presents a customer churn prediction analysis in the telecommunications sector, achieving an ROC-AUC of approximately 0.86 using statistically validated features and interpretable AI models. Key churn drivers identified include the number of products held, customer age, and geographic location. Ensemble models, such as Random Forest and Gradient Boosting, provided the highest predictive performance. Ethical AI principles were applied to ensure fairness, transparency, privacy, and accountability. Business insights derived from the analysis inform targeted retention strategies, prioritising multi-product users, specific age groups, and geographic segments. Deployment recommendations include the tuned Random Forest model with ongoing monitoring, governance, and future …


Customer Response To Marketing Campaigns., Alexandru Enoiu Jan 2025

Customer Response To Marketing Campaigns., Alexandru Enoiu

ICT

This project applies machine learning to predict customer responses to marketing campaigns, aiming to enhance customer engagement and optimise resource allocation. Using a dataset containing demographic, lifestyle, and purchase behaviour data, a Random Forest classifier was developed to identify potential responders to marketing initiatives. The project follows a structured methodology including data exploration, preprocessing, model building, and evaluation. By accurately predicting customer behaviour, businesses can improve campaign targeting, design personalised marketing strategies, and allocate resources more efficiently. The results provide actionable insights to support customer segmentation and strategic decision-making, promoting business growth and customer satisfaction.


Using Machine Learning To Predict Credit Card Fraud, Pedro Henrique Das Chagas Morais Jan 2025

Using Machine Learning To Predict Credit Card Fraud, Pedro Henrique Das Chagas Morais

ICT

Credit card fraud poses a significant challenge to financial institutions, leading to substantial financial losses and declining customer trust. This project develops and evaluates machine learning models to detect fraudulent credit card transactions using a large, realistic synthetic dataset. Following data preprocessing, exploratory analysis, and class-balancing using SMOTE, four supervised models—Logistic Regression, Decision Tree, Random Forest, and XGBoost—were trained and compared. Performance was assessed using metrics suited to imbalanced classification, including AUC, Recall, Precision, F1-score, and Average Precision. Results show that XGBoost, particularly after hyperparameter optimisation, delivered the strongest performance (AUC 0.99, Recall 0.83, AP 0.70), outperforming other models and …


Chess Evaluation And Player Profiling Using Convolutional Neural Networks (Cnns) And Spatial Recognition., Joel D’Mello Jan 2025

Chess Evaluation And Player Profiling Using Convolutional Neural Networks (Cnns) And Spatial Recognition., Joel D’Mello

ICT

This thesis explores the feasibility of employing data analytics techniques in chess, with the purpose of profiling player styles and building a comprehensive chess analytics platform. The data set consists of over 20,000 anonymized games, and therefore, the study involved feature engineering, classification and visual analytics, in order to gain more insight into player decision making in chess. The data pre-processing part of analysis involved parsing Portable Game Notation (PGN) files, and feature engineering positional characteristics - material imbalance, pawn structure, king safety, and piece mobility - along with quantifying the sample using Average Centipawn Loss (ACPL) using Stockfish. ACPL …


Exploring The Determinants Of Life Expectancy At Birth: Predicting And Forecasting Global Health Trends Using Statistical, Machine Learning, And Deep Learning Models., Emma Rath Jan 2025

Exploring The Determinants Of Life Expectancy At Birth: Predicting And Forecasting Global Health Trends Using Statistical, Machine Learning, And Deep Learning Models., Emma Rath

ICT

Accurate life expectancy forecasting is essential for health policy planning, yet research comparing statistical, machine learning, and deep learning approaches under real-world constraints remains limited. This study evaluates ARIMA/ARIMAX, tree-based, and neural network models using Irish and global datasets, considering small samples, missing data, and COVID-19 shocks. ARIMAX with lagged socioeconomic variables outperformed LSTM and other ML/DL methods. Income-based stratification improved predictive accuracy and interpretability, with SHAP analysis highlighting GDP per capita for developed countries and school enrolment and trade indicators for developing contexts. Results provide practical guidance for policymakers and establish limits for model complexity under constrained health data.


Forecasting Hourly Police Call For Service Volumes: A Comparative Analysis Of Statistical, Machine Learning And Neural Network Models For Operational Planning., Patrick Duggan Jan 2025

Forecasting Hourly Police Call For Service Volumes: A Comparative Analysis Of Statistical, Machine Learning And Neural Network Models For Operational Planning., Patrick Duggan

ICT

Accurate demand forecasting is critical in operational settings where resource allocation and planning decisions depend on anticipated service volumes. Transactional systems that capture timestamped records provide valuable data sources for developing demand forecasts. This study examines hourly call volume forecasting using New Orleans police calls for service data, comparing the performance of statistical models, tree-based methods, and recurrent neural networks.

The research evaluates four primary modelling approaches: ARIMA models representing traditional statistical methods, XGBoost and Random Forest as a tree-based ensemble technique, and Gated Recurrent Units (GRUs) as deep learning alternatives. A naive seasonal model serves as the baseline benchmark. …


Predicting S&P Corporate Credit Ratings Using Financial Ratios And Machine Learning: An Analysis Of European Non-Financial Companies., Gabriele Frattaroli Jan 2025

Predicting S&P Corporate Credit Ratings Using Financial Ratios And Machine Learning: An Analysis Of European Non-Financial Companies., Gabriele Frattaroli

ICT

This study investigates the prediction of multi-class S&P corporate credit ratings for European non-financial firms from 2010 to 2024 using a machine learning framework grounded in financial fundamentals. To ensure robustness and generalizability, the analysis excluded the Year variable, which was identified as a source of data leakage. After this correction, non-linear ensemble models demonstrated a clear advantage over linear baselines. The top-performing Random Forest model achieved a weighted F1-score of approximately 0.60, more than doubling the performance of the Logistic Regression benchmark used as a baseline (0.26), with most misclassifications concentrated in adjacent rating categories. This indicates that while …


Spatial Analysis And Machine Learning Integration For Nutritional Status Mapping Using Ann And Random Forest Models, Desi Anis Anggraini, Fachrul Kurniawan, Fresy Nugroho, Meidya Koeshardianto, Mohammad Iqbal Bachtiar Jan 2025

Spatial Analysis And Machine Learning Integration For Nutritional Status Mapping Using Ann And Random Forest Models, Desi Anis Anggraini, Fachrul Kurniawan, Fresy Nugroho, Meidya Koeshardianto, Mohammad Iqbal Bachtiar

Knowledge Engineering and Data Science

Nutritional problems among children under five remain a major public health challenge. This research seeks to create a spatially oriented system for evaluating and mapping nutritional status utilizing Artificial Neural Network (ANN) and Random Forest (RF) algorithms. Data obtained from the Sumenep District Health Office included age, weight, height, and gender variables. Both models were trained using a 70:30 data ratio and evaluated with accuracy, precision, recall, and F1-score metrics. The ANN model achieved an accuracy of 95.8%, while the RF model reached 97.7%. Classification results were visualized through a Geographic Information System (GIS) to illustrate spatial distribution and identify …


Classification Of Indonesian Sign Language (Sibi) Using Data Mining Algorithms K-Nearest Neighbor And Random Forest, Muhammad Zaki Wirawan, Achmad Afif, Anik Nur Handayani, Imanuel Hitipeuw, Osamu Fukuda Jan 2025

Classification Of Indonesian Sign Language (Sibi) Using Data Mining Algorithms K-Nearest Neighbor And Random Forest, Muhammad Zaki Wirawan, Achmad Afif, Anik Nur Handayani, Imanuel Hitipeuw, Osamu Fukuda

Knowledge Engineering and Data Science

This study aims to address the communication hallenges faced by the Indonesian deaf community by developing an automatic classification model for Sistem Bahasa Isyarat Indonesia (SIBI) using data mining techniques. The main objective is to identify a practical algorithm for recognizing SIBI hand gestures to enhance accessibility and inclusiveness in digital communication. A comprehensive dataset consisting of 32,850 gesture samples representing SIBI alphabet signs was collected and processed through feature extraction, data cleaning, and normalization using Z-Transform and Min-Max methods. Two classification algorithms, K-Nearest Neighbor (KNN) and Random Forest, were implemented and evaluated using metrics such as accuracy, precision, recall, …


Enhancing Imputation Accuracy: A Multi-Faceted Approach For Missing Data In Chicago Arrest Records, Steve Bramhall, Jae Chung, Nicholas Mueller Sep 2024

Enhancing Imputation Accuracy: A Multi-Faceted Approach For Missing Data In Chicago Arrest Records, Steve Bramhall, Jae Chung, Nicholas Mueller

SMU Data Science Review

This paper introduces a novel approach to enhance the imputation process for missing data, utilizing crime records from Chicago with arrests as the target feature. Robust imputation techniques are crucial in the era of burgeoning datasets for generating reliable insights. Our core objective is to present an innovative method that improves imputation techniques, augmenting model performance and bolstering the reliability of analytical outcomes. Leveraging numeric crime data, we establish a Gradient Boosting (GBM) baseline model, then introduce ensemble methods including Random Forest and Decision Trees for further refinement. By systematically exploring multiple imputation processes, we establish a baseline for comparative …


Classification Models Using Python In Industrial/Organizational Psychology, Beyza Ceylan Jan 2024

Classification Models Using Python In Industrial/Organizational Psychology, Beyza Ceylan

Williams Honors College, Honors Research Projects

Companies, industries, and places of business use artificial intelligence and statistics to predict the characteristics of their employees and staff. Data collected from these individuals is also used to make decisions about them regarding their work life, such as promotions, salaries, or within the hiring process. Two models that are commonly used throughout the field of psychology and specifically in industrial/organizational psychology are the linear regression and the logistic regression. Examining different classification models using Python shows the potential that there may be different models that are more accurate in their predictions of employee success, including a Random Forest model …


Eeg Classification While Listening To Murottal Al-Quran And Classical Music Using Random Forest Method, Heni Sumarti, Fahira Septiani, Agus Sudarmanto, Wahyu Caesarendra, Rizki Edmi Edison Dec 2023

Eeg Classification While Listening To Murottal Al-Quran And Classical Music Using Random Forest Method, Heni Sumarti, Fahira Septiani, Agus Sudarmanto, Wahyu Caesarendra, Rizki Edmi Edison

Knowledge Engineering and Data Science

This study is aimed to classify the brain activity of adolescents associated with audio stimuli; murottal Al-Quran and classical music. The raw data were filtered using Independent Component Analisys (ICA) and followed by band-pass filter in Python on the Google Colab Extraction was processed with Power Spectral Density (PSD) and the Random Forest Method in Weka Machine Learning was used for classification. The research results showed the same results between the two types of stimulation, namely the order of brain waves from highest to lowest were delta, alpha, theta and beta. The average brain waves of teenagers when given murottal …


Optimizing Random Forest Algorithm To Classify Player's Memorisation Via In-Game Data, Akmal Vrisna Alzuhdi, Harits Ar Rosyid, Mohammad Yasser Chuttur, Shah Nazir Jul 2023

Optimizing Random Forest Algorithm To Classify Player's Memorisation Via In-Game Data, Akmal Vrisna Alzuhdi, Harits Ar Rosyid, Mohammad Yasser Chuttur, Shah Nazir

Knowledge Engineering and Data Science

Assessment of a player's knowledge in game education has been around for some time. Traditional evaluation in and around a gaming session may disrupt the players' immersion. This research uses an optimized Random Forest to construct a non-invasive prediction of a game education player's Memorization via in-game data. Firstly, we obtained the dataset from a 3-month survey to record in-game data of 50 players who play 4-15 game stages of the Chem Fight (a test case game). Next, we generated three variants of datasets via the preprocessing stages: resampling method (SMOTE), normalization (min-max), and a combination of resampling and normalization. …


Distance Correlation Based Feature Selection In Random Forest, Jose Munoz-Lopez May 2023

Distance Correlation Based Feature Selection In Random Forest, Jose Munoz-Lopez

Electronic Theses, Projects, and Dissertations

The Pearson correlation coefficient is a commonly used measure of correlation, but it has limitations as it only measures the linear relationship between two numerical variables. In 2007, Szekely et al. introduced the distance correlation, which measures all types of dependencies between random vectors X and Y in arbitrary dimensions, not just the linear ones. In this thesis, we propose a filter method that utilizes distance correlation as a criterion for feature selection in Random Forest regression. We conduct extensive simulation studies to evaluate its performance compared to existing methods under various data settings, in terms of the prediction mean …


Analysis Of Chemical Elements In Basalts Using Mislabeled Data, A Machine Learning Approach, Jenifer Vivar Jan 2023

Analysis Of Chemical Elements In Basalts Using Mislabeled Data, A Machine Learning Approach, Jenifer Vivar

Dissertations and Theses

Scientists use basalt chemistry to discriminate among different tectonic settings. There are well-known chemical elements used to classify tectonic settings. An exploration of new features is done using Logistic Regression and Random Forest to discover any new elements of interest. The models were used with other tools, such as recursive feature elimination and permutations, to increase reliability. Among the scarcely explored chemical elements are Terbium (Tb), Holmium (Ho), Samarium (Sm), and Erbium (Er). The data used for the exploration contained many outliers. Therefore, an ensemble model was created to explore the location and composition of such outliers. The ensemble was …


Classification Of Pixel Tracks To Improve Track Reconstruction From Proton-Proton Collisions, Kebur Fantahun, Jobin Joseph, Halle Purdom, Nibhrat Lohia Sep 2022

Classification Of Pixel Tracks To Improve Track Reconstruction From Proton-Proton Collisions, Kebur Fantahun, Jobin Joseph, Halle Purdom, Nibhrat Lohia

SMU Data Science Review

In this paper, machine learning techniques are used to reconstruct particle collision pathways. CERN (Conseil européen pour la recherche nucléaire) uses a massive underground particle collider, called the Large Hadron Collider or LHC, to produce particle collisions at extremely high speeds. There are several layers of detectors in the collider that track the pathways of particles as they collide. The data produced from collisions contains an extraneous amount of background noise, i.e., decays from known particle collisions produce fake signal. Particularly, in the first layer of the detector, the pixel tracker, there is an overwhelming amount of background noise that …


Analysis Of The Electric Power Outage Data And Prediction Of Electric Power Outage For Major Metropolitan Areas In Texas Using Machine Learning And Time Series Methods, Renfeng Wang, Venkata Leela 'Mg' Vanga, Zachary B. Zaiken, Jonathan Bennett Jun 2022

Analysis Of The Electric Power Outage Data And Prediction Of Electric Power Outage For Major Metropolitan Areas In Texas Using Machine Learning And Time Series Methods, Renfeng Wang, Venkata Leela 'Mg' Vanga, Zachary B. Zaiken, Jonathan Bennett

SMU Data Science Review

With growing energy usage, power outages affect millions of households. This case study focuses on gathering power outage historical data, modifying the data to attach weather attributes, and gathering ERCOT energy market conditions for Dallas-Fort Worth and Houston metropolitan areas of Texas. The transformed data is then analyzed using machine learning algorithms including, but not limited to, Regression, Random Forests and XGBoost to consider current weather and ERCOT features and predict power outage percentage for locations. The transformed data is also trained using time series models and serially correlated models including Autoregression and Vector Autoregression. This study also focuses on …


Urban Traffic Simulation: Network And Demand Representation Impacts On Congestion Metrics, Aaron Faltesek, Balasubramaniam Dakshinamoorthi, Sreeni Prabhala, Akbar Thobani, Anu Kuncheria, Jane Macfarlane Dec 2021

Urban Traffic Simulation: Network And Demand Representation Impacts On Congestion Metrics, Aaron Faltesek, Balasubramaniam Dakshinamoorthi, Sreeni Prabhala, Akbar Thobani, Anu Kuncheria, Jane Macfarlane

SMU Data Science Review

Traffic simulations are often used by city planners as a basis for predicting the impact of policies, plans, and operations. The complexities underpinning traffic simulations are often not described in detail yet can significantly impact the simulation outcome. Conflating underlying data for simulations is complex and hinders the interest in this type of exploration. This paper aims to elucidate critical features of traffic simulations that drive the generated metrics of the modeled urban environment. Specifically, this paper examines differences in two road graph networks for the metropolitan region of Houston, TX: a reduced network composed of 45,675 road links and …


Identifying Vacant Lots To Reduce Violent Crime In Dallas, Texas, Laura Lazarescou, Andrew Mejia, Tina Pai, Sabrina Purvis, Robert Slater, Owen Wilson-Chavez Dec 2021

Identifying Vacant Lots To Reduce Violent Crime In Dallas, Texas, Laura Lazarescou, Andrew Mejia, Tina Pai, Sabrina Purvis, Robert Slater, Owen Wilson-Chavez

SMU Data Science Review

Vacant lots have been associated with community violence for many years. Researchers have confirmed a positive correlation between vacant lots and vacant buildings with increased violence in urban and rural geographies. However, identifying vacant lots has been a challenge, and modeling methods were largely manual and time-intensive. This prevented cities and non-profit organizations from acting on the information since it was expensive and high-risk to develop remediation programs without clearly understanding where or how many vacant lots existed.

The primary objective of this study was to provide a predictive model that accelerates and improves the accuracy of prior land classification …


Forecasting The Daily Percentage Of Delayed Flights Based On The National Weather Data, Parto Mahmoudi Jan 2021

Forecasting The Daily Percentage Of Delayed Flights Based On The National Weather Data, Parto Mahmoudi

Graduate Student Theses, Dissertations, & Professional Papers

Flight delays cost airlines and affect passenger’s satisfaction. In this research work, we predicted the daily percentage of delayed flights based on the national weather data using the multiple linear regression and the random forest models. We extracted the passenger flight on-time performance data from the Bureau of Transportation Statistics and the weather dataset from NOAA National Centers for Environmental Information for the years from 2015 to 2019. We used the flight dataset for Seattle airport as the origin. We predicted the daily percentage of delayed flights for the Seattle-originated flights based on the features such as weather conditions of …


Classifying Imbalanced Financial Fraud Data Utilizing Enhanced Random Forest Algorithm, Charles Gardner Dec 2020

Classifying Imbalanced Financial Fraud Data Utilizing Enhanced Random Forest Algorithm, Charles Gardner

Master of Science in Computer Science Theses

Imbalanced datasets have been a unique challenge for machine learning, requiring specialized approaches to correctly classify the minority class. Financial fraud detection involves using highly imbalanced datasets with a class imbalance of up to .01% frauds to 99.99% regular transactions. It is essential to identify all frauds in financial fraud detection, even if some classifications' precision is low. I developed a random forest assembly that separates fraudulent transactions into tiers of precision. With this approach, 96% of fraudulent transactions are identified, showing an 8% increase in recall when compared to standard approaches. 59% of fraud classifications' precision increases by 10% …


Comparing Variable Importance In Prediction Of Silence Behaviours Between Random Forest And Conditional Inference Forest Models., Stephen Barrett Dr, Geraldine Gray Dr, Colm Mcguinness Dr, Michael Knoll Dr. Oct 2020

Comparing Variable Importance In Prediction Of Silence Behaviours Between Random Forest And Conditional Inference Forest Models., Stephen Barrett Dr, Geraldine Gray Dr, Colm Mcguinness Dr, Michael Knoll Dr.

Articles

This paper explores variable importance metrics of Conditional Inference Trees (CIT) and classical Classification And Regression Trees (CART) based Random Forests. The paper compares both algorithms variable importance rankings and highlights why CIT should be used when dealing with data with different levels of aggregation. The models analysed explored the role of cultural factors at individual and societal level when predicting Organisational Silence behaviours.


Machine Learning Approaches For Improving Prediction Performance Of Structure-Activity Relationship Models, Gabriel Idakwo Aug 2020

Machine Learning Approaches For Improving Prediction Performance Of Structure-Activity Relationship Models, Gabriel Idakwo

Dissertations

In silico bioactivity prediction studies are designed to complement in vivo and in vitro efforts to assess the activity and properties of small molecules. In silico methods such as Quantitative Structure-Activity/Property Relationship (QSAR) are used to correlate the structure of a molecule to its biological property in drug design and toxicological studies. In this body of work, I started with two in-depth reviews into the application of machine learning based approaches and feature reduction methods to QSAR, and then investigated solutions to three common challenges faced in machine learning based QSAR studies.

First, to improve the prediction accuracy of learning …


Analysis On Suicidal Ideation Among Adolescents (12-17 Years) In The Usa, Himani Raturi Jul 2020

Analysis On Suicidal Ideation Among Adolescents (12-17 Years) In The Usa, Himani Raturi

Electronic Theses, Projects, and Dissertations

Suicide is one of the leading health concerns in United States among adolescents and the presence of suicidal ideation (SI) is quite high, with ~20-30% of adolescents reporting it at some point. Though we have seen growth and development in the prevention of suicide, there is limited research on the ability to identify the adolescents which might be at risk for SI. The objective behind the project is to identify adolescents with SI using machine learning.

The project shows statistics from different articles on adolescents in the U.S. For this study, adolescent data was taken from NSDUH 2018. Moreover, detailed …


An Application Of Machine Learning To Explore Relationships Between Factors Of Organisational Silence And Culture, With Specific Focus On Predicting Silence Behaviours, Stephen Barrett Dr May 2020

An Application Of Machine Learning To Explore Relationships Between Factors Of Organisational Silence And Culture, With Specific Focus On Predicting Silence Behaviours, Stephen Barrett Dr

Articles

Research indicates that there are many individual reasons why people do not speak up when confronted with situations that may concern them within their working environment. One of the areas that requires more focused research is the role culture plays in why a person may remain silent when such situations arise. The purpose of this study is to use data science techniques to explore the patterns in a data set that would lead a person to engage in organisational silence. The main research question the thesis asks is: Is Machine Learning a tool that Social Scientists can use with respect …


Automated Morgan Keenan Classification Of Observed Stellar Spectra Collected By The Sloan Digital Sky Survey Using A Single Classifier, Michael J. Brice, Răzvan Andonie Oct 2019

Automated Morgan Keenan Classification Of Observed Stellar Spectra Collected By The Sloan Digital Sky Survey Using A Single Classifier, Michael J. Brice, Răzvan Andonie

All Faculty Scholarship for the College of the Sciences

The classification of stellar spectra is a fundamental task in stellar astrophysics. Stellar spectra from the Sloan Digital Sky Survey are applied to standard classification methods, k-nearest neighbors and random forest, to automatically classify the spectra. Stellar spectra are high dimensional data and the dimensionality is reduced using astronomical knowledge because classifiers work in low dimensional space. These methods are utilized to classify the stellar spectra into a complete Morgan Keenan classification (spectral and luminosity) using a single classifier. The motion of stars (radial velocity) causes machine-learning complications through the feature matrix when classifying stellar spectra. Due to the nature …