Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Computer Sciences (4)
- Statistics and Probability (4)
- Computer Engineering (3)
- Data Storage Systems (3)
- Electrical and Computer Engineering (3)
-
- Engineering (3)
- Systems and Communications (3)
- Artificial Intelligence and Robotics (2)
- Business (2)
- Mathematics (2)
- Medicine and Health Sciences (2)
- Physics (2)
- Quantum Physics (2)
- Social and Behavioral Sciences (2)
- Applied Mathematics (1)
- Astrophysics and Astronomy (1)
- Biochemistry (1)
- Biochemistry, Biophysics, and Structural Biology (1)
- Business Administration, Management, and Operations (1)
- Business Intelligence (1)
- Chemistry (1)
- Earth Sciences (1)
- Elementary Particles and Fields and String Theory (1)
- Emergency Medicine (1)
- Geochemistry (1)
- Geology (1)
- Industrial and Organizational Psychology (1)
- Life Sciences (1)
- Institution
- Publication
-
- ICT (7)
- SMU Data Science Review (6)
- Knowledge Engineering and Data Science (4)
- Electronic Theses, Projects, and Dissertations (3)
- Articles (2)
-
- All Faculty Scholarship for the College of the Sciences (1)
- Dissertations (1)
- Dissertations and Theses (1)
- Dissertations, Theses, and Capstone Projects (1)
- Graduate Student Theses, Dissertations, & Professional Papers (1)
- Master of Science in Computer Science Theses (1)
- Mathematics, Statistics, and Computer Science Honors Projects (1)
- Williams Honors College, Honors Research Projects (1)
- Publication Type
Articles 1 - 30 of 30
Full-Text Articles in Data Science
Machine Learning-Based Regression For Magnetic Field Prediction From Odmr Spectral Data, Jesse B. Hernandez
Machine Learning-Based Regression For Magnetic Field Prediction From Odmr Spectral Data, Jesse B. Hernandez
Electronic Theses, Projects, and Dissertations
Optically Detected Magnetic Resonance (ODMR) using nitrogen-vacancy (NV) centers in diamond enables sensitive, room-temperature magnetic field sensing, but real ODMR spectra are often noisy and difficult to analyze with traditional peak-fitting methods. This thesis investigates whether machine learning can reliably predict magnetic field strength directly from ODMR spectra, and compares four model families under a single regression task: a random forest, an artificial neural network (ANN), a one-dimensional convolutional neural network (1D-CNN), and a Transformer.
Training data were generated from an NV-ensemble simulation calibrated to real measurements provided by the Ulsan National Institute of Science and Technology (UNIST), spanning 0 …
Identifying Textual Predictors Of Early Termination In Clinical Trials In Medicine: An Explainable Machine-Learning Study, Rohan Ramnarain
Identifying Textual Predictors Of Early Termination In Clinical Trials In Medicine: An Explainable Machine-Learning Study, Rohan Ramnarain
Dissertations, Theses, and Capstone Projects
About one in five clinical trials in medicine ends early, wasting valuable resources and reducing the evidence available for developing life-saving medical treatments. This project uses a method called Trial2Vec, which is a self-supervised machine-learning method that converts clinical trial documents into dense numerical representations that capture their key design and clinical characteristics, to turn each proposed clinical trial’s written protocol into a compact numerical profile (a process referred to as embedding). These profiles are then paired with a predictive machine learning models to identify the words and phrases in the trial documents that can signal a higher risk of …
Understanding Delays In Emergency Department Care: A National Analysis Of Wait Times, Gregory Forsberg
Understanding Delays In Emergency Department Care: A National Analysis Of Wait Times, Gregory Forsberg
Mathematics, Statistics, and Computer Science Honors Projects
Emergency department (ED) wait times remain a persistent bottleneck in the United States healthcare system, impacting patient outcomes, hospital efficiency, and equitable access to care. This study analyzes nationally representative data from the National Hospital Ambulatory Medical Care Survey (NHAMCS), a complex, multi-stage probability sample. Using survey-weighted analyses and predictive modeling, we examine the effects of patient characteristics, triage acuity, and visit timing. Results indicate that operational and system-level factors, including hospital capacity, geographic region, and temporal variation, are among the most influential predictors of ED wait times
Enhancing Network Security Through Dual-Layer Log Analysis: Integrating Machine Learning Classifiers With Large Language Models For Intelligent Anomaly Detection, Anthony Burton-Cordova, O'Neil Gray, Mohammad Al Rousan
Enhancing Network Security Through Dual-Layer Log Analysis: Integrating Machine Learning Classifiers With Large Language Models For Intelligent Anomaly Detection, Anthony Burton-Cordova, O'Neil Gray, Mohammad Al Rousan
SMU Data Science Review
This paper presents an innovative approach to enhancing network security by integrating machine learning algorithms with fine-tuned large language models (LLMs) to provide an expert assistant querying. The proposed method utilizes machine learning for efficient preprocessing and feature extraction from log data, followed by the application of a fine-tuned LLM to analyze and interpret anomalies with greater accuracy. This dual-layer detection system is designed to improve the identification of subtle and sophisticated security threats. The research team’s extensive evaluation using real-world log datasets indicates that the combined approach increases detection rates and communicates results in an understandable manner, demonstrating its …
Predictive Analytics For Customer Churns In Financial Services., Thant Thiha
Predictive Analytics For Customer Churns In Financial Services., Thant Thiha
ICT
This project presents a customer churn prediction analysis in the telecommunications sector, achieving an ROC-AUC of approximately 0.86 using statistically validated features and interpretable AI models. Key churn drivers identified include the number of products held, customer age, and geographic location. Ensemble models, such as Random Forest and Gradient Boosting, provided the highest predictive performance. Ethical AI principles were applied to ensure fairness, transparency, privacy, and accountability. Business insights derived from the analysis inform targeted retention strategies, prioritising multi-product users, specific age groups, and geographic segments. Deployment recommendations include the tuned Random Forest model with ongoing monitoring, governance, and future …
Customer Response To Marketing Campaigns., Alexandru Enoiu
Customer Response To Marketing Campaigns., Alexandru Enoiu
ICT
This project applies machine learning to predict customer responses to marketing campaigns, aiming to enhance customer engagement and optimise resource allocation. Using a dataset containing demographic, lifestyle, and purchase behaviour data, a Random Forest classifier was developed to identify potential responders to marketing initiatives. The project follows a structured methodology including data exploration, preprocessing, model building, and evaluation. By accurately predicting customer behaviour, businesses can improve campaign targeting, design personalised marketing strategies, and allocate resources more efficiently. The results provide actionable insights to support customer segmentation and strategic decision-making, promoting business growth and customer satisfaction.
Using Machine Learning To Predict Credit Card Fraud, Pedro Henrique Das Chagas Morais
Using Machine Learning To Predict Credit Card Fraud, Pedro Henrique Das Chagas Morais
ICT
Credit card fraud poses a significant challenge to financial institutions, leading to substantial financial losses and declining customer trust. This project develops and evaluates machine learning models to detect fraudulent credit card transactions using a large, realistic synthetic dataset. Following data preprocessing, exploratory analysis, and class-balancing using SMOTE, four supervised models—Logistic Regression, Decision Tree, Random Forest, and XGBoost—were trained and compared. Performance was assessed using metrics suited to imbalanced classification, including AUC, Recall, Precision, F1-score, and Average Precision. Results show that XGBoost, particularly after hyperparameter optimisation, delivered the strongest performance (AUC 0.99, Recall 0.83, AP 0.70), outperforming other models and …
Chess Evaluation And Player Profiling Using Convolutional Neural Networks (Cnns) And Spatial Recognition., Joel D’Mello
Chess Evaluation And Player Profiling Using Convolutional Neural Networks (Cnns) And Spatial Recognition., Joel D’Mello
ICT
This thesis explores the feasibility of employing data analytics techniques in chess, with the purpose of profiling player styles and building a comprehensive chess analytics platform. The data set consists of over 20,000 anonymized games, and therefore, the study involved feature engineering, classification and visual analytics, in order to gain more insight into player decision making in chess. The data pre-processing part of analysis involved parsing Portable Game Notation (PGN) files, and feature engineering positional characteristics - material imbalance, pawn structure, king safety, and piece mobility - along with quantifying the sample using Average Centipawn Loss (ACPL) using Stockfish. ACPL …
Exploring The Determinants Of Life Expectancy At Birth: Predicting And Forecasting Global Health Trends Using Statistical, Machine Learning, And Deep Learning Models., Emma Rath
ICT
Accurate life expectancy forecasting is essential for health policy planning, yet research comparing statistical, machine learning, and deep learning approaches under real-world constraints remains limited. This study evaluates ARIMA/ARIMAX, tree-based, and neural network models using Irish and global datasets, considering small samples, missing data, and COVID-19 shocks. ARIMAX with lagged socioeconomic variables outperformed LSTM and other ML/DL methods. Income-based stratification improved predictive accuracy and interpretability, with SHAP analysis highlighting GDP per capita for developed countries and school enrolment and trade indicators for developing contexts. Results provide practical guidance for policymakers and establish limits for model complexity under constrained health data.
Forecasting Hourly Police Call For Service Volumes: A Comparative Analysis Of Statistical, Machine Learning And Neural Network Models For Operational Planning., Patrick Duggan
ICT
Accurate demand forecasting is critical in operational settings where resource allocation and planning decisions depend on anticipated service volumes. Transactional systems that capture timestamped records provide valuable data sources for developing demand forecasts. This study examines hourly call volume forecasting using New Orleans police calls for service data, comparing the performance of statistical models, tree-based methods, and recurrent neural networks.
The research evaluates four primary modelling approaches: ARIMA models representing traditional statistical methods, XGBoost and Random Forest as a tree-based ensemble technique, and Gated Recurrent Units (GRUs) as deep learning alternatives. A naive seasonal model serves as the baseline benchmark. …
Predicting S&P Corporate Credit Ratings Using Financial Ratios And Machine Learning: An Analysis Of European Non-Financial Companies., Gabriele Frattaroli
Predicting S&P Corporate Credit Ratings Using Financial Ratios And Machine Learning: An Analysis Of European Non-Financial Companies., Gabriele Frattaroli
ICT
This study investigates the prediction of multi-class S&P corporate credit ratings for European non-financial firms from 2010 to 2024 using a machine learning framework grounded in financial fundamentals. To ensure robustness and generalizability, the analysis excluded the Year variable, which was identified as a source of data leakage. After this correction, non-linear ensemble models demonstrated a clear advantage over linear baselines. The top-performing Random Forest model achieved a weighted F1-score of approximately 0.60, more than doubling the performance of the Logistic Regression benchmark used as a baseline (0.26), with most misclassifications concentrated in adjacent rating categories. This indicates that while …
Spatial Analysis And Machine Learning Integration For Nutritional Status Mapping Using Ann And Random Forest Models, Desi Anis Anggraini, Fachrul Kurniawan, Fresy Nugroho, Meidya Koeshardianto, Mohammad Iqbal Bachtiar
Spatial Analysis And Machine Learning Integration For Nutritional Status Mapping Using Ann And Random Forest Models, Desi Anis Anggraini, Fachrul Kurniawan, Fresy Nugroho, Meidya Koeshardianto, Mohammad Iqbal Bachtiar
Knowledge Engineering and Data Science
Nutritional problems among children under five remain a major public health challenge. This research seeks to create a spatially oriented system for evaluating and mapping nutritional status utilizing Artificial Neural Network (ANN) and Random Forest (RF) algorithms. Data obtained from the Sumenep District Health Office included age, weight, height, and gender variables. Both models were trained using a 70:30 data ratio and evaluated with accuracy, precision, recall, and F1-score metrics. The ANN model achieved an accuracy of 95.8%, while the RF model reached 97.7%. Classification results were visualized through a Geographic Information System (GIS) to illustrate spatial distribution and identify …
Classification Of Indonesian Sign Language (Sibi) Using Data Mining Algorithms K-Nearest Neighbor And Random Forest, Muhammad Zaki Wirawan, Achmad Afif, Anik Nur Handayani, Imanuel Hitipeuw, Osamu Fukuda
Classification Of Indonesian Sign Language (Sibi) Using Data Mining Algorithms K-Nearest Neighbor And Random Forest, Muhammad Zaki Wirawan, Achmad Afif, Anik Nur Handayani, Imanuel Hitipeuw, Osamu Fukuda
Knowledge Engineering and Data Science
This study aims to address the communication hallenges faced by the Indonesian deaf community by developing an automatic classification model for Sistem Bahasa Isyarat Indonesia (SIBI) using data mining techniques. The main objective is to identify a practical algorithm for recognizing SIBI hand gestures to enhance accessibility and inclusiveness in digital communication. A comprehensive dataset consisting of 32,850 gesture samples representing SIBI alphabet signs was collected and processed through feature extraction, data cleaning, and normalization using Z-Transform and Min-Max methods. Two classification algorithms, K-Nearest Neighbor (KNN) and Random Forest, were implemented and evaluated using metrics such as accuracy, precision, recall, …
Enhancing Imputation Accuracy: A Multi-Faceted Approach For Missing Data In Chicago Arrest Records, Steve Bramhall, Jae Chung, Nicholas Mueller
Enhancing Imputation Accuracy: A Multi-Faceted Approach For Missing Data In Chicago Arrest Records, Steve Bramhall, Jae Chung, Nicholas Mueller
SMU Data Science Review
This paper introduces a novel approach to enhance the imputation process for missing data, utilizing crime records from Chicago with arrests as the target feature. Robust imputation techniques are crucial in the era of burgeoning datasets for generating reliable insights. Our core objective is to present an innovative method that improves imputation techniques, augmenting model performance and bolstering the reliability of analytical outcomes. Leveraging numeric crime data, we establish a Gradient Boosting (GBM) baseline model, then introduce ensemble methods including Random Forest and Decision Trees for further refinement. By systematically exploring multiple imputation processes, we establish a baseline for comparative …
Classification Models Using Python In Industrial/Organizational Psychology, Beyza Ceylan
Classification Models Using Python In Industrial/Organizational Psychology, Beyza Ceylan
Williams Honors College, Honors Research Projects
Companies, industries, and places of business use artificial intelligence and statistics to predict the characteristics of their employees and staff. Data collected from these individuals is also used to make decisions about them regarding their work life, such as promotions, salaries, or within the hiring process. Two models that are commonly used throughout the field of psychology and specifically in industrial/organizational psychology are the linear regression and the logistic regression. Examining different classification models using Python shows the potential that there may be different models that are more accurate in their predictions of employee success, including a Random Forest model …
Eeg Classification While Listening To Murottal Al-Quran And Classical Music Using Random Forest Method, Heni Sumarti, Fahira Septiani, Agus Sudarmanto, Wahyu Caesarendra, Rizki Edmi Edison
Eeg Classification While Listening To Murottal Al-Quran And Classical Music Using Random Forest Method, Heni Sumarti, Fahira Septiani, Agus Sudarmanto, Wahyu Caesarendra, Rizki Edmi Edison
Knowledge Engineering and Data Science
This study is aimed to classify the brain activity of adolescents associated with audio stimuli; murottal Al-Quran and classical music. The raw data were filtered using Independent Component Analisys (ICA) and followed by band-pass filter in Python on the Google Colab Extraction was processed with Power Spectral Density (PSD) and the Random Forest Method in Weka Machine Learning was used for classification. The research results showed the same results between the two types of stimulation, namely the order of brain waves from highest to lowest were delta, alpha, theta and beta. The average brain waves of teenagers when given murottal …
Optimizing Random Forest Algorithm To Classify Player's Memorisation Via In-Game Data, Akmal Vrisna Alzuhdi, Harits Ar Rosyid, Mohammad Yasser Chuttur, Shah Nazir
Optimizing Random Forest Algorithm To Classify Player's Memorisation Via In-Game Data, Akmal Vrisna Alzuhdi, Harits Ar Rosyid, Mohammad Yasser Chuttur, Shah Nazir
Knowledge Engineering and Data Science
Assessment of a player's knowledge in game education has been around for some time. Traditional evaluation in and around a gaming session may disrupt the players' immersion. This research uses an optimized Random Forest to construct a non-invasive prediction of a game education player's Memorization via in-game data. Firstly, we obtained the dataset from a 3-month survey to record in-game data of 50 players who play 4-15 game stages of the Chem Fight (a test case game). Next, we generated three variants of datasets via the preprocessing stages: resampling method (SMOTE), normalization (min-max), and a combination of resampling and normalization. …
Distance Correlation Based Feature Selection In Random Forest, Jose Munoz-Lopez
Distance Correlation Based Feature Selection In Random Forest, Jose Munoz-Lopez
Electronic Theses, Projects, and Dissertations
The Pearson correlation coefficient is a commonly used measure of correlation, but it has limitations as it only measures the linear relationship between two numerical variables. In 2007, Szekely et al. introduced the distance correlation, which measures all types of dependencies between random vectors X and Y in arbitrary dimensions, not just the linear ones. In this thesis, we propose a filter method that utilizes distance correlation as a criterion for feature selection in Random Forest regression. We conduct extensive simulation studies to evaluate its performance compared to existing methods under various data settings, in terms of the prediction mean …
Analysis Of Chemical Elements In Basalts Using Mislabeled Data, A Machine Learning Approach, Jenifer Vivar
Analysis Of Chemical Elements In Basalts Using Mislabeled Data, A Machine Learning Approach, Jenifer Vivar
Dissertations and Theses
Scientists use basalt chemistry to discriminate among different tectonic settings. There are well-known chemical elements used to classify tectonic settings. An exploration of new features is done using Logistic Regression and Random Forest to discover any new elements of interest. The models were used with other tools, such as recursive feature elimination and permutations, to increase reliability. Among the scarcely explored chemical elements are Terbium (Tb), Holmium (Ho), Samarium (Sm), and Erbium (Er). The data used for the exploration contained many outliers. Therefore, an ensemble model was created to explore the location and composition of such outliers. The ensemble was …
Classification Of Pixel Tracks To Improve Track Reconstruction From Proton-Proton Collisions, Kebur Fantahun, Jobin Joseph, Halle Purdom, Nibhrat Lohia
Classification Of Pixel Tracks To Improve Track Reconstruction From Proton-Proton Collisions, Kebur Fantahun, Jobin Joseph, Halle Purdom, Nibhrat Lohia
SMU Data Science Review
In this paper, machine learning techniques are used to reconstruct particle collision pathways. CERN (Conseil européen pour la recherche nucléaire) uses a massive underground particle collider, called the Large Hadron Collider or LHC, to produce particle collisions at extremely high speeds. There are several layers of detectors in the collider that track the pathways of particles as they collide. The data produced from collisions contains an extraneous amount of background noise, i.e., decays from known particle collisions produce fake signal. Particularly, in the first layer of the detector, the pixel tracker, there is an overwhelming amount of background noise that …
Analysis Of The Electric Power Outage Data And Prediction Of Electric Power Outage For Major Metropolitan Areas In Texas Using Machine Learning And Time Series Methods, Renfeng Wang, Venkata Leela 'Mg' Vanga, Zachary B. Zaiken, Jonathan Bennett
Analysis Of The Electric Power Outage Data And Prediction Of Electric Power Outage For Major Metropolitan Areas In Texas Using Machine Learning And Time Series Methods, Renfeng Wang, Venkata Leela 'Mg' Vanga, Zachary B. Zaiken, Jonathan Bennett
SMU Data Science Review
With growing energy usage, power outages affect millions of households. This case study focuses on gathering power outage historical data, modifying the data to attach weather attributes, and gathering ERCOT energy market conditions for Dallas-Fort Worth and Houston metropolitan areas of Texas. The transformed data is then analyzed using machine learning algorithms including, but not limited to, Regression, Random Forests and XGBoost to consider current weather and ERCOT features and predict power outage percentage for locations. The transformed data is also trained using time series models and serially correlated models including Autoregression and Vector Autoregression. This study also focuses on …
Urban Traffic Simulation: Network And Demand Representation Impacts On Congestion Metrics, Aaron Faltesek, Balasubramaniam Dakshinamoorthi, Sreeni Prabhala, Akbar Thobani, Anu Kuncheria, Jane Macfarlane
Urban Traffic Simulation: Network And Demand Representation Impacts On Congestion Metrics, Aaron Faltesek, Balasubramaniam Dakshinamoorthi, Sreeni Prabhala, Akbar Thobani, Anu Kuncheria, Jane Macfarlane
SMU Data Science Review
Traffic simulations are often used by city planners as a basis for predicting the impact of policies, plans, and operations. The complexities underpinning traffic simulations are often not described in detail yet can significantly impact the simulation outcome. Conflating underlying data for simulations is complex and hinders the interest in this type of exploration. This paper aims to elucidate critical features of traffic simulations that drive the generated metrics of the modeled urban environment. Specifically, this paper examines differences in two road graph networks for the metropolitan region of Houston, TX: a reduced network composed of 45,675 road links and …
Identifying Vacant Lots To Reduce Violent Crime In Dallas, Texas, Laura Lazarescou, Andrew Mejia, Tina Pai, Sabrina Purvis, Robert Slater, Owen Wilson-Chavez
Identifying Vacant Lots To Reduce Violent Crime In Dallas, Texas, Laura Lazarescou, Andrew Mejia, Tina Pai, Sabrina Purvis, Robert Slater, Owen Wilson-Chavez
SMU Data Science Review
Vacant lots have been associated with community violence for many years. Researchers have confirmed a positive correlation between vacant lots and vacant buildings with increased violence in urban and rural geographies. However, identifying vacant lots has been a challenge, and modeling methods were largely manual and time-intensive. This prevented cities and non-profit organizations from acting on the information since it was expensive and high-risk to develop remediation programs without clearly understanding where or how many vacant lots existed.
The primary objective of this study was to provide a predictive model that accelerates and improves the accuracy of prior land classification …
Forecasting The Daily Percentage Of Delayed Flights Based On The National Weather Data, Parto Mahmoudi
Forecasting The Daily Percentage Of Delayed Flights Based On The National Weather Data, Parto Mahmoudi
Graduate Student Theses, Dissertations, & Professional Papers
Flight delays cost airlines and affect passenger’s satisfaction. In this research work, we predicted the daily percentage of delayed flights based on the national weather data using the multiple linear regression and the random forest models. We extracted the passenger flight on-time performance data from the Bureau of Transportation Statistics and the weather dataset from NOAA National Centers for Environmental Information for the years from 2015 to 2019. We used the flight dataset for Seattle airport as the origin. We predicted the daily percentage of delayed flights for the Seattle-originated flights based on the features such as weather conditions of …
Classifying Imbalanced Financial Fraud Data Utilizing Enhanced Random Forest Algorithm, Charles Gardner
Classifying Imbalanced Financial Fraud Data Utilizing Enhanced Random Forest Algorithm, Charles Gardner
Master of Science in Computer Science Theses
Imbalanced datasets have been a unique challenge for machine learning, requiring specialized approaches to correctly classify the minority class. Financial fraud detection involves using highly imbalanced datasets with a class imbalance of up to .01% frauds to 99.99% regular transactions. It is essential to identify all frauds in financial fraud detection, even if some classifications' precision is low. I developed a random forest assembly that separates fraudulent transactions into tiers of precision. With this approach, 96% of fraudulent transactions are identified, showing an 8% increase in recall when compared to standard approaches. 59% of fraud classifications' precision increases by 10% …
Comparing Variable Importance In Prediction Of Silence Behaviours Between Random Forest And Conditional Inference Forest Models., Stephen Barrett Dr, Geraldine Gray Dr, Colm Mcguinness Dr, Michael Knoll Dr.
Comparing Variable Importance In Prediction Of Silence Behaviours Between Random Forest And Conditional Inference Forest Models., Stephen Barrett Dr, Geraldine Gray Dr, Colm Mcguinness Dr, Michael Knoll Dr.
Articles
This paper explores variable importance metrics of Conditional Inference Trees (CIT) and classical Classification And Regression Trees (CART) based Random Forests. The paper compares both algorithms variable importance rankings and highlights why CIT should be used when dealing with data with different levels of aggregation. The models analysed explored the role of cultural factors at individual and societal level when predicting Organisational Silence behaviours.
Machine Learning Approaches For Improving Prediction Performance Of Structure-Activity Relationship Models, Gabriel Idakwo
Machine Learning Approaches For Improving Prediction Performance Of Structure-Activity Relationship Models, Gabriel Idakwo
Dissertations
In silico bioactivity prediction studies are designed to complement in vivo and in vitro efforts to assess the activity and properties of small molecules. In silico methods such as Quantitative Structure-Activity/Property Relationship (QSAR) are used to correlate the structure of a molecule to its biological property in drug design and toxicological studies. In this body of work, I started with two in-depth reviews into the application of machine learning based approaches and feature reduction methods to QSAR, and then investigated solutions to three common challenges faced in machine learning based QSAR studies.
First, to improve the prediction accuracy of learning …
Analysis On Suicidal Ideation Among Adolescents (12-17 Years) In The Usa, Himani Raturi
Analysis On Suicidal Ideation Among Adolescents (12-17 Years) In The Usa, Himani Raturi
Electronic Theses, Projects, and Dissertations
Suicide is one of the leading health concerns in United States among adolescents and the presence of suicidal ideation (SI) is quite high, with ~20-30% of adolescents reporting it at some point. Though we have seen growth and development in the prevention of suicide, there is limited research on the ability to identify the adolescents which might be at risk for SI. The objective behind the project is to identify adolescents with SI using machine learning.
The project shows statistics from different articles on adolescents in the U.S. For this study, adolescent data was taken from NSDUH 2018. Moreover, detailed …
An Application Of Machine Learning To Explore Relationships Between Factors Of Organisational Silence And Culture, With Specific Focus On Predicting Silence Behaviours, Stephen Barrett Dr
An Application Of Machine Learning To Explore Relationships Between Factors Of Organisational Silence And Culture, With Specific Focus On Predicting Silence Behaviours, Stephen Barrett Dr
Articles
Research indicates that there are many individual reasons why people do not speak up when confronted with situations that may concern them within their working environment. One of the areas that requires more focused research is the role culture plays in why a person may remain silent when such situations arise. The purpose of this study is to use data science techniques to explore the patterns in a data set that would lead a person to engage in organisational silence. The main research question the thesis asks is: Is Machine Learning a tool that Social Scientists can use with respect …
Automated Morgan Keenan Classification Of Observed Stellar Spectra Collected By The Sloan Digital Sky Survey Using A Single Classifier, Michael J. Brice, Răzvan Andonie
Automated Morgan Keenan Classification Of Observed Stellar Spectra Collected By The Sloan Digital Sky Survey Using A Single Classifier, Michael J. Brice, Răzvan Andonie
All Faculty Scholarship for the College of the Sciences
The classification of stellar spectra is a fundamental task in stellar astrophysics. Stellar spectra from the Sloan Digital Sky Survey are applied to standard classification methods, k-nearest neighbors and random forest, to automatically classify the spectra. Stellar spectra are high dimensional data and the dimensionality is reduced using astronomical knowledge because classifiers work in low dimensional space. These methods are utilized to classify the stellar spectra into a complete Morgan Keenan classification (spectral and luminosity) using a single classifier. The motion of stars (radial velocity) causes machine-learning complications through the feature matrix when classifying stellar spectra. Due to the nature …