Open Access. Powered by Scholars. Published by Universities.®

Digital Commons Network

Open Access. Powered by Scholars. Published by Universities.®

PDF

University of Central Florida

Honors Undergraduate Theses

Physical Sciences and Mathematics

Machine Learning

Publication Year

Articles 1 - 2 of 2

Full-Text Articles in Entire DC Network

Biomarker Identification For Breast Cancer Types Using Feature Selection And Explainable Ai Methods, David E. La Rosa Giraud Jan 2023

Biomarker Identification For Breast Cancer Types Using Feature Selection And Explainable Ai Methods, David E. La Rosa Giraud

Honors Undergraduate Theses

This paper investigates the impact the LASSO, mRMR, SHAP, and Reinforcement Feature Selection techniques on random forest models for the breast cancer subtypes markers ER, HER2, PR, and TN as well as identifying a small subset of biomarkers that could potentially cause the disease and explain them using explainable AI techniques. This is important because in areas such as healthcare understanding why the model makes a specific decision is important it is a diagnostic of an individual which requires reliable AI. Another contribution is using feature selection methods to identify a small subset of biomarkers capable of predicting if a …


Genetic Algorighm Representation Selection Impact On Binary Classification Problems, Stephen V. Maldonado Jan 2022

Genetic Algorighm Representation Selection Impact On Binary Classification Problems, Stephen V. Maldonado

Honors Undergraduate Theses

In this thesis, we explore the impact of problem representation on the ability for the genetic algorithms (GA) to evolve a binary prediction model to predict whether a physical therapist is paid above or below the median amount from Medicare. We explore three different problem representations, the vector GA (VGA), the binary GA (BGA), and the proportional GA (PGA). We find that all three representations can produce models with high accuracy and low loss that are better than Scikit-Learn’s logistic regression model and that all three representations select the same features; however, the PGA representation tends to create lower weights …