Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Statistics and Probability (37)
- Computer Sciences (34)
- Social and Behavioral Sciences (23)
- Artificial Intelligence and Robotics (18)
- Applied Statistics (16)
-
- Statistical Models (16)
- Business (14)
- Medicine and Health Sciences (14)
- Engineering (13)
- Applied Mathematics (10)
- Longitudinal Data Analysis and Time Series (10)
- Theory and Algorithms (7)
- Categorical Data Analysis (6)
- Statistical Methodology (6)
- Biostatistics (5)
- Electrical and Computer Engineering (5)
- Finance and Financial Management (5)
- Life Sciences (5)
- Mathematics (5)
- Other Computer Sciences (5)
- Diseases (4)
- Education (4)
- Environmental Sciences (4)
- Information Security (4)
- Law (4)
- Numerical Analysis and Computation (4)
- Other Public Health (4)
- Public Affairs, Public Policy and Public Administration (4)
- Keyword
-
- Machine Learning (17)
- NLP (15)
- Data Science (14)
- Machine learning (12)
- CNN (10)
-
- Deep Learning (9)
- Natural language processing (9)
- Statistics (8)
- Time series (8)
- Deep learning (7)
- Neural Networks (6)
- Random Forest (6)
- COVID-19 (5)
- Classification (5)
- LSTM (5)
- ARIMA (4)
- Clustering (4)
- Computer Science (4)
- Computer vision (4)
- GAN (4)
- LLM (4)
- LLMs (4)
- Natural Language Processing (4)
- XGBoost (4)
- AI (3)
- American Community Survey (3)
- BERT (3)
- Bias (3)
- Biostatistics (3)
- Data science (3)
- Publication
- Publication Type
Articles 91 - 120 of 144
Full-Text Articles in Data Science
Examining Bias In Jury Selection For Criminal Trials In Dallas County, Megan Ball, Brandon Birmingham, Matt Farrow, Katherine Mitchell, Bivin Sadler, Lynne Stokes
Examining Bias In Jury Selection For Criminal Trials In Dallas County, Megan Ball, Brandon Birmingham, Matt Farrow, Katherine Mitchell, Bivin Sadler, Lynne Stokes
SMU Data Science Review
One of the hallmarks of the American judicial system is the concept of trial by jury, and for said trial to consist of an impartial jury of your peers. Several landmark legal cases in the history of the United States have challenged this notion of equal representation by jury—most notably Batson v. Kentucky, 476 U.S. 79 (1986). Most of the previous research, focus, and legal precedence has centered around peremptory challenges and attempting to prove if bias was suspected in excluding certain jurors from serving. Few studies, however, focus on examining challenges for cause based on self-reported biases from the …
Application Of Probabilistic Ranking Systems On Women’S Junior Division Beach Volleyball, Cameron Stewart, Michael Mazel, Bivin Sadler
Application Of Probabilistic Ranking Systems On Women’S Junior Division Beach Volleyball, Cameron Stewart, Michael Mazel, Bivin Sadler
SMU Data Science Review
Women’s beach volleyball is one of the fastest growing collegiate sports today. The increase in popularity has come with an increase in valuable scholarship opportunities across the country. With thousands of athletes to sort through, college scouts depend on websites that aggregate tournament results and rank players nationally. This project partnered with the company Volleyball Life, who is the current market leader in the ranking space of junior beach volleyball players. Utilizing the tournament information provided by Volleyball Life, this study explored replacements to the current ranking systems, which are designed to aggregate player points from recent tournament placements. Three …
Market Segmentation And Recency Frequency Monetary Value Analysis For A Freemium Mobile Game, Satvik Ajmera, Taylor Bonar, Dylan Scott, Carol Miu, Alana Manuel
Market Segmentation And Recency Frequency Monetary Value Analysis For A Freemium Mobile Game, Satvik Ajmera, Taylor Bonar, Dylan Scott, Carol Miu, Alana Manuel
SMU Data Science Review
Bricks ‘N Balls is a freemium game that relies on in-app purchases and ad monetization from users to be profitable at no upfront cost to the players. This study explores how in-game data analytics and purchase data can be used to segment players. Features taken into consideration for segmentation include past purchasing habits along with the players interactions within the missions. This study uses the Recency Frequency Monetary Value (RFM) framework to extract insights on player purchasing behavior to segment players into clusters and predict how much users will spend in the future.
Using Hospital Bed Capacity Prediction During Covid-19 To Determine Feature Importance, Helene Barrera, Justin Ehly, Blake Freeman, Chris Papesh, Brad Blanchard
Using Hospital Bed Capacity Prediction During Covid-19 To Determine Feature Importance, Helene Barrera, Justin Ehly, Blake Freeman, Chris Papesh, Brad Blanchard
SMU Data Science Review
The COVID-19 pandemic has exacerbated existing hospital capacity limitations in the United States, causing hospitals in certain regions to hit maximum capacity. The purpose of this study is to investigate key features of COVID-19 related admissions to help create a higher level of public understanding and help guide healthcare management professionals and governments when considering preventive measures. The introduction of preventative measures and new regulations during the pandemic have led to the generation of multiple types of models and feature selection methods in the field of Machine Learning that are increasingly complicated. This study focuses on the exploration of feature …
A Machine Learning Approach To Revenue Generation Within The Professional Hair Care Industry, Alexander K. Sepenu, Linda Eliasen
A Machine Learning Approach To Revenue Generation Within The Professional Hair Care Industry, Alexander K. Sepenu, Linda Eliasen
SMU Data Science Review
The cosmetic and beauty industry continues to grow and evolve to satisfy its patrons. In the United States, the industry is heavily science-driven, innovative, and fast-paced, suggesting that to remain productive and profitable, companies must seek smart alternatives to their current modus operandi or risk losing out on this multi-billion-dollar industry to fierce competition. In this paper, the authors seek to utilize machine learning models such as clustering and regression to improve the efficiency of current sales and customer segmentation models to help HairCo (pseudonym for confidentiality), a professional hair products manufacturer, strategize their marketing and sales efforts for revenue …
Analysis Of The Electric Power Outage Data And Prediction Of Electric Power Outage For Major Metropolitan Areas In Texas Using Machine Learning And Time Series Methods, Renfeng Wang, Venkata Leela 'Mg' Vanga, Zachary B. Zaiken, Jonathan Bennett
Analysis Of The Electric Power Outage Data And Prediction Of Electric Power Outage For Major Metropolitan Areas In Texas Using Machine Learning And Time Series Methods, Renfeng Wang, Venkata Leela 'Mg' Vanga, Zachary B. Zaiken, Jonathan Bennett
SMU Data Science Review
With growing energy usage, power outages affect millions of households. This case study focuses on gathering power outage historical data, modifying the data to attach weather attributes, and gathering ERCOT energy market conditions for Dallas-Fort Worth and Houston metropolitan areas of Texas. The transformed data is then analyzed using machine learning algorithms including, but not limited to, Regression, Random Forests and XGBoost to consider current weather and ERCOT features and predict power outage percentage for locations. The transformed data is also trained using time series models and serially correlated models including Autoregression and Vector Autoregression. This study also focuses on …
Web Page Multiclass Classification, Brian Gaither, Antonio Debouse, Catherine Huang
Web Page Multiclass Classification, Brian Gaither, Antonio Debouse, Catherine Huang
SMU Data Science Review
As the internet age evolves, the volume of content hosted on the Web is rapidly expanding. With this ever-expanding content, the capability to accurately categorize web pages is a current challenge to serve many use cases. This paper proposes a variation in the approach to text preprocessing pipeline whereby noun phrase extraction is performed first followed by lemmatization, contraction expansion, removing special characters, removing extra white space, lower casing, and removal of stop words. The first step of noun phrase extraction is aimed at reducing the set of terms to those that best describe what the web pages are about …
Anomaly Detection Methods To Improve Supply Chain Data Quality And Operations, Ana E. Glaser, Jake P. Harrison, David Josephs
Anomaly Detection Methods To Improve Supply Chain Data Quality And Operations, Ana E. Glaser, Jake P. Harrison, David Josephs
SMU Data Science Review
Supply chain operations drive the planning, manufacture, and distribution of billions of semiconductors a year, spanning thousands of products across many supply chain configurations. The customizations span from wafer technology to die stacking and chip feature enablement. Data quality drives efficiency in these processes and anomalies in data can be very disruptive, and at times, consequential. Developing preventative measures that automate the detection of anomalies before they reach downstream execution systems would result in significant efficiency gain for the organization. The purpose of this research is to identify an effective, actionable, and computationally efficient approach to highlight anomalies in a …
Adjusting Community Survey Data Benchmarks For External Factors, Allen Miller, Nicole M. Norelli, Robert Slater, Mingyang N. Yu
Adjusting Community Survey Data Benchmarks For External Factors, Allen Miller, Nicole M. Norelli, Robert Slater, Mingyang N. Yu
SMU Data Science Review
Abstract. Using U.S. resident survey data from the National Community Survey in combination with public data from the U.S. Census and additional sources, a Voting Regressor Model was developed to establish fair benchmark values for city performance. These benchmarks were adjusted for characteristics the city cannot easily influence that contribute to confidence in local government, such as population size, demographics, and income. This adjustment allows for a more meaningful comparison and interpretation of survey results among individual cities. Methods explored for the benchmark adjustment included cluster analysis, anomaly detection, and a variety of regression techniques, including random forest, ridge, decision …
Aspect-Based Sentiment Analysis Of Movie Reviews, Samuel Onalaja, Eric Romero, Bosang Yun
Aspect-Based Sentiment Analysis Of Movie Reviews, Samuel Onalaja, Eric Romero, Bosang Yun
SMU Data Science Review
This study investigates a comparison of classification models used to determine aspect based separated text sentiment and predict binary sentiments of movie reviews with genre and aspect specific driving factors. To gain a broader classification analysis, five machine and deep learning algorithms were compared: Logistic Regression (LR), Naive Bayes (NB), Support Vector Machine (SVM), and Recurrent Neural Network Long-Short-Term Memory (RNN LSTM). The various movie aspects that are utilized to separate the sentences are determined through aggregating aspect words from lexicon-base, supervised and unsupervised learning. The driving factors are randomly assigned to various movie aspects and their impact tied to …
Reading Level Identification Using Natural Language Processing Techniques, William Arnost, Ellen Lull, Joseph Schueder, Joseph Engler
Reading Level Identification Using Natural Language Processing Techniques, William Arnost, Ellen Lull, Joseph Schueder, Joseph Engler
SMU Data Science Review
This paper investigates using the Bidirectional Encoder Representations from Transformers (BERT) algorithm and lexical-syntactic features to measure readability. Readability is important in many disciplines, for functions such as selecting passages for school children, assessing the complexity of publications, and writing documentation. Text at an appropriate reading level will help make communication clear and effective. Readability is primarily measured using well-established statistical methods. Recent advances in Natural Language Processing (NLP) have had mixed success incorporating higher-level text features in a way that consistently beats established metrics. This paper contributes a readability method using a modern transformer technique and compares the results …
Predicting Power Using Time Series Analysis Of Power Generation And Consumption In Texas, Joshua Eysenbach, Bodie Franklin, Andrew J. Larsen, Joel Lindsey
Predicting Power Using Time Series Analysis Of Power Generation And Consumption In Texas, Joshua Eysenbach, Bodie Franklin, Andrew J. Larsen, Joel Lindsey
SMU Data Science Review
Due to the recent power events in Texas, power forecasting has been brought national attention. Accurate demand forecasting is necessary to be sure that there is adequate power supply to meet consumer's needs. While Texas has a forecasting model created by the Electricity Reliability Council of Texas (ERCOT), constant efforts are required to ensure that the model stays at the state-of-the-art and is producing the most reliable forecasts possible. This research seeks to provide improved short- and medium-term forecasting models, bringing in state-of-the-art deep learning models to compare to ERCOT’s forecasts. A model that is more accurate than ERCOT’s own …
Emotion Integrated Music Recommendation System Using Generative Adversarial Networks, Mrinmoy Bhaumik, Patrica U. Attah, Faizan Javed
Emotion Integrated Music Recommendation System Using Generative Adversarial Networks, Mrinmoy Bhaumik, Patrica U. Attah, Faizan Javed
SMU Data Science Review
Music can stimulate emotions within us; hence is often called the “language of emotion.” This study explores emotion as an additional feature in generating a playlist with a deep learning model to improve the current music recommendation system. This study will sample emotions from certain subjects for each song in a sample of the data. Since the effect of music on emotion is subjective and is different person to person, this study would need a considerable number of subjects to reduce subjectivity. Due to the limited resources, a portion of the data will be labeled with emotion from subjects and …
Alternative Methods For Deriving Emotion Metrics In The Spotify® Recommendation Algorithm, Ronald M. Sherga Jr., David Wei, Neil Benson, Faizan Javed
Alternative Methods For Deriving Emotion Metrics In The Spotify® Recommendation Algorithm, Ronald M. Sherga Jr., David Wei, Neil Benson, Faizan Javed
SMU Data Science Review
Spotify's® recommendation algorithm tailors music offerings to create a unique listening experience for each user. Though what this recommender does is highly impressive, there is always room for improvement given that these techniques are not fully prescient. This study posits that in addition to creating certain features based on audio analysis, incorporating new features derived from album art color as well as lyrical sentiment analysis may provide additional value to the end user. This team did not find that a significant difference existed between color valence and Spotify® valence; however, all other comparisons resulted in statistically significant difference of means …
Clinical Diagnosis Support With Convolutional Neural Network By Transfer Learning, Spencer Fogleman, Jeremy Otsap, Sangrae Cho
Clinical Diagnosis Support With Convolutional Neural Network By Transfer Learning, Spencer Fogleman, Jeremy Otsap, Sangrae Cho
SMU Data Science Review
Breast cancer is prevalent among women in the United States. Breast cancer screening is standard but requires a radiologist to review screening images to make a diagnosis. Diagnosis through the traditional screening method of mammography currently has an accuracy of about 78% for women of all ages and demographics. A more recent and precise technique called Digital Breast Tomosynthesis (DBT) has shown to be more promising but is less well studied. A machine learning model trained on DBT images has the potential to increase the success of identifying breast cancer and reduce the time it takes to diagnose a patient, …
Covid-19 - A Graph Network Approach, Nibhrat Lohia, Rajesh Satluri, Suchismita Moharana, Venkat Kasarla
Covid-19 - A Graph Network Approach, Nibhrat Lohia, Rajesh Satluri, Suchismita Moharana, Venkat Kasarla
SMU Data Science Review
The effects of COVID-19 and its spreads are attributed to various factors. This study uses CDC open-source data on COVID-19 effected population with features ranging from location to ethnicity, to create a Knowledge Graph to measure the similarity between COVID-19 cases and estimate the risk for people likely affected by COVID-19. This data could be used to find correlations between distinct factors, like ethnicity and pre-existing health conditions, to find the vulnerability of a given COVID-19 patient. Using the Jaccard similarity coefficient, in the knowledge graph, we are able to identify and explore relationships between COVID-19 cases as well as …
Rocket Learn, Daanesh Ibrahim, Jules Stacy, David Stroud, Yusi Zhang
Rocket Learn, Daanesh Ibrahim, Jules Stacy, David Stroud, Yusi Zhang
SMU Data Science Review
Abstract. This paper covers the development, testing, and implementation of Reinforcement Learning methods designed to autonomously learn and optimize Rocket League play. This study aims to analyze and benchmark model frameworks commonly used in Reinforcement Learning applications. These models can be applied to tasks ranging in difficulty from simple to superhumanly complex, and this study will begin with and build upon simple models performing simple tasks. It will result in complex models performing difficult tasks. Models will be allowed to train autonomously on the game using mass parallelization to expedite training times with the goal of maximizing reward function scores. …
Pokégan: P2p (Pet To Pokémon) Stylizer, Michael B. Hedge, Morgan Nelson, Thomas Pengilly, Michael Weatherford
Pokégan: P2p (Pet To Pokémon) Stylizer, Michael B. Hedge, Morgan Nelson, Thomas Pengilly, Michael Weatherford
SMU Data Science Review
This paper covers the development, testing, and implementation of an automatic framework for converting common images of pets into a Pokémon cartoon with the style of a Pokémon trading card. The technique will first implement object detection for common animals to facilitate image segmentation and apply the appropriate style transfer model to ensure the most aesthetic stylization. It explores various methods to address artifacts in the results of common neural style transfer techniques using Generative Adversarial Networks (GANs). This research sets up a framework to create an app that converts user-submitted pet pictures to Pokémon styled images using the most …
Machine Learning Approach To Distinguish Ulcerative Colitis And Crohn’S Disease Using Smote (Synthetic Minority Oversampling Technique) Methods, Kris Ghimire, Walter Lai, Yasser Omar, Thad Schwebke, Jamie Vo
Machine Learning Approach To Distinguish Ulcerative Colitis And Crohn’S Disease Using Smote (Synthetic Minority Oversampling Technique) Methods, Kris Ghimire, Walter Lai, Yasser Omar, Thad Schwebke, Jamie Vo
SMU Data Science Review
Irritable Bowel Disease (IBD) affects a sizable portion of the US population, causing symptoms such as vomiting, abdominal pain, and diarrhea. Despite the disease’s prevalence, the precise cause is not fully understood. This study consists of endoscopic and histological data from patients diagnosed with IBD and a control population for reference. The machine learning models' focus is to classify patients into IBD types. Several models were analyzed, including decision trees, logistic regression, and k-nearest neighbors. In addition, various methods of SMOTE were applied to determine the most effective transformation and ensuring that the dataset is balanced. The best model with …
Urban Traffic Simulation: Network And Demand Representation Impacts On Congestion Metrics, Aaron Faltesek, Balasubramaniam Dakshinamoorthi, Sreeni Prabhala, Akbar Thobani, Anu Kuncheria, Jane Macfarlane
Urban Traffic Simulation: Network And Demand Representation Impacts On Congestion Metrics, Aaron Faltesek, Balasubramaniam Dakshinamoorthi, Sreeni Prabhala, Akbar Thobani, Anu Kuncheria, Jane Macfarlane
SMU Data Science Review
Traffic simulations are often used by city planners as a basis for predicting the impact of policies, plans, and operations. The complexities underpinning traffic simulations are often not described in detail yet can significantly impact the simulation outcome. Conflating underlying data for simulations is complex and hinders the interest in this type of exploration. This paper aims to elucidate critical features of traffic simulations that drive the generated metrics of the modeled urban environment. Specifically, this paper examines differences in two road graph networks for the metropolitan region of Houston, TX: a reduced network composed of 45,675 road links and …
Intelligent Investment Portfolio Management Using Time-Series Analytics And Deep Reinforcement Learning, Sachin Chavan, Pradeep Kumar, Tom Gianelle
Intelligent Investment Portfolio Management Using Time-Series Analytics And Deep Reinforcement Learning, Sachin Chavan, Pradeep Kumar, Tom Gianelle
SMU Data Science Review
Abstract. – With globalization, the capital markets have exploded in size and value, making them exceedingly difficult to predict. These days the public has access to real-time data of the market-leading to more participation. As a positive step, this might lead to better wealth distribution in the society, and it also adds to the random nature of the market, making it more unpredictable. The portfolio accounts consisting of stocks and bonds are considered serious investment assets. They can make or break a person’s future. It is also a way of shielding one against market risk or rising inflation. These accounts, …
Identifying Vacant Lots To Reduce Violent Crime In Dallas, Texas, Laura Lazarescou, Andrew Mejia, Tina Pai, Sabrina Purvis, Robert Slater, Owen Wilson-Chavez
Identifying Vacant Lots To Reduce Violent Crime In Dallas, Texas, Laura Lazarescou, Andrew Mejia, Tina Pai, Sabrina Purvis, Robert Slater, Owen Wilson-Chavez
SMU Data Science Review
Vacant lots have been associated with community violence for many years. Researchers have confirmed a positive correlation between vacant lots and vacant buildings with increased violence in urban and rural geographies. However, identifying vacant lots has been a challenge, and modeling methods were largely manual and time-intensive. This prevented cities and non-profit organizations from acting on the information since it was expensive and high-risk to develop remediation programs without clearly understanding where or how many vacant lots existed.
The primary objective of this study was to provide a predictive model that accelerates and improves the accuracy of prior land classification …
Identification And Characterization Of Forest Fire Risk Zones Leveraging Machine Learning Methods, Joshua Balson, Matt Chinchilla, Cam Lu, Jeff Washburn, Nibhrat Lohia
Identification And Characterization Of Forest Fire Risk Zones Leveraging Machine Learning Methods, Joshua Balson, Matt Chinchilla, Cam Lu, Jeff Washburn, Nibhrat Lohia
SMU Data Science Review
Across the United States, record numbers of wildfires are observed costing billions of dollars in property damage, polluting the environment, and putting lives at risk. The ability of emergency management professionals, city planners, and private entities such as insurance companies to determine if an area is at higher risk of a fire breaking out has never been greater. This paper proposes a novel methodology for identifying and characterizing zones with increased risks of forest fires. Methods involving machine learning techniques use the widely available and recorded data, thus making it possible to implement the tool quickly.
Qualitative Leveraging Natural Language Processing To Establish Judge Incrimination Statistics To Educate Voters In Re-Elections, Aurian Ghaemmaghami, Paul Huggins, Grace Lang, Julia Layne, Robert Slater
Qualitative Leveraging Natural Language Processing To Establish Judge Incrimination Statistics To Educate Voters In Re-Elections, Aurian Ghaemmaghami, Paul Huggins, Grace Lang, Julia Layne, Robert Slater
SMU Data Science Review
The prevalence of data has given consumers the power to make informed choices based off reviews, ratings, and descriptive statistics. However, when a local judge is coming up for re-election there is not any available data that aids voters in making data-driven decision on their vote. Currently court docket data is stored in text or PDFs with very little uniformity. Scaling the collection of this information could prove to be complicated and tiresome. There is a demand for an automated, intelligent system that can extract and organize useful information from the datasets. This paper covers the process of web scraping …
Fast Multipole Methods For Wave And Charge Source Interactions In Layered Media And Deep Neural Network Algorithms For High-Dimensional Pdes, Wenzhong Zhang
Fast Multipole Methods For Wave And Charge Source Interactions In Layered Media And Deep Neural Network Algorithms For High-Dimensional Pdes, Wenzhong Zhang
Mathematics Theses and Dissertations
In this dissertation, we develop fast algorithms for large scale numerical computations, including the fast multipole method (FMM) in layered media, and the forward-backward stochastic differential equation (FBSDE) based deep neural network (DNN) algorithms for high-dimensional parabolic partial differential equations (PDEs), addressing the issues of real-world challenging computational problems in various computation scenarios.
We develop the FMM in layered media, by first studying analytical and numerical properties of the Green's functions in layered media for the 2-D and 3-D Helmholtz equation, the linearized Poisson--Boltzmann equation, the Laplace's equation, and the tensor Green's functions for the time-harmonic Maxwell's equations and the …
Electricity Market Operations With Massive Renewable Integration: New Designs, Shengfei Yin
Electricity Market Operations With Massive Renewable Integration: New Designs, Shengfei Yin
Electrical Engineering Theses and Dissertations
Electricity market has been transitioning from a conventional and deterministic operation to a stochastic operation under the increasing penetration of renewable energy. Industry-level solutions toward the future electricity market operation ask for both accuracy and efficiency while maintaining model interpretability. Hence, reliable stochastic optimization techniques come to the first place for such a complex and dynamic problem.
This work starts at proposing a solution strategy for the uncertainty-based power system planning problem, which acts as a preliminary and instructs the electricity market operation. Considering 100% renewable penetration in the future, it analyzes the cost-effectiveness of renewable energy from a long-term …
Using Machine Learning Methods To Predict The Movement Trajectories Of The Louisiana Black Bear, Daniel Clark, David Shaw, Armando Vela, Shane Weinstock, John Santerre, Joseph D. Clark
Using Machine Learning Methods To Predict The Movement Trajectories Of The Louisiana Black Bear, Daniel Clark, David Shaw, Armando Vela, Shane Weinstock, John Santerre, Joseph D. Clark
SMU Data Science Review
In 1992, the Louisiana black bear (Ursus americanus luteolus) was placed on the U.S. Endangered Species List. This was due to bear populations in Louisiana being small and isolated enough where their populations couldn’t intersect with other populations to grow. Interchange of individuals between subpopulations of bears in Louisiana is critical to maintain genetic diversity and avoid inbreeding effects. Utilizing GPS (Global Positioning System) data gathered from 31 radio-collared bears from 2010 through 2012, this research will investigate how bears traverse the landscape, which has implications for gene exchange. This paper will leverage machine learning tools to improve upon existing …
Analyzing Empirical Quality Metrics Of Deep Learning Models For Antimicrobial Resistance, Huy H. Nguyen, Sanjay Pillay, Allison Roderick, Hao Wang, John Santerre
Analyzing Empirical Quality Metrics Of Deep Learning Models For Antimicrobial Resistance, Huy H. Nguyen, Sanjay Pillay, Allison Roderick, Hao Wang, John Santerre
SMU Data Science Review
Antimicrobial Resistance (AMR) is a growing concern in the medical field. Over-prescription of antibiotics as well as bacterial mutations have caused some once lifesaving drugs to become ineffective against bacteria. However, the problem of AMR might be addressed using Machine Learning (ML) thanks to increased availability of genomic data and large computing resources. The Pathosystems Resource Integration Center (PATRIC) has genomic data of various bacterial genera with sample isolates that are either resistant or susceptible to certain antibiotics. Past research has used this database to use ML algorithms to model AMR with successful results, including accuracies over 80%. To better …
Analysis Of Individual Player Performances And Their Effect On Winning In College Soccer, Angelo Bravo, Thomas Karba, Sean Mcwhirter, Billy Nayden
Analysis Of Individual Player Performances And Their Effect On Winning In College Soccer, Angelo Bravo, Thomas Karba, Sean Mcwhirter, Billy Nayden
SMU Data Science Review
This study describes the process of modernizing the approach of the Southern Methodist University (SMU) Men's Soccer coaching staff through the use of location and tracking data from their matches in the 2019 season. This study utilizes a variety of modeling and analysis techniques to explore and categorize the data and use it to evaluate the types of plays that are most often correlated with victories. This study's contribution to college soccer analytics includes the implementation of a model to determine individual players' performance, the production of team-level metrics, and visualizations to increase the efficiency of the coaching staff's efforts. …
Machine Learning In The Health Industry: Predicting Congestive Heart Failure And Impactors, Alexandra Norman, James Harding, Daria Zhukova
Machine Learning In The Health Industry: Predicting Congestive Heart Failure And Impactors, Alexandra Norman, James Harding, Daria Zhukova
SMU Data Science Review
Cardiovascular diseases, Congestive Heart Failure in particular, are a leading cause of deaths worldwide. Congestive Heart Failure has high mortality and morbidity rates. The key to decreasing the morbidity and mortality rates associated with Congestive Heart Failure is determining a method to detect high-risk individuals prior to the development of this often-fatal disease. Providing high-risk individuals with advanced knowledge of risk factors that could potentially lead to Congestive Heart Failure, enhances the likelihood of preventing the disease through implementation of lifestyle changes for healthy living. When dealing with healthcare and patient data, there are restrictions that led to difficulties accessing …