Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- Old Dominion University (12)
- Claremont Colleges (11)
- East Tennessee State University (7)
- Southern Methodist University (5)
- University of New Mexico (5)
-
- Ursinus College (5)
- Chapman University (4)
- Florida Institute of Technology (4)
- Illinois State University (4)
- University of Missouri, St. Louis (4)
- West Virginia University (4)
- Kennesaw State University (3)
- Montclair State University (3)
- Murray State University (3)
- New Jersey Institute of Technology (3)
- University of Central Florida (3)
- Utah State University (3)
- Binghamton University (2)
- California Polytechnic State University, San Luis Obispo (2)
- City University of New York (CUNY) (2)
- Clemson University (2)
- Embry-Riddle Aeronautical University (2)
- Georgia Southern University (2)
- Marshall University (2)
- Purdue University (2)
- South Dakota State University (2)
- Southeastern University (2)
- University of Arkansas, Fayetteville (2)
- University of Connecticut (2)
- University of Kentucky (2)
- Keyword
-
- Machine Learning (13)
- Machine learning (9)
- Data Science (7)
- Deep learning (6)
- Data science (5)
-
- Graph Theory (5)
- Neural networks (5)
- Statistics (5)
- Algorithms (4)
- Artificial Intelligence (4)
- COVID-19 (4)
- Neural Networks (4)
- Clustered data (3)
- Deep Learning (3)
- Informative cluster size (3)
- Mathematics (3)
- Reinforcement Learning (3)
- Sentiment (3)
- Time series (3)
- ARIMA (2)
- Algebraic topology (2)
- Analytics (2)
- Artificial intelligence (2)
- Big data (2)
- Bioinformatics (2)
- CIE (2)
- Central Limit Theorem (2)
- Classification (2)
- Computer Vision (2)
- Computer vision (2)
- Publication Year
- Publication
-
- Mathematics & Statistics Faculty Publications (10)
- Electronic Theses and Dissertations (9)
- Theses and Dissertations (8)
- Mathematics, Computer Science & Statistics Presentations (5)
- SMU Data Science Review (4)
-
- Theses (4)
- Annual Symposium on Biomathematics and Ecology Education and Research (3)
- Data Science and Data Mining (3)
- Dissertations (3)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (3)
- Journal of Humanistic Mathematics (3)
- All Dissertations (2)
- Branch Mathematics and Statistics Faculty and Staff Publications (2)
- CMC Senior Theses (2)
- CODEE Journal (2)
- College of Graduate Studies: Theses & Dissertations (2)
- Data Science Undergraduate Honors Theses (2)
- Department of Computer Science Faculty Scholarship and Creative Works (2)
- Discovery Day - Daytona Beach (2)
- Honors College Theses (2)
- Honors Projects (2)
- Honors Theses and Capstones (2)
- MPP Published Research (2)
- Master's Theses (2)
- Mathematics & Statistics ETDs (2)
- Northeast Journal of Complex Systems (NEJCS) (2)
- Pitzer Senior Theses (2)
- SDSU Data Science Symposium (2)
- Scripps Senior Theses (2)
- Selected Honors Theses (2)
- Publication Type
- File Type
Articles 31 - 60 of 144
Full-Text Articles in Data Science
Car Price Prediction Using Machine Learning: Analyzing The Dvm-Car Dataset, Yaman Abu Ghareebaih
Car Price Prediction Using Machine Learning: Analyzing The Dvm-Car Dataset, Yaman Abu Ghareebaih
Electronic Theses and Dissertations
The objective of this study is to predict car prices using machine learning models and the DVM-CAR dataset, which includes over 1.4 million images and car specifi- cations from 899 car models. Key factors such as mileage, engine power, and year of registration were analyzed for their correlation with car prices. Extensive data cleaning was performed, including filling missing values, identifying outliers, and normalizing numerical variables. Discrete variables like car make and body type were encoded using one-hot encoding. Linear relationships were analyzed with Multiple Logistic Regression, and Random Forest models were used for nonlinear patterns. Model performance was evaluated …
Analyzing The Sentiment Of Feminist And Non-Feminist Works, Jasmine Borie, Megan G. Falschlehner
Analyzing The Sentiment Of Feminist And Non-Feminist Works, Jasmine Borie, Megan G. Falschlehner
Mathematics, Computer Science & Statistics Presentations
This presentation focuses on a group of texts that advocate for a change in the current belief system. These texts are the Feminist Manifesto, Sojourner Truth: Ain’t I a Woman?, and Civilization and Its Discontents. These first two texts advocate for women’s rights, while Freud’s book is focused on civilization’s decline and how our understanding of community can affect this. Through our presentation, we want to examine the differences in sentiment and language between the feminist texts and Freud’s texts to pinpoint whether or not sentiment changes when advocating for different beliefs.
Analyzing Cie Texts Through History Using R, Rachel A. Hart, Aaron Ditto
Analyzing Cie Texts Through History Using R, Rachel A. Hart, Aaron Ditto
Mathematics, Computer Science & Statistics Presentations
In this presentation, we analyzed three separate CIE texts from different time periods. First, “The Allegory of the Cave” from 380 BC, then “The Declaration of Independence” from 1776, and lastly “The Lottery” from 1948. We compared them using tidy text techniques like sentiment lexicons, creating word clouds, and bigram analysis to see if the types of words and sentiments used have changed over time in these short texts.
A Statistical Comparison Of Selected Old Testament And New Testament Books, Branden F. Stahl, Kevin Guan, Adam Denn
A Statistical Comparison Of Selected Old Testament And New Testament Books, Branden F. Stahl, Kevin Guan, Adam Denn
Mathematics, Computer Science & Statistics Presentations
The purpose of this project was to discover similarities between sentiments in Old Testament and New Testament books of the Bible, track emotional valence and find the most common words and sentiments in the books. Text analysis was performed on Genesis, Exodus, Matthew and Luke. Word clouds were also created for these texts.
Data Science For Engineers, Heidi Moulton
Data Science For Engineers, Heidi Moulton
Student Research Symposium
30% of USU undergraduate students participate in some sort of research, and for engineering students this often means generating large amounts of data.
Data Science for Engineers is a series of four modules that introduce students to data processing, visualization, and graphing in the Python programming language using Pandas DataFrames and Juypter Notebooks.
The modules are intended for students with a basic understanding of programming in Python, specifically those who have taken CS 1400 Introduction to Computer Science.
Analysis Of Systematic Trade-Offs Between Military And Healthcare Expenditure Alongside Gdp Growth Of Select Asian And Western Exporting Economies In The 21st Century, Rahul Balamurugan, Carlos Gershenson, Preethi Nanjundan, Hiroki Sayama
Analysis Of Systematic Trade-Offs Between Military And Healthcare Expenditure Alongside Gdp Growth Of Select Asian And Western Exporting Economies In The 21st Century, Rahul Balamurugan, Carlos Gershenson, Preethi Nanjundan, Hiroki Sayama
Northeast Journal of Complex Systems (NEJCS)
This study explores the complexity in the trade-offs between military expenditure, healthcare expenditure, and GDP growth across select Asian nations and major weapon-exporting countries, examining how nations allocate finite resources between national security and human well-being over the past two decades. Using a systems science approach, the research integrates Granger causality testing to analyze temporal and directional relationships among GDP growth, military expenditure, and healthcare expenditure, uncovering their dynamic interdependencies. The methodology includes trend and slope analysis, Granger causality testing, outlier detection, and clustering to identify heterogeneity in resource allocation strategies. Developed, weapon-exporting nations exhibit complementary trends, with strong causality …
Optimized Hiv/Aids Resource Allocation In Ohio: A Linear Programming Approach, Godfred Ahenkroa Kesse
Optimized Hiv/Aids Resource Allocation In Ohio: A Linear Programming Approach, Godfred Ahenkroa Kesse
Data Science and Data Mining
This study employs a linear and integer programming approach to optimize HIV resource allocation in Ohio, aiming to minimize new infections and enhance the impact of limited resources. With the advances in HIV prevention and treatment, Ohio faces challenges in addressing disparities in access to healthcare, particularly among high-risk populations. The proposed model integrates data on infection rates, transmission patterns, demographic factors, and cost-effectiveness to provide a decision-support framework for policymakers. Using epidemiological data and equity constraints, the model prioritizes high-risk regions and populations while ensuring fair resource distribution. Results indicate that increased funding allocations significantly enhance the potential to …
Kroger Post-Pandemic Customer Segmentation, Mario Mata, Joey Truitt, Renn Spigelmyer, Dhanuja Kasturiratna, Lisa Holden, Nitish Baidya, Hanna Tafari
Kroger Post-Pandemic Customer Segmentation, Mario Mata, Joey Truitt, Renn Spigelmyer, Dhanuja Kasturiratna, Lisa Holden, Nitish Baidya, Hanna Tafari
Posters-at-the-Capitol
The grocery retail industry landscape has changed greatly in the wake of the pandemic. Specifically, delivery and pickup services have become more popular and customer buying habits have evolved. At the same time, improvements in data collection and analysis have allowed grocery marketing strategies to become highly individualized.
We worked with 84.51, an analytics firm, to identify customer segments for the Kroger Company based on data from 2023. Using clustering techniques, we organized customers into groups, or segments, based on similar characteristics. We identified and profiled four distinct groups of customers. Three segments were characterized by high frequency and spending …
Random Graph Models For Dual Graphs, Anne Friedman
Random Graph Models For Dual Graphs, Anne Friedman
Scripps Senior Theses
This paper aims to better characterize dual graphs derived from state districting maps by developing random graph models that replicate their structural properties. Dual graphs provide a simplified way to represent districting maps, making it computationally feasible to analyze their structure. These representations enable researchers, legislators, and courts to assess district compactness, detect signs of gerrymandering, and generate alternative districting plans. A deeper understanding of the structural patterns of these dual graphs can help researchers choose or design more effective algorithms for redistricting analysis. The random graph models developed in this study serve as testbeds for evaluating algorithmic approaches to …
Analyzing Patterns In Chicago Motor Vehicle Crashes Using Time-Series Techniques, Christina Trotta
Analyzing Patterns In Chicago Motor Vehicle Crashes Using Time-Series Techniques, Christina Trotta
Senior Honors Theses and Projects
This project explores time series forecasting of daily traffic crash rates in Chicago from 2018 to 2024, with a focus on understanding how past crash patterns and external conditions influence future risk. The primary research question asks: To what extent does yesterday’s crash rate help predict today’s? Using a combination of Holt-Winters exponential smoothing, Prophet forecasting, and SARIMAX models, we assess the role of autoregression, seasonality, and exogenous variables such as weather and roadway conditions. Daily crash data was cleaned, aggregated, and enriched with engineered features including holiday indicators, weather metrics from O’Hare and Midway airports, and binary flags for …
Optimizing Decision-Making In A Cerebral Palsy Model Using Reinforcement Learning, Richard Ampah
Optimizing Decision-Making In A Cerebral Palsy Model Using Reinforcement Learning, Richard Ampah
Pitzer Senior Theses
This study presents an original interdisciplinary investigation into how reinforcement learning (RL) can model motor and cognitive defects and potentially improve motor and cognitive functions in individuals with cerebral palsy (CP), a non-progressive neurological disorder that impairs movement and adaptability. Integrating computational neuroscience and machine learning, the research applies policy gradient methods and Markov Decision Processes (MDPs) to simulate adaptive learning in agents with and without CP-related constraints.
The central aim is to compare the cumulative rewards of optimal policies, derived from value iteration, and human-like learning policies using the REINFORCE algorithm, both with and without the Bellman baseline. The …
Empirical Analysis Of Political Districting Splitability Via Uniform Spanning Trees In Polynomial Time, Brooke C. Feinberg
Empirical Analysis Of Political Districting Splitability Via Uniform Spanning Trees In Polynomial Time, Brooke C. Feinberg
Scripps Senior Theses
This work expands a recently proven conjecture that a polynomial fraction of all uniform spanning trees (USTs) are splittable into k balanced partitions on grid graphs to real-world political districting plans. We investigate whether similar structural properties hold for the planar dual graphs of U.S. counties (cnty) and tracts (t), using Wilson’s algorithm to generate uniform random spanning trees and Breadth- First Search (BFS) to check for splitability into balanced partitions. Our empirical findings suggest that real-world districting plans can be split into 2-balanced, connected partitions in a fraction of polynomial time. This result highlights the potential for scalable redistricting …
Analyzing Political Sentiment On Micro-Blogging Data: A Lexicon And Machine Learning Approach To The 2024 U.S. Presidential Election, Ava Grey
CMC Senior Theses
This paper explores the trends in sentiment towards U.S. presidential candidates Kamala Harris and Donald Trump through micro-blogging social media text during the five months leading up to the election. Two datasets of varying sizes and origins were used to contextualize and validate analysis findings. The analyses include both a lexicon-based approach and a machine learning predictive method. Common sentiment analysis techniques like term frequency, term frequency inverse, various lexicons, and n-grams were utilized during the lexicon approach. During the modeling, a random forest was utilized in addition to the methods used during the lexicon approach. Results showed that overall …
Centralized Deep Reinforcement Learning For Homogeneous Multi-Component Maintenance Optimization, Joseph W. Wittrock
Centralized Deep Reinforcement Learning For Homogeneous Multi-Component Maintenance Optimization, Joseph W. Wittrock
Theses and Dissertations
This thesis explores an application of reinforcement learning (RL) in maintenance optimization. Recent advances in hardware-accelerated computation and deep learning have made RL a powerful tool for solving optimization problems which are too complex for traditional methods. Maintenance optimization involves improving the efficiency and effectiveness of maintenance activities through data-driven approaches, ultimately reducing costs and increasing asset availability. Making informed maintenance decisions is crucial to long-term sustainability.
A desirable maintenance policy maximizes a utility signal while minimizing the cost of maintenance. Techniques in sequential decision making such as dynamic programming (DP) and RL have found success in optimizing these maintenance …
Machine Learning Methods For Intrusion Detection And Response In Network Security, Ayomide Oyemaja
Machine Learning Methods For Intrusion Detection And Response In Network Security, Ayomide Oyemaja
College of Graduate Studies: Theses & Dissertations
Intrusion Detection Systems (IDS) play a crucial role in computer network security by identifying malicious activities and potential cyberattacks. This thesis combines machine learning and cybersecurity by applying Reinforcement Learning (RL) in intrusion detection and response using the NSL-KDD dataset.
We designed and implemented a Q-learning framework where an agent learns to classify network traffic over time by interacting with the environment and receiving rewards based on detection accuracy. We also look at the importance of feature selection and classification techniques and how effective they are in improving model performance, reducing the complexity of computation, and producing more desirable results. …
Model-Free Organization Of Patient Reported Outcomes Data: Geometrical Rep-Resentation Of The Modified Compartmen-Talization Method, Manasi Sheth, N. Rao Chaganty
Model-Free Organization Of Patient Reported Outcomes Data: Geometrical Rep-Resentation Of The Modified Compartmen-Talization Method, Manasi Sheth, N. Rao Chaganty
Mathematics & Statistics Faculty Publications
There is a recent advancement in the field of mathematics and statistics to understand the geometry or connectedness of the data due to the massive amounts of data being generated. The data provided for analyses are usually very large and need to be organized and minimized in order to make it more useful and meaningful. In biostatistics or medical field, it is important for patients to have access to high-quality, safe and effective and/ or efficacious medical products. It is quite necessary to ascertain that the patients and their care-partners stay at the center of the regulatory decision-making process. In …
Calculation And Statistical Analysis Of Wins Above Replacement, Joshua Taylor
Calculation And Statistical Analysis Of Wins Above Replacement, Joshua Taylor
Departmental Honors & Graduate Capstone Projects
The Wins Above Replacement (WAR) statistic in Major League Baseball is a prominent metric used to estimate player value by quantifying all aspects of play in terms of wins added to a baseball team. We will use R to calculate WAR for all players from 1871 to 2012 and use data from those years to construct multivariate predictive models to attempt to estimate WAR for players from 2013 to 2024. We find strong correlations between predicted and actual WAR values for most models, with the exception of the polynomial predictive model for non-qualified pitchers.
Applications Of Neural Networks In Parkinson’S Disease Diagnosis, Saladin Minhaaj
Applications Of Neural Networks In Parkinson’S Disease Diagnosis, Saladin Minhaaj
Theses
Parkinson's disease (PD) is a complex and debilitating neurodegenerative disorder that affects millions of people worldwide. Early and accurate diagnosis is crucial for effective treatment and management of PD. This thesis explores the application of neural networks in PD diagnosis, leveraging their ability to learn patterns from large datasets and make accurate predictions.
Thesis provides an overview of PD, including its symptoms, diagnosis, and current challenges in diagnosis. We then delve into the fundamentals of neural networks, including supervised learning, mathematical interpretations, and parametric models. This research focuses on the development of neural network models that can accurately diagnose PD …
Computational Representation, Analysis And Verification Of Requirements In Engineering Design And Systems Engineering, Chandan Kumar Sahu
Computational Representation, Analysis And Verification Of Requirements In Engineering Design And Systems Engineering, Chandan Kumar Sahu
All Dissertations
Systems are developed to satisfy a set of requirements derived from stakeholders’ needs, defining the problem space for which the system is created as a feasible solution. The system design process begins with eliciting these requirements and concludes with validating whether the created system meets them. Requirements engineering (RE) encompasses elicitation, representation, analysis, documentation, verification, and validation. However, challenges in RE, such as imprecision in natural language (NL), proprietary restrictions, and a lack of standardized quality metrics, hinder the creation of well-formed and comprehensive requirements. These challenges complicate formalization and analysis of requirements.
This dissertation addresses these challenges by proposing …
Decoding Neural Networks: An Information-Theoretic Guide To Interpretability, Error Analysis And Efficiency, Mackenzie J. Meni
Decoding Neural Networks: An Information-Theoretic Guide To Interpretability, Error Analysis And Efficiency, Mackenzie J. Meni
Theses and Dissertations
This dissertation addresses critical challenges in neural network design by leveraging entropy-based techniques to improve model efficiency, interpretability, and bias reduction. Focusing on the unique demands of computer vision applications, particularly object detection and classification for real-time systems, this work introduces a series of innovative methods centered on information theory. At the core of these methods is the Probabilistic Explanations of Entropic Knowledge (PEEK) framework, a tool developed to analyze and visualize entropy distributions across feature maps. PEEK offers insights into information flow within neural networks, making it possible to pinpoint layers that contribute meaningfully to decision-making or identify those …
A Machine Learning Approach For Survival Analysis Of Transplanted Kidneys Based On Donors’ And Recipients’ Factors., Alain Edward Despeignes
A Machine Learning Approach For Survival Analysis Of Transplanted Kidneys Based On Donors’ And Recipients’ Factors., Alain Edward Despeignes
Theses and Dissertations
Over seven thousand people on average die each year in the United States waiting for an organ transplant due to the shortage of donated organs. With this alarming concern, efforts from the health organizations like the United Network Organ Sharing (UNOS) and government officials have considered avenues to remedy this distress, one of which is to investigate the characteristics among donors and recipients that affects the longevity of donated organs. The goal of this project is to investigate the survival time of transplanted kidneys from 1987 to 2018 with regards to the donors’ and the recipients’ characteristics including gender, ethnicity, …
Discrete Time Series Forecasting Of Hive Weight, In-Hive Temperature, And Hive Entrance Traffic In Non-Invasive Monitoring Of Managed Honey Bee Colonies: Part I, Vladimir A. Kulyukin, Daniel Coster, Aleksey V. Kulyukin, William Meikle, Milagra Weiss
Discrete Time Series Forecasting Of Hive Weight, In-Hive Temperature, And Hive Entrance Traffic In Non-Invasive Monitoring Of Managed Honey Bee Colonies: Part I, Vladimir A. Kulyukin, Daniel Coster, Aleksey V. Kulyukin, William Meikle, Milagra Weiss
Computer Science Faculty and Staff Publications
From June to October, 2022, we recorded the weight, the internal temperature, and the hive entrance video traffic of ten managed honey bee (Apis mellifera) colonies at a research apiary of the Carl Hayden Bee Research Center in Tucson, AZ, USA. The weight and temperature were recorded every five minutes around the clock. The 30 s videos were recorded every five minutes daily from 7:00 to 20:55. We curated the collected data into a dataset of 758,703 records (208,760–weight; 322,570–temperature; 155,373–video). A principal objective of Part I of our investigation was to use the curated dataset to investigate …
Mathematical Modeling, Analysis, And Simulation Of Patient Addiction Journey, Adan Baca, Diego Gonzalez, Alonso G. Ogueda, Holly C. Matto, Padmanabhan Seshaiyer
Mathematical Modeling, Analysis, And Simulation Of Patient Addiction Journey, Adan Baca, Diego Gonzalez, Alonso G. Ogueda, Holly C. Matto, Padmanabhan Seshaiyer
CODEE Journal
This paper aims to develop a mathematical model to study the dynamics of addiction as individuals go through their detox journey. The motivation for this work is three fold. First, there has been a significant increase in drug overdose and drug addiction following the COVID-19 pandemic, and addiction may be interpreted as a infectious disease. Secondly, the dynamics of infectious disease could be modeled via compartmental models described by differential equations and one can therefore leverage the existing analytical and numerical methods to model addiction as a disease. Finally, the work helps to inform how mathematical models governed by differential …
Book Review: How To Expect The Unexpected: The Science Of Making Predictions -- And The Art Of Knowing When Not To By Kit Yates, Mark Huber
Journal of Humanistic Mathematics
Humans think about the future all the time. Prediction is a part of how we prepare for the coming of both good and bad events in our lives. Kit Yates' book, How to expect the unexpected, concentrates primarily on the question of why prediction is difficult, and what mental shortcuts people take in prediction that can lead to incorrect results. Unfortunately, a lack of concern for details and several omissions undermine the quality of the book.
Murmurations And Root Numbers, Alexey Pozdnyakov
Murmurations And Root Numbers, Alexey Pozdnyakov
University Scholar Projects
We report on a machine learning investigation of large datasets of elliptic curves and L-functions. This leads to the discovery of murmurations, an unexpected correlation between the root numbers and Dirichlet coefficients of L-functions. We provide a formal definition of murmurations, describe the connection with 1-level density, and provide three examples for which the murmuration phenomenon has been rigorously proven. Using our understanding of murmurations, we then build new machine learning models in search of a polynomial time algorithm for predicting root numbers. Based on our models and several heuristic arguments, we conclude that it is unlikely for …
Spatiotemporal Negative Inventory Outlier Decomposition For Supply Chain Applications In Consumer-Packaged Goods (Cpg), Hayden Mcdonald
Spatiotemporal Negative Inventory Outlier Decomposition For Supply Chain Applications In Consumer-Packaged Goods (Cpg), Hayden Mcdonald
Data Science Undergraduate Honors Theses
Coca-Cola is a popular soft drink brand with sales occurring in every Walmart store across the world, which generates large quantities of data and requires a robust supply chain system. However, the company does not currently have a sophisticated, automated, and/or prescriptive system for detecting where, when, and why inventory outages occur and applying preventative measures to avoid loss of revenue from the absence of inventory on store shelves. This thesis proposes and applies a novel, prescriptive system for this purpose. An inventory outage can be seen as a ‘negative’ statistical outlier in a time series of inventory for an …
Plumbing The Depths Of The Shallow End: Exploring Persistent Homology Using Small Data, R. Anne Flynn
Plumbing The Depths Of The Shallow End: Exploring Persistent Homology Using Small Data, R. Anne Flynn
All NMU Master's Theses
Persistent homology is a prominent tool in topological data analysis. This thesis is designed to be an introduction and guide to a beginner in persistent homology. This comprehensive overview discusses the math used behind it, the code needed to apply it, and its current place in the field. We explain and demonstrate the algebraic topology which fuels persistent homology. Homotopies inspire homology groups, which are able to determine how many holes a shape has. By visualizing data as a shape, persistent homology determines what type of holes are present.
We demonstrate this by using the package TDA in the manipulation …
Representation Learning For Generative Models With Applications To Healthcare, Astronautics, And Aviation, Van Minh Nguyen
Representation Learning For Generative Models With Applications To Healthcare, Astronautics, And Aviation, Van Minh Nguyen
Theses and Dissertations
This dissertation explores applications of representation learning and generative models to challenges in healthcare, astronautics, and aviation.
The first part investigates the use of Generative Adversarial Networks (GANs) to synthesize realistic electronic health record (EHR) data. An initial attempt at training a GAN on the MIMIC-IV dataset encountered stability and convergence issues, motivating a deeper study of 1-Lipschitz regularization techniques for Auxiliary Classifier GANs (AC-GANs). An extensive ablation study on the CIFAR-10 dataset found that Spectral Normalization is key for AC-GAN stability and performance, while Weight Clipping fails to converge without Spectral Normalization. Analysis of the training dynamics provided further …
An Investigation Into The Causes Of Home Field Advantage In Professional Soccer, Paige E. Tomer
An Investigation Into The Causes Of Home Field Advantage In Professional Soccer, Paige E. Tomer
Mathematics, Statistics, and Computer Science Honors Projects
Home-field advantage is the sporting phenomenon in which the home team outperforms the away team. Despite its widespread occurrence across sports, the underlying reasons for home-field advantage remain uncertain. In this paper, we employ a range of statistical methods to explore the causal relationships of potential determinants of home-field advantage. We measure home-field advantage using match outcomes and differential metrics (e.g., differences in yellow cards received). In an attempt to narrow the research disparity between men’s and women’s sports, we utilize data from the National Women’s Soccer League (NWSL) and the English Premier League (EPL) to investigate potential causes of …
Mathematically Rigorous Deep Learning Paradigms For Data-Driven Scientific Modeling, Owen Nicholas Davis
Mathematically Rigorous Deep Learning Paradigms For Data-Driven Scientific Modeling, Owen Nicholas Davis
Mathematics & Statistics ETDs
This dissertation explores the crucial role of data-driven modeling in science and engineering, with a focus on developing surrogate models to accelerate large-scale computational tasks, aiding in both outer-loop functions like uncertainty quantification and expensive inner-loop tasks within broader computational frameworks. Challenges arise with increased problem dimension and sparse, noisy training data, particularly significant when constructing surrogates for very expensive computational models where acquiring sufficient high-fidelity training data is unfeasible. In such scenarios, training surrogates from an ensemble of multifidelity information sources of varying accuracy and cost becomes essential. We emphasize neural network-based modeling paradigms, which are flexible in integrating …