Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Computer Sciences (22)
- Statistics and Probability (13)
- Social and Behavioral Sciences (12)
- Engineering (9)
- Education (7)
-
- Applied Statistics (6)
- Artificial Intelligence and Robotics (6)
- Databases and Information Systems (6)
- Business (5)
- Mathematics (5)
- Medicine and Health Sciences (5)
- Life Sciences (4)
- Statistical Models (4)
- Arts and Humanities (3)
- Business Analytics (3)
- Health Information Technology (3)
- Programming Languages and Compilers (3)
- Public Health (3)
- Sociology (3)
- Applied Mathematics (2)
- Categorical Data Analysis (2)
- Computer Engineering (2)
- Criminology (2)
- Epidemiology (2)
- Higher Education (2)
- Information Security (2)
- Interprofessional Education (2)
- Library and Information Science (2)
- Institution
-
- City University of New York (CUNY) (5)
- Belmont University (3)
- Minnesota State University, Mankato (3)
- Southern Methodist University (3)
- University of Arkansas, Fayetteville (3)
-
- Claremont Colleges (2)
- Dartmouth College (2)
- Mississippi State University (2)
- Old Dominion University (2)
- University of Louisville (2)
- Virginia Commonwealth University (2)
- Aga Khan University (1)
- Air Force Institute of Technology (1)
- Binghamton University (1)
- California Polytechnic State University, San Luis Obispo (1)
- Case Western Reserve University (1)
- Chapman University (1)
- Clemson University (1)
- DePaul University (1)
- Embry-Riddle Aeronautical University (1)
- Georgia Southern University (1)
- Gettysburg College (1)
- Kennesaw State University (1)
- Kutztown University (1)
- Louisiana State University (1)
- Loyola Marymount University and Loyola Law School (1)
- Marshall University (1)
- Northern Michigan University (1)
- Portland State University (1)
- Purdue University (1)
- Publication Year
- Publication
-
- All Graduate Theses, Dissertations, and Other Capstone Projects (3)
- Data Science Undergraduate Honors Theses (3)
- SPARK Symposium Presentations (3)
- Capstone Projects (2)
- Graduate Research Posters (2)
-
- Journal of Humanistic Mathematics (2)
- Publications and Research (2)
- SMU Data Science Review (2)
- All Dissertations (1)
- All NMU Master's Theses (1)
- Arts & Sciences Faculty Publications (1)
- College of Computing and Digital Media Dissertations (1)
- College of Engineering Summer Undergraduate Research Program (1)
- College of Graduate Studies: Theses & Dissertations (1)
- Computational and Data Sciences (PhD) Dissertations (1)
- Computer Science and Information Technology Faculty (1)
- Dartmouth College Master’s Theses (1)
- Dartmouth Scholarship (1)
- Dissertations and Theses (1)
- Dissertations, Theses, and Capstone Projects (1)
- Electronic Theses and Dissertations (1)
- Engineering Technology Faculty Publications (1)
- Faculty Scholarship (1)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (1)
- Honors Theses and Capstones (1)
- Honors Thesis (1)
- Information Technology & Decision Sciences Faculty Publications (1)
- Institute for Human Development, East Africa (1)
- International Conference on Gambling & Risk Taking (1)
- Journal on Empowering Teaching Excellence (1)
- Publication Type
Articles 1 - 30 of 58
Full-Text Articles in Data Science
Data-Driven Characterization Of Counties In The Prison Industrial Complex Using Clustering Analysis, Riley N. Tuccio
Data-Driven Characterization Of Counties In The Prison Industrial Complex Using Clustering Analysis, Riley N. Tuccio
Capstone Projects
This project investigates the complex relationship between counties that house prisons in the United States and the rurality associated with them. The central research question explores how both county characteristics, such as variables corresponding to cost of living and demographics of a county, and prison characteristics, such as programming available to inmates and staffing levels, differ across the census-designated rural-urban distinctions. Furthermore, the study examines whether modern data science methods can more accurately define and distinguish these characteristics, providing a nuanced understanding of the Prison Industrial Complex (PIC) and its manifestation across various American communities. The motivation for this research …
Mycelial Modeling: Teaching Biology Students Statistical Modeling With Mushrooms, Colette Wolf
Mycelial Modeling: Teaching Biology Students Statistical Modeling With Mushrooms, Colette Wolf
University Honors Theses
This paper summarizes and describes the development of a set of learning materials that were created to educate students and professionals from other fields in statistical modeling techniques. These materials are primarily aimed at biology students, but are still intended to be useful for anyone who is interested in incorporating decision trees and random forest models into their personal research in the future. By directing the reader towards the JMP software, these materials navigate around the statistical knowledge base and coding implementation practices that otherwise would serve as a barrier to learning statistical modeling techniques, and instead focus on the …
Labeling And Describing Objects In A Photogrammetry-Based 3d Environment Using Computer Vision And Artificial Intelligence, Colin Donald Moschella
Labeling And Describing Objects In A Photogrammetry-Based 3d Environment Using Computer Vision And Artificial Intelligence, Colin Donald Moschella
Capstone Projects
Cataloging and digitizing the objects inside a building manually is a task that is often impractical at scale. This project therefore automates the process, using a custom-made system. Using a photogrammetry-based 3D reconstruction of a room, this system is applied to sequences of 2D images used to make the 3D models. The system applies object detection, image segmentation, and image-text models to identify and describe objects, using CNN based models such as YOLO and OpenCLIP. Each analyzed object is then stored in a structured database with spatial coordinates from the 3D scanning, descriptive attributes from the image-text models, and other …
Lifestyle Factors' Effect On Political And Religious Behavior, Madison A. Price
Lifestyle Factors' Effect On Political And Religious Behavior, Madison A. Price
SPARK Symposium Presentations
There is a general understanding of an individuals’ political/religious beliefs when you analyze predictors like support of same sex marriage or views on abortion, but there is not a ton of research on how everyday lifestyle factors might affect how someone aligns themselves politically or religiously. Researchers have begun to expand the horizons of political and religious research by investigating how income, health, and politics affect religion (Francis-Tan & Tian, 2022), but there is a need to expand into more specific lifestyle factors. Which brings reason to question a few things: what lifestyle and demographic factors predict how religious someone …
A Data-Driven Framework For Automation Readiness In Minnesota State University, Mankato Course Scheduling, Prisca Bongu Payanzo Maba
A Data-Driven Framework For Automation Readiness In Minnesota State University, Mankato Course Scheduling, Prisca Bongu Payanzo Maba
All Graduate Theses, Dissertations, and Other Capstone Projects
University course scheduling is one of the most complex optimization problems in higher education institutions. With universities growing in size and offering a broad spectrum of majors and disciplines, the number of possible course scheduling combinations increases exponentially, rendering traditional ways of scheduling ineffective.
Although operations research has extensively studied automated scheduling algorithms, there has been limited investigations into the organization readiness of academic departments to implement such systems. This paper offers a hybrid data science framework that assesses departmental readiness for scheduling automation.
The study combines qualitative Zoom interview data from 19 academic departments with institutional scheduling rules from …
Predicting Criminal Behavior In Major Us Cities, Madison A. Price
Predicting Criminal Behavior In Major Us Cities, Madison A. Price
SPARK Symposium Presentations
In recent years, especially post pandemic, there has been a decrease in crime in the United States. Unfortunately, the country’s violent crime rates are still significantly higher compared to similar high-income countries, so what predicts crime in major American cities? There is tons of research to support the idea that demographics can offer some insight into predicting crime. There are countless online resources that seek to identify major crime centrals in the United States (Petrino, 2025). In the late 1990s, researchers noticed that crime rates in cities had a downward slope due to an important contributor: demographic change (Fox & …
Data-Driven Climate Damage Functions For Capital Formation: Estimating The Climate Penalty Using Maching Learning, Pramudya Wicaksono
Data-Driven Climate Damage Functions For Capital Formation: Estimating The Climate Penalty Using Maching Learning, Pramudya Wicaksono
All Graduate Theses, Dissertations, and Other Capstone Projects
Traditional integrated assessment models assume parametric climate damage functions that may miss nonlinearities, heterogeneity, and dynamic effects on investment. This thesis develops a data-driven climate damage function for capital formation by estimating the predictive relationship between climate conditions and future gross fixed capital formation (% GDP) across 125 countries over 1982–2019. We construct a panel dataset by combining daily ERA5 climate reanalysis data (accessed via the Copernicus Climate Data Store API and aggregated to yearly country-level variables including temperature anomalies, extreme heat days, frost days, precipitation, and solar radiation) with economic indicators from the World Bank World Development Indicators and …
Reverse (Bio)Engineering: A Machine Learning Approach To Optimize Baseball Pitcher Health And Performance, Robert C. Moore
Reverse (Bio)Engineering: A Machine Learning Approach To Optimize Baseball Pitcher Health And Performance, Robert C. Moore
All Dissertations
Ball tracking systems are becoming ubiquitous in sport, creating an unprecedented opportunity for big data applications to optimize human health and performance. These applications are especially common in baseball, a sport known for analyzing ball flight data to quantify performance. Analysts routinely use ball flight data to identify the attributes of top performing pitchers, finding that the best pitchers throw with optimal combinations of release speed and spin to precise locations. However, for certain pitchers, the throwing motion required to produce optimal ball flight places exceedingly high biomechanical load on the elbow, and consequently injury rates continue to rise. This …
Computational Data Analysis, Kathryn S. Biles
Computational Data Analysis, Kathryn S. Biles
LSU Master's Theses
Data science has emerged as a cornerstone of innovation, shaping an ever-expanding range
of professional careers. As technology advances and the volume of data expands expo-
nentially, the ability to extract meaningful insights from data has become indispensable
across industries. Far from representing a single career path, data science enables profes-
sionals in nearly every domain to make informed decisions, optimize systems, and drive
innovation. Yet, many high school students and incoming college freshmen have limited
exposure to data science fundamentals or the career opportunities they unlock. This is
the gap that Computational Data Analysis, a high school-level curriculum I …
Mapping Food Justice: Urban Farms And The Examination Of Equitable Food Access, Aaron Avila, Mark Ayiah, Jake Stavely, Marc T. Sager, Maximilian K. Sherard, Anthony J. Petrosino
Mapping Food Justice: Urban Farms And The Examination Of Equitable Food Access, Aaron Avila, Mark Ayiah, Jake Stavely, Marc T. Sager, Maximilian K. Sherard, Anthony J. Petrosino
SMU Journal of Undergraduate Research
In this project, we worked alongside members from an urban farm in South Dallas to learn about issues related to food justice, urban farming, and food deserts. Using participatory design research methods, we created data visualizations showing how society can reduce inequities relating to food access produced in historically underserved neighborhoods. The research goals guiding this study are: a) to identify food deserts and urban farms in the Dallas-Fort Worth metropolitan region (DFW) and b) to determine which urban farms service the needs of these food deserts. To identify food deserts, we took two steps: First, we used open-access data …
Mat 301 - Applied Statistics And Data Analysis, Eric Aragundi
Mat 301 - Applied Statistics And Data Analysis, Eric Aragundi
Open Educational Resources
Data analysis using standard statistical methods and relevant computer software. Emphasis on real-world data, interpretation, and misinterpretation of computer output.
This syllabus contains open source notebook about data analysis content.
Enhancing The University Of Arkansas' Operations Through Data Science, Aura L. Pinto-Avelar
Enhancing The University Of Arkansas' Operations Through Data Science, Aura L. Pinto-Avelar
Data Science Undergraduate Honors Theses
As universities navigate financial constraints and resource allocation challenges, data driven financial analysis has become increasingly important. Universities employ various methods to assess financial efficiency, predict future expenditures, and optimize student credit hour distribution. However, the approaches to financial analysis vary widely, with some institutions leveraging advanced predictive modeling and business intelligence tools, while others rely on traditional budgeting techniques and manual forecasting.
This thesis examines how the University of Arkansas' (“Uark”) financial analysis methods compare to those of other institutions and alternative data-driven approaches. Using four years of financial and student credit hour data, this study evaluates cost trends …
Enhancing Animal Shelter Operations With Time Series And Machine Learning, Sakava L. Kiv, Donald L. Anderson, Shivam Negi, Jacquelyn Cheun
Enhancing Animal Shelter Operations With Time Series And Machine Learning, Sakava L. Kiv, Donald L. Anderson, Shivam Negi, Jacquelyn Cheun
SMU Data Science Review
Enhancing animal shelter operations through machine learning involves employing a variety of advanced techniques aimed at increasing efficiency, promoting animal welfare, and optimizing resource allocation. This paper explores predictive analytics for adoption rates using regression models to estimate the likelihood of adoption based on historical data, encompassing variables such as breed, health status, and previous adoption trends. Additionally, classification algorithms are utilized to categorize animals by adoption probability, facilitating better resources and marketing prioritization. Clustering algorithms are employed to group animals according to behavior patterns and/or physical health, enabling tailored medical care and enrichment activities that improve their mental and …
Spotify Recommender Using Content Filtering, Colin M. Anderson
Spotify Recommender Using Content Filtering, Colin M. Anderson
SPARK Symposium Presentations
This project is a music recommender system that analyzes Spotify data to establish relationships between songs by analyzing their musical components. A user can receive recommendations through two methods. Firstly, the user can select a song from the database that they enjoy, and the system will provide them with a list of recommendations, as well as predict the genre of the song that they have entered. Secondly, the user can set personalized values for various musical features (i.e. energy, danceability). In either case, by selecting a number 1 to 5 recommendations that the user would like, they will get that …
Book Review: How To Expect The Unexpected: The Science Of Making Predictions -- And The Art Of Knowing When Not To By Kit Yates, Mark Huber
Journal of Humanistic Mathematics
Humans think about the future all the time. Prediction is a part of how we prepare for the coming of both good and bad events in our lives. Kit Yates' book, How to expect the unexpected, concentrates primarily on the question of why prediction is difficult, and what mental shortcuts people take in prediction that can lead to incorrect results. Unfortunately, a lack of concern for details and several omissions undermine the quality of the book.
Identifying The O’Connell Effect In Eclipsing Binary Stars, Nicholas Paolella
Identifying The O’Connell Effect In Eclipsing Binary Stars, Nicholas Paolella
Computer Science and Information Technology Faculty
Data science techniques have wide-ranging applications throughout scientific explorations. One, is filtering astronomical data to better understand specific populations, such as binary stars. Specifically, binary stars that exhibit the O’Connell effect are worthy of study as this phenomenon is still not well understood. The O’Connell effect can be defined as the asymmetry of maxima in the light curves, as captured by the instrument, while observing the eclipsing binary system in question. There is significant data captured by NASA and curated by Villanova University, which enabled the investigation of eclipsing binary stars and the attributes of which may help identify the …
Interpretable Learning In Multivariate Big Data Analysis For Network Monitoring, José Camacho, Katarzyna Wasielewska, Rasmus Bro, David Kotz
Interpretable Learning In Multivariate Big Data Analysis For Network Monitoring, José Camacho, Katarzyna Wasielewska, Rasmus Bro, David Kotz
Dartmouth Scholarship
There is an increasing interest in the development of new data-driven models useful to assess the performance of communication networks. For many applications, like network monitoring and troubleshooting, a data model is of little use if it cannot be interpreted by a human operator. In this paper, we present an extension of the Multivariate Big Data Analysis (MBDA) methodology, a recently proposed interpretable data analysis tool. In this extension, we propose a solution to the automatic derivation of features, a cornerstone step for the application of MBDA when the amount of data is massive. The resulting network monitoring approach allows …
A Symbolic Approach To Nonlinear Time Series Analysis, Ranjan Karki, Nibhrat Lohia, Michael B. Schulte
A Symbolic Approach To Nonlinear Time Series Analysis, Ranjan Karki, Nibhrat Lohia, Michael B. Schulte
SMU Data Science Review
Current nonlinear time series methods such as neural networks forecast well. However, they act as a black box and are difficult to interpret, leaving the researchers and the audience with little insight into why the forecasts are the way they are. There is a need for a method that forecasts accurately while also being easy to interpret. This paper aims to develop a method to build an interpretable model for univariate and multivariate nonlinear time series data using wavelets and symbolic regression. The final method relies on multilayer perceptron (MLP) neural networks as a form of dimensionality reduction and the …
A Spatiotemporal Analysis Of Violent Crime In Little Rock, Arkansas From 1999-2022, Nicole Rogers
A Spatiotemporal Analysis Of Violent Crime In Little Rock, Arkansas From 1999-2022, Nicole Rogers
Data Science Undergraduate Honors Theses
Little Rock, Arkansas is not only the capital and largest city in Arkansas, but it has one of the highest crime rates amongst cities with over 100,000 people in the country. According to the US Census in 2020, Little Rock had a population of 202,591. In the same year, Little Rock Police Department recorded 3,567 cases of violent crime, leading to a violent crime rate of 1,805 violent crime occurrences per 100,000 people. For perspective, Chicago’s violent crime rate was approximately half of that of Little Rock during the same time period. Crime, like other social phenomena is unevenly distributed …
Plumbing The Depths Of The Shallow End: Exploring Persistent Homology Using Small Data, R. Anne Flynn
Plumbing The Depths Of The Shallow End: Exploring Persistent Homology Using Small Data, R. Anne Flynn
All NMU Master's Theses
Persistent homology is a prominent tool in topological data analysis. This thesis is designed to be an introduction and guide to a beginner in persistent homology. This comprehensive overview discusses the math used behind it, the code needed to apply it, and its current place in the field. We explain and demonstrate the algebraic topology which fuels persistent homology. Homotopies inspire homology groups, which are able to determine how many holes a shape has. By visualizing data as a shape, persistent homology determines what type of holes are present.
We demonstrate this by using the package TDA in the manipulation …
Data Engineering: Building Software Efficiency In Medium To Large Organizations, Alessandro De La Torre
Data Engineering: Building Software Efficiency In Medium To Large Organizations, Alessandro De La Torre
Whittier Scholars Program
The introduction of PoetHQ, a mobile application, offers an economical strategy for colleges, potentially ushering in significant cost savings. These savings could be redirected towards enhancing academic programs and services, enriching the educational landscape for students. PoetHQ aims to democratize access to crucial software, effectively removing financial barriers and facilitating a richer educational experience. By providing an efficient software solution that reduces organizational overhead while maximizing accessibility for students, the project highlights the essential role of equitable education and resource optimization within academic institutions.
On Intrinsic Dimensionality Of Data Sets And Neural Networks, Ori Chachmo
On Intrinsic Dimensionality Of Data Sets And Neural Networks, Ori Chachmo
Theses and Dissertations
The concept of Intrinsic Dimensionality (ID) is of special interest in the field of Neural Networks (NNs) since it promotes both (a) a deeper understanding of the underlying mechanisms, and (b) embraces parsimonious modeling (that is, building the right-sized model for the task) with associated benefits to processing speed and storage requirements. This thesis explores the concept of ID via two separate, but related, questions. First, we study the potential of NN ID prediction by exploiting easily obtained quantities measured on the data. We then explore NN ID as an independent concept by comparing the results of different methods for …
Data Science In Finance: Challenges And Opportunities, Xianrong Zheng, Elizabeth Gildea, Sheng Chai, Tongxiao Zhang, Shuxi Wang
Data Science In Finance: Challenges And Opportunities, Xianrong Zheng, Elizabeth Gildea, Sheng Chai, Tongxiao Zhang, Shuxi Wang
Information Technology & Decision Sciences Faculty Publications
Data science has become increasingly popular due to emerging technologies, including generative AI, big data, deep learning, etc. It can provide insights from data that are hard to determine from a human perspective. Data science in finance helps to provide more personal and safer experiences for customers and develop cutting-edge solutions for a company. This paper surveys the challenges and opportunities in applying data science to finance. It provides a state-of-the-art review of financial technologies, algorithmic trading, and fraud detection. Also, the paper identifies two research topics. One is how to use generative AI in algorithmic trading. The other is …
Intelligent Traffic Management Systems, Mohammad Mazhar
Intelligent Traffic Management Systems, Mohammad Mazhar
All Graduate Theses, Dissertations, and Other Capstone Projects
With the increase in population and in particular urban population. The traffic and travel times in between cities and inside cities has increased due to more and more people using private means of transportation. Due to this need arose for tackling the increase in traffic by managing it using various means. For this we look towards The Intelligent Traffic Management System (ITMS). ITMS is an AI-powered solution designed to optimize traffic flow, reduce congestion, and improve overall road safety. The system will monitor real-time traffic data using a combination of cameras and sensors, identify traffic jams, and send alerts to …
Teaching Reproducibility To First Year College Students: Reflections From An Introductory Data Science Course, Brennan L. Bean
Teaching Reproducibility To First Year College Students: Reflections From An Introductory Data Science Course, Brennan L. Bean
Journal on Empowering Teaching Excellence
Access the online Pressbooks version of this article here.
Modern technology threatens traditional modes of classroom assessment by providing students with automated ways to write essays and take exams. At the same time, modern technology continues to expand the accessibility of computational tools that promise to increase the potential scope and quality of class projects. This paper presents a case study where students are asked to complete a “reproducible” final project in an introductory data science course using the R programming language. A reproducible project is one where an instructor can easily regenerate the results and conclusions from the submitted …
Promoting Data Harmonization To Evaluate Vaccine Hesitancy In Lmics: Approach And Applications, Ryan Rego, Yuri Zhukov, Kyrani Reneau, Amy Pienta, Kristina L. Rice, Patrick Brady, Geoffrey Siwo, Peninah Wachira, Amina Abubakar, Ken Kollman
Promoting Data Harmonization To Evaluate Vaccine Hesitancy In Lmics: Approach And Applications, Ryan Rego, Yuri Zhukov, Kyrani Reneau, Amy Pienta, Kristina L. Rice, Patrick Brady, Geoffrey Siwo, Peninah Wachira, Amina Abubakar, Ken Kollman
Institute for Human Development, East Africa
Background: Factors influencing the health of populations are subjects of interdisciplinary study. However, datasets relevant to public health often lack interdisciplinary breath. It is difficult to combine data on health outcomes with datasets on potentially important contextual factors, like political violence or development, due to incompatible levels of geographic support; differing data formats and structures; differences in sampling procedures and wording; and the stability of temporal trends. We present a computational package to combine spatially misaligned datasets, and provide an illustrative analysis of multi-dimensional factors in health outcomes.
Methods: We rely on a new software toolkit, Sub-National Geospatial Data Archive …
Digital Scholarship And Data Science Intersect In Libraries: A Needs Assessment Report, Halie Kerns
Digital Scholarship And Data Science Intersect In Libraries: A Needs Assessment Report, Halie Kerns
Library Created Resources
The following report summarized the results of a needs assessment completed in the fall of 2023 at Binghamton University by the Libraries’ Digital Scholarship team. The aim was to understand how data science-focused programming, as part of the digital scholarship’s offerings, would be utilized on campus. The report evaluates existing literature, summarizes findings from twenty-eight interviews done across campus, and lays out an action plan for the Digital Scholarship team’s future planning.
Dei: Exploring Academic Reflections Using Natural Language Processing To Create A Roadmap Of Student Success And Foster Inclusive Engineering Education, Rajvir H. Vyas, Nidhi Raviprasad
Dei: Exploring Academic Reflections Using Natural Language Processing To Create A Roadmap Of Student Success And Foster Inclusive Engineering Education, Rajvir H. Vyas, Nidhi Raviprasad
College of Engineering Summer Undergraduate Research Program
Every year, the College of Engineering (CENG) students and faculty reach out to admitted students through “Text-a-Thon” programs to answer their questions about being a student at Cal Poly. In order to improve CENG outreach efforts, we analyzed these text conversations to predict the likelihood of an admitted student accepting an offer of admission from Cal Poly. Through our research, we discovered key factors that play a role in a student committing to Cal Poly through data-based insights. Additionally, we successfully used a human-on-the-loop system to help create Machine Learning (ML) models that predict satisfaction of response by way of …
Effects Of Weight Initialization Methods On Ffn's, Ida K. Karem
Effects Of Weight Initialization Methods On Ffn's, Ida K. Karem
The Cardinal Edge
Weight initialization is the method of determining starting values of weights in a neural network. The way this method is done can have massive effects on the network[2, 3, 6, 9] and can halt training if not handled properly. On the other hand, if initialization is chosen tactfully it can improve training and accuracy greatly. The initialization method usually called Normalized Xavier will be referred to as Nox in this paper to avoid confusion with the Xavier initialization method. This study analyzes five methods of weight initialization(Nox, He, Xavier, Plutonian, and Self-Root), two of them …
Responsible Data Science For Genocide Prevention, Victor Piercey
Responsible Data Science For Genocide Prevention, Victor Piercey
Journal of Humanistic Mathematics
The term "genocide" emerged out of an effort to describe mass atrocities committed in the first half of the 20th century. Despite a convention of the United Nations outlawing genocide as a matter of international law, the problem persists. Some organizations (including the United Nations) are developing indicator frameworks and “early-warning” systems that leverage data science to produce risk assessments of countries where conflict is present. These tools raise questions about responsible data use, specifically regarding the data sources and social biases built into algorithms through their training data. This essay seeks to engage mathematicians in discussing these concerns.