Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Computer Sciences (99)
- Social and Behavioral Sciences (45)
- Engineering (37)
- Statistics and Probability (32)
- Databases and Information Systems (31)
-
- Artificial Intelligence and Robotics (29)
- Medicine and Health Sciences (26)
- Electrical and Computer Engineering (20)
- Life Sciences (20)
- Library and Information Science (19)
- Computer Engineering (18)
- Data Storage Systems (14)
- Other Computer Sciences (14)
- Theory and Algorithms (14)
- Business (13)
- Software Engineering (13)
- Systems and Communications (13)
- Applied Statistics (11)
- Applied Mathematics (10)
- Bioinformatics (10)
- Numerical Analysis and Scientific Computing (10)
- Education (9)
- Information Security (9)
- Arts and Humanities (8)
- Collection Development and Management (8)
- Medical Specialties (8)
- Biomedical Informatics (7)
- Earth Sciences (7)
- Institution
-
- Southern Methodist University (16)
- Singapore Management University (14)
- Universitas Negeri Malang (12)
- University of Kentucky (10)
- City University of New York (CUNY) (8)
-
- Technological University Dublin (8)
- The Texas Medical Center Library (8)
- University of Malaya (8)
- University of Rhode Island (8)
- Dakota State University (7)
- Smith College (6)
- University of Nebraska - Lincoln (6)
- Kennesaw State University (5)
- San Jose State University (5)
- California Polytechnic State University, San Luis Obispo (4)
- Edith Cowan University (4)
- Louisiana State University (4)
- Purdue University (4)
- The University of Southern Mississippi (4)
- West Virginia University (4)
- Chapman University (3)
- Dartmouth College (3)
- DePaul University (3)
- Harrisburg University of Science and Technology (3)
- Illinois State University (3)
- Munster Technological University (3)
- Rowan University (3)
- SIT Graduate Institute/SIT Study Abroad (3)
- University of Georgia School of Law (3)
- Western Michigan University (3)
- Keyword
-
- Machine Learning (21)
- Machine learning (18)
- Deep learning (11)
- Deep Learning (9)
- Library science (8)
-
- Big Data (7)
- Big data (7)
- Classification (7)
- COVID-19 (5)
- Computer science (5)
- Data (5)
- Data Science (5)
- Humans (5)
- Natural Language Processing (5)
- Random Forest (5)
- Analytics (4)
- Data Analysis (4)
- Prediction (4)
- Sentiment analysis (4)
- Artificial Intelligence (3)
- Bias (3)
- Cancer (3)
- Computer Sciences (3)
- Convolutional Neural Network (3)
- Data analysis (3)
- Data mining (3)
- Data visualization (3)
- Electronic Health Records (3)
- Library Impact Statement, Faculty Senate, Data Science, Collection Development (3)
- MITB student (3)
- Publication
-
- Research Collection School Of Computing and Information Systems (14)
- SMU Data Science Review (13)
- Knowledge Engineering and Data Science (12)
- Library Impact Statements (8)
- Student Works (2020-2029) (8)
-
- Dissertations (7)
- Faculty, Staff and Student Publications (6)
- Articles (5)
- Masters Theses & Doctoral Dissertations (5)
- Publications and Research (5)
- Statistical and Data Sciences: Faculty Publications (5)
- Theses and Dissertations (5)
- Conference papers (4)
- Master's Theses (4)
- Research outputs 2014 to 2021 (4)
- Student Papers in Public Policy (4)
- Annual Symposium on Biomathematics and Ecology Education and Research (3)
- College of Science & Mathematics Departmental Research (3)
- Electronic Theses and Dissertations (3)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (3)
- Independent Study Project (ISP) Collection (3)
- LSU Doctoral Dissertations (3)
- Library Philosophy and Practice (e-journal) (3)
- Master of Science in Computer Science Theses (3)
- Presentations (3)
- College of Computing and Digital Media Dissertations (2)
- Computational and Data Sciences (PhD) Dissertations (2)
- Dartmouth Scholarship (2)
- Department of Computer Science Faculty Scholarship and Creative Works (2)
- Dissertations and Theses (2)
- Publication Type
- File Type
Articles 1 - 30 of 235
Full-Text Articles in Data Science
Distributed Load Testing By Modeling And Simulating User Behavior, Chester Ira Parrott
Distributed Load Testing By Modeling And Simulating User Behavior, Chester Ira Parrott
LSU Doctoral Dissertations
Modern human-machine systems such as microservices rely upon agile engineering practices which require changes to be tested and released more frequently than classically engineered systems. A critical step in the testing of such systems is the generation of realistic workloads or load testing. Generated workload emulates the expected behaviors of users and machines within a system under test in order to find potentially unknown failure states. Typical testing tools rely on static testing artifacts to generate realistic workload conditions. Such artifacts can be cumbersome and costly to maintain; however, even model-based alternatives can prevent adaptation to changes in a system …
Data: The Good, The Bad And The Ethical, John D. Kelleher, Filipe Cabral Pinto, Luis M. Cortesao
Data: The Good, The Bad And The Ethical, John D. Kelleher, Filipe Cabral Pinto, Luis M. Cortesao
Articles
It is often the case with new technologies that it is very hard to predict their long-term impacts and as a result, although new technology may be beneficial in the short term, it can still cause problems in the longer term. This is what happened with oil by-products in different areas: the use of plastic as a disposable material did not take into account the hundreds of years necessary for its decomposition and its related long-term environmental damage. Data is said to be the new oil. The message to be conveyed is associated with its intrinsic value. But as in …
Improving The Maximum Power Point Tracking Efficiency Of Photovoltaic Arrays Via Machine Learning And Deep Learning, Sumedha Inamdar
Improving The Maximum Power Point Tracking Efficiency Of Photovoltaic Arrays Via Machine Learning And Deep Learning, Sumedha Inamdar
Master of Science in Computer Science Theses
Under partial shading conditions, photovoltaic (PV) modules in a solar array experience varying irradiance. A Global Maximum (GM) and multiple Local Maximums (LMs) can originate on the Power-Voltage (P-V) curve under nonuniform irradiance conditions. There are many maximum power point tracking (MPPT) algorithms developed to detect the true maximum power point (MPP) of a PV array. However, in the real-world environment, limited samples of power-voltage (P-V) data might be available to quickly and accurately predict the position of the global maximum point. Since the change of environmental conditions are dynamic, limited time is available to locate the global peak. Machine …
Bayesian Semi-Supervised Keyphrase Extraction And Jackknife Empirical Likelihood For Assessing Heterogeneity In Meta-Analysis, Guanshen Wang
Bayesian Semi-Supervised Keyphrase Extraction And Jackknife Empirical Likelihood For Assessing Heterogeneity In Meta-Analysis, Guanshen Wang
Statistical Science Theses and Dissertations
This dissertation investigates: (1) A Bayesian Semi-supervised Approach to Keyphrase Extraction with Only Positive and Unlabeled Data, (2) Jackknife Empirical Likelihood Confidence Intervals for Assessing Heterogeneity in Meta-analysis of Rare Binary Events.
In the big data era, people are blessed with a huge amount of information. However, the availability of information may also pose great challenges. One big challenge is how to extract useful yet succinct information in an automated fashion. As one of the first few efforts, keyphrase extraction methods summarize an article by identifying a list of keyphrases. Many existing keyphrase extraction methods focus on the unsupervised setting, …
Analysis Of Github Pull Requests, Canon Ellis
Analysis Of Github Pull Requests, Canon Ellis
Computer Science and Engineering Theses and Dissertations
The popularity of the software repository site GitHub has created a rise in the Pull Based Development Models' use. An essential portion of pull-based development is the creation of Pull Requests. Pull Requests often have to be reviewed by an individual to be approved and accepted into the Master branch of a software repository. The reviewing process can often be time-consuming and introduce a relatively high level of lost development time. This paper examines thousands of pull requests to understand the most valuable metadata of pull requests. We then introduce metrics in comparing the metadata of pull requests to understand …
Machine Learning Model Selection For Predicting Global Bathymetry, Nicholas P. Moran
Machine Learning Model Selection For Predicting Global Bathymetry, Nicholas P. Moran
LSU New Orleans Theses and Dissertations
This work is concerned with the viability of Machine Learning (ML) in training models for predicting global bathymetry, and whether there is a best fit model for predicting that bathymetry. The desired result is an investigation of the ability for ML to be used in future prediction models and to experiment with multiple trained models to determine an optimum selection. Ocean features were aggregated from a set of external studies and placed into two minute spatial grids representing the earth's oceans. A set of regression models, classification models, and a novel classification model were then fit to this data and …
Classifying Imbalanced Financial Fraud Data Utilizing Enhanced Random Forest Algorithm, Charles Gardner
Classifying Imbalanced Financial Fraud Data Utilizing Enhanced Random Forest Algorithm, Charles Gardner
Master of Science in Computer Science Theses
Imbalanced datasets have been a unique challenge for machine learning, requiring specialized approaches to correctly classify the minority class. Financial fraud detection involves using highly imbalanced datasets with a class imbalance of up to .01% frauds to 99.99% regular transactions. It is essential to identify all frauds in financial fraud detection, even if some classifications' precision is low. I developed a random forest assembly that separates fraudulent transactions into tiers of precision. With this approach, 96% of fraudulent transactions are identified, showing an 8% increase in recall when compared to standard approaches. 59% of fraud classifications' precision increases by 10% …
Reading Pdfs Using Adversarially Trained Convolutional Neural Network Based Optical Character Recognition, Michael B. Brewer, Michael Catalano, Yat Leung, David Stroud
Reading Pdfs Using Adversarially Trained Convolutional Neural Network Based Optical Character Recognition, Michael B. Brewer, Michael Catalano, Yat Leung, David Stroud
SMU Data Science Review
A common problem that has plagued companies for years is digitizing documents and making use of the data contained within. Optical Character Recognition (OCR) technology has flooded the market, but companies still face challenges productionizing these solutions at scale. Although these technologies can identify and recognize the text on the page, they fail to classify the data to the appropriate datatype in an automated system that uses OCR technology as its data mining process. The research contained in this paper presents a novel framework for the identification of datapoints on check stub images by utilizing generative adversarial networks (GANs) to …
Data Science In The Time Of Covid-19, Tony Breitzman
Data Science In The Time Of Covid-19, Tony Breitzman
College of Science & Mathematics Departmental Research
No abstract provided.
Principal Component Analysis For Predicting The Party Of The Legislators, Afsana Mimi
Principal Component Analysis For Predicting The Party Of The Legislators, Afsana Mimi
Publications and Research
In Spring 2020, I did a project, "Decision Tree Predicting the Party of Legislators," and construct a decision tree model to predict legislators' parties' based on their votes. We also use this model to identify legislators who frequently voted against their parties. We used the legislators' roll call votes, Office of Clerk U.S. House of Representatives Data Sets (Categorical values) collected in 2018 and 2019. In this new project, We study the 2018 and 2019 vote data using Principal Component Analysis (PCA). The goal is to find a (compressed) model using unsupervised learning to distinguish the legislators' parties, and PCA …
Introduction To Data Science Lti 110, Joanna Burkhardt
Introduction To Data Science Lti 110, Joanna Burkhardt
Library Impact Statements
No abstract provided.
Introduction To Data Science, Joanna Burkhardt
Introduction To Data Science, Joanna Burkhardt
Library Impact Statements
No abstract provided.
Spatial Frequency Implications For Global And Local Processing In Autistic Children, Riya Mody, Ayra Tusneem, Louanne Boyd, Vincent Berardi
Spatial Frequency Implications For Global And Local Processing In Autistic Children, Riya Mody, Ayra Tusneem, Louanne Boyd, Vincent Berardi
Student Scholar Symposium Abstracts and Posters
Visual processing in humans is done by integrating and updating multiple streams of global and local sensory input. Interaction between these two systems can be disrupted in individuals with ASD and other learning disabilities. When this integration is not done smoothly, it becomes difficult to see the “big picture”, which has been found to have implications on emotion recognition, social skills, and conversation skills. An example of this phenomenon is local interference, which is when local details are prioritized over the global features. Previous research in this field has aimed to decrease local interference by developing and evaluating a filter …
Factors Affecting Computer Science Research Productivity And Impact In Nigeria: A Bibliometric Evidence, Azubuike Ezenwoke
Factors Affecting Computer Science Research Productivity And Impact In Nigeria: A Bibliometric Evidence, Azubuike Ezenwoke
Library Philosophy and Practice (e-journal)
Computer science is a burgeoning research field and has the potential to accelerate the rate of industrialisation and subsequently, economic development. Using bibliometric data obtained from Scopus, this study employed a 15-year bibliometric analysis to highlight Nigeria’s productivity and impact trends in the computer science research landscape. Our findings are summarised as follows: First, Nigeria’s computer science research contribution and citations are meager in comparison to the global output. Secondly, international collaboration is generally weak as most collaborations are national in scope. Third, Nigeria’s computer science-related research is published in low-quality outlets, as Scopus has discontinued the indexing of most …
Evaluating The Reproducibility Of Physiological Stress Detection Models, Varun Mishra, Sougata Sen, Grace Chen, Tian Hao, Jeffrey Rogers, Ching-Hua Chen, David Kotz
Evaluating The Reproducibility Of Physiological Stress Detection Models, Varun Mishra, Sougata Sen, Grace Chen, Tian Hao, Jeffrey Rogers, Ching-Hua Chen, David Kotz
Dartmouth Scholarship
Recent advances in wearable sensor technologies have led to a variety of approaches for detecting physiological stress. Even with over a decade of research in the domain, there still exist many significant challenges, including a near-total lack of reproducibility across studies. Researchers often use some physiological sensors (custom-made or off-the-shelf), conduct a study to collect data, and build machine-learning models to detect stress. There is little effort to test the applicability of the model with similar physiological data collected from different devices, or the efficacy of the model on data collected from different studies, populations, or demographics.
This paper takes …
Open Data, Collaborative Working Platforms, And Interdisciplinary Collaboration: Building An Early Career Scientist Community Of Practice To Leverage Ocean Observatories Initiative Data To Address Critical Questions In Marine Science, Robert M. Levine, Kristen E. Fogaren, Johna E. Rudzin, Christopher J. Russoniello, Dax C. Soule, Justine M. Whitaker
Open Data, Collaborative Working Platforms, And Interdisciplinary Collaboration: Building An Early Career Scientist Community Of Practice To Leverage Ocean Observatories Initiative Data To Address Critical Questions In Marine Science, Robert M. Levine, Kristen E. Fogaren, Johna E. Rudzin, Christopher J. Russoniello, Dax C. Soule, Justine M. Whitaker
Publications and Research
Ocean observing systems are well-recognized as platforms for long-term monitoring of near-shore and remote locations in the global ocean. High-quality observatory data is freely available and accessible to all members of the global oceanographic community—a democratization of data that is particularly useful for early career scientists (ECS), enabling ECS to conduct research independent of traditional funding models or access to laboratory and field equipment. The concurrent collection of distinct data types with relevance for oceanographic disciplines including physics, chemistry, biology, and geology yields a unique incubator for cutting-edge, timely, interdisciplinary research. These data are both an opportunity and an incentive …
Healthcare Regulation And Governance: Big Data Analytics And Healthcare Data Protection, Xuejuan Zhang
Healthcare Regulation And Governance: Big Data Analytics And Healthcare Data Protection, Xuejuan Zhang
School of Continuing and Professional Studies Student Papers
No abstract provided.
Simple Modification For An Apriori Algorithm With Combination Reduction And Iteration Limitation Technique, Adie Wahyudi Oktavia Gama, Ni Made Widnyani
Simple Modification For An Apriori Algorithm With Combination Reduction And Iteration Limitation Technique, Adie Wahyudi Oktavia Gama, Ni Made Widnyani
Knowledge Engineering and Data Science
Apriori algorithm is one of the methods with regard to association rules in data mining. This algorithm uses knowledge from an itemset previously formed with frequent occurrence frequencies to form the next itemset. An a priori algorithm generates a combination by iteration methods that are using repeated database scanning process, pairing one product with another product and then recording the number of occurrences of the combination with the minimum limit of support and confidence values. The a priori algorithm will slow down to an expanding database in the process of finding frequent itemset to form association rules. Modification techniques are …
Segmentation Method For Face Modelling In Thermal Images, Albar Albar, Hendrick Hendrick, Rahmat Hidayat
Segmentation Method For Face Modelling In Thermal Images, Albar Albar, Hendrick Hendrick, Rahmat Hidayat
Knowledge Engineering and Data Science
Face detection is mostly applied in RGB images. The object detection usually applied the Deep Learning method for model creation. One method face spoofing is by using a thermal camera. The famous object detection methods are Yolo, Fast Region Based Convolutional Neural Networks (RCNN), Faster RCNN, SSD, and Mask RCNN. We proposed a segmentation Mask RCNN method to create a face model from thermal images. This model was able to locate the face area in images. The dataset was established using 1600 images. The images were created from direct capturing and collecting from the online dataset. The Mask RCNN was …
Generating Javanese Stopwords List Using K-Means Clustering Algorithm, Aji Prasetya Wibawa, Hidayah Kariima Fithri, Ilham Ari Elbaith Zaeni, Andrew Nafalski
Generating Javanese Stopwords List Using K-Means Clustering Algorithm, Aji Prasetya Wibawa, Hidayah Kariima Fithri, Ilham Ari Elbaith Zaeni, Andrew Nafalski
Knowledge Engineering and Data Science
Stopword removal necessary in Information Retrieval. It can remove frequently appeared and general words to reduce memory storage. The algorithm eliminates each word that is precisely the same as the word in the stopword list. However, generating the list could be time-consuming. The words in a specific language and domain must be collected and validated by specialists. This research aims to develop a new way to generate a stop word list using the K-means Clustering method. The proposed approach groups words based on their frequency. The confusion matrix calculates the difference between the findings with a valid stopword list created …
Efficient Scheduling Of Plantation Company Workers Using Genetic Algorithm, Wayan Firdaus Mahmudy, Andreas Pardede, Agus Wahyu Widodo, Muh Arif Rahman
Efficient Scheduling Of Plantation Company Workers Using Genetic Algorithm, Wayan Firdaus Mahmudy, Andreas Pardede, Agus Wahyu Widodo, Muh Arif Rahman
Knowledge Engineering and Data Science
Workers at large plantation companies have various activities. These activities include caring for plants, regularly applying fertilizers according to schedule, and crop harvesting activities. The density of worker activities must be balanced with efficient and fair work scheduling. A good schedule will minimize worker dissatisfaction while also maintaining their physical health. This study aims to optimize workers' schedules using a genetic algorithm. An efficient chromosome representation is designed to produce a good schedule in a reasonable amount of time. The mutation method is used in combination with reciprocal mutation and exchange mutation, while the type of crossover used is one …
A Review Of Accessing Big Data With Significant Ontologies, Jumah Y.J Sleeman, Jehad A.H Hammad
A Review Of Accessing Big Data With Significant Ontologies, Jumah Y.J Sleeman, Jehad A.H Hammad
Knowledge Engineering and Data Science
Ontology Based Data Access (OBDA) is a recently proposed approach which is able to provide a conceptual view on relational data sources. It addresses the problem of the direct access to big data through providing end-users with an ontology that goes between users and sources in which the ontology is connected to the data via mappings. We introduced the languages used to represent the ontologies and the mapping assertions technique that derived the query answering from sources. Query answering is divided into two steps: (i) Ontology rewriting, in which the query is rewritten with respect to the ontology into new …
Convolutional Neural Network On Tanned And Synthetic Leather Textures, Faadihilah Ahnaf Faiz, Ahmad Azhari
Convolutional Neural Network On Tanned And Synthetic Leather Textures, Faadihilah Ahnaf Faiz, Ahmad Azhari
Knowledge Engineering and Data Science
Tanned leather is an output from complex processes called tanning. Leather tanning is an important step that used to protect the fiber or protein structure of animal’s skin. Another reason of tanning process is to prevent the animal’s skin from any defect or rot. After the tanning is complete, the leather can be applied to produce a wide variety of leather products. Thus, the leather prices usually more expensive because it takes longer time in process. Another way to get cheaper price is make non-animal leather that usually known as synthetic or imitation leather. The purpose of this paper is …
Computational Behavioral Analytics: Estimating Psychological Traits In Foreign Languages., Kristopher Wayne Reese
Computational Behavioral Analytics: Estimating Psychological Traits In Foreign Languages., Kristopher Wayne Reese
Electronic Theses and Dissertations
The rise of technology proliferating into the workplace has increased the threat of loss of intellectual property, classified, and proprietary information for companies, governments, and academics. This can cause economic damage to the creators of new IP, companies, and whole economies. This technology proliferation has also assisted terror groups and lone wolf actors in pushing their message to a larger audience or finding similar tribal groups that share common, sometimes flawed, beliefs across various social media platforms. These types of challenges have created numerous studies in psycholinguistics, as well as commercial tools, that look to assist in identifying potential threats …
Detecting Hacker Threats: Performance Of Word And Sentence Embedding Models In Identifying Hacker Communications, Susan Mckeever, Brian Keegan, Andrei Quieroz
Detecting Hacker Threats: Performance Of Word And Sentence Embedding Models In Identifying Hacker Communications, Susan Mckeever, Brian Keegan, Andrei Quieroz
Conference papers
Abstract—Cyber security is striving to find new forms of protection against hacker attacks. An emerging approach nowadays is the investigation of security-related messages exchanged on deep/dark web and even surface web channels. This approach can be supported by the use of supervised machine learning models and text mining techniques. In our work, we compare a variety of machine learning algorithms, text representations and dimension reduction approaches for the detection accuracies of software-vulnerability-related communications. Given the imbalanced nature of the three public datasets used, we investigate appropriate sampling approaches to boost detection accuracies of our models. In addition, we examine how …
Exploring Information For Quantum Machine Learning Models, Michael Telahun
Exploring Information For Quantum Machine Learning Models, Michael Telahun
Electronic Theses and Dissertations
Quantum computing performs calculations by using physical phenomena and quantum mechanics principles to solve problems. This form of computation theoretically has been shown to provide speed ups to some problems of modern-day processing. With much anticipation the utilization of quantum phenomena in the field of Machine Learning has become apparent. The work here develops models from two software frameworks: TensorFlow Quantum (TFQ) and PennyLane for machine learning purposes. Both developed models utilize an information encoding technique amplitude encoding for preparation of states in a quantum learning model. This thesis explores both the capacity for amplitude encoding to provide enriched state …
Hierarchical Aggregation Of Multidimensional Data For Efficient Data Mining, Safaa Khalil Alwajidi
Hierarchical Aggregation Of Multidimensional Data For Efficient Data Mining, Safaa Khalil Alwajidi
Dissertations
Big data analysis is essential for many smart applications in areas such as connected healthcare, intelligent transportation, human activity recognition, environment, and climate change monitoring. Traditional data mining algorithms do not scale well to big data due to the enormous number of data points and the velocity of their generation. Mining and learning from big data need time and memory efficiency techniques, albeit the cost of possible loss in accuracy. This research focuses on the mining of big data using aggregated data as input. We developed a data structure that is to be used to aggregate data at multiple resolutions. …
Data And Assessment Management In Collegiate Recreation, Jeana Carow
Data And Assessment Management In Collegiate Recreation, Jeana Carow
Graduate Theses and Dissertations
Collegiate recreation programs and centers typically provide traditional programming space in addition to a range of physical activity spaces and resources, as a valuable part of the student experience. The external pressures of identifying and communicating departmental value and impact on the campus community has resulted in collegiate recreation departments’ use of data to communicate the effectiveness and impact of their work. The purpose of the study was to identify the data collection and assessment management practices of collegiate recreation departments, particularly focusing on the organization of data and assessment strategies as well as data collection, storage, reporting, analyzing, and …
Incorporating Shear Resistance Into Debris Flow Triggering Model Statistics, Noah J. Lyman
Incorporating Shear Resistance Into Debris Flow Triggering Model Statistics, Noah J. Lyman
Master's Theses
Several regions of the Western United States utilize statistical binary classification models to predict and manage debris flow initiation probability after wildfires. As the occurrence of wildfires and large intensity rainfall events increase, so has the frequency in which development occurs in the steep and mountainous terrain where these events arise. This resulting intersection brings with it an increasing need to derive improved results from existing models, or develop new models, to reduce the economic and human impacts that debris flows may bring. Any development or change to these models could also theoretically increase the ease of collection, processing, and …
Creating Optimal Conditions For Reproducible Data Analysis In R With ‘Fertile’, Audrey M. Bertin, Benjamin Baumer
Creating Optimal Conditions For Reproducible Data Analysis In R With ‘Fertile’, Audrey M. Bertin, Benjamin Baumer
Statistical and Data Sciences: Faculty Publications
The advancement of scientific knowledge increasingly depends on ensuring that data-driven research is reproducible: that two people with the same data obtain the same results. However, while the necessity of reproducibility is clear, there are significant behavioral and technical challenges that impede its widespread implementation and no clear consensus on standards of what constitutes reproducibility in published research. We present fertile, an R package that focuses on a series of common mistakes programmers make while conducting data science projects in R, primarily through the RStudio integrated development environment. fertile operates in two modes: proactively, to prevent reproducibility mistakes from happening …