Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

3,235 Full-Text Articles 9,310 Authors 1,316,836 Downloads 221 Institutions

All Articles in Data Science

Faceted Search

3,235 full-text articles. Page 135 of 155.

Classifying Imbalanced Financial Fraud Data Utilizing Enhanced Random Forest Algorithm, Charles Gardner 2020 Kennesaw State University

Classifying Imbalanced Financial Fraud Data Utilizing Enhanced Random Forest Algorithm, Charles Gardner

Master of Science in Computer Science Theses

Imbalanced datasets have been a unique challenge for machine learning, requiring specialized approaches to correctly classify the minority class. Financial fraud detection involves using highly imbalanced datasets with a class imbalance of up to .01% frauds to 99.99% regular transactions. It is essential to identify all frauds in financial fraud detection, even if some classifications' precision is low. I developed a random forest assembly that separates fraudulent transactions into tiers of precision. With this approach, 96% of fraudulent transactions are identified, showing an 8% increase in recall when compared to standard approaches. 59% of fraud classifications' precision increases by 10% …


Reading Pdfs Using Adversarially Trained Convolutional Neural Network Based Optical Character Recognition, Michael B. Brewer, Michael Catalano, Yat Leung, David Stroud 2020 Southern Methodist University

Reading Pdfs Using Adversarially Trained Convolutional Neural Network Based Optical Character Recognition, Michael B. Brewer, Michael Catalano, Yat Leung, David Stroud

SMU Data Science Review

A common problem that has plagued companies for years is digitizing documents and making use of the data contained within. Optical Character Recognition (OCR) technology has flooded the market, but companies still face challenges productionizing these solutions at scale. Although these technologies can identify and recognize the text on the page, they fail to classify the data to the appropriate datatype in an automated system that uses OCR technology as its data mining process. The research contained in this paper presents a novel framework for the identification of datapoints on check stub images by utilizing generative adversarial networks (GANs) to …


Data Science In The Time Of Covid-19, Tony Breitzman 2020 Rowan University

Data Science In The Time Of Covid-19, Tony Breitzman

College of Science & Mathematics Departmental Research

No abstract provided.


Principal Component Analysis For Predicting The Party Of The Legislators, Afsana Mimi 2020 CUNY New York City College of Technology

Principal Component Analysis For Predicting The Party Of The Legislators, Afsana Mimi

Publications and Research

In Spring 2020, I did a project, "Decision Tree Predicting the Party of Legislators," and construct a decision tree model to predict legislators' parties' based on their votes. We also use this model to identify legislators who frequently voted against their parties. We used the legislators' roll call votes, Office of Clerk U.S. House of Representatives Data Sets (Categorical values) collected in 2018 and 2019. In this new project, We study the 2018 and 2019 vote data using Principal Component Analysis (PCA). The goal is to find a (compressed) model using unsupervised learning to distinguish the legislators' parties, and PCA …


Introduction To Data Science Lti 110, Joanna Burkhardt 2020 University of Rhode Island

Introduction To Data Science Lti 110, Joanna Burkhardt

Library Impact Statements

No abstract provided.


Introduction To Data Science, Joanna Burkhardt 2020 University of Rhode Island

Introduction To Data Science, Joanna Burkhardt

Library Impact Statements

No abstract provided.


Spatial Frequency Implications For Global And Local Processing In Autistic Children, Riya Mody, Ayra Tusneem, LouAnne Boyd, Vincent Berardi 2020 Chapman University

Spatial Frequency Implications For Global And Local Processing In Autistic Children, Riya Mody, Ayra Tusneem, Louanne Boyd, Vincent Berardi

Student Scholar Symposium Abstracts and Posters

Visual processing in humans is done by integrating and updating multiple streams of global and local sensory input. Interaction between these two systems can be disrupted in individuals with ASD and other learning disabilities. When this integration is not done smoothly, it becomes difficult to see the “big picture”, which has been found to have implications on emotion recognition, social skills, and conversation skills. An example of this phenomenon is local interference, which is when local details are prioritized over the global features. Previous research in this field has aimed to decrease local interference by developing and evaluating a filter …


Factors Affecting Computer Science Research Productivity And Impact In Nigeria: A Bibliometric Evidence, Azubuike Ezenwoke 2020 Covenant University

Factors Affecting Computer Science Research Productivity And Impact In Nigeria: A Bibliometric Evidence, Azubuike Ezenwoke

Library Philosophy and Practice (e-journal)

Computer science is a burgeoning research field and has the potential to accelerate the rate of industrialisation and subsequently, economic development. Using bibliometric data obtained from Scopus, this study employed a 15-year bibliometric analysis to highlight Nigeria’s productivity and impact trends in the computer science research landscape. Our findings are summarised as follows: First, Nigeria’s computer science research contribution and citations are meager in comparison to the global output. Secondly, international collaboration is generally weak as most collaborations are national in scope. Third, Nigeria’s computer science-related research is published in low-quality outlets, as Scopus has discontinued the indexing of most …


Hierarchical Aggregation Of Multidimensional Data For Efficient Data Mining, Safaa Khalil Alwajidi 2020 Western Michigan University

Hierarchical Aggregation Of Multidimensional Data For Efficient Data Mining, Safaa Khalil Alwajidi

Dissertations

Big data analysis is essential for many smart applications in areas such as connected healthcare, intelligent transportation, human activity recognition, environment, and climate change monitoring. Traditional data mining algorithms do not scale well to big data due to the enormous number of data points and the velocity of their generation. Mining and learning from big data need time and memory efficiency techniques, albeit the cost of possible loss in accuracy. This research focuses on the mining of big data using aggregated data as input. We developed a data structure that is to be used to aggregate data at multiple resolutions. …


Evaluating The Reproducibility Of Physiological Stress Detection Models, Varun Mishra, Sougata Sen, Grace Chen, Tian Hao, Jeffrey Rogers, Ching-Hua Chen, David Kotz 2020 Dartmouth College

Evaluating The Reproducibility Of Physiological Stress Detection Models, Varun Mishra, Sougata Sen, Grace Chen, Tian Hao, Jeffrey Rogers, Ching-Hua Chen, David Kotz

Dartmouth Scholarship

Recent advances in wearable sensor technologies have led to a variety of approaches for detecting physiological stress. Even with over a decade of research in the domain, there still exist many significant challenges, including a near-total lack of reproducibility across studies. Researchers often use some physiological sensors (custom-made or off-the-shelf), conduct a study to collect data, and build machine-learning models to detect stress. There is little effort to test the applicability of the model with similar physiological data collected from different devices, or the efficacy of the model on data collected from different studies, populations, or demographics.

This paper takes …


Open Data, Collaborative Working Platforms, And Interdisciplinary Collaboration: Building An Early Career Scientist Community Of Practice To Leverage Ocean Observatories Initiative Data To Address Critical Questions In Marine Science, Robert M. Levine, Kristen E. Fogaren, Johna E. Rudzin, Christopher J. Russoniello, Dax C. Soule, Justine M. Whitaker 2020 University of Washington - Seattle Campus

Open Data, Collaborative Working Platforms, And Interdisciplinary Collaboration: Building An Early Career Scientist Community Of Practice To Leverage Ocean Observatories Initiative Data To Address Critical Questions In Marine Science, Robert M. Levine, Kristen E. Fogaren, Johna E. Rudzin, Christopher J. Russoniello, Dax C. Soule, Justine M. Whitaker

Publications and Research

Ocean observing systems are well-recognized as platforms for long-term monitoring of near-shore and remote locations in the global ocean. High-quality observatory data is freely available and accessible to all members of the global oceanographic community—a democratization of data that is particularly useful for early career scientists (ECS), enabling ECS to conduct research independent of traditional funding models or access to laboratory and field equipment. The concurrent collection of distinct data types with relevance for oceanographic disciplines including physics, chemistry, biology, and geology yields a unique incubator for cutting-edge, timely, interdisciplinary research. These data are both an opportunity and an incentive …


Healthcare Regulation And Governance: Big Data Analytics And Healthcare Data Protection, Xuejuan Zhang 2020 Thomas Jefferson University

Healthcare Regulation And Governance: Big Data Analytics And Healthcare Data Protection, Xuejuan Zhang

School of Continuing and Professional Studies Student Papers

No abstract provided.


Simple Modification For An Apriori Algorithm With Combination Reduction And Iteration Limitation Technique, Adie Wahyudi Oktavia Gama, Ni Made Widnyani 2020 Department of Information Technology, Universitas Pendidikan Nasional, Indonesia

Simple Modification For An Apriori Algorithm With Combination Reduction And Iteration Limitation Technique, Adie Wahyudi Oktavia Gama, Ni Made Widnyani

Knowledge Engineering and Data Science

Apriori algorithm is one of the methods with regard to association rules in data mining. This algorithm uses knowledge from an itemset previously formed with frequent occurrence frequencies to form the next itemset. An a priori algorithm generates a combination by iteration methods that are using repeated database scanning process, pairing one product with another product and then recording the number of occurrences of the combination with the minimum limit of support and confidence values. The a priori algorithm will slow down to an expanding database in the process of finding frequent itemset to form association rules. Modification techniques are …


Segmentation Method For Face Modelling In Thermal Images, Albar Albar, Hendrick Hendrick, Rahmat Hidayat 2020 Department of Electrical Engineering, Politeknik Negeri Padang, Indonesia

Segmentation Method For Face Modelling In Thermal Images, Albar Albar, Hendrick Hendrick, Rahmat Hidayat

Knowledge Engineering and Data Science

Face detection is mostly applied in RGB images. The object detection usually applied the Deep Learning method for model creation. One method face spoofing is by using a thermal camera. The famous object detection methods are Yolo, Fast Region Based Convolutional Neural Networks (RCNN), Faster RCNN, SSD, and Mask RCNN. We proposed a segmentation Mask RCNN method to create a face model from thermal images. This model was able to locate the face area in images. The dataset was established using 1600 images. The images were created from direct capturing and collecting from the online dataset. The Mask RCNN was …


Generating Javanese Stopwords List Using K-Means Clustering Algorithm, Aji Prasetya Wibawa, Hidayah Kariima Fithri, Ilham Ari Elbaith Zaeni, Andrew Nafalski 2020 Electrical Engineering Department, Universitas Negeri Malang, Indonesia

Generating Javanese Stopwords List Using K-Means Clustering Algorithm, Aji Prasetya Wibawa, Hidayah Kariima Fithri, Ilham Ari Elbaith Zaeni, Andrew Nafalski

Knowledge Engineering and Data Science

Stopword removal necessary in Information Retrieval. It can remove frequently appeared and general words to reduce memory storage. The algorithm eliminates each word that is precisely the same as the word in the stopword list. However, generating the list could be time-consuming. The words in a specific language and domain must be collected and validated by specialists. This research aims to develop a new way to generate a stop word list using the K-means Clustering method. The proposed approach groups words based on their frequency. The confusion matrix calculates the difference between the findings with a valid stopword list created …


Efficient Scheduling Of Plantation Company Workers Using Genetic Algorithm, Wayan Firdaus Mahmudy, Andreas Pardede, Agus Wahyu Widodo, Muh Arif Rahman 2020 Faculty of Computer Science, Brawijaya University, Indonesia

Efficient Scheduling Of Plantation Company Workers Using Genetic Algorithm, Wayan Firdaus Mahmudy, Andreas Pardede, Agus Wahyu Widodo, Muh Arif Rahman

Knowledge Engineering and Data Science

Workers at large plantation companies have various activities. These activities include caring for plants, regularly applying fertilizers according to schedule, and crop harvesting activities. The density of worker activities must be balanced with efficient and fair work scheduling. A good schedule will minimize worker dissatisfaction while also maintaining their physical health. This study aims to optimize workers' schedules using a genetic algorithm. An efficient chromosome representation is designed to produce a good schedule in a reasonable amount of time. The mutation method is used in combination with reciprocal mutation and exchange mutation, while the type of crossover used is one …


A Review Of Accessing Big Data With Significant Ontologies, Jumah Y.J Sleeman, Jehad A.H Hammad 2020 Department of Computer Information Systems, Al-Quds Open University, Palestine

A Review Of Accessing Big Data With Significant Ontologies, Jumah Y.J Sleeman, Jehad A.H Hammad

Knowledge Engineering and Data Science

Ontology Based Data Access (OBDA) is a recently proposed approach which is able to provide a conceptual view on relational data sources. It addresses the problem of the direct access to big data through providing end-users with an ontology that goes between users and sources in which the ontology is connected to the data via mappings. We introduced the languages used to represent the ontologies and the mapping assertions technique that derived the query answering from sources. Query answering is divided into two steps: (i) Ontology rewriting, in which the query is rewritten with respect to the ontology into new …


Convolutional Neural Network On Tanned And Synthetic Leather Textures, Faadihilah Ahnaf Faiz, Ahmad Azhari 2020 Department of Informatics, Universitas Ahmad Dahlan, Indonesia

Convolutional Neural Network On Tanned And Synthetic Leather Textures, Faadihilah Ahnaf Faiz, Ahmad Azhari

Knowledge Engineering and Data Science

Tanned leather is an output from complex processes called tanning. Leather tanning is an important step that used to protect the fiber or protein structure of animal’s skin. Another reason of tanning process is to prevent the animal’s skin from any defect or rot. After the tanning is complete, the leather can be applied to produce a wide variety of leather products. Thus, the leather prices usually more expensive because it takes longer time in process. Another way to get cheaper price is make non-animal leather that usually known as synthetic or imitation leather. The purpose of this paper is …


Computational Behavioral Analytics: Estimating Psychological Traits In Foreign Languages., Kristopher Wayne Reese 2020 University of Louisville

Computational Behavioral Analytics: Estimating Psychological Traits In Foreign Languages., Kristopher Wayne Reese

Electronic Theses and Dissertations

The rise of technology proliferating into the workplace has increased the threat of loss of intellectual property, classified, and proprietary information for companies, governments, and academics. This can cause economic damage to the creators of new IP, companies, and whole economies. This technology proliferation has also assisted terror groups and lone wolf actors in pushing their message to a larger audience or finding similar tribal groups that share common, sometimes flawed, beliefs across various social media platforms. These types of challenges have created numerous studies in psycholinguistics, as well as commercial tools, that look to assist in identifying potential threats …


Detecting Hacker Threats: Performance Of Word And Sentence Embedding Models In Identifying Hacker Communications, Susan McKeever, Brian Keegan, Andrei Quieroz 2020 Technological University Dublin

Detecting Hacker Threats: Performance Of Word And Sentence Embedding Models In Identifying Hacker Communications, Susan Mckeever, Brian Keegan, Andrei Quieroz

Conference papers

Abstract—Cyber security is striving to find new forms of protection against hacker attacks. An emerging approach nowadays is the investigation of security-related messages exchanged on deep/dark web and even surface web channels. This approach can be supported by the use of supervised machine learning models and text mining techniques. In our work, we compare a variety of machine learning algorithms, text representations and dimension reduction approaches for the detection accuracies of software-vulnerability-related communications. Given the imbalanced nature of the three public datasets used, we investigate appropriate sampling approaches to boost detection accuracies of our models. In addition, we examine how …


Digital Commons powered by bepress