Classifying Imbalanced Financial Fraud Data Utilizing Enhanced Random Forest Algorithm,
2020
Kennesaw State University
Classifying Imbalanced Financial Fraud Data Utilizing Enhanced Random Forest Algorithm, Charles Gardner
Master of Science in Computer Science Theses
Imbalanced datasets have been a unique challenge for machine learning, requiring specialized approaches to correctly classify the minority class. Financial fraud detection involves using highly imbalanced datasets with a class imbalance of up to .01% frauds to 99.99% regular transactions. It is essential to identify all frauds in financial fraud detection, even if some classifications' precision is low. I developed a random forest assembly that separates fraudulent transactions into tiers of precision. With this approach, 96% of fraudulent transactions are identified, showing an 8% increase in recall when compared to standard approaches. 59% of fraud classifications' precision increases by 10% …
Reading Pdfs Using Adversarially Trained Convolutional Neural Network Based Optical Character Recognition,
2020
Southern Methodist University
Reading Pdfs Using Adversarially Trained Convolutional Neural Network Based Optical Character Recognition, Michael B. Brewer, Michael Catalano, Yat Leung, David Stroud
SMU Data Science Review
A common problem that has plagued companies for years is digitizing documents and making use of the data contained within. Optical Character Recognition (OCR) technology has flooded the market, but companies still face challenges productionizing these solutions at scale. Although these technologies can identify and recognize the text on the page, they fail to classify the data to the appropriate datatype in an automated system that uses OCR technology as its data mining process. The research contained in this paper presents a novel framework for the identification of datapoints on check stub images by utilizing generative adversarial networks (GANs) to …
Data Science In The Time Of Covid-19,
2020
Rowan University
Data Science In The Time Of Covid-19, Tony Breitzman
College of Science & Mathematics Departmental Research
No abstract provided.
Principal Component Analysis For Predicting The Party Of The Legislators,
2020
CUNY New York City College of Technology
Principal Component Analysis For Predicting The Party Of The Legislators, Afsana Mimi
Publications and Research
In Spring 2020, I did a project, "Decision Tree Predicting the Party of Legislators," and construct a decision tree model to predict legislators' parties' based on their votes. We also use this model to identify legislators who frequently voted against their parties. We used the legislators' roll call votes, Office of Clerk U.S. House of Representatives Data Sets (Categorical values) collected in 2018 and 2019. In this new project, We study the 2018 and 2019 vote data using Principal Component Analysis (PCA). The goal is to find a (compressed) model using unsupervised learning to distinguish the legislators' parties, and PCA …
Introduction To Data Science Lti 110,
2020
University of Rhode Island
Introduction To Data Science Lti 110, Joanna Burkhardt
Library Impact Statements
No abstract provided.
Introduction To Data Science,
2020
University of Rhode Island
Introduction To Data Science, Joanna Burkhardt
Library Impact Statements
No abstract provided.
Spatial Frequency Implications For Global And Local Processing In Autistic Children,
2020
Chapman University
Spatial Frequency Implications For Global And Local Processing In Autistic Children, Riya Mody, Ayra Tusneem, Louanne Boyd, Vincent Berardi
Student Scholar Symposium Abstracts and Posters
Visual processing in humans is done by integrating and updating multiple streams of global and local sensory input. Interaction between these two systems can be disrupted in individuals with ASD and other learning disabilities. When this integration is not done smoothly, it becomes difficult to see the “big picture”, which has been found to have implications on emotion recognition, social skills, and conversation skills. An example of this phenomenon is local interference, which is when local details are prioritized over the global features. Previous research in this field has aimed to decrease local interference by developing and evaluating a filter …
Factors Affecting Computer Science Research Productivity And Impact In Nigeria: A Bibliometric Evidence,
2020
Covenant University
Factors Affecting Computer Science Research Productivity And Impact In Nigeria: A Bibliometric Evidence, Azubuike Ezenwoke
Library Philosophy and Practice (e-journal)
Computer science is a burgeoning research field and has the potential to accelerate the rate of industrialisation and subsequently, economic development. Using bibliometric data obtained from Scopus, this study employed a 15-year bibliometric analysis to highlight Nigeria’s productivity and impact trends in the computer science research landscape. Our findings are summarised as follows: First, Nigeria’s computer science research contribution and citations are meager in comparison to the global output. Secondly, international collaboration is generally weak as most collaborations are national in scope. Third, Nigeria’s computer science-related research is published in low-quality outlets, as Scopus has discontinued the indexing of most …
Hierarchical Aggregation Of Multidimensional Data For Efficient Data Mining,
2020
Western Michigan University
Hierarchical Aggregation Of Multidimensional Data For Efficient Data Mining, Safaa Khalil Alwajidi
Dissertations
Big data analysis is essential for many smart applications in areas such as connected healthcare, intelligent transportation, human activity recognition, environment, and climate change monitoring. Traditional data mining algorithms do not scale well to big data due to the enormous number of data points and the velocity of their generation. Mining and learning from big data need time and memory efficiency techniques, albeit the cost of possible loss in accuracy. This research focuses on the mining of big data using aggregated data as input. We developed a data structure that is to be used to aggregate data at multiple resolutions. …
Evaluating The Reproducibility Of Physiological Stress Detection Models,
2020
Dartmouth College
Evaluating The Reproducibility Of Physiological Stress Detection Models, Varun Mishra, Sougata Sen, Grace Chen, Tian Hao, Jeffrey Rogers, Ching-Hua Chen, David Kotz
Dartmouth Scholarship
Recent advances in wearable sensor technologies have led to a variety of approaches for detecting physiological stress. Even with over a decade of research in the domain, there still exist many significant challenges, including a near-total lack of reproducibility across studies. Researchers often use some physiological sensors (custom-made or off-the-shelf), conduct a study to collect data, and build machine-learning models to detect stress. There is little effort to test the applicability of the model with similar physiological data collected from different devices, or the efficacy of the model on data collected from different studies, populations, or demographics.
This paper takes …
Open Data, Collaborative Working Platforms, And Interdisciplinary Collaboration: Building An Early Career Scientist Community Of Practice To Leverage Ocean Observatories Initiative Data To Address Critical Questions In Marine Science,
2020
University of Washington - Seattle Campus
Open Data, Collaborative Working Platforms, And Interdisciplinary Collaboration: Building An Early Career Scientist Community Of Practice To Leverage Ocean Observatories Initiative Data To Address Critical Questions In Marine Science, Robert M. Levine, Kristen E. Fogaren, Johna E. Rudzin, Christopher J. Russoniello, Dax C. Soule, Justine M. Whitaker
Publications and Research
Ocean observing systems are well-recognized as platforms for long-term monitoring of near-shore and remote locations in the global ocean. High-quality observatory data is freely available and accessible to all members of the global oceanographic community—a democratization of data that is particularly useful for early career scientists (ECS), enabling ECS to conduct research independent of traditional funding models or access to laboratory and field equipment. The concurrent collection of distinct data types with relevance for oceanographic disciplines including physics, chemistry, biology, and geology yields a unique incubator for cutting-edge, timely, interdisciplinary research. These data are both an opportunity and an incentive …
Healthcare Regulation And Governance: Big Data Analytics And Healthcare Data Protection,
2020
Thomas Jefferson University
Healthcare Regulation And Governance: Big Data Analytics And Healthcare Data Protection, Xuejuan Zhang
School of Continuing and Professional Studies Student Papers
No abstract provided.
Simple Modification For An Apriori Algorithm With Combination Reduction And Iteration Limitation Technique,
2020
Department of Information Technology, Universitas Pendidikan Nasional, Indonesia
Simple Modification For An Apriori Algorithm With Combination Reduction And Iteration Limitation Technique, Adie Wahyudi Oktavia Gama, Ni Made Widnyani
Knowledge Engineering and Data Science
Apriori algorithm is one of the methods with regard to association rules in data mining. This algorithm uses knowledge from an itemset previously formed with frequent occurrence frequencies to form the next itemset. An a priori algorithm generates a combination by iteration methods that are using repeated database scanning process, pairing one product with another product and then recording the number of occurrences of the combination with the minimum limit of support and confidence values. The a priori algorithm will slow down to an expanding database in the process of finding frequent itemset to form association rules. Modification techniques are …
Segmentation Method For Face Modelling In Thermal Images,
2020
Department of Electrical Engineering, Politeknik Negeri Padang, Indonesia
Segmentation Method For Face Modelling In Thermal Images, Albar Albar, Hendrick Hendrick, Rahmat Hidayat
Knowledge Engineering and Data Science
Face detection is mostly applied in RGB images. The object detection usually applied the Deep Learning method for model creation. One method face spoofing is by using a thermal camera. The famous object detection methods are Yolo, Fast Region Based Convolutional Neural Networks (RCNN), Faster RCNN, SSD, and Mask RCNN. We proposed a segmentation Mask RCNN method to create a face model from thermal images. This model was able to locate the face area in images. The dataset was established using 1600 images. The images were created from direct capturing and collecting from the online dataset. The Mask RCNN was …
Generating Javanese Stopwords List Using K-Means Clustering Algorithm,
2020
Electrical Engineering Department, Universitas Negeri Malang, Indonesia
Generating Javanese Stopwords List Using K-Means Clustering Algorithm, Aji Prasetya Wibawa, Hidayah Kariima Fithri, Ilham Ari Elbaith Zaeni, Andrew Nafalski
Knowledge Engineering and Data Science
Stopword removal necessary in Information Retrieval. It can remove frequently appeared and general words to reduce memory storage. The algorithm eliminates each word that is precisely the same as the word in the stopword list. However, generating the list could be time-consuming. The words in a specific language and domain must be collected and validated by specialists. This research aims to develop a new way to generate a stop word list using the K-means Clustering method. The proposed approach groups words based on their frequency. The confusion matrix calculates the difference between the findings with a valid stopword list created …
Efficient Scheduling Of Plantation Company Workers Using Genetic Algorithm,
2020
Faculty of Computer Science, Brawijaya University, Indonesia
Efficient Scheduling Of Plantation Company Workers Using Genetic Algorithm, Wayan Firdaus Mahmudy, Andreas Pardede, Agus Wahyu Widodo, Muh Arif Rahman
Knowledge Engineering and Data Science
Workers at large plantation companies have various activities. These activities include caring for plants, regularly applying fertilizers according to schedule, and crop harvesting activities. The density of worker activities must be balanced with efficient and fair work scheduling. A good schedule will minimize worker dissatisfaction while also maintaining their physical health. This study aims to optimize workers' schedules using a genetic algorithm. An efficient chromosome representation is designed to produce a good schedule in a reasonable amount of time. The mutation method is used in combination with reciprocal mutation and exchange mutation, while the type of crossover used is one …
A Review Of Accessing Big Data With Significant Ontologies,
2020
Department of Computer Information Systems, Al-Quds Open University, Palestine
A Review Of Accessing Big Data With Significant Ontologies, Jumah Y.J Sleeman, Jehad A.H Hammad
Knowledge Engineering and Data Science
Ontology Based Data Access (OBDA) is a recently proposed approach which is able to provide a conceptual view on relational data sources. It addresses the problem of the direct access to big data through providing end-users with an ontology that goes between users and sources in which the ontology is connected to the data via mappings. We introduced the languages used to represent the ontologies and the mapping assertions technique that derived the query answering from sources. Query answering is divided into two steps: (i) Ontology rewriting, in which the query is rewritten with respect to the ontology into new …
Convolutional Neural Network On Tanned And Synthetic Leather Textures,
2020
Department of Informatics, Universitas Ahmad Dahlan, Indonesia
Convolutional Neural Network On Tanned And Synthetic Leather Textures, Faadihilah Ahnaf Faiz, Ahmad Azhari
Knowledge Engineering and Data Science
Tanned leather is an output from complex processes called tanning. Leather tanning is an important step that used to protect the fiber or protein structure of animal’s skin. Another reason of tanning process is to prevent the animal’s skin from any defect or rot. After the tanning is complete, the leather can be applied to produce a wide variety of leather products. Thus, the leather prices usually more expensive because it takes longer time in process. Another way to get cheaper price is make non-animal leather that usually known as synthetic or imitation leather. The purpose of this paper is …
Computational Behavioral Analytics: Estimating Psychological Traits In Foreign Languages.,
2020
University of Louisville
Computational Behavioral Analytics: Estimating Psychological Traits In Foreign Languages., Kristopher Wayne Reese
Electronic Theses and Dissertations
The rise of technology proliferating into the workplace has increased the threat of loss of intellectual property, classified, and proprietary information for companies, governments, and academics. This can cause economic damage to the creators of new IP, companies, and whole economies. This technology proliferation has also assisted terror groups and lone wolf actors in pushing their message to a larger audience or finding similar tribal groups that share common, sometimes flawed, beliefs across various social media platforms. These types of challenges have created numerous studies in psycholinguistics, as well as commercial tools, that look to assist in identifying potential threats …
Detecting Hacker Threats: Performance Of Word And Sentence Embedding Models In Identifying Hacker Communications,
2020
Technological University Dublin
Detecting Hacker Threats: Performance Of Word And Sentence Embedding Models In Identifying Hacker Communications, Susan Mckeever, Brian Keegan, Andrei Quieroz
Conference papers
Abstract—Cyber security is striving to find new forms of protection against hacker attacks. An emerging approach nowadays is the investigation of security-related messages exchanged on deep/dark web and even surface web channels. This approach can be supported by the use of supervised machine learning models and text mining techniques. In our work, we compare a variety of machine learning algorithms, text representations and dimension reduction approaches for the detection accuracies of software-vulnerability-related communications. Given the imbalanced nature of the three public datasets used, we investigate appropriate sampling approaches to boost detection accuracies of our models. In addition, we examine how …
