Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons™

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 2701 - 2730 of 3244

Full-Text Articles in Data Science

Simple Modification For An Apriori Algorithm With Combination Reduction And Iteration Limitation Technique, Adie Wahyudi Oktavia Gama, Ni Made Widnyani Dec 2020

Simple Modification For An Apriori Algorithm With Combination Reduction And Iteration Limitation Technique, Adie Wahyudi Oktavia Gama, Ni Made Widnyani

Knowledge Engineering and Data Science

Apriori algorithm is one of the methods with regard to association rules in data mining. This algorithm uses knowledge from an itemset previously formed with frequent occurrence frequencies to form the next itemset. An a priori algorithm generates a combination by iteration methods that are using repeated database scanning process, pairing one product with another product and then recording the number of occurrences of the combination with the minimum limit of support and confidence values. The a priori algorithm will slow down to an expanding database in the process of finding frequent itemset to form association rules. Modification techniques are …


Segmentation Method For Face Modelling In Thermal Images, Albar Albar, Hendrick Hendrick, Rahmat Hidayat Dec 2020

Segmentation Method For Face Modelling In Thermal Images, Albar Albar, Hendrick Hendrick, Rahmat Hidayat

Knowledge Engineering and Data Science

Face detection is mostly applied in RGB images. The object detection usually applied the Deep Learning method for model creation. One method face spoofing is by using a thermal camera. The famous object detection methods are Yolo, Fast Region Based Convolutional Neural Networks (RCNN), Faster RCNN, SSD, and Mask RCNN. We proposed a segmentation Mask RCNN method to create a face model from thermal images. This model was able to locate the face area in images. The dataset was established using 1600 images. The images were created from direct capturing and collecting from the online dataset. The Mask RCNN was …


Generating Javanese Stopwords List Using K-Means Clustering Algorithm, Aji Prasetya Wibawa, Hidayah Kariima Fithri, Ilham Ari Elbaith Zaeni, Andrew Nafalski Dec 2020

Generating Javanese Stopwords List Using K-Means Clustering Algorithm, Aji Prasetya Wibawa, Hidayah Kariima Fithri, Ilham Ari Elbaith Zaeni, Andrew Nafalski

Knowledge Engineering and Data Science

Stopword removal necessary in Information Retrieval. It can remove frequently appeared and general words to reduce memory storage. The algorithm eliminates each word that is precisely the same as the word in the stopword list. However, generating the list could be time-consuming. The words in a specific language and domain must be collected and validated by specialists. This research aims to develop a new way to generate a stop word list using the K-means Clustering method. The proposed approach groups words based on their frequency. The confusion matrix calculates the difference between the findings with a valid stopword list created …


Efficient Scheduling Of Plantation Company Workers Using Genetic Algorithm, Wayan Firdaus Mahmudy, Andreas Pardede, Agus Wahyu Widodo, Muh Arif Rahman Dec 2020

Efficient Scheduling Of Plantation Company Workers Using Genetic Algorithm, Wayan Firdaus Mahmudy, Andreas Pardede, Agus Wahyu Widodo, Muh Arif Rahman

Knowledge Engineering and Data Science

Workers at large plantation companies have various activities. These activities include caring for plants, regularly applying fertilizers according to schedule, and crop harvesting activities. The density of worker activities must be balanced with efficient and fair work scheduling. A good schedule will minimize worker dissatisfaction while also maintaining their physical health. This study aims to optimize workers' schedules using a genetic algorithm. An efficient chromosome representation is designed to produce a good schedule in a reasonable amount of time. The mutation method is used in combination with reciprocal mutation and exchange mutation, while the type of crossover used is one …


A Review Of Accessing Big Data With Significant Ontologies, Jumah Y.J Sleeman, Jehad A.H Hammad Dec 2020

A Review Of Accessing Big Data With Significant Ontologies, Jumah Y.J Sleeman, Jehad A.H Hammad

Knowledge Engineering and Data Science

Ontology Based Data Access (OBDA) is a recently proposed approach which is able to provide a conceptual view on relational data sources. It addresses the problem of the direct access to big data through providing end-users with an ontology that goes between users and sources in which the ontology is connected to the data via mappings. We introduced the languages used to represent the ontologies and the mapping assertions technique that derived the query answering from sources. Query answering is divided into two steps: (i) Ontology rewriting, in which the query is rewritten with respect to the ontology into new …


Convolutional Neural Network On Tanned And Synthetic Leather Textures, Faadihilah Ahnaf Faiz, Ahmad Azhari Dec 2020

Convolutional Neural Network On Tanned And Synthetic Leather Textures, Faadihilah Ahnaf Faiz, Ahmad Azhari

Knowledge Engineering and Data Science

Tanned leather is an output from complex processes called tanning. Leather tanning is an important step that used to protect the fiber or protein structure of animal’s skin. Another reason of tanning process is to prevent the animal’s skin from any defect or rot. After the tanning is complete, the leather can be applied to produce a wide variety of leather products. Thus, the leather prices usually more expensive because it takes longer time in process. Another way to get cheaper price is make non-animal leather that usually known as synthetic or imitation leather. The purpose of this paper is …


Computational Behavioral Analytics: Estimating Psychological Traits In Foreign Languages., Kristopher Wayne Reese Dec 2020

Computational Behavioral Analytics: Estimating Psychological Traits In Foreign Languages., Kristopher Wayne Reese

Electronic Theses and Dissertations

The rise of technology proliferating into the workplace has increased the threat of loss of intellectual property, classified, and proprietary information for companies, governments, and academics. This can cause economic damage to the creators of new IP, companies, and whole economies. This technology proliferation has also assisted terror groups and lone wolf actors in pushing their message to a larger audience or finding similar tribal groups that share common, sometimes flawed, beliefs across various social media platforms. These types of challenges have created numerous studies in psycholinguistics, as well as commercial tools, that look to assist in identifying potential threats …


Exploring Information For Quantum Machine Learning Models, Michael Telahun Dec 2020

Exploring Information For Quantum Machine Learning Models, Michael Telahun

Electronic Theses and Dissertations

Quantum computing performs calculations by using physical phenomena and quantum mechanics principles to solve problems. This form of computation theoretically has been shown to provide speed ups to some problems of modern-day processing. With much anticipation the utilization of quantum phenomena in the field of Machine Learning has become apparent. The work here develops models from two software frameworks: TensorFlow Quantum (TFQ) and PennyLane for machine learning purposes. Both developed models utilize an information encoding technique amplitude encoding for preparation of states in a quantum learning model. This thesis explores both the capacity for amplitude encoding to provide enriched state …


Data And Assessment Management In Collegiate Recreation, Jeana Carow Dec 2020

Data And Assessment Management In Collegiate Recreation, Jeana Carow

Graduate Theses and Dissertations

Collegiate recreation programs and centers typically provide traditional programming space in addition to a range of physical activity spaces and resources, as a valuable part of the student experience. The external pressures of identifying and communicating departmental value and impact on the campus community has resulted in collegiate recreation departments’ use of data to communicate the effectiveness and impact of their work. The purpose of the study was to identify the data collection and assessment management practices of collegiate recreation departments, particularly focusing on the organization of data and assessment strategies as well as data collection, storage, reporting, analyzing, and …


Incorporating Shear Resistance Into Debris Flow Triggering Model Statistics, Noah J. Lyman Dec 2020

Incorporating Shear Resistance Into Debris Flow Triggering Model Statistics, Noah J. Lyman

Master's Theses

Several regions of the Western United States utilize statistical binary classification models to predict and manage debris flow initiation probability after wildfires. As the occurrence of wildfires and large intensity rainfall events increase, so has the frequency in which development occurs in the steep and mountainous terrain where these events arise. This resulting intersection brings with it an increasing need to derive improved results from existing models, or develop new models, to reduce the economic and human impacts that debris flows may bring. Any development or change to these models could also theoretically increase the ease of collection, processing, and …


Hierarchical Aggregation Of Multidimensional Data For Efficient Data Mining, Safaa Khalil Alwajidi Dec 2020

Hierarchical Aggregation Of Multidimensional Data For Efficient Data Mining, Safaa Khalil Alwajidi

Dissertations

Big data analysis is essential for many smart applications in areas such as connected healthcare, intelligent transportation, human activity recognition, environment, and climate change monitoring. Traditional data mining algorithms do not scale well to big data due to the enormous number of data points and the velocity of their generation. Mining and learning from big data need time and memory efficiency techniques, albeit the cost of possible loss in accuracy. This research focuses on the mining of big data using aggregated data as input. We developed a data structure that is to be used to aggregate data at multiple resolutions. …


Detecting Hacker Threats: Performance Of Word And Sentence Embedding Models In Identifying Hacker Communications, Susan Mckeever, Brian Keegan, Andrei Quieroz Dec 2020

Detecting Hacker Threats: Performance Of Word And Sentence Embedding Models In Identifying Hacker Communications, Susan Mckeever, Brian Keegan, Andrei Quieroz

Conference papers

Abstract—Cyber security is striving to find new forms of protection against hacker attacks. An emerging approach nowadays is the investigation of security-related messages exchanged on deep/dark web and even surface web channels. This approach can be supported by the use of supervised machine learning models and text mining techniques. In our work, we compare a variety of machine learning algorithms, text representations and dimension reduction approaches for the detection accuracies of software-vulnerability-related communications. Given the imbalanced nature of the three public datasets used, we investigate appropriate sampling approaches to boost detection accuracies of our models. In addition, we examine how …


Creating Optimal Conditions For Reproducible Data Analysis In R With ‘Fertile’, Audrey M. Bertin, Benjamin Baumer Nov 2020

Creating Optimal Conditions For Reproducible Data Analysis In R With ‘Fertile’, Audrey M. Bertin, Benjamin Baumer

Statistical and Data Sciences: Faculty Publications

The advancement of scientific knowledge increasingly depends on ensuring that data-driven research is reproducible: that two people with the same data obtain the same results. However, while the necessity of reproducibility is clear, there are significant behavioral and technical challenges that impede its widespread implementation and no clear consensus on standards of what constitutes reproducibility in published research. We present fertile, an R package that focuses on a series of common mistakes programmers make while conducting data science projects in R, primarily through the RStudio integrated development environment. fertile operates in two modes: proactively, to prevent reproducibility mistakes from happening …


Secure Unlinkability Schemes For Privacy Preserving Data Publishing In Weighted Social Networks, Chong Kah Meng Nov 2020

Secure Unlinkability Schemes For Privacy Preserving Data Publishing In Weighted Social Networks, Chong Kah Meng

Student Works (2020-2029)

Preserving privacy of users has been one of the important research issues in social networks. Social networks contain sensitive personal information that are often released for business and research purposes. The privacy of a user can be breached if the data are not released in an anonymized form. In this thesis, we address edge weight disclosure, link disclosure and identity disclosure problems in publishing weighted network data. To counter these privacy risks while preserving high utility of the published data, we define two key privacy properties, namely edge weight unlinkability and node unlinkability. We design two novel anonymization schemes namely …


Development Of Reduced Order Models Using Reservoir Simulation And Physics Informed Machine Learning Techniques, Mark V. Behl Jr Nov 2020

Development Of Reduced Order Models Using Reservoir Simulation And Physics Informed Machine Learning Techniques, Mark V. Behl Jr

LSU Master's Theses

Reservoir simulation is the industry standard for prediction and characterization of processes in the subsurface. However, simulation is computationally expensive and time consuming. This study explores reduced order models (ROMs) as an appropriate alternative. ROMs that use neural networks effectively capture nonlinear dependencies, and only require available operational data as inputs. Neural networks are a black box and difficult to interpret, however. Physics informed neural networks (PINNs) provide a potential solution to these shortcomings, but have not yet been applied extensively in petroleum engineering.

A mature black-oil simulation model from Volve public data release was used to generate training data …


Data Analysis To Evaluate The Performance Of Breathing Masks Used For Filtering Nano-Level Particles At Manufacturing Sites, Gracia M. Dardano Nov 2020

Data Analysis To Evaluate The Performance Of Breathing Masks Used For Filtering Nano-Level Particles At Manufacturing Sites, Gracia M. Dardano

Honors College Theses

The work performed in this research aims to evaluate the performance of commercially available breathing masks in filtering airborne nanoparticles at manufacturing sites. Nanoparticles are found virtually anywhere, from dust in a worksite to a simple sneeze. Therefore, they pose a substantial threat to human health as their velocity and volatility are high. This research analyzes if current efforts of breathing masks to hinder nanoparticles are effective, especially at manufacturing sites. Data has been collected in order to analyze the behavior of nanoparticles and to measure nanoparticle levels at manufacturing sites and its working environment. Data is statistical in nature …


Application Of Tda Mapper To Water Data And Bird Data, Wako Bungula Nov 2020

Application Of Tda Mapper To Water Data And Bird Data, Wako Bungula

Annual Symposium on Biomathematics and Ecology Education and Research

No abstract provided.


A Study Of Sentiment Of Covid-19 Related Tweets In The Usa, Jack Luu, Rosangela Follmann Nov 2020

A Study Of Sentiment Of Covid-19 Related Tweets In The Usa, Jack Luu, Rosangela Follmann

Annual Symposium on Biomathematics and Ecology Education and Research

No abstract provided.


Stochastic Modeling Of Ovarian Follicle Growth In Adult Female Rats, Zhaozhi Li Nov 2020

Stochastic Modeling Of Ovarian Follicle Growth In Adult Female Rats, Zhaozhi Li

Annual Symposium on Biomathematics and Ecology Education and Research

No abstract provided.


Viral Data, Agnieszka Leszczynski, Matthew Zook Nov 2020

Viral Data, Agnieszka Leszczynski, Matthew Zook

Geography Faculty Publications

We are experiencing a historical moment characterized by unprecedented conditions of virality: a viral pandemic, the viral diffusion of misinformation and conspiracy theories, the viral momentum of ongoing Hong Kong protests, and the viral spread of #BlackLivesMatter demonstrations and related efforts to defund policing. These co-articulations of crises, traumas, and virality both implicate and are implicated by big data practices occurring in a present that is pervasively mediated by data materialities, deeply rooted dataist ideologies that entrench processes of datafication as granting objective access to truth and attendant practices of tracking, data analytics, algorithmic prediction, and data-driven targeting of individuals …


Ensemble Labeling Towards Scientific Information Extraction (Elsie), Erin Murphy Nov 2020

Ensemble Labeling Towards Scientific Information Extraction (Elsie), Erin Murphy

College of Computing and Digital Media Dissertations

Extracting scientific facts from unstructured text is difficult due to challenges specific to the ambiguity of the language, the complexity of the scientific named entities and relations to be extracted. This problem is well illustrated through the extraction of polymer names and their properties. Even in the cases where the property is a temperature, identifying the polymer name associated with the temperature may require expertise due to the use of acronyms, synonyms, complicated naming conventions and by the fact that new polymer names are being “introduced” to the vernacular as polymer science advances. While there exist domain-specific machine learning toolkits …


An Analysis Of Technological Components In Relation To Privacy In A Smart City, Kayla Rutherford, Ben Lands, A. J. Stiles Nov 2020

An Analysis Of Technological Components In Relation To Privacy In A Smart City, Kayla Rutherford, Ben Lands, A. J. Stiles

James Madison Undergraduate Research Journal (JMURJ)

A smart city is an interconnection of technological components that store, process, and wirelessly transmit information to enhance the efficiency of applications and the individuals who use those applications. Over the course of the 21st century, it is expected that an overwhelming majority of the world’s population will live in urban areas and that the number of wireless devices will increase. The resulting increase in wireless data transmission means that the privacy of data will be increasingly at risk. This paper uses a holistic problem-solving approach to evaluate the security challenges posed by the technological components that make up a …


Cash Flow Forecasting Using Probabilistic Neural Networks, Marwan Ashour Nov 2020

Cash Flow Forecasting Using Probabilistic Neural Networks, Marwan Ashour

Journal of the Arab American University مجلة الجامعة العربية الامريكية للبحوث

This paper aimed to compare the modern methods of cash flow forecasting with the traditional ones. In other words, the researcher compared between the Probabilistic Neural Networks and Transfer Function. It is worth mentioning that cash flow forecasting , nowadays, is very important and helps the upper management plan, control, assess the performance and make decisions. More specifically, in this paper, the Artificial Neural networks were used to diagnose the nature of the cash flow for the next period of time and then forecast the cash flow. The experiment was conducted in The General company for Electricity Distribution in Baghdad. …


Lifespan Analysis Of Earth Satellites, Venkata Jaipal Reddy Batthula Nov 2020

Lifespan Analysis Of Earth Satellites, Venkata Jaipal Reddy Batthula

Theses and Dissertations

Different countries have their own satellites for their various needs like communication, weather forecast, and security. The first satellite was launched in 1957 into space. Thousands of satellite lifetimes have already ended but they are still in orbit. The present world has more advanced technology when compared with previous technology. So, the technology for satellites is improving compared with the past. We need to understand trends in improvements to satellites related to lifespans better, using a new dataset that has not been available before, as well as datasets that we have worked with before, and that is the purpose of …


Lis Online Graduate Certificate In Data Science, Joanna Burkhardt Nov 2020

Lis Online Graduate Certificate In Data Science, Joanna Burkhardt

Library Impact Statements

No abstract provided.


Efficient And Fair Data Valuation For Horizontal Federated Learning, Shuyue Wei, Yongxin Tong, Zimu Zhou, Tianshu Song Nov 2020

Efficient And Fair Data Valuation For Horizontal Federated Learning, Shuyue Wei, Yongxin Tong, Zimu Zhou, Tianshu Song

Research Collection School Of Computing and Information Systems

Availability of big data is crucial for modern machine learning applications and services. Federated learning is an emerging paradigm to unite different data owners for machine learning on massive data sets without worrying about data privacy. Yet data owners may still be reluctant to contribute unless their data sets are fairly valuated and paid. In this work, we adapt Shapley value, a widely used data valuation metric to valuating data providers in federated learning. Prior data valuation schemes for machine learning incur high computation cost because they require training of extra models on all data set combinations. For efficient data …


Using Data Analytics To Predict Students Score, Nang Laik Ma, Gim Hong Chua Nov 2020

Using Data Analytics To Predict Students Score, Nang Laik Ma, Gim Hong Chua

Research Collection School Of Computing and Information Systems

Education is very important to Singapore, and the government has continued to invest heavily in our education system to become one of the world-class systems today. A strong foundation of Science, Technology, Engineering, and Mathematics (STEM) was what underpinned Singapore's development over the past 50 years. PISA is a triennial international survey that evaluates education systems worldwide by testing the skills and knowledge of 15-year-old students who are nearing the end of compulsory education. In this paper, the authors used the PISA data from 2012 and 2015 and developed machine learning techniques to predictive the students' scores and understand the …


A New Efficient Method To Detect Genetic Interactions For Lung Cancer Gwas, Jennifer Luyapan, Xuemei Ji, Siting Li, Xiangjun Xiao, Dakai Zhu, Eric J. Duell, David C. Christiani, Matthew B. Schabath, Susanne M. Arnold, Shanbeh Zienolddiny, Hans Brunnström, Olle Melander, Mark D. Thornquist, Todd A. Mackenzie, Christopher I. Amos, Jiang Gui Oct 2020

A New Efficient Method To Detect Genetic Interactions For Lung Cancer Gwas, Jennifer Luyapan, Xuemei Ji, Siting Li, Xiangjun Xiao, Dakai Zhu, Eric J. Duell, David C. Christiani, Matthew B. Schabath, Susanne M. Arnold, Shanbeh Zienolddiny, Hans Brunnström, Olle Melander, Mark D. Thornquist, Todd A. Mackenzie, Christopher I. Amos, Jiang Gui

Markey Cancer Center Faculty Publications

BACKGROUND: Genome-wide association studies (GWAS) have proven successful in predicting genetic risk of disease using single-locus models; however, identifying single nucleotide polymorphism (SNP) interactions at the genome-wide scale is limited due to computational and statistical challenges. We addressed the computational burden encountered when detecting SNP interactions for survival analysis, such as age of disease-onset. To confront this problem, we developed a novel algorithm, called the Efficient Survival Multifactor Dimensionality Reduction (ES-MDR) method, which used Martingale Residuals as the outcome parameter to estimate survival outcomes, and implemented the Quantitative Multifactor Dimensionality Reduction method to identify significant interactions associated with age of …


Comparing Variable Importance In Prediction Of Silence Behaviours Between Random Forest And Conditional Inference Forest Models., Stephen Barrett Dr, Geraldine Gray Dr, Colm Mcguinness Dr, Michael Knoll Dr. Oct 2020

Comparing Variable Importance In Prediction Of Silence Behaviours Between Random Forest And Conditional Inference Forest Models., Stephen Barrett Dr, Geraldine Gray Dr, Colm Mcguinness Dr, Michael Knoll Dr.

Articles

This paper explores variable importance metrics of Conditional Inference Trees (CIT) and classical Classification And Regression Trees (CART) based Random Forests. The paper compares both algorithms variable importance rankings and highlights why CIT should be used when dealing with data with different levels of aggregation. The models analysed explored the role of cultural factors at individual and societal level when predicting Organisational Silence behaviours.


Towards High Performance Stock Market Prediction Methods, Warren M. Landis, Sangwhan Cha Oct 2020

Towards High Performance Stock Market Prediction Methods, Warren M. Landis, Sangwhan Cha

Other Student Works

Stock markets of today, and will continue to in the future, rely on the metrics of timeliness and efficiency to reach optimal profits. A way stock investors have continued to strive for the best of these two factors of the business is through the use of predictive machine learning systems to help aid in their decision making. However, among the many systems currently in use, it could be said that the myriad of data that they are based on may not be sufficient. In an effort to devise an ensemble learning predictive system that will utilize an array of big …