Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

3,233 Full-Text Articles 9,308 Authors 1,316,836 Downloads 221 Institutions

All Articles in Data Science

Faceted Search

3,233 full-text articles. Page 136 of 155.

Incorporating Shear Resistance Into Debris Flow Triggering Model Statistics, Noah J. Lyman 2020 California Polytechnic State University, San Luis Obispo

Incorporating Shear Resistance Into Debris Flow Triggering Model Statistics, Noah J. Lyman

Master's Theses

Several regions of the Western United States utilize statistical binary classification models to predict and manage debris flow initiation probability after wildfires. As the occurrence of wildfires and large intensity rainfall events increase, so has the frequency in which development occurs in the steep and mountainous terrain where these events arise. This resulting intersection brings with it an increasing need to derive improved results from existing models, or develop new models, to reduce the economic and human impacts that debris flows may bring. Any development or change to these models could also theoretically increase the ease of collection, processing, and …


Creating Optimal Conditions For Reproducible Data Analysis In R With ‘Fertile’, Audrey M. Bertin, Benjamin Baumer 2020 Smith College

Creating Optimal Conditions For Reproducible Data Analysis In R With ‘Fertile’, Audrey M. Bertin, Benjamin Baumer

Statistical and Data Sciences: Faculty Publications

The advancement of scientific knowledge increasingly depends on ensuring that data-driven research is reproducible: that two people with the same data obtain the same results. However, while the necessity of reproducibility is clear, there are significant behavioral and technical challenges that impede its widespread implementation and no clear consensus on standards of what constitutes reproducibility in published research. We present fertile, an R package that focuses on a series of common mistakes programmers make while conducting data science projects in R, primarily through the RStudio integrated development environment. fertile operates in two modes: proactively, to prevent reproducibility mistakes from happening …


Secure Unlinkability Schemes For Privacy Preserving Data Publishing In Weighted Social Networks, Chong Kah Meng 2020 Universiti Malaya

Secure Unlinkability Schemes For Privacy Preserving Data Publishing In Weighted Social Networks, Chong Kah Meng

Student Works (2020-2029)

Preserving privacy of users has been one of the important research issues in social networks. Social networks contain sensitive personal information that are often released for business and research purposes. The privacy of a user can be breached if the data are not released in an anonymized form. In this thesis, we address edge weight disclosure, link disclosure and identity disclosure problems in publishing weighted network data. To counter these privacy risks while preserving high utility of the published data, we define two key privacy properties, namely edge weight unlinkability and node unlinkability. We design two novel anonymization schemes namely …


Development Of Reduced Order Models Using Reservoir Simulation And Physics Informed Machine Learning Techniques, Mark V. Behl Jr 2020 Louisiana State University and Agricultural and Mechanical College

Development Of Reduced Order Models Using Reservoir Simulation And Physics Informed Machine Learning Techniques, Mark V. Behl Jr

LSU Master's Theses

Reservoir simulation is the industry standard for prediction and characterization of processes in the subsurface. However, simulation is computationally expensive and time consuming. This study explores reduced order models (ROMs) as an appropriate alternative. ROMs that use neural networks effectively capture nonlinear dependencies, and only require available operational data as inputs. Neural networks are a black box and difficult to interpret, however. Physics informed neural networks (PINNs) provide a potential solution to these shortcomings, but have not yet been applied extensively in petroleum engineering.

A mature black-oil simulation model from Volve public data release was used to generate training data …


Data Analysis To Evaluate The Performance Of Breathing Masks Used For Filtering Nano-Level Particles At Manufacturing Sites, Gracia M. Dardano 2020 Georgia Southern University

Data Analysis To Evaluate The Performance Of Breathing Masks Used For Filtering Nano-Level Particles At Manufacturing Sites, Gracia M. Dardano

Honors College Theses

The work performed in this research aims to evaluate the performance of commercially available breathing masks in filtering airborne nanoparticles at manufacturing sites. Nanoparticles are found virtually anywhere, from dust in a worksite to a simple sneeze. Therefore, they pose a substantial threat to human health as their velocity and volatility are high. This research analyzes if current efforts of breathing masks to hinder nanoparticles are effective, especially at manufacturing sites. Data has been collected in order to analyze the behavior of nanoparticles and to measure nanoparticle levels at manufacturing sites and its working environment. Data is statistical in nature …


Application Of Tda Mapper To Water Data And Bird Data, Wako Bungula 2020 University of Wisconsin - La Crosse

Application Of Tda Mapper To Water Data And Bird Data, Wako Bungula

Annual Symposium on Biomathematics and Ecology Education and Research

No abstract provided.


A Study Of Sentiment Of Covid-19 Related Tweets In The Usa, Jack Luu, Rosangela Follmann 2020 Illinois State University

A Study Of Sentiment Of Covid-19 Related Tweets In The Usa, Jack Luu, Rosangela Follmann

Annual Symposium on Biomathematics and Ecology Education and Research

No abstract provided.


Stochastic Modeling Of Ovarian Follicle Growth In Adult Female Rats, Zhaozhi Li 2020 Illinois State University

Stochastic Modeling Of Ovarian Follicle Growth In Adult Female Rats, Zhaozhi Li

Annual Symposium on Biomathematics and Ecology Education and Research

No abstract provided.


Viral Data, Agnieszka Leszczynski, Matthew Zook 2020 Western University, Canada

Viral Data, Agnieszka Leszczynski, Matthew Zook

Geography Faculty Publications

We are experiencing a historical moment characterized by unprecedented conditions of virality: a viral pandemic, the viral diffusion of misinformation and conspiracy theories, the viral momentum of ongoing Hong Kong protests, and the viral spread of #BlackLivesMatter demonstrations and related efforts to defund policing. These co-articulations of crises, traumas, and virality both implicate and are implicated by big data practices occurring in a present that is pervasively mediated by data materialities, deeply rooted dataist ideologies that entrench processes of datafication as granting objective access to truth and attendant practices of tracking, data analytics, algorithmic prediction, and data-driven targeting of individuals …


Ensemble Labeling Towards Scientific Information Extraction (Elsie), Erin Murphy 2020 DePaul University

Ensemble Labeling Towards Scientific Information Extraction (Elsie), Erin Murphy

College of Computing and Digital Media Dissertations

Extracting scientific facts from unstructured text is difficult due to challenges specific to the ambiguity of the language, the complexity of the scientific named entities and relations to be extracted. This problem is well illustrated through the extraction of polymer names and their properties. Even in the cases where the property is a temperature, identifying the polymer name associated with the temperature may require expertise due to the use of acronyms, synonyms, complicated naming conventions and by the fact that new polymer names are being “introduced” to the vernacular as polymer science advances. While there exist domain-specific machine learning toolkits …


An Analysis Of Technological Components In Relation To Privacy In A Smart City, Kayla Rutherford, Ben Lands, A. J. Stiles 2020 James Madison University

An Analysis Of Technological Components In Relation To Privacy In A Smart City, Kayla Rutherford, Ben Lands, A. J. Stiles

James Madison Undergraduate Research Journal (JMURJ)

A smart city is an interconnection of technological components that store, process, and wirelessly transmit information to enhance the efficiency of applications and the individuals who use those applications. Over the course of the 21st century, it is expected that an overwhelming majority of the world’s population will live in urban areas and that the number of wireless devices will increase. The resulting increase in wireless data transmission means that the privacy of data will be increasingly at risk. This paper uses a holistic problem-solving approach to evaluate the security challenges posed by the technological components that make up a …


Cash Flow Forecasting Using Probabilistic Neural Networks, Marwan Ashour 2020 University of Baghdad - Iraq

Cash Flow Forecasting Using Probabilistic Neural Networks, Marwan Ashour

Journal of the Arab American University مجلة الجامعة العربية الامريكية للبحوث

This paper aimed to compare the modern methods of cash flow forecasting with the traditional ones. In other words, the researcher compared between the Probabilistic Neural Networks and Transfer Function. It is worth mentioning that cash flow forecasting , nowadays, is very important and helps the upper management plan, control, assess the performance and make decisions. More specifically, in this paper, the Artificial Neural networks were used to diagnose the nature of the cash flow for the next period of time and then forecast the cash flow. The experiment was conducted in The General company for Electricity Distribution in Baghdad. …


Lifespan Analysis Of Earth Satellites, Venkata Jaipal Reddy Batthula 2020 University of Arkansas Little Rock

Lifespan Analysis Of Earth Satellites, Venkata Jaipal Reddy Batthula

Theses and Dissertations

Different countries have their own satellites for their various needs like communication, weather forecast, and security. The first satellite was launched in 1957 into space. Thousands of satellite lifetimes have already ended but they are still in orbit. The present world has more advanced technology when compared with previous technology. So, the technology for satellites is improving compared with the past. We need to understand trends in improvements to satellites related to lifespans better, using a new dataset that has not been available before, as well as datasets that we have worked with before, and that is the purpose of …


Lis Online Graduate Certificate In Data Science, Joanna Burkhardt 2020 University of Rhode Island

Lis Online Graduate Certificate In Data Science, Joanna Burkhardt

Library Impact Statements

No abstract provided.


Efficient And Fair Data Valuation For Horizontal Federated Learning, Shuyue WEI, Yongxin TONG, Zimu ZHOU, Tianshu SONG 2020 Singapore Management University

Efficient And Fair Data Valuation For Horizontal Federated Learning, Shuyue Wei, Yongxin Tong, Zimu Zhou, Tianshu Song

Research Collection School Of Computing and Information Systems

Availability of big data is crucial for modern machine learning applications and services. Federated learning is an emerging paradigm to unite different data owners for machine learning on massive data sets without worrying about data privacy. Yet data owners may still be reluctant to contribute unless their data sets are fairly valuated and paid. In this work, we adapt Shapley value, a widely used data valuation metric to valuating data providers in federated learning. Prior data valuation schemes for machine learning incur high computation cost because they require training of extra models on all data set combinations. For efficient data …


Using Data Analytics To Predict Students Score, Nang Laik MA, Gim Hong CHUA 2020 Singapore University of Social Sciences

Using Data Analytics To Predict Students Score, Nang Laik Ma, Gim Hong Chua

Research Collection School Of Computing and Information Systems

Education is very important to Singapore, and the government has continued to invest heavily in our education system to become one of the world-class systems today. A strong foundation of Science, Technology, Engineering, and Mathematics (STEM) was what underpinned Singapore's development over the past 50 years. PISA is a triennial international survey that evaluates education systems worldwide by testing the skills and knowledge of 15-year-old students who are nearing the end of compulsory education. In this paper, the authors used the PISA data from 2012 and 2015 and developed machine learning techniques to predictive the students' scores and understand the …


A New Efficient Method To Detect Genetic Interactions For Lung Cancer Gwas, Jennifer Luyapan, Xuemei Ji, Siting Li, Xiangjun Xiao, Dakai Zhu, Eric J. Duell, David C. Christiani, Matthew B. Schabath, Susanne M. Arnold, Shanbeh Zienolddiny, Hans Brunnström, Olle Melander, Mark D. Thornquist, Todd A. MacKenzie, Christopher I. Amos, Jiang Gui 2020 Dartmouth College

A New Efficient Method To Detect Genetic Interactions For Lung Cancer Gwas, Jennifer Luyapan, Xuemei Ji, Siting Li, Xiangjun Xiao, Dakai Zhu, Eric J. Duell, David C. Christiani, Matthew B. Schabath, Susanne M. Arnold, Shanbeh Zienolddiny, Hans Brunnström, Olle Melander, Mark D. Thornquist, Todd A. Mackenzie, Christopher I. Amos, Jiang Gui

Markey Cancer Center Faculty Publications

BACKGROUND: Genome-wide association studies (GWAS) have proven successful in predicting genetic risk of disease using single-locus models; however, identifying single nucleotide polymorphism (SNP) interactions at the genome-wide scale is limited due to computational and statistical challenges. We addressed the computational burden encountered when detecting SNP interactions for survival analysis, such as age of disease-onset. To confront this problem, we developed a novel algorithm, called the Efficient Survival Multifactor Dimensionality Reduction (ES-MDR) method, which used Martingale Residuals as the outcome parameter to estimate survival outcomes, and implemented the Quantitative Multifactor Dimensionality Reduction method to identify significant interactions associated with age of …


Comparing Variable Importance In Prediction Of Silence Behaviours Between Random Forest And Conditional Inference Forest Models., Stephen Barrett Dr, Geraldine Gray Dr, Colm McGuinness Dr, Michael Knoll Dr. 2020 Technological University Dublin

Comparing Variable Importance In Prediction Of Silence Behaviours Between Random Forest And Conditional Inference Forest Models., Stephen Barrett Dr, Geraldine Gray Dr, Colm Mcguinness Dr, Michael Knoll Dr.

Articles

This paper explores variable importance metrics of Conditional Inference Trees (CIT) and classical Classification And Regression Trees (CART) based Random Forests. The paper compares both algorithms variable importance rankings and highlights why CIT should be used when dealing with data with different levels of aggregation. The models analysed explored the role of cultural factors at individual and societal level when predicting Organisational Silence behaviours.


Towards High Performance Stock Market Prediction Methods, Warren M. Landis, Sangwhan Cha 2020 Harrisburg University of Science and Technology

Towards High Performance Stock Market Prediction Methods, Warren M. Landis, Sangwhan Cha

Other Student Works

Stock markets of today, and will continue to in the future, rely on the metrics of timeliness and efficiency to reach optimal profits. A way stock investors have continued to strive for the best of these two factors of the business is through the use of predictive machine learning systems to help aid in their decision making. However, among the many systems currently in use, it could be said that the myriad of data that they are based on may not be sufficient. In an effort to devise an ensemble learning predictive system that will utilize an array of big …


Tapping Twitter Data For Analyzing And Visualizing Public Sentiments On Censorship, Naveen Kumar Yadav, Akhilesh K.S. Yadav 2020 CFEES, Defence Research Development Organisation, New Delhi, India.

Tapping Twitter Data For Analyzing And Visualizing Public Sentiments On Censorship, Naveen Kumar Yadav, Akhilesh K.S. Yadav

Library Philosophy and Practice (e-journal)

The main objective of this research study is to analyse and visualize Twitter data with tags “#Censorship”. A connection was established with twitter using Twitter API, and receiving the tweets on Google Spreadsheets. Data visualization was performed using various tools such as Voyant Tools, Tableau, Google Spreadsheet and Orange in order to generate different visualizations based upon, language, geographical areas, retweets etc. The sentiment analysis was performed for the sentiments that were attached to the given set of data by the public in their respective tweets. The 23680 tweets were retrieved during the data collection time and there were 13,771 …


Digital Commons powered by bepress