Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

3,233 Full-Text Articles 9,308 Authors 1,316,836 Downloads 221 Institutions

All Articles in Data Science

Faceted Search

3,233 full-text articles. Page 139 of 155.

How Port Logistics Competitiveness Evolves Among Major Ports In China And Europe (1998-2018), Jiawei Wang 2020 World Maritime University

How Port Logistics Competitiveness Evolves Among Major Ports In China And Europe (1998-2018), Jiawei Wang

World Maritime University Dissertations

No abstract provided.


Data Visualization And Infographics Design Art404g/Dsp Xxx, Harrison Dekker 2020 University of Rhode Island

Data Visualization And Infographics Design Art404g/Dsp Xxx, Harrison Dekker

Library Impact Statements

No abstract provided.


Multi‑View Clustering For Multi‑Omics Data Using Unifed Embedding, Mohammed Hasanuzzaman, Sayantan Mitra, Sriparna Saha 2020 ADAPT Centre, Cork Institute of Technology, Cork, Ireland.

Multi‑View Clustering For Multi‑Omics Data Using Unifed Embedding, Mohammed Hasanuzzaman, Sayantan Mitra, Sriparna Saha

Articles

In real world applications, data sets are often comprised of multiple views, which provide consensus and complementary information to each other. Embedding learning is an effective strategy for nearest neighbour search and dimensionality reduction in large data sets. This paper attempts to learn a unified probability distribution of the points across different views and generates a unified embedding in a low-dimensional space to optimally preserve neighbourhood identity. Probability distributions generated for each point for each view are combined by conflation method to create a single unified distribution. The goal is to approximate this unified distribution as much as possible when …


Multi‑View Clustering For Multi‑Omics Data Using Unifed Embedding, Mohammed Hasanuzzaman, Sayantan Mitra, Sriparna Saha 2020 ADAPT Centre, Cork Institute of Technology, Cork, Ireland.

Multi‑View Clustering For Multi‑Omics Data Using Unifed Embedding, Mohammed Hasanuzzaman, Sayantan Mitra, Sriparna Saha

Department of Computer Science Publications

In real world applications, data sets are often comprised of multiple views, which provide consensus and complementary information to each other. Embedding learning is an effective strategy for nearest neighbour search and dimensionality reduction in large data sets. This paper attempts to learn a unified probability distribution of the points across different views and generates a unified embedding in a low-dimensional space to optimally preserve neighbourhood identity. Probability distributions generated for each point for each view are combined by conflation method to create a single unified distribution. The goal is to approximate this unified distribution as much as possible when …


Big Data Analytics Applied To Healthcare, Xuejuan Zhang, Boris Vishnevsky 2020 Thomas Jefferson University

Big Data Analytics Applied To Healthcare, Xuejuan Zhang, Boris Vishnevsky

School of Continuing and Professional Studies Student Papers

In this paper, we review the recent literature related to Big Data Analytics (BDA). We also discuss ways of applying BDA in Healthcare. In Section 1, we discuss the definition of Big Data Analytics and its characteristics. In Section 2, we discuss the healthcare ecosystem's main stakeholders and the data of each main stakeholder. Section 3 discusses the challenges and opportunities of leveraging Big Data Analytics by healthcare stakeholders.


Cell Assembly Detection In Low Firing-Rate Spike Train Data, Phan Minh Duc Truong 2020 Southern Methodist University

Cell Assembly Detection In Low Firing-Rate Spike Train Data, Phan Minh Duc Truong

Mathematics Theses and Dissertations

Cell assemblies, defined as groups of neurons forming temporal spike coordination, are thought to be fundamental units supporting major cognitive functions. However, detecting cell assemblies is challenging since they can occur at a range of time scales and with a range of precisions, from synchronous spikes to co-variations in firing rate. In this dissertation, we use a recently published cell assembly detection (CAD) algorithm that is capable of detecting assemblies at a range of time scales and precisions. We first showed that the CAD method can be applied to sparser spike train data than what have previously been reported. This …


Iba Newsletter [August 2020], Communications Department, Office of the Registrar 2020 Institute of Business Administration

Iba Newsletter [August 2020], Communications Department, Office Of The Registrar

IBA News

No abstract provided.


A Study Of Information Bots And Knowledge Bots, Amartya Hatua 2020 University of Southern Mississippi

A Study Of Information Bots And Knowledge Bots, Amartya Hatua

Dissertations

In this dissertation, a study of different aspects of information bots and knowledge bots is done. The research contributes to a better understanding of the various characteristics of information bots as well as the different patterns and factors responsible for the information diffusion in a social network. This research also shows how these factors can be used to predict information diffusion for a particular topic in a social network. The second part of the research is focused on strategies for improving the knowledge base of knowledge bots, where two different approaches are studied. In the first approach, knowledge is transferred …


Variable Compact Multi-Point Upscaling Schemes For Anisotropic Diffusion Problems In Three-Dimensions, James Quinlan 2020 The University of Southern Mississippi

Variable Compact Multi-Point Upscaling Schemes For Anisotropic Diffusion Problems In Three-Dimensions, James Quinlan

Dissertations

Simulation is a useful tool to mitigate risk and uncertainty in subsurface flow models that contain geometrically complex features and in which the permeability field is highly heterogeneous. However, due to the level of detail in the underlying geocellular description, an upscaling procedure is needed to generate a coarsened model that is computationally feasible to perform simulations. These procedures require additional attention when coefficients in the system exhibit full-tensor anisotropy due to heterogeneity or not aligned with the computational grid. In this thesis, we generalize a multi-point finite volume scheme in several ways and benchmark it against the industry-standard routines. …


Empirical Studies Of Deep Learning On Information Diffusion On Social Networks And Collective Task Learning For Swarm Robotics, Trung T. Nguyen 2020 University of Southern Mississippi

Empirical Studies Of Deep Learning On Information Diffusion On Social Networks And Collective Task Learning For Swarm Robotics, Trung T. Nguyen

Dissertations

Researchers in multiple disciplines have recently adopted deep learning because of its ability of high accuracy representation learning from big and complex data. My research goal in this thesis is developing deep learning models for information diffusion analysis on social networks and collective tasks learning in swarm robotics. Firstly, the information diffusion on social networks is modeled as a multivariate time series in three dimensions with ten features. Then, we applied time-series clustering algorithms with Dynamic Time Warping to discover different patterns of our models. Then, we build a prediction model based on LSTM, which outperforms traditional time-series prediction methods. …


Machine Learning Approaches For Improving Prediction Performance Of Structure-Activity Relationship Models, Gabriel Idakwo 2020 The University of Southern Mississippi

Machine Learning Approaches For Improving Prediction Performance Of Structure-Activity Relationship Models, Gabriel Idakwo

Dissertations

In silico bioactivity prediction studies are designed to complement in vivo and in vitro efforts to assess the activity and properties of small molecules. In silico methods such as Quantitative Structure-Activity/Property Relationship (QSAR) are used to correlate the structure of a molecule to its biological property in drug design and toxicological studies. In this body of work, I started with two in-depth reviews into the application of machine learning based approaches and feature reduction methods to QSAR, and then investigated solutions to three common challenges faced in machine learning based QSAR studies.

First, to improve the prediction accuracy of learning …


A Novel Correction For The Adjusted Box-Pierce Test — New Risk Factors For Emergency Department Return Visits Within 72 Hours For Children With Respiratory Conditions — General Pediatric Model For Understanding And Predicting Prolonged Length Of Stay, Sidy Danioko 2020 Chapman University

A Novel Correction For The Adjusted Box-Pierce Test — New Risk Factors For Emergency Department Return Visits Within 72 Hours For Children With Respiratory Conditions — General Pediatric Model For Understanding And Predicting Prolonged Length Of Stay, Sidy Danioko

Computational and Data Sciences (PhD) Dissertations

This thesis represents the results of three research projects that underline the breadth and depth of my interests.

Firstly, I devoted some efforts to the well-known Box-Pierce goodness-of-fit tests for time series models which has been an important research topic over the last few decades. All previously proposed tests are focused on changes of the test statistics. Instead, I adopted a different approach that takes the best performing test and modifying the rejection region. Thus, I developed a semiparametric correction of the Adjusted Box-Pierce test that attains the best I error rates for all sample sizes and lags and outperforms …


Shotgun Metagenomic Sequencing Data Of Sunflower Rhizosphere Microbial Community In South Africa, Olubukola Oluranti Babalola, Temitayo Tosin Alawiye, Carlos M. Rodriguez Lopez, Ayansina Segun Ayangbenro 2020 North-West University, South Africa

Shotgun Metagenomic Sequencing Data Of Sunflower Rhizosphere Microbial Community In South Africa, Olubukola Oluranti Babalola, Temitayo Tosin Alawiye, Carlos M. Rodriguez Lopez, Ayansina Segun Ayangbenro

Horticulture Faculty Publications

This dataset presents shotgun metagenomic sequencing of sunflower rhizosphere microbiome in Bloemhof, South Africa. Data were collected to decipher the structure and function in the sunflower microbial community. Illumina HiSeq platform using next generation sequencing of the DNA was carried out. The metagenome comprised 8,991,566 sequences totaling 1,607,022,279 bp size and 66% GC content. The metagenome was deposited into the NCBI database and can be accessed with the SRA accession number SRR10418054. An online metagenome server (MG RAST) using the subsystem database revealed bacteria had the highest taxonomical representation with 98.47%, eukaryote at 1.23%, and archaea at 0.20%. The most …


A Unified Framework For Sparse Online Learning, Peilin ZHAO, Dayong WONG, Pengcheng WU, Steven C. H. HOI 2020 Tencent AL Lab

A Unified Framework For Sparse Online Learning, Peilin Zhao, Dayong Wong, Pengcheng Wu, Steven C. H. Hoi

Research Collection School Of Computing and Information Systems

The amount of data in our society has been exploding in the era of big data. This article aims to address several open challenges in big data stream classification. Many existing studies in data mining literature follow the batch learning setting, which suffers from low efficiency and poor scalability. To tackle these challenges, we investigate a unified online learning framework for the big data stream classification task. Different from the existing online data stream classification techniques, we propose a unified Sparse Online Classification (SOC) framework. Based on SOC, we derive a second-order online learning algorithm and a cost-sensitive sparse online …


Maia And Admonita: Mandatory Integrity Control Language And Dynamic Trust Framework For Arbitrary Structured Data, Wassnaa Al-Mawee 2020 Western Michigan University

Maia And Admonita: Mandatory Integrity Control Language And Dynamic Trust Framework For Arbitrary Structured Data, Wassnaa Al-Mawee

Dissertations

The expansion of attacks against information systems of companies that operate nuclear power stations and other energy facilities in the United States and other countries, are noticeable with potential catastrophic real-world implications. Data integrity is a fundamental component of information security. It refers to the accuracy and the trustworthiness of data or resources. Data integrity within information systems becomes an important factor of security protection as the data becomes more integrated and crucial to decision-making. The security threats brought by human errors whether, malicious or unintentional, such as viruses, hacking, and many other cybersecurity threats, are dangerous and require mandatory …


Gaining Computational Insight Into Psychological Data: Applications Of Machine Learning With Eating Disorders And Autism Spectrum Disorder, Natalia Rosenfield 2020 Chapman University

Gaining Computational Insight Into Psychological Data: Applications Of Machine Learning With Eating Disorders And Autism Spectrum Disorder, Natalia Rosenfield

Computational and Data Sciences (PhD) Dissertations

Over the past 100 years, assessment tools have been developed that allow us to explore mental and behavioral processes that could not be measured before. However, conventional statistical models used for psychological data are lacking in thoroughness and predictability. This provides a perfect opportunity to use machine learning to study the data in a novel way. In this paper, we present examples of using machine learning techniques with data in three areas: eating disorders, body satisfaction, and Autism Spectrum Disorder (ASD). We explore clustering algorithms as well as virtual reality (VR).

Our first study employs the k-means clustering algorithm to …


Human Trafficking In Nepal: Can Big Data Help?, Shushant Khanal 2020 University of Nebraska at Kearney

Human Trafficking In Nepal: Can Big Data Help?, Shushant Khanal

Undergraduate Research Journal

This paper provides an overview of human trafficking in Nepal, identifies strategies implemented by the government of the country to handle the problem and possibilities of using big data as a solution to the problem of human trafficking in Nepal. Big data, may be defined as the collection of a large volume of data from the past that is processed using machine learning and artificial intelligence to find a common pattern. The use of big data in tackling the problem of human trafficking is not new in developed countries like the United States but it is still a foreign idea …


Applications Of Artificial Intelligence And Graphy Theory To Cyberbullying, Jesse D. Simpson 2020 Missouri State University

Applications Of Artificial Intelligence And Graphy Theory To Cyberbullying, Jesse D. Simpson

Graduate Theses/Dissertations

Cyberbullying is an ongoing and devastating issue in today's online social media. Abusive users engage in cyber-harassment by utilizing social media to send posts, private messages, tweets, or pictures to innocent social media users. Detecting and preventing cases of cyberbullying is crucial. In this work, I analyze multiple machine learning, deep learning, and graph analysis algorithms and explore their applicability and performance in pursuit of a robust system for detecting cyberbullying. First, I evaluate the performance of the machine learning algorithms Support Vector Machine, Naïve Bayes, Random Forest, Decision Tree, and Logistic Regression. This yielded positive results and obtained upwards …


Colleague To Banner Migration: Data Conversion Guide For Institutional Research, Laura Osborn 2020 Dakota State University

Colleague To Banner Migration: Data Conversion Guide For Institutional Research, Laura Osborn

Masters Theses & Doctoral Dissertations

When the SDBOR decided to migrate their current student information system into a shared system with HR and Finance, adjustments needed to be made to accommodate for current Banner settings and work around tables that were already populated with HRFIS data. The change in data type of the student identifier from that of a 7-digit numeric field to a 9- digit alpha-numeric field poses problems for running aggregate data calculations. Additional complications include having some information such as first-generation status that was not migrated between the systems, and cases such as college coding where tables that were designed for student …


Understanding Spatial Language In Radiology: Representation Framework, Annotation, And Spatial Relation Extraction From Chest X-Ray Reports Using Deep Learning, Surabhi Datta, Yuqi Si, Laritza Rodriguez, Sonya E Shooshan, Dina Demner-Fushman, Kirk Roberts 2020 University of Texas Health Science Center at Houston, School of Health Information Sciences, Houston TX, USA

Understanding Spatial Language In Radiology: Representation Framework, Annotation, And Spatial Relation Extraction From Chest X-Ray Reports Using Deep Learning, Surabhi Datta, Yuqi Si, Laritza Rodriguez, Sonya E Shooshan, Dina Demner-Fushman, Kirk Roberts

Faculty, Staff and Student Publications

Radiology reports contain a radiologist's interpretations of images, and these images frequently describe spatial relations. Important radiographic findings are mostly described in reference to an anatomical location through spatial prepositions. Such spatial relationships are also linked to various differential diagnoses and often described through uncertainty phrases. Structured representation of this clinically significant spatial information has the potential to be used in a variety of downstream clinical informatics applications. Our focus is to extract these spatial representations from the reports. For this, we first define a representation framework based on the Spatial Role Labeling (SpRL) scheme, which we refer to as …


Digital Commons powered by bepress