Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 901 - 930 of 1156

Full-Text Articles in Data Science

Explainable Recommendation With Comparative Constraints On Product Aspects, Trung-Hoang Le, Hady W. Lauw Mar 2021

Explainable Recommendation With Comparative Constraints On Product Aspects, Trung-Hoang Le, Hady W. Lauw

Research Collection School Of Computing and Information Systems

To aid users in choice-making, explainable recommendation models seek to provide not only accurate recommendations but also accompanying explanations that help to make sense of those recommendations. Most of the previous approaches rely on evaluative explanations, assessing the quality of an individual item along some aspects of interest to the user. In this work, we are interested in comparative explanations, the less studied problem of assessing a recommended item in comparison to another reference item.

In particular, we propose to anchor reference items on the previously adopted items in a user's history. Not only do we aim at providing comparative …


Unsupervised Data Mining Technique For Clustering Library In Indonesia, Robbi Rahim, Joseph Teguh Santoso, Sri Jumini, Gita Widi Bhawika, Daniel Susilo, Danny Wibowo Feb 2021

Unsupervised Data Mining Technique For Clustering Library In Indonesia, Robbi Rahim, Joseph Teguh Santoso, Sri Jumini, Gita Widi Bhawika, Daniel Susilo, Danny Wibowo

Library Philosophy and Practice (e-journal)

Organizing school libraries not only keeps library materials, but helps students and teachers in completing tasks in the teaching process so that national development goals are in order to improve community welfare by producing quality and competitive human resources. The purpose of this study is to analyze the Unsupervised Learning technique in conducting cluster mapping of the number of libraries at education levels in Indonesia. The data source was obtained from the Ministry of Education and Culture which was processed by the Central Statistics Agency (abbreviated as BPS) with url: bps.go.id/. The data consisted of 34 records where the attribute …


A New Feature Selection Method Based On Class Association Rule, Sami A. Al-Dhaheri Feb 2021

A New Feature Selection Method Based On Class Association Rule, Sami A. Al-Dhaheri

Dissertations, Theses, and Capstone Projects

Feature selection is a key process for supervised learning algorithms. It involves discarding irrelevant attributes from the training dataset from which the models are derived. One of the vital feature selection approaches is Filtering, which often uses mathematical models to compute the relevance for each feature in the training dataset and then sorts the features into descending order based on their computed scores. However, most Filtering methods face several challenges including, but not limited to, merely considering feature-class correlation when defining a feature’s relevance; additionally, not recommending which subset of features to retain. Leaving this decision to the end-user may …


Exploring Media Portrayals Of People With Mental Disorders Using Nlp, Swapna Gottipati, Mark Chong, Andrew Wei Kiat Lim, Benny Haryanto Kawidiredjo Feb 2021

Exploring Media Portrayals Of People With Mental Disorders Using Nlp, Swapna Gottipati, Mark Chong, Andrew Wei Kiat Lim, Benny Haryanto Kawidiredjo

Research Collection School Of Computing and Information Systems

Media plays an important role in creating an impact in society. Several studies show that news media and entertainment channels, at times may create overwhelming images of the mental illness that emphasize criminality and dangerousness. The consequences of such negative impact may impact the audience with stigma and on the other hand, they impair the self-esteem and help-seeking behavior of the people with mental disorders. This is the first study to examine the Singapore media’s portrayal of persons with mental disorders (MDs) using text analytics and natural language processing. To date, most studies on media portrayal of people with MDs …


Deep Learning For Multi-Tissue Cancer Classification Of Gene Expressions, Tarek Khorshed Jan 2021

Deep Learning For Multi-Tissue Cancer Classification Of Gene Expressions, Tarek Khorshed

Theses and Dissertations

We contribute in saving the lives of cancer patients through early detection and diagnosis, since one of the major challenges in cancer treatment is that patients are diagnosed at very late stages when appropriate medical interventions become less effective and full curative treatment is no longer achievable. Cancer classification using gene expressions is extremely challenging given the complexity and high dimensionality of the data. Current classification methods typically rely on samples collected from a single tissue type and perform a prerequisite of gene feature selection to avoid processing the full set of genes. These methods fall short in taking advantage …


Delineating Knowledge Domains In Scientific Domains In Scientific Literature Using Machine Learning (Ml), Abhay Maurya, Smarajit Paul Choudhury Mr., Kshitij Jaiswal Mr. Jan 2021

Delineating Knowledge Domains In Scientific Domains In Scientific Literature Using Machine Learning (Ml), Abhay Maurya, Smarajit Paul Choudhury Mr., Kshitij Jaiswal Mr.

Library Philosophy and Practice (e-journal)

The recent years have witnessed an upsurge in the number of published documents. Organizations are showing an increased interest in text classification for effective use of the information. Manual procedures for text classification can be fruitful for a handful of documents, but the same lack in credibility when the number of documents increases besides being laborious and time-consuming. Text mining techniques facilitate assigning text strings to categories rendering the process of classification fast, accurate, and hence reliable. This paper classifies chemistry documents using machine learning and statistical methods. The procedure of text classification has been described in chronological order like …


Multi-Modal Classification Using Images And Text, Stuart J. Miller, Justin Howard, Paul Adams, Mel Schwan, Robert Slater Jan 2021

Multi-Modal Classification Using Images And Text, Stuart J. Miller, Justin Howard, Paul Adams, Mel Schwan, Robert Slater

SMU Data Science Review

This paper proposes a method for the integration of natural language understanding in image classification to improve classification accuracy by making use of associated metadata. Traditionally, only image features have been used in the classification process; however, metadata accompanies images from many sources. This study implemented a multi-modal image classification model that combines convolutional methods with natural language understanding of descriptions, titles, and tags to improve image classification. The novelty of this approach was to learn from additional external features associated with the images using natural language understanding with transfer learning. It was found that the combination of ResNet-50 image …


Sars-Cov-2 Pandemic Analytical Overview With Machine Learning Predictability, Anthony Tanaydin, Jingchen Liang, Daniel W. Engels Jan 2021

Sars-Cov-2 Pandemic Analytical Overview With Machine Learning Predictability, Anthony Tanaydin, Jingchen Liang, Daniel W. Engels

SMU Data Science Review

Understanding diagnostic tests and examining important features of novel coronavirus (COVID-19) infection are essential steps for controlling the current pandemic of 2020. In this paper, we study the relationship between clinical diagnosis and analytical features of patient blood panels from the US, Mexico, and Brazil. Our analysis confirms that among adults, the risk of severe illness from COVID-19 increases with pre-existing conditions such as diabetes and immunosuppression. Although more than eight months into pandemic, more data have become available to indicate that more young adults were getting infected. In addition, we expand on the definition of COVID-19 test and discuss …


Challenges When Identifying Migration From Geo-Located Twitter Data, Caitrin Armstrong, Ate Poorthuis, Matthew Zook, Derek Ruths, Thomas Soehl Jan 2021

Challenges When Identifying Migration From Geo-Located Twitter Data, Caitrin Armstrong, Ate Poorthuis, Matthew Zook, Derek Ruths, Thomas Soehl

Geography Faculty Publications

Given the challenges in collecting up-to-date, comparable data on migrant populations the potential of digital trace data to study migration and migrants has sparked considerable interest among researchers and policy makers. In this paper we assess the reliability of one such data source that is heavily used within the research community: geolocated tweets. We assess strategies used in previous work to identify migrants based on their geolocation histories. We apply these approaches to infer the travel history of a set of Twitter users who regularly posted geolocated tweets between July 2012 and June 2015. In a second step we hand-code …


Adverse Health Effects Of Kratom: An Analysis Of Social Media Data, Abdullah Wahbeh, Tareq Nasralah, Omar El-Gayar, Mohammad A. Al-Ramahi, Ahmed El Noshokaty Jan 2021

Adverse Health Effects Of Kratom: An Analysis Of Social Media Data, Abdullah Wahbeh, Tareq Nasralah, Omar El-Gayar, Mohammad A. Al-Ramahi, Ahmed El Noshokaty

Computer Information Systems Faculty Publications (Archived)

This study investigates the adverse healthcare effects associated with the use of kratom. Using machine learning techniques, we analyzed a total of 36,516 users’ posts related to kratom. The results and analysis showed that social media could help identify important insights related to the use of kratom. The sentiment and emotion analyses showed that the kratom experience was negative and largely associated with anger, fear, disgust, and sadness. The results from. topic modeling showed that kratom is associated with a number of healthcare issues such as rashes and itching, urination, constipation, loss of appetite/weight, dry mouth, seizures, nausea, heartburn, dehydration, …


Applications Of Machine Learning To Facilitate Software Engineering And Scientific Computing, Natalie Best Jan 2021

Applications Of Machine Learning To Facilitate Software Engineering And Scientific Computing, Natalie Best

Computational and Data Sciences (PhD) Dissertations

The use of machine learning has risen in recent years, though many areas remain unexplored due to lack of data or lack of computational tools. This dissertation explores machine learning approaches in case studies involving image classification and natural language processing. In addition, a software library in the form of two-way bridge connecting deep learning models in Keras with ones available in the Fortran programming language is also presented.

In Chapter 2, we explore the applicability of transfer learning utilizing models pre-trained on non-software engineering data applied to the problem of classifying software unified modeling language diagrams where data is …


Proposed Data Governance Framework For Small And Medium Scale Enterprises (Smes), Rejoice Okoro Jan 2021

Proposed Data Governance Framework For Small And Medium Scale Enterprises (Smes), Rejoice Okoro

All Graduate Theses, Dissertations, and Other Capstone Projects

Data governance is not a one size fits all, instead, it should be an evolutionary process that can be started small and measurable along the way. This research aims at proposing a data governance framework by ensuring data management processes, data security and control are compliant with laws and policies. This article also presents the first results of a comparative analysis between three data privacy laws and outlines five components which together form a data governance framework for SMEs. The data governance model documents data quality roles and their type of interaction with data quality management activities exploring how data …


Semantics Of The Black-Box: Can Knowledge Graphs Help Make Deep Learning Systems More Interpretable And Explainable?, Manas Gaur, Keyur Faldu, Amit Sheth Jan 2021

Semantics Of The Black-Box: Can Knowledge Graphs Help Make Deep Learning Systems More Interpretable And Explainable?, Manas Gaur, Keyur Faldu, Amit Sheth

Publications

The recent series of innovations in deep learning (DL) have shown enormous potential to impact individuals and society, both positively and negatively. The DL models utilizing massive computing power and enormous datasets have significantly outperformed prior historical benchmarks on increasingly difficult, well-defined research tasks across technology domains such as computer vision, natural language processing, signal processing, and human-computer interactions. However, the Black-Box nature of DL models and their over-reliance on massive amounts of data condensed into labels and dense representations poses challenges for interpretability and explainability of the system. Furthermore, DLs have not yet been proven in their ability to …


A Hybrid Gene Selection Strategy Based On Fisher And Ant Colony Optimization Algorithm For Breast Cancer Classification, Mohammed Hamim, Ismail El Moudden, Mohan D. Pant, Hicham Moutachaouik, Mustapha Hain Jan 2021

A Hybrid Gene Selection Strategy Based On Fisher And Ant Colony Optimization Algorithm For Breast Cancer Classification, Mohammed Hamim, Ismail El Moudden, Mohan D. Pant, Hicham Moutachaouik, Mustapha Hain

EVMS School of Health Professions Faculty Publications

Breast cancer poses the greatest threat to human life and especially to women's life. Despite the progress made in data mining technology in recent years, the ability to predict and diagnose such fatal diseases based on gene expression data still reveals a limited prediction performance, which may not be surprising since most of the genes in expression data are believed to be irrelevant or redundant. The dimensionality reduction process may be considered as a crucial step to analyze gene expression data, as it can reduce the high dimensionality of the breast cancer datasets, which may result into a better prediction …


Convolutional Audio Source Separation Applied To Drum Signal Separation, Marius Orehovschi Jan 2021

Convolutional Audio Source Separation Applied To Drum Signal Separation, Marius Orehovschi

Honors Theses

This study examined the task of drum signal separation from full music mixes via both classical methods (Independent Component Analysis) and a combination of Time-Frequency Binary Masking and Convolutional Neural Networks. The results indicate that classical methods relying on predefined computations do not achieve any meaningful results, while convolutional neural networks can achieve imperfect but musically useful results. Furthermore, neural network performance can be improved by data augmentation via transposition – a technique that can only be applied in the context of drum signal separation.


Information Architecture For A Chemical Modeling Knowledge Graph, Adam R. Luxon Jan 2021

Information Architecture For A Chemical Modeling Knowledge Graph, Adam R. Luxon

Theses and Dissertations

Machine learning models for chemical property predictions are high dimension design challenges spanning multiple disciplines. Free and open-source software libraries have streamlined the model implementation process, but the design complexity remains. In order better navigate and understand the machine learning design space, model information needs to be organized and contextualized. In this work, instances of chemical property models and their associated parameters were stored in a Neo4j property graph database. Machine learning model instances were created with permutations of dataset, learning algorithm, molecular featurization, data scaling, data splitting, hyperparameters, and hyperparameter optimization techniques. The resulting graph contains over 83,000 nodes …


Reliable And Interpretable Machine Learning For Modeling Physical And Cyber Systems, Daniel L. Marino Lizarazo Jan 2021

Reliable And Interpretable Machine Learning For Modeling Physical And Cyber Systems, Daniel L. Marino Lizarazo

Theses and Dissertations

Over the past decade, Machine Learning (ML) research has predominantly focused on building extremely complex models in order to improve predictive performance. The idea was that performance can be improved by adding complexity to the models. This approach proved to be successful in creating models that can approximate highly complex relationships while taking advantage of large datasets. However, this approach led to extremely complex black-box models that lack reliability and are difficult to interpret. By lack of reliability, we specifically refer to the lack of consistent (unpredictable) behavior in situations outside the training data. Lack of interpretability refers to the …


Big Data Management In The Shipping Industry : Examining Strengths Vs Weaknesses And Highlighting Relevant Business Opportunities, Dimitrios Dalaklis, Georgios Vaitsos, Nikitas Nikitakos, Dimitrios Papachristos, Angelo Dalaklis, Esslam Hassan Jan 2021

Big Data Management In The Shipping Industry : Examining Strengths Vs Weaknesses And Highlighting Relevant Business Opportunities, Dimitrios Dalaklis, Georgios Vaitsos, Nikitas Nikitakos, Dimitrios Papachristos, Angelo Dalaklis, Esslam Hassan

Conference Papers

History testifies that there is a dialectic relationship between humans and technology. Especially during the last couple of decades, the shipping industry has benefitted from a very extended number of advanced technology innovations. Today, all systems supporting the conduct of navigation and the various information technology (IT) applications related to ship management activities are heavily reliant upon (almost) real-time information to safely/effectively fulfil their allocated tasks. As a result, truly vast quantities of data -which are often described as “Big Data” in the wider literature- are created and the issue of how to effectively manage all the associated information is …


Public Interest Technology – Exploring Covid-19 Health Data, Sarah Zelikovitz Jan 2021

Public Interest Technology – Exploring Covid-19 Health Data, Sarah Zelikovitz

Open Educational Resources

This module is part of a Introduction to Data Science course that covers the different parts of the data science process: data acquisition, cleaning, exploratory data analysis, and modeling. The COVID-19 pandemic has created much interest in public health data, as well as interest in visualization of all types of data. Public health data has a set of challenges that is unique to health data, with HIPAA laws, and real time collection of data. With COVID-19, the challenges are particularly amplified, as data collection and statistics collected are constantly changing in response to feedback from labs, hospitals, drug companies, and …


Question Answering By Bert, Suman Karanjit Jan 2021

Question Answering By Bert, Suman Karanjit

Student Academic Conference

No abstract provided.


Multi-Stream Longitudinal Data Analysis Using Deep Learning, Sajjad Fouladvand Jan 2021

Multi-Stream Longitudinal Data Analysis Using Deep Learning, Sajjad Fouladvand

Theses and Dissertations--Computer Science

Longitudinal healthcare data encompasses all tasks where patients information are collected at multiple follow-up times. Analyzing this data is critical in addressing many real world problems in healthcare such as disease prediction and prevention. In this thesis, technical challenges in analyzing longitudinal administrative claims data are addressed and novel deep learning based models are proposed for multi-stream data analysis and disease prediction tasks. These algorithms and frameworks are assessed mainly on substance use disorders prediction tasks and specifically designed to tackled these disorders. Substance use disorder is a public health crisis costing the US an estimated $740 billion annually in …


Fast And Memory-Efficient Tfidf Calculation For Text Analysis Of Large Datasets, Samah Senbel Jan 2021

Fast And Memory-Efficient Tfidf Calculation For Text Analysis Of Large Datasets, Samah Senbel

School of Computer Science & Engineering Faculty Publications

Term frequency – Inverse Document Frequency (TFIDF) is a vital first step in text analytics for information retrieval and machine learning applications. It is a memory-intensive and complex task due to the need to create and process a large sparse matrix of term frequencies, with the documents as rows and the term as columns and populate it with the term frequency of each word in each document.

The standard method of storing the sparse array is the “Compressed Sparse Row” (CSR), which stores the sparse array as three one-dimensional arrays for the row id, column id, and term frequencies. We …


Neural Representations Of Concepts And Texts For Biomedical Information Retrieval, Jiho Noh Jan 2021

Neural Representations Of Concepts And Texts For Biomedical Information Retrieval, Jiho Noh

Theses and Dissertations--Computer Science

Information retrieval (IR) methods are an indispensable tool in the current landscape of exponentially increasing textual data, especially on the Web. A typical IR task involves fetching and ranking a set of documents (from a large corpus) in terms of relevance to a user's query, which is often expressed as a short phrase. IR methods are the backbone of modern search engines where additional system-level aspects including fault tolerance, scale, user interfaces, and session maintenance are also addressed. In addition to fetching documents, modern search systems may also identify snippets within the documents that are potentially most relevant to the …


Revisiting Absolute Pose Regression, Hunter Blanton Jan 2021

Revisiting Absolute Pose Regression, Hunter Blanton

Theses and Dissertations--Computer Science

Images provide direct evidence for the position and orientation of the camera in space, known as camera pose. Traditionally, the problem of estimating the camera pose requires reference data for determining image correspondence and leveraging geometric relationships between features in the image. Recent advances in deep learning have led to a new class of methods that regress the pose directly from a single image.

This thesis proposes methods for absolute camera pose regression. Absolute pose regression estimates the pose of a camera from a single image as the output of a fixed computation pipeline. These methods have many practical benefits …


Requirements Engineering Education Slr Data Set 1988-2020, Marian Daun, Alicia M. Grubb, Bastian Tenbergen Jan 2021

Requirements Engineering Education Slr Data Set 1988-2020, Marian Daun, Alicia M. Grubb, Bastian Tenbergen

Data

Requirements Engineering (RE) has established itself as a core software engineering discipline. It is well acknowledged that good RE leads to higher quality software and considerably reduces the risk of failure or exceeding budgets of software development projects. Therefore, it is of vital importance to train future software engineers in RE and educate future requirements engineers to adequately manage requirements in various projects. However, to date there exists no central dataset for RE Education articles. To lay the foundation for this important mission, we conducted a systematic literature review. In this dataset, we present 152 articles from the Requirements Engineering …


"Who Can Help Me?'': Knowledge Infused Matching Of Support Seekers And Support Providers During Covid-19 On Reddit, Manas Gaur, Kaushik Roy, Aditya Sharma, Biplav Srivastava, Amit Sheth Jan 2021

"Who Can Help Me?'': Knowledge Infused Matching Of Support Seekers And Support Providers During Covid-19 On Reddit, Manas Gaur, Kaushik Roy, Aditya Sharma, Biplav Srivastava, Amit Sheth

Publications

During the ongoing COVID-19 crisis, subreddits on Reddit, such as r/Coronavirus saw a rapid growth in user's requests for help (support seekers - SSs) including individuals with varying professions and experiences with diverse perspectives on care (support providers - SPs). Currently, knowledgeable human moderators match an SS with a user with relevant experience, i.e, an SP on these subreddits. This unscalable process defers timely care. We present a medical knowledge-infused approach to efficient matching of SS and SPs validated by experts for the users affected by anxiety and depression, in the context of with COVID-19. After matching, each SP to …


Uncovering Object Categories In Infant Views, Naiti S. Bhatt Jan 2021

Uncovering Object Categories In Infant Views, Naiti S. Bhatt

Scripps Senior Theses

While adults recognize objects in a near-instant, infants must learn how to categorize the objects in their visual environments. Recent work has shown that egocentric head-mounted camera videos contain rich data that illuminate the infant experience (Clerkin et al., 2017; Franchak et al., 2011; Yoshida & Smith, 2008). While past work has focused on the social information in view, in this work, we aim to characterize the objects in infants’ at-home visual environments by modifying modern computer vision models for the infant view. To do so, we collected manual annotations of objects that infants seemed to be interacting within a …


Improving Space Efficiency Of Deep Neural Networks, Aliakbar Panahi Jan 2021

Improving Space Efficiency Of Deep Neural Networks, Aliakbar Panahi

Theses and Dissertations

Language models employ a very large number of trainable parameters. Despite being highly overparameterized, these networks often achieve good out-of-sample test performance on the original task and easily fine-tune to related tasks. Recent observations involving, for example, intrinsic dimension of the objective landscape and the lottery ticket hypothesis, indicate that often training actively involves only a small fraction of the parameter space. Thus, a question remains how large a parameter space needs to be in the first place — the evidence from recent work on model compression, parameter sharing, factorized representations, and knowledge distillation increasingly shows that models can be …


A Multi-Resolution Graph Convolution Network For Contiguous Epitope Prediction, Lisa Oh Jan 2021

A Multi-Resolution Graph Convolution Network For Contiguous Epitope Prediction, Lisa Oh

Dartmouth College Master’s Theses

Computational methods for predicting binding interfaces between antigens and antibodies (epitopes and paratopes) are faster and cheaper than traditional experimental structure determination methods. A sufficiently reliable computational predictor that could scale to large sets of available antibody sequence data could thus inform and expedite many biomedical pursuits, such as better understanding immune responses to vaccination and natural infection and developing better drugs and vaccines. However, current state-of-the-art predictors produce discontiguous predictions, e.g., predicting the epitope in many different spots on an antigen, even though in reality they typically comprise a single localized region. We seek to produce contiguous predicted epitopes, …


Automatic Hierarchy Expansion For Improved Structure And Chord Evaluation, Katherine M. Kinnaird, Brian Mcfee Jan 2021

Automatic Hierarchy Expansion For Improved Structure And Chord Evaluation, Katherine M. Kinnaird, Brian Mcfee

Statistical and Data Sciences: Faculty Publications

No abstract provided.