Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

3,235 Full-Text Articles 9,310 Authors 1,358,113 Downloads 221 Institutions

All Articles in Data Science

Faceted Search

3,235 full-text articles. Page 124 of 155.

An Integrated Magneto-Electrochemical Device For The Rapid Profiling Of Tumour Extracellular Vesicles From Blood Plasma, Jongmin Park, Jun Seok Park, Chen Han Huang, Ala Jo, Kaitlyn Cook, Rui Wang, Hsing Ying Lin, Jan Van Deun, Huiyan Li, Jouha Min, Lan Wang, Ghilsuk Yoon, Bob S. Carter, Leonora Balaj, Gyu Seog Choi, Cesar M. Castro, Ralph Weissleder, Hakho Lee 2021 Harvard Medical School

An Integrated Magneto-Electrochemical Device For The Rapid Profiling Of Tumour Extracellular Vesicles From Blood Plasma, Jongmin Park, Jun Seok Park, Chen Han Huang, Ala Jo, Kaitlyn Cook, Rui Wang, Hsing Ying Lin, Jan Van Deun, Huiyan Li, Jouha Min, Lan Wang, Ghilsuk Yoon, Bob S. Carter, Leonora Balaj, Gyu Seog Choi, Cesar M. Castro, Ralph Weissleder, Hakho Lee

Statistical and Data Sciences: Faculty Publications

Assays for cancer diagnosis via the analysis of biomarkers on circulating extracellular vesicles (EVs) typically have lengthy sample workups, limited throughput or insufficient sensitivity, or do not use clinically validated biomarkers. Here we report the development and performance of a 96-well assay that integrates the enrichment of EVs by antibody-coated magnetic beads and the electrochemical detection, in less than one hour of total assay time, of EV-bound proteins after enzymatic amplification. By using the assay with a combination of antibodies for clinically relevant tumour biomarkers (EGFR, EpCAM, CD24 and GPA33) of colorectal cancer (CRC), we classified plasma samples from 102 …


Privacy-Preserving Cloud-Assisted Data Analytics, Wei Bao 2021 University of Arkansas, Fayetteville

Privacy-Preserving Cloud-Assisted Data Analytics, Wei Bao

Graduate Theses and Dissertations

Nowadays industries are collecting a massive and exponentially growing amount of data that can be utilized to extract useful insights for improving various aspects of our life. Data analytics (e.g., via the use of machine learning) has been extensively applied to make important decisions in various real world applications. However, it is challenging for resource-limited clients to analyze their data in an efficient way when its scale is large. Additionally, the data resources are increasingly distributed among different owners. Nonetheless, users' data may contain private information that needs to be protected.

Cloud computing has become more and more popular in …


Variational Learning From Implicit Bandit Feedback, Quoc Tuan TRUONG, Hady W. LAUW 2021 Singapore Management University

Variational Learning From Implicit Bandit Feedback, Quoc Tuan Truong, Hady W. Lauw

Research Collection School Of Computing and Information Systems

Recommendations are prevalent in Web applications (e.g., search ranking, item recommendation, advertisement placement). Learning from bandit feedback is challenging due to the sparsity of feedback limited to system-provided actions. In this work, we focus on batch learning from logs of recommender systems involving both bandit and organic feedbacks. We develop a probabilistic framework with a likelihood function for estimating not only explicit positive observations but also implicit negative observations inferred from the data. Moreover, we introduce a latent variable model for organic-bandit feedbacks to robustly capture user preference distributions. Next, we analyze the behavior of the new likelihood under two …


Exploring Cross-Modality Utilization In Recommender Systems, Quoc Tuan TRUONG, Aghiles SALAH, Thanh-Binh TRAN, Jingyao GUO, Hady W. LAUW 2021 Singapore Management University

Exploring Cross-Modality Utilization In Recommender Systems, Quoc Tuan Truong, Aghiles Salah, Thanh-Binh Tran, Jingyao Guo, Hady W. Lauw

Research Collection School Of Computing and Information Systems

Multimodal recommender systems alleviate the sparsity of historical user-item interactions. They are commonly catalogued based on the type of auxiliary data (modality) they leverage, such as preference data plus user-network (social), user/item texts (textual), or item images (visual) respectively. One consequence of this categorization is the tendency for virtual walls to arise between modalities. For instance, a study involving images would compare to only baselines ostensibly designed for images. However, a closer look at existing models' statistical assumptions about any one modality would reveal that many could work just as well with other modalities. Therefore, we pursue a systematic investigation …


Pediatric Asthma – Another Negative Outcome Of Recurrent Flooding, Odu Researchers Find, News @ ODU 2021 Old Dominion University

Pediatric Asthma – Another Negative Outcome Of Recurrent Flooding, Odu Researchers Find, News @ Odu

News Items

No abstract provided.


Counting And Sampling Small Structures In Graph And Hypergraph Data Streams, Themistoklis Haris 2021 Dartmouth College

Counting And Sampling Small Structures In Graph And Hypergraph Data Streams, Themistoklis Haris

Dartmouth College Undergraduate Theses

In this thesis, we explore the problem of approximating the number of elementary substructures called simplices in large k-uniform hypergraphs. The hypergraphs are assumed to be too large to be stored in memory, so we adopt a data stream model, where the hypergraph is defined by a sequence of hyperedges.

First we propose an algorithm that (ε, δ)-estimates the number of simplices using O(m1+1/k / T) bits of space. In addition, we prove that no constant-pass streaming algorithm can (ε, δ)- approximate the number of simplices using less than O( m 1+1/k / T ) bits of space. Thus …


A Configurable Social Network For Running Irb-Approved Experiments, Mihovil Mandic 2021 Dartmouth College

A Configurable Social Network For Running Irb-Approved Experiments, Mihovil Mandic

Dartmouth College Undergraduate Theses

Our world has never been more connected, and the size of the social media landscape draws a great deal of attention from academia. However, social networks are also a growing challenge for the Institutional Review Boards concerned with the subjects’ privacy. These networks contain a monumental variety of personal information of almost 4 billion people, allow for precise social profiling, and serve as a primary news source for many users. They are perfect environments for influence operations that are becoming difficult to defend against. Motivated to study online social influence via IRB-approved experiments, we designed and implemented a flexible, scalable, …


Lexical Complexity Prediction With Assembly Models, Aadil Islam 2021 Dartmouth College

Lexical Complexity Prediction With Assembly Models, Aadil Islam

Dartmouth College Undergraduate Theses

Tuning the complexity of one's writing is essential to presenting ideas in a logical, intuitive manner to audiences. This paper describes a system submitted by team BigGreen to LCP 2021 for predicting the lexical complexity of English words in a given context. We assemble a feature engineering-based model and a deep neural network model with an underlying Transformer architecture based on BERT. While BERT itself performs competitively, our feature engineering-based model helps in extreme cases, eg. separating instances of easy and neutral difficulty. Our handcrafted features comprise a breadth of lexical, semantic, syntactic, and novel phonetic measures. Visualizations of BERT …


Fine-Grained Detection Of Hate Speech Using Bertoxic, Yakoob Khan 2021 Dartmouth College

Fine-Grained Detection Of Hate Speech Using Bertoxic, Yakoob Khan

Dartmouth College Undergraduate Theses

This thesis describes our approach towards the fine-grained detection of hate speech using deep learning. We leverage the transformer encoder architecture to propose BERToxic, a system that fine-tunes a pre-trained BERT model to locate toxic text spans in a given text and utilizes additional post-processing steps to refine the prediction boundaries. The post-processing steps involve (1) labeling character offsets between consecutive toxic tokens as toxic and (2) assigning a toxic label to words that have at least one token labeled as toxic. Through experiments, we show that these two post-processing steps improve the performance of our model by 4.16% on …


Improving Existing Methods For Calculating Embodied Carbon Emissions In Trade Through Feature Discovery: An Information Theoretic Approach, Sam Morton 2021 Dartmouth College

Improving Existing Methods For Calculating Embodied Carbon Emissions In Trade Through Feature Discovery: An Information Theoretic Approach, Sam Morton

Dartmouth College Undergraduate Theses

The continued societal and ecological risks posed by climate change have spurred renewed interest in quantitative tools that can improve policy aimed at climate mitigation. In 2008, international trade accounted for up to 26\% of global anthropogenic emissions, and therefore trade has garnered increased attention from policymakers seeking carbon mitigation. The concept of embodied carbon emissions in trade (EET) quantifies overall carbon emitted in the production and transport of goods for the purposes of trade. EET in theory could prove an indispensable tool to climate-concerned policymakers, but current implementations and data availability limit EET calculation to annual snapshots that extend …


Exploring The Long Tail, Joseph H. Hajjar 2021 Dartmouth College

Exploring The Long Tail, Joseph H. Hajjar

Dartmouth College Undergraduate Theses

The migration of datasets online has created a near-infinite inventory for big name retailers such as Amazon and Netflix, giving rise to recommendation systems to assist users in navigating the massive catalog. This has also allowed for the possibility of retailers storing much less popular, uncommon items which would not appear in a more traditional brick-and-mortar setting due to the cost of storage. Nevertheless, previous work has highlighted the profit potential which lies in the so-called "long tail'' of niche, unpopular items. Unfortunately, due to the limited amount of data in this subset of the inventory, recommendation systems often struggle …


Investigating Daily Fantasy Baseball: An Approach To Automated Lineup Generation, Ryan Smith 2021 California Polytechnic State University, San Luis Obispo

Investigating Daily Fantasy Baseball: An Approach To Automated Lineup Generation, Ryan Smith

Master's Theses

A recent trend among sports fans along both sides of the letterman jacket is that of Daily Fantasy Sports (DFS). The DFS industry has been under legal scrutiny recently, due to the view that daily sports data is too random to make its prediction skillful. Therefore, a common view is that it constitutes online gambling. This thesis proves that DFS, as it pertains to Baseball, is significantly more predictable than random chance, and thus does not constitute gambling.

We propose a system which generates daily lists of lineups for Fanduel Daily Fantasy Baseball contests. The system consists of two components: …


Exploring The Use Of Social Media To Infer Relationships Between Demographics, Psychographics And Vaccine Hesitancy, Abhimanyu Kapur 2021 Dartmouth College

Exploring The Use Of Social Media To Infer Relationships Between Demographics, Psychographics And Vaccine Hesitancy, Abhimanyu Kapur

Computer Science Senior Theses

The growing popularity of social media as a platform to obtain information and share one's opinions on various topics makes it a rich source of information for research. In this study, we aimed to develop a framework to infer relationships between demographic and psychographic characteristics of a user and their opinion on a specific narrative - in this case, their stance on taking the COVID-19 vaccine. Twitter was the chosen platform due to the large USA user base and easily available data. Demographic traits included Race, Age, Gender, and Human-vs-Organization Status. Psychographic traits included the Big Five personality traits (Conscientiousness, …


Advancing The Ability To Predict Cognitive Decline And Alzheimer’S Disease Based On Genetic Variants Beyond Amyloid-Beta And Tau, Naveen Rawat 2021 San Jose State University

Advancing The Ability To Predict Cognitive Decline And Alzheimer’S Disease Based On Genetic Variants Beyond Amyloid-Beta And Tau, Naveen Rawat

Master's Projects

A growing amount of neurodegenerative R&D is focused on identifying genomic- based explanations of AD that are beyond Amyloid-b and Tau. The proposed effort involves identifying some of the genomic variations, such as single nucleotide polymorphisms (SNPs), allele , chromosome, epigenetic contributors to MCI and AD that are beyond Aβ and Tau.

The project involves building a prediction model based on a support vector machine (SVM) classifier that takes into account the genomic variations and epigenetic factors to predict the early stage of mild cognitive impairment (MCI) and Alzheimer disease (AD). To achieve this, picking up important feature sets which …


Learn Biologically Meaningful Representation With Transfer Learning, Di He 2021 CUNY Graduate Center

Learn Biologically Meaningful Representation With Transfer Learning, Di He

Dissertations, Theses, and Capstone Projects

Machine learning has made significant contributions to bioinformatics and computational biol­ogy. In particular, supervised learning approaches have been widely used in solving problems such as bio­marker identification, drug response prediction, and so on. However, because of the limited availability of comprehensively labeled and clean data, constructing predictive models in super­ vised settings is not always desirable or possible, especially when using data­hunger, red­hot learning paradigms such as deep learning methods. Hence, there are urgent needs to develop new approaches that could leverage more readily available unlabeled data in driving successful machine learning ap­ plications in this area.

In my dissertation, …


Do Animated Line Graphs Increase Risk Inferences?, Junghan KIM, Arun LAKSHMANAN 2021 Singapore Management University

Do Animated Line Graphs Increase Risk Inferences?, Junghan Kim, Arun Lakshmanan

Research Collection Lee Kong Chian School Of Business

This article shows that animated display of time-varying data (e.g., stock or commodity prices) enhances risk judgments. We outline a process whereby animated display enhances the visual salience of transitions in a trajectory (i.e., successive changes in data values), which leads to transitions being utilized more to form cognitive inferences about risk. In turn, this leads to inflated risk judgments. The studies reported in this article provide converging evidence via eye tracking (Study 1), serial mediation analyses (Studies 2 and 3), and experimental manipulations of transition salience (graph type; Study 3) and utilization of transitions (global trend; Study 4 and …


Soarnet, Deep Learning Thermal Detection For Free Flight, Jake T. Tallman 2021 California Polytechnic State University, San Luis Obispo

Soarnet, Deep Learning Thermal Detection For Free Flight, Jake T. Tallman

Master's Theses

Thermals are regions of rising hot air formed on the ground through the warming of the surface by the sun. Thermals are commonly used by birds and glider pilots to extend flight duration, increase cross-country distance, and conserve energy. This kind of powerless flight using natural sources of lift is called soaring. Once a thermal is encountered, the pilot flies in circles to keep within the thermal, so gaining altitude before flying off to the next thermal and towards the destination. A single thermal can net a pilot thousands of feet of elevation gain, however estimating thermal locations is not …


Modeling The Spread Of Covid-19 Over Varied Contact Networks, Ryan L. Solorzano 2021 California Polytechnic State University, San Luis Obispo

Modeling The Spread Of Covid-19 Over Varied Contact Networks, Ryan L. Solorzano

Master's Theses

When attempting to mitigate the spread of an epidemic without the use of a vaccine, many measures may be made to dampen the spread of the disease such as physically distancing and wearing masks. The implementation of an effective test and quarantine strategy on a population has the potential to make a large impact on the spread of the disease as well. Testing and quarantining strategies become difficult when a portion of the population are asymptomatic spreaders of the disease. Additionally, a study has shown that randomly testing a portion of a population for asymptomatic individuals makes a small impact …


A Performance Survey Of Text-Based Sentiment Analysis Methods For Automating Usability Evaluations, Kelsi Van Damme 2021 California Polytechnic State University, San Luis Obispo

A Performance Survey Of Text-Based Sentiment Analysis Methods For Automating Usability Evaluations, Kelsi Van Damme

Master's Theses

Usability testing, or user experience (UX) testing, is increasingly recognized as an important part of the user interface design process. However, evaluating usability tests can be expensive in terms of time and resources and can lack consistency between human evaluators. This makes automation an appealing expansion or alternative to conventional usability techniques.

Early usability automation focused on evaluating human behavior through quantitative metrics but the explosion of opinion mining and sentiment analysis applications in recent decades has led to exciting new possibilities for usability evaluation methods.

This paper presents a survey of modern, open-source sentiment analyzers’ usefulness in extracting and …


Per-Pixel Cloud Cover Classification Of Multispectral Landsat-8 Data, Salome E. Carrasco, Torrey J. Wagner, Brent T. Langhals 2021 Riverside Research

Per-Pixel Cloud Cover Classification Of Multispectral Landsat-8 Data, Salome E. Carrasco, Torrey J. Wagner, Brent T. Langhals

Faculty Publications

Random forest and neural network algorithms are applied to identify cloud cover using 10 of the wavelength bands available in Landsat 8 imagery. The methods classify each pixel into 4 different classes: clear, cloud shadow, light cloud, or cloud. The first method is based on a fully connected neural network with ten input neurons, two hidden layers of 8 and 10 neurons respectively, and a single-neuron output for each class. This type of model is considered with and without L2 regularization applied to the kernel weighting. The final model type is a random forest classifier created from an ensemble of …


Digital Commons powered by bepress