Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

3,235 Full-Text Articles 9,310 Authors 1,358,113 Downloads 221 Institutions

All Articles in Data Science

Faceted Search

3,235 full-text articles. Page 120 of 155.

Representation Learning On Multi-Layered Heterogeneous Network, Delvin Ce ZHANG, Hady W. LAUW 2021 Singapore Management University

Representation Learning On Multi-Layered Heterogeneous Network, Delvin Ce Zhang, Hady W. Lauw

Research Collection School Of Computing and Information Systems

Network data can often be represented in a multi-layered structure with rich semantics. One example is e-commerce data, containing user-user social network layer and item-item context layer, with cross-layer user-item interactions. Given the dual characters of homogeneity within each layer and heterogeneity across layers, we seek to learn node representations from such a multi-layered heterogeneous network while jointly preserving structural information and network semantics. In contrast, previous works on network embedding mainly focus on single-layered or homogeneous networks with one type of nodes and links. In this paper we propose intra- and cross-layer proximity concepts. Intra-layer proximity simulates propagation along …


Analytical Models For Traffic Congestion And Accident Analysis, Hongrui Liu, Rahul Ramachandra Shetty 2021 San Jose State University

Analytical Models For Traffic Congestion And Accident Analysis, Hongrui Liu, Rahul Ramachandra Shetty

Mineta Transportation Institute

In the US, over 38,000 people die in road crashes each year, and 2.35 million are injured or disabled, according to the statistics report from the Association for Safe International Road Travel (ASIRT) in 2020. In addition, traffic congestion keeping Americans stuck on the road wastes millions of hours and billions of dollars each year. Using statistical techniques and machine learning algorithms, this research developed accurate predictive models for traffic congestion and road accidents to increase understanding of the complex causes of these challenging issues. The research used US Accidents data consisting of 49 variables describing 4.2 million accident records …


The Negative Effects Of Cyberbullying Among Secondary School Adolescents In Tanzania, Hezron ZACHARIA Onditi, Jennifer Shapka 2021 University of British Columbia

The Negative Effects Of Cyberbullying Among Secondary School Adolescents In Tanzania, Hezron Zacharia Onditi, Jennifer Shapka

Journal of Humanities and Social Sciences

Cyberbullying and its associated consequences on children and adolescents has recently become a problem of global concern. Using phenomenological research design, this study explores the negative effects of cyberbullying on Tanzanian secondary school female and male adolescents. A total of 20 adolescents (50% female) in secondary schools (Form I to Form IV) who identified themselves as victims of cyberbullying were involved in the interview to share their lived experiences on the negative effects of cyberbullying. From thematic analysis, results indicated that, similar to their counterpart adolescents in the developed countries, Tanzanian female and male adolescents have experienced emotional, social, cognitive, …


Automated Data Processing: Making Community Indicators Possible For Lafayette, Indiana, Jace T. Newell, Eli W. Coltin, Eric D. Flaningam 2021 Purdue University

Automated Data Processing: Making Community Indicators Possible For Lafayette, Indiana, Jace T. Newell, Eli W. Coltin, Eric D. Flaningam

The Journal of Purdue Undergraduate Research

No abstract provided.


Facilitating Team-Based Data Science: Lessons Learned From The Dsc-Wav Project, Chelsey Legacy, Andrew Zieffler, Benjamin S. Baumer, Valerie Barr, Nicholas J. Horton 2021 University of Minnesota

Facilitating Team-Based Data Science: Lessons Learned From The Dsc-Wav Project, Chelsey Legacy, Andrew Zieffler, Benjamin S. Baumer, Valerie Barr, Nicholas J. Horton

Statistical and Data Sciences: Faculty Publications

While coursework provides undergraduate data science students with some relevant analytic skills, many are not given the rich experiences with data and computing they need to be successful in the workplace. Additionally, students often have limited exposure to team-based data science and the principles and tools of collaboration that are encountered outside of school. In this paper, we describe the DSC-WAV program, an NSF-funded data science workforce development project in which teams of undergraduate sophomores and juniors work with a local non-profit organization on a data-focused problem. To help students develop a sense of agency and improve confidence in their …


Concept Drift Adaptation With Incremental–Decremental Svm, Honorius Gâlmeanu, Răzvan Andonie 2021 Transilvania University of Braşov

Concept Drift Adaptation With Incremental–Decremental Svm, Honorius Gâlmeanu, Răzvan Andonie

Computer Science Faculty Scholarship

Data classification in streams where the underlying distribution changes over time is known to be difficult. This problem—known as concept drift detection—involves two aspects: (i) detecting the concept drift and (ii) adapting the classifier. Online training only considers the most recent samples; they form the so-called shifting window. Dynamic adaptation to concept drift is performed by varying the width of the window. Defining an online Support Vector Machine (SVM) classifier able to cope with concept drift by dynamically changing the window size and avoiding retraining from scratch is currently an open problem. We introduce the Adaptive Incremental–Decremental SVM (AIDSVM), a …


Addressing The Learning Loss During The Covid-19 Pandemic Through The Adaptation Of Virtual Platforms, Nazrul I. Khandaker, Anika Nawar Mayeesha, Violeta Escandon Correa, Toralv Munro, Andrew Singh, Matthew Khargie, Ality Aghedo, Jasmin Budhan, Krishna Mahabir, Belal A. Sayeed 2021 CUNY York College

Addressing The Learning Loss During The Covid-19 Pandemic Through The Adaptation Of Virtual Platforms, Nazrul I. Khandaker, Anika Nawar Mayeesha, Violeta Escandon Correa, Toralv Munro, Andrew Singh, Matthew Khargie, Ality Aghedo, Jasmin Budhan, Krishna Mahabir, Belal A. Sayeed

Publications and Research

The York College-hosted NASA MAA (MUREP AEROSPACE ACADEMY) has always played a pivotal role in minimizing the learning loss during the summer months, which was heightened during the pandemic. Support from AT&T, Con Edison and NASA enabled the MAA program at York College to offer a virtual STEM education with an earth science concentration to 1000 plus underserved K1-12 students from the community last summer, including 160 high school students. Two factors made this endeavor fruitful: allowing additional time to engage in STEM lessons and increasing self-motivation to successfully accomplish assigned tasks. Students built partnerships and resolved technical issues with …


Leveraging The Popularity Of Virtual Conferencing Due To The Covid-19 Pandemic To Create New Opportunities For Stem Education, Andrew Singh, Nazrul I. Khandaker, Violeta Escandon Correa, Omadevi Singh, Ariel Skobelsky, Farhan Tanvir, Brian Sukhnandan, Matthew Khargie, Elton Selby, Masud Ahmed 2021 CUNY York College

Leveraging The Popularity Of Virtual Conferencing Due To The Covid-19 Pandemic To Create New Opportunities For Stem Education, Andrew Singh, Nazrul I. Khandaker, Violeta Escandon Correa, Omadevi Singh, Ariel Skobelsky, Farhan Tanvir, Brian Sukhnandan, Matthew Khargie, Elton Selby, Masud Ahmed

Publications and Research

Due to the COVID-19 pandemic, virtual learning has become a necessity for K9-16 education. Virtual classwork has been administered through platforms such as Google Classroom, Clever, and iReady. During the summer of 2021, the City University of New York (C.U.N.Y) York College campus hosted its NASA MAA MUREP (Minority University Research and Education Project Aerospace Academy) program virtually using a combination of Zoom, Google Docs, and even Canva, which some students requested as a more intuitive alternative to Microsoft PowerPoint. Students were mentored to use the scientific method to explore their interests in the STEM field, with a geoscience or …


Feature Engineering Vs Feature Selection Vs Hyperparameter Optimization In The Spotify Song Popularity Dataset, Alan Cueva Mora, Brendan Tierney 2021 Technological University Dublin

Feature Engineering Vs Feature Selection Vs Hyperparameter Optimization In The Spotify Song Popularity Dataset, Alan Cueva Mora, Brendan Tierney

Conference Papers

Research in Featuring Engineering has been part of the data pre-processing phase of machine learning projects for many years. It can be challenging for new people working with machine learning to understand its importance along with various approaches to find an optimized model. This work uses the Spotify Song Popularity dataset to compare and evaluate Feature Engineering, Feature Selection and Hyperparameter Optimization. The result of this work will demonstrate Feature Engineering has a greater effect on model efficiency when compared to the alternative approaches.


Crest Or Trough? How Research Libraries Used Emerging Technologies To Survive The Pandemic, So Far, Scout Calvert 2021 University of Nebraska-Lincoln

Crest Or Trough? How Research Libraries Used Emerging Technologies To Survive The Pandemic, So Far, Scout Calvert

University of Nebraska-Lincoln Libraries: Faculty Publications

Introduction

In the first months of the COVID-19 pandemic, it was impossible to tell if we were at the crest of a wave of new transmissions, or a trough of a much larger wave, still yet to peak. As of this writing, as colleges and universities prepare for mostly in-person fall 2021 semesters, case counts in the United States are increasing again after a decline that coincided with easier access to the COVID vaccine. Plans for a return to campus made with confidence this spring may be in doubt, as we climb the curve of what is already the second …


Identification Of Factors Associated With Fume Events Using Text Mining And Data Mining Methods, Mary B. O'Connor 2021 Embry-Riddle Aeronautical University

Identification Of Factors Associated With Fume Events Using Text Mining And Data Mining Methods, Mary B. O'Connor

Doctoral Dissertations and Master's Theses

Pilots, flight attendants, and passengers can be exposed to toxic compounds when the bleed air that supplies the cabin and flight deck is contaminated with pyrolyzed hydraulic fluid or oil from turbine jet engines. These fume events occur sporadically and can result in acute or chronic exposure in air crews and can have catastrophic consequences if flight crew members become impaired or incapacitated. The purpose of this research was to explore unstructured textual data and identify important factors associated with these events. Models using machine learning algorithms were developed and tested using variables gleaned from the text mining process and …


Science Is For Everybody: A Resource For Understanding Glaciers, Climate, And Modeling, Emma Watson 2021 SIT Study Abroad

Science Is For Everybody: A Resource For Understanding Glaciers, Climate, And Modeling, Emma Watson

Independent Study Project (ISP) Collection

Climate change threatens the existence of glaciers worldwide. In order to properly interact with these changing systems, we must first understand them. Glacial models provide an excellent way to do this; however, the language and mathematical concepts used in their creation is generally inaccessible to a common audience. This project presents an online resource for a general audience to interact with climate science, glaciology, and glacial modeling. Long term goals for the project include the incorporation of a glacial model of Drangajökull, Vestfirðir, NW Iceland. As such, focus for the project includes a literature review of glaciers, Drangajökull in particular, …


The Labyrinth Of Data Collection For Humanitarian Project Funding And Implementation, Maria Alejandra Pulido 2021 SIT Study Abroad

The Labyrinth Of Data Collection For Humanitarian Project Funding And Implementation, Maria Alejandra Pulido

Independent Study Project (ISP) Collection

My research concentrates on four NGOs: IOM, IDMC, JIPS, and OCHA which use different tools to collect data and translate the information into evidence for data-driven decision making (DDDM) for the implementation of humanitarian assistance projects. I focus on the importance, advantages, and various data collection tools which help ameliorate the humanitarian sector since it does not have a current professionalized path to enter the workforce. I incorporated four interviews, attended two conferences and analyzed multiple online sources during my project.


The Classification Of Basket Neural Cells In The Mammalian Neocortex, Sreya Pudi 2021 University of South Carolina

The Classification Of Basket Neural Cells In The Mammalian Neocortex, Sreya Pudi

Senior Theses

Basket neuronal cells of the mammalian neocortex have been classically categorized into two or more groups. Originally, it was thought that the large and small types are the naturally occurring groups that emerge from reasons that relate to neurobiological function and anatomical position. Later, a study based on anatomical and physiological features of these neurons introduced a third type, the net basket cell which is intermediate in size as compared to the large and small types. In this study, multivariate analysis was used to test the hypothesis that the large and small types are morphologically distinct groups. The results of …


Towards Source-Aligned Variational Models For Cross-Domain Recommendation, Aghiles SALAH, Thanh-Binh TRAN, Hady W. LAUW 2021 Singapore Management University

Towards Source-Aligned Variational Models For Cross-Domain Recommendation, Aghiles Salah, Thanh-Binh Tran, Hady W. Lauw

Research Collection School Of Computing and Information Systems

Data sparsity is a long-standing challenge in recommender systems. Among existing approaches to alleviate this problem, cross-domain recommendation consists in leveraging knowledge from a source domain or category (e.g., Movies) to improve item recommendation in a target domain (e.g., Books). In this work, we advocate a probabilistic approach to cross-domain recommendation and rely on variational autoencoders (VAEs) as our latent variable models. More precisely, we assume that we have access to a VAE trained on the source domain that we seek to leverage to improve preference modeling in the target domain. To this end, we propose a model which learns …


Data Analysis Of The “2021 Covid, Equity And Social Justice Showcase”, Cristo Leon, James Lipuma 2021 New Jersey Institute of Technology

Data Analysis Of The “2021 Covid, Equity And Social Justice Showcase”, Cristo Leon, James Lipuma

STEM Month

During the “2021 STEM for All Video Showcase” (NSF, 2021) funded by the National Science Foundation, 287 short videos showcasing federally funded projects aimed at improving STEM and CS education were presented.

The videos highlight strategies to engage students during COVID-19 and address educational inequities.


Deep Fakes: The Algorithms That Create And Detect Them And The National Security Risks They Pose, Nick Dunard 2021 James Madison University

Deep Fakes: The Algorithms That Create And Detect Them And The National Security Risks They Pose, Nick Dunard

James Madison Undergraduate Research Journal (JMURJ)

The dissemination of deep fakes for nefarious purposes poses significant national security risks to the United States, requiring an urgent development of technologies to detect their use and strategies to mitigate their effects. Deep fakes are images and videos created by or with the assistance of AI algorithms in which a person’s likeness, actions, or words have been replaced by someone else’s to deceive an audience. Often created with the help of generative adversarial networks, deep fakes can be used to blackmail, harass, exploit, and intimidate individuals and businesses; in large-scale disinformation campaigns, they can incite political tensions around the …


Infer: An R Package For Tidyverse-Friendly Statistical Inference, Simon P. Couch, Andrew P. Bray, Chester Ismay, Evgeni Chasnovski, B. Baumer, Mine Cetinkaya-Rundel 2021 Johns Hopkins University

Infer: An R Package For Tidyverse-Friendly Statistical Inference, Simon P. Couch, Andrew P. Bray, Chester Ismay, Evgeni Chasnovski, B. Baumer, Mine Cetinkaya-Rundel

Statistical and Data Sciences: Faculty Publications

infer implements an expressive grammar to perform statistical inference that adheres to the tidyverse design framework (Wickham et al., 2019). Rather than providing methods for specific statistical tests, this package consolidates the principles that are shared among common hypothesis tests and confidence intervals into a set of four main verbs (functions), supplemented with many utilities to visualize and extract value from their outputs.


The Concurrence Of Dna Methylation And Demethylation Is Associated With Transcription Regulation, Jiejun Shi, Jianfeng Xu, Yiling Elaine Chen, Jason Sheng Li, Ya Cui, Lanlan Shen, Jingyi Jessica Li, Wei Li 2021 The Texas Medical Center Library

The Concurrence Of Dna Methylation And Demethylation Is Associated With Transcription Regulation, Jiejun Shi, Jianfeng Xu, Yiling Elaine Chen, Jason Sheng Li, Ya Cui, Lanlan Shen, Jingyi Jessica Li, Wei Li

Children’s Nutrition Research Center Staff Publications

The mammalian DNA methylome is formed by two antagonizing processes, methylation by DNA methyltransferases (DNMT) and demethylation by ten-eleven translocation (TET) dioxygenases. Although the dynamics of either methylation or demethylation have been intensively studied in the past decade, the direct effects of their interaction on gene expression remain elusive. Here, we quantify the concurrence of DNA methylation and demethylation by the percentage of unmethylated CpGs within a partially methylated read from bisulfite sequencing. After verifying 'methylation concurrence' by its strong association with the co-localization of DNMT and TET enzymes, we observe that methylation concurrence is strongly correlated with gene expression. …


Comprehensive Characterization Of Covid-19 Patients With Repeatedly Positive Sars-Cov-2 Tests Using A Large Us Electronic Health Record Database, Xiao Dong, Yujia Zhou, Xiao-Ou Shu, Elmer V Bernstam, Rebecca Stern, David M Aronoff, Hua Xu, Loren Lipworth 2021 University of Texas Health Science Center at Houston, School of Health Information Sciences, Houston TX, USA

Comprehensive Characterization Of Covid-19 Patients With Repeatedly Positive Sars-Cov-2 Tests Using A Large Us Electronic Health Record Database, Xiao Dong, Yujia Zhou, Xiao-Ou Shu, Elmer V Bernstam, Rebecca Stern, David M Aronoff, Hua Xu, Loren Lipworth

Faculty, Staff and Student Publications

In the absence of genome sequencing, two positive molecular tests for severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) separated by negative tests, prolonged time, and symptom resolution remain the best surrogate measure of possible reinfection. Using a large electronic health record database, we characterized clinical and testing data for 23 patients with repeatedly positive SARS-CoV-2 PCR test results ≥60 days apart, separated by ≥2 consecutive negative test results. The prevalence of chronic medical conditions, symptoms, and severe outcomes related to coronavirus disease 19 (COVID-19) illness were ascertained. The median age of patients was 64.5 years, 40% were Black, and 39% …


Digital Commons powered by bepress