Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

2022

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 271 - 300 of 418

Full-Text Articles in Data Science

Fusionai, A Dna-Sequence-Based Deep Learning Protocol Reduces The False Positives Of Human Fusion Gene Prediction, Pora Kim, Hua Tan, Jiajia Liu, Himansu Kumar, Xiaobo Zhou Mar 2022

Fusionai, A Dna-Sequence-Based Deep Learning Protocol Reduces The False Positives Of Human Fusion Gene Prediction, Pora Kim, Hua Tan, Jiajia Liu, Himansu Kumar, Xiaobo Zhou

Faculty, Staff and Student Publications

Even though there were many tool developments of fusion gene prediction from NGS data, too many false positives are still an issue. Wise use of the genomic features around the fusion gene breakpoints will be helpful to identify reliable fusion genes efficiently. For this aim, we developed FusionAI, a deep learning pipeline predicting human fusion gene breakpoints from DNA sequence. FusionAI is freely available via https://compbio.uth.edu/FusionGDB2/FusionAI. For complete details on the use and execution of this protocol, please refer to Kim et al. (2021b).


How Graduate Student Fellows Enhance What A Center For Digital Scholarship Does, Ben B. Chiewphasa Mar 2022

How Graduate Student Fellows Enhance What A Center For Digital Scholarship Does, Ben B. Chiewphasa

Transforming Libraries for Graduate Students

Multiple disciplines are increasingly embracing data science and digital scholarship. However, insufficient training for digital and computational methodologies within subject/departmental silos means that these needs often get overlooked. Opportunities for learning how to teach technical concepts (i.e., how to handle troubleshooting, live participatory coding, etc.) are also rare or non-existent via departmental offerings. To respond to these needs, the Navari Family Center for Digital Scholarship launched its Pedagogy Fellowship Program in Fall 2021 where Notre Dame PhD students/candidates build their instructional expertise and experience related to digital scholarship with an added bonus of enhancing their competitiveness on the job market. …


Split Classification Model For Complex Clustered Data, Katherine Gerot Mar 2022

Split Classification Model For Complex Clustered Data, Katherine Gerot

Honors Program: Senior Projects (Public)

Classification in high-dimensional data has generated tremendous interest in a multitude of fields. Data in higher dimensions often tend to reside in non-Euclidean metric space. This prevents Euclidean-based classification methodologies, such as regression, from reliably modeling the data. Many proposed models rely on computationally-complex embedding to convert the data to a more usable format. Others, namely the Support Vector Machine, rely on kernel manipulation to implicitly describe the "feature space" to arrive at a non-linear decision boundary. The proposed methodology in this paper seeks to classify complex data in a relatively computationally-simple and explainable manner.


Bcse: Blockchain-Based Trusted Service Evaluation Model Over Big Data, Fengyin Li, Xinying Yu, Rui Ge, Yanli Wang, Yang Cui, Huiyu Zhou Mar 2022

Bcse: Blockchain-Based Trusted Service Evaluation Model Over Big Data, Fengyin Li, Xinying Yu, Rui Ge, Yanli Wang, Yang Cui, Huiyu Zhou

Big Data Mining and Analytics

The blockchain, with its key characteristics of decentralization, persistence, anonymity, and auditability, has become a solution to overcome the overdependence and lack of trust for a traditional public key infrastructure on third-party institutions. Because of these characteristics, the blockchain is suitable for solving certain open problems in the service-oriented social network, where the unreliability of submitted reviews of service vendors can cause serious security problems. To solve the unreliability problems of submitted reviews, this paper first proposes a blockchain-based identity authentication scheme and a new trusted service evaluation model by introducing the scheme into a service evaluation model. The new …


Big Data With Cloud Computing: Discussions And Challenges, Amanpreet Kaur Sandhu Mar 2022

Big Data With Cloud Computing: Discussions And Challenges, Amanpreet Kaur Sandhu

Big Data Mining and Analytics

With the recent advancements in computer technologies, the amount of data available is increasing day by day. However, excessive amounts of data create great challenges for users. Meanwhile, cloud computing services provide a powerful environment to store large volumes of data. They eliminate various requirements, such as dedicated space and maintenance of expensive computer hardware and software. Handling big data is a time-consuming task that requires large computational clusters to ensure successful data storage and processing. In this work, the definition, classification, and characteristics of big data are discussed, along with various cloud services, such as Microsoft Azure, Google Cloud, …


Exploiting More Associations Between Slots For Multi-Domain Dialog State Tracking, Hui Bai, Yan Yang, Jie Wang Mar 2022

Exploiting More Associations Between Slots For Multi-Domain Dialog State Tracking, Hui Bai, Yan Yang, Jie Wang

Big Data Mining and Analytics

Dialog State Tracking (DST) aims to extract the current state from the conversation and plays an important role in dialog systems. Existing methods usually predict the value of each slot independently and do not consider the correlations among slots, which will exacerbate the data sparsity problem because of the increased number of candidate values. In this paper, we propose a multi-domain DST model that integrates slot-relevant information. In particular, certain connections may exist among slots in different domains, and their corresponding values can be obtained through explicit or implicit reasoning. Therefore, we use the graph adjacency matrix to determine the …


Sampling With Prior Knowledge For High-Dimensional Gravitational Wave Data Analysis, He Wang, Zhoujian Cao, Yue Zhou, Zong-Kuan Guo, Zhixiang Ren Mar 2022

Sampling With Prior Knowledge For High-Dimensional Gravitational Wave Data Analysis, He Wang, Zhoujian Cao, Yue Zhou, Zong-Kuan Guo, Zhixiang Ren

Big Data Mining and Analytics

Extracting knowledge from high-dimensional data has been notoriously difficult, primarily due to the so-called "curse of dimensionality" and the complex joint distributions of these dimensions. This is a particularly profound issue for high-dimensional gravitational wave data analysis where one requires to conduct Bayesian inference and estimate joint posterior distributions. In this study, we incorporate prior physical knowledge by sampling from desired interim distributions to develop the training dataset. Accordingly, the more relevant regions of the high-dimensional feature space are covered by additional data points, such that the model can learn the subtle but important details. We adapt the normalizing flow …


Toward Intelligent Financial Advisors For Identifying Potential Clients: A Multitask Perspective, Qixiang Shao, Runlong Yu, Hongke Zhao, Chunli Liu, Mengyi Zhang, Hongmei Song, Qi Liu Mar 2022

Toward Intelligent Financial Advisors For Identifying Potential Clients: A Multitask Perspective, Qixiang Shao, Runlong Yu, Hongke Zhao, Chunli Liu, Mengyi Zhang, Hongmei Song, Qi Liu

Big Data Mining and Analytics

Intelligent Financial Advisors (IFAs) in online financial applications (apps) have brought new life to personal investment by providing appropriate and high-quality portfolios for users. In real-world scenarios, identifying potential clients is a crucial issue for IFAs, i.e., identifying users who are willing to purchase the portfolios. Thus, extracting useful information from various characteristics of users and further predicting their purchase inclination are urgent. However, two critical problems encountered in real practice make this prediction task challenging, i.e., sample selection bias and data sparsity. In this study, we formalize a potential conversion relationship, i.e., user→activated user→client and decompose this relationship into …


A Comparison Of Computational Approaches For Intron Retention Detection, Jiantao Zheng, Cuixiang Lin, Zhenpeng Wu, Hong-Dong Li Mar 2022

A Comparison Of Computational Approaches For Intron Retention Detection, Jiantao Zheng, Cuixiang Lin, Zhenpeng Wu, Hong-Dong Li

Big Data Mining and Analytics

Intron Retention (IR) is an alternative splicing mode through which introns are retained in mature RNAs rather than being spliced in most cases. IR has been gaining increasing attention in recent years because of its recognized association with gene expression regulation and complex diseases. Continuous efforts have been dedicated to the development of IR detection methods. These methods differ in their metrics to quantify retention propensity, performance to detect IR events, functional enrichment of detected IRs, and computational speed. A systematic experimental comparison would be valuable to the selection and use of existing methods. In this work, we conduct an …


Ingredient Classification Using Food Ontology, Ricky Flores Mar 2022

Ingredient Classification Using Food Ontology, Ricky Flores

UNO Student Research and Creative Activity Fair

A food label provides some of the most crucial information for a food product. The food label is a key resource for many health-conscious consumers for understanding ingredients. It is also vital for individuals to avoid food allergens or help patients follow dietary recommendations. While the food labels in the United States are regulated by the Food and Drug Administration (FDA) many labels contain additional information or statements that are not regulated. Moreover, the food label may be complex or contain terminology that the layperson may not understand. Evidence has indicated that consumers often find nutrition labels confusing, especially when …


The Mathematics Of Risk: An Introduction To Guaranteed Data De-Identification, Kristi Thompson Mar 2022

The Mathematics Of Risk: An Introduction To Guaranteed Data De-Identification, Kristi Thompson

Western Libraries Presentations

This webinar is devoted to the mathematical and theoretical underpinnings of guaranteed data anonymization. Topics covered include an overview of identifiers and quasi-identifiers, an introduction to k-anonymity, a look at some cases where k-anonymity breaks down, and anonymization hierarchies. The presenter will describe a method to assess a survey dataset for anonymization using standard statistical software and consider the question of "anonymization overkill". Much of the academic material looking at data anonymization is quite abstract and aimed at computer scientists, while material aimed at data curators does not always consider recent developments. This webinar is intended to help bridge the …


Autonomous, Long-Range, Sensor Emplacement Using Unmanned Aircraft Systems, Adam Plowcha, Justin Bradley, Jacob Hoberg, Thomas Ammon, Mark Nail, Brittany Duncan, Carrick Detweiler Mar 2022

Autonomous, Long-Range, Sensor Emplacement Using Unmanned Aircraft Systems, Adam Plowcha, Justin Bradley, Jacob Hoberg, Thomas Ammon, Mark Nail, Brittany Duncan, Carrick Detweiler

School of Computing: Faculty Publications

Automated, in-ground sensor emplacement can significantly improve remote, terrestrial, data collection capabilities. Utilizing a multicopter, unmanned aircraft system (UAS) for this purpose allows sensor insertion with minimal disturbance to the target site or surrounding area. However, developing an emplacement mechanism for a small multicopter, autonomy to manage the target selection and implantation process, as well as long-range deployment are challenging to address. We have developed an autonomous, multicopter UAS that can implant subsurface sensor devices. We enhanced the UAS autopilot with autonomy for target and landing zone selection, as well as ensuring the sensor is implanted properly in the ground. …


Generalized Robust Feature Selection, Bradford L. Lott Mar 2022

Generalized Robust Feature Selection, Bradford L. Lott

Theses and Dissertations

Feature selection may be summarized as identifying salient features to a given response. Understanding which features affect the response enables, in the future, only collecting consequential data; hence, the feature selection algorithm may lead to saving effort spent collecting data, storage resources, as well as computational resources for making predictions. We propose a generalized approach to select the salient features of data sets. Our approach may also be applied to unsupervised datasets to understand which data streams provide unique information. We contend our approach identifies salient features robust to the sub-sequent predictive model applied. The proposed algorithm considers all provided …


Directional Pairwise Class Confusion Bias And Its Mitigation, Sudhashree Sayenju, Ramazan Aygun Phd, Jonathan Boardman, Duleep Prasanna Rathgamage Don, Yifan Zhang Phd, Bill Franks, Sereres Johnston Phd, George Lee, Dan Sullivan, Girish Modgil Phd Mar 2022

Directional Pairwise Class Confusion Bias And Its Mitigation, Sudhashree Sayenju, Ramazan Aygun Phd, Jonathan Boardman, Duleep Prasanna Rathgamage Don, Yifan Zhang Phd, Bill Franks, Sereres Johnston Phd, George Lee, Dan Sullivan, Girish Modgil Phd

Published and Grey Literature from PhD Candidates

Recent advances in Natural Language Processing have led to powerful and sophisticated models like BERT (Bidirectional Encoder Representations from Transformers) that have bias. These models are mostly trained on text corpora that deviate in important ways from the text encountered by a chatbot in a problem-specific context. While a lot of research in the past has focused on measuring and mitigating bias with respect to protected attributes (stereotyping like gender, race, ethnicity, etc.), there is lack of research in model bias with respect to classification labels. We investigate whether a classification model hugely favors one class with respect to another. …


The Clock Modulator Nobiletin Mitigates Astrogliosis-Associated Neuroinflammation And Disease Hallmarks In An Alzheimer’S Disease Model, Marvin Wirianto, Chih-Yen Wang, Eunju Kim, Nobuya Koike, Ruben Gomez-Gutierrez, Kazunari Nohara, Gabriel Escobedo, Jong Min Choi, Chorong Han, Kazuhiro Yagita, Sung Yun Jung, Claudio Soto, Hyun Kyoung Lee, Rodrigo Morales, Seung-Hee Yoo, Zheng Chen Mar 2022

The Clock Modulator Nobiletin Mitigates Astrogliosis-Associated Neuroinflammation And Disease Hallmarks In An Alzheimer’S Disease Model, Marvin Wirianto, Chih-Yen Wang, Eunju Kim, Nobuya Koike, Ruben Gomez-Gutierrez, Kazunari Nohara, Gabriel Escobedo, Jong Min Choi, Chorong Han, Kazuhiro Yagita, Sung Yun Jung, Claudio Soto, Hyun Kyoung Lee, Rodrigo Morales, Seung-Hee Yoo, Zheng Chen

Faculty, Staff and Student Publications

Alzheimer's disease (AD) is a devastating neurodegenerative disorder, and there is a pressing need to identify disease-modifying factors and devise interventional strategies. The circadian clock, our intrinsic biological timer, orchestrates various cellular and physiological processes including gene expression, sleep, and neuroinflammation; conversely, circadian dysfunctions are closely associated with and/or contribute to AD hallmarks. We previously reported that the natural compound Nobiletin (NOB) is a clock-enhancing modulator that promotes physiological health and healthy aging. In the current study, we treated the double transgenic AD model mice, APP/PS1, with NOB-containing diets. NOB significantly alleviated β-amyloid burden in both the hippocampus and the …


Use Of The Deep Learning Approach To Measure Alveolar Bone Level, Chun-Teh Lee, Tanjida Kabir, Jiman Nelson, Sally Sheng, Hsiu-Wan Meng, Thomas E Van Dyke, Muhammad F Walji, Xiaoqian Jiang, Shayan Shams Mar 2022

Use Of The Deep Learning Approach To Measure Alveolar Bone Level, Chun-Teh Lee, Tanjida Kabir, Jiman Nelson, Sally Sheng, Hsiu-Wan Meng, Thomas E Van Dyke, Muhammad F Walji, Xiaoqian Jiang, Shayan Shams

Faculty, Staff and Student Publications

AIM: The goal was to use a deep convolutional neural network to measure the radiographic alveolar bone level to aid periodontal diagnosis.

MATERIALS AND METHODS: A deep learning (DL) model was developed by integrating three segmentation networks (bone area, tooth, cemento-enamel junction) and image analysis to measure the radiographic bone level and assign radiographic bone loss (RBL) stages. The percentage of RBL was calculated to determine the stage of RBL for each tooth. A provisional periodontal diagnosis was assigned using the 2018 periodontitis classification. RBL percentage, staging, and presumptive diagnosis were compared with the measurements and diagnoses made by the …


Counterfactual Analysis Of Differential Comorbidity Risk Factors In Alzheimer’S Disease And Related Dementias, Yejin Kim, Kai Zhang, Sean I Savitz, Luyao Chen, Paul E Schulz, Xiaoqian Jiang Mar 2022

Counterfactual Analysis Of Differential Comorbidity Risk Factors In Alzheimer’S Disease And Related Dementias, Yejin Kim, Kai Zhang, Sean I Savitz, Luyao Chen, Paul E Schulz, Xiaoqian Jiang

Faculty, Staff and Student Publications

Alzheimer’s disease and related dementias (ADRD) is a multifactorial disease that involves several different etiologic mechanisms with various comorbidities. There is also significant heterogeneity in the prevalence of ADRD across diverse demographics groups. Association studies on such heterogeneous comorbidity risk factors are limited in their ability to determine causation. We aim to compare counterfactual treatment effects of various comorbidity in ADRD in different racial groups (African Americans and Caucasians). We used 138,026 ADRD and 1:1 matched older adults without ADRD from nationwide electronic health records, which extensively cover a large population’s long medical history in breadth. We matched African Americans …


Leveraging Machine Learning For Large Scale Analysis Of Publicly-Available Data For Gnss Interference Events, David K. Stamper Mar 2022

Leveraging Machine Learning For Large Scale Analysis Of Publicly-Available Data For Gnss Interference Events, David K. Stamper

Theses and Dissertations

This research documents architecture and implementation of an enhanced interference detection and classification analysis system, using both a database and storage solution utilizing machine learning algorithms to detect changes in Carrier-to-Noise strength over multiple GNSS sites. The system uses publicly-available government supported receivers to detect interference, and built using FOSS packaged as a programming library through Python. Two algorithms are discussed in terms of enhancing interference detection using both non-machine learning and machine learning approaches. Two algorithms are also discussed which are used for classification of events. In addition, an approach to Large Scale data analytics is demonstrated via a …


Constructing Prediction Intervals With Neural Networks: An Empirical Evaluation Of Bootstrapping And Conformal Inference Methods, Alexander N. Contarino Mar 2022

Constructing Prediction Intervals With Neural Networks: An Empirical Evaluation Of Bootstrapping And Conformal Inference Methods, Alexander N. Contarino

Theses and Dissertations

Artificial neural networks (ANNs) are popular tools for accomplishing many machine learning tasks, including predicting continuous outcomes. However, the general lack of confidence measures provided with ANN predictions limit their applicability, especially in military settings where accuracy is paramount. Supplementing point predictions with prediction intervals (PIs) is common for other learning algorithms, but the complex structure and training of ANNs renders constructing PIs difficult. This work provides the network design choices and inferential methods for creating better performing PIs with ANNs to enable their adaptation for military use. A two-step experiment is executed across 11 datasets, including an imaged-based dataset. …


Telemetry Data Mining For Unmanned Aircraft Systems, Li Yu Mar 2022

Telemetry Data Mining For Unmanned Aircraft Systems, Li Yu

Theses and Dissertations

With ever more data becoming available to the US Air Force, it is vital to develop effective methods to leverage this strategic asset. Machine learning (ML) techniques present a means of meeting this challenge, as these tools have demonstrated successful use in commercial applications. For this research, three ML methods were applied to a unmanned aircraft system (UAS) telemetry dataset with the aim of extracting useful insight related to phases of flight. It was shown that ML provides an advantage in exploratory data analysis and as well as classification of phases. Neural network models demonstrated the best performance with over …


Online Masters In Data Science, Joanna Burkhardt Feb 2022

Online Masters In Data Science, Joanna Burkhardt

Library Impact Statements

No abstract provided.


Integrating Web Applications Into Popular Survey Platforms For Online Experiments, Benjamin Carter, Alessandro Del Ponte Feb 2022

Integrating Web Applications Into Popular Survey Platforms For Online Experiments, Benjamin Carter, Alessandro Del Ponte

Political Science Faculty Articles and Research

Research using custom-made web applications is burgeoning as scholars increasingly conduct their experiments online. We show how researchers can integrate their web applications into popular survey software such as Qualtrics in five simple steps and provide the full JavaScript code and screenshots. This procedure allows participants to seamlessly switch from Qualtrics to their web applications without leaving the survey platform. This integration has two benefits: (1) it eliminates the risk that participants inadvertently drop out of the survey while switching from the survey software to the web application and vice versa; and (2) it saves researchers the fees charged by …


Covid-19 Pandemic Analysis By The Volterra Integral Equation Models: A Preliminary Study Of Brazil, Italy, And South Africa, Yajni Warnapala, Emma Dehetre, Kate Gilbert Feb 2022

Covid-19 Pandemic Analysis By The Volterra Integral Equation Models: A Preliminary Study Of Brazil, Italy, And South Africa, Yajni Warnapala, Emma Dehetre, Kate Gilbert

Arts & Sciences Faculty Publications

The COVID-19 pandemic has affected many people throughout the world. The objective of this research project was to find numerical solutions through the Gaussian Quadrature Method for the Volterra Integral Equation Model. The non-homogenous Volterra Integral Equation of the second kind is used to capture a broader range of disease distributions. Volterra Integral equation models are used in the context of applied mathematics, public health, and evolutionary biology. The mathematical models of this integral equation gave valid convergence results for the COVID-19 data for 3 countries Italy, South Africa and Brazil. The modeling of these countries was done using the …


Learning Latent Causal Dynamics, Weiran Yao, Guangyi Chen, Kun Zhang Feb 2022

Learning Latent Causal Dynamics, Weiran Yao, Guangyi Chen, Kun Zhang

Machine Learning Faculty Publications

One critical challenge of time-series modeling is how to learn and quickly correct the model under unknown distribution shifts. In this work, we propose a principled framework, called LiLY, to first recover time-delayed latent causal variables and identify their relations from measured temporal data under different distribution shifts. The correction step is then formulated as learning the low-dimensional change factors with a few samples from the new environment, leveraging the identified causal structure. Specifically, the framework factorizes unknown distribution shifts into transition distribution changes caused by fixed dynamics and time-varying latent causal relations, and by global changes in observation. We …


Iseeq: Information Seeking Question Generation Using Dynamic Meta-Information Retrieval And Knowledge Graphs, Manas Gaur, Kalpa Gunaratna, Vijay Srinivasan, Hongxia Jin Feb 2022

Iseeq: Information Seeking Question Generation Using Dynamic Meta-Information Retrieval And Knowledge Graphs, Manas Gaur, Kalpa Gunaratna, Vijay Srinivasan, Hongxia Jin

Publications

Conversational Information Seeking (CIS) is a relatively new research area within conversational AI that attempts to seek information from end-users in order to understand and satisfy users’ needs. If realized, such a system has far-reaching benefits in the real world; for example, a CIS system can assist clinicians in pre-screening or triaging patients in healthcare. A key open sub-problem in CIS that remains unaddressed in the literature is generating Information Seeking Questions (ISQs) based on a short initial query from the end user. To address this open problem, we propose Information SEEking Question generator (ISEEQ), a novel approach for generating …


Data Analytics And Visualization Dsp 562, Harrison Dekker Feb 2022

Data Analytics And Visualization Dsp 562, Harrison Dekker

Collection Development Reports and Documents

No abstract provided.


Advanced Topics In Machine Learning Dsp 566, Harrison Dekker Feb 2022

Advanced Topics In Machine Learning Dsp 566, Harrison Dekker

Collection Development Reports and Documents

No abstract provided.


Introduction To Statistical Computing Dsp 565, Harrison Dekker Feb 2022

Introduction To Statistical Computing Dsp 565, Harrison Dekker

Collection Development Reports and Documents

No abstract provided.


Applications Of Data Science In Biological Science Dsp 569, Harrison Dekker Feb 2022

Applications Of Data Science In Biological Science Dsp 569, Harrison Dekker

Collection Development Reports and Documents

No abstract provided.


Mathematical Foundations For Data Science Ams/Dsp 563, Harrison Dekker Feb 2022

Mathematical Foundations For Data Science Ams/Dsp 563, Harrison Dekker

Collection Development Reports and Documents

No abstract provided.