Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Computer Sciences (141)
- Medicine and Health Sciences (130)
- Life Sciences (117)
- Bioinformatics (99)
- Biomedical Informatics (96)
-
- Social and Behavioral Sciences (71)
- Engineering (67)
- Statistics and Probability (65)
- Artificial Intelligence and Robotics (63)
- Medical Sciences (39)
- Medical Specialties (35)
- Computer Engineering (32)
- Electrical and Computer Engineering (30)
- Applied Mathematics (23)
- Applied Statistics (23)
- Statistical Models (23)
- Diseases (20)
- Other Computer Sciences (20)
- Environmental Sciences (19)
- Public Health (19)
- Categorical Data Analysis (18)
- Medical Genetics (18)
- Business (17)
- Databases and Information Systems (16)
- Mathematics (16)
- Data Storage Systems (15)
- Public Affairs, Public Policy and Public Administration (15)
- Systems and Communications (15)
- Institution
-
- The Texas Medical Center Library (97)
- Southern Methodist University (19)
- City University of New York (CUNY) (16)
- Old Dominion University (14)
- Universitas Negeri Malang (14)
-
- Kennesaw State University (12)
- Chapman University (10)
- Smith College (10)
- Technological University Dublin (9)
- Air Force Institute of Technology (8)
- University of Rhode Island (8)
- Virginia Commonwealth University (8)
- West Virginia University (8)
- University of Louisville (7)
- Chinese Academy of Sciences (6)
- Embry-Riddle Aeronautical University (6)
- Tsinghua University Press (6)
- Bryant University (5)
- New Jersey Institute of Technology (5)
- University of Kentucky (5)
- University of South Carolina (5)
- Western University (5)
- Central Bank of Nigeria (4)
- Claremont Colleges (4)
- The University of Akron (4)
- University of New Mexico (4)
- Bowling Green State University (3)
- California Polytechnic State University, San Luis Obispo (3)
- Central Washington University (3)
- Clemson University (3)
- Keyword
-
- Humans (52)
- Machine learning (40)
- Machine Learning (36)
- Deep learning (16)
- COVID-19 (14)
-
- Data science (12)
- Deep Learning (12)
- Natural language processing (11)
- Algorithms (9)
- Neural Networks (9)
- Artificial intelligence (8)
- Data (8)
- Library Impact Statement, Faculty Senate, Data Science, Collection Development (8)
- Library science (8)
- Classification (7)
- Data Science (7)
- Privacy (7)
- Statistics (7)
- Adult (6)
- Artificial Intelligence (6)
- CNN (6)
- Computer (6)
- Computer science (6)
- Data analysis (6)
- Genome-Wide Association Study (6)
- Mathematics (6)
- Natural Language Processing (6)
- Prediction (6)
- Retrospective Studies (6)
- Sentiment analysis (6)
- Publication
-
- Faculty, Staff and Student Publications (95)
- Theses and Dissertations (20)
- SMU Data Science Review (19)
- Knowledge Engineering and Data Science (14)
- Dissertations, Theses, and Capstone Projects (12)
-
- Electronic Theses and Dissertations (10)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (8)
- Collection Development Reports and Documents (7)
- Statistical and Data Sciences: Faculty Publications (7)
- Big Data Mining and Analytics (6)
- Bulletin of Chinese Academy of Sciences (Chinese Version) (6)
- Publications (6)
- Articles (5)
- Dissertations (5)
- Honors Projects (5)
- Honors Projects in Data Science (5)
- CBN Journal of Applied Statistics (JAS) (4)
- Doctor of Data Science and Analytics Dissertations (4)
- Published and Grey Literature from PhD Candidates (4)
- Theses (4)
- Williams Honors College, Honors Research Projects (4)
- College of Graduate Studies: Theses & Dissertations (3)
- Electrical & Computer Engineering Faculty Publications (3)
- Electronic Theses and Dissertations, 2020-2023 (3)
- Graduate Student Theses, Dissertations, & Professional Papers (3)
- LSU New Orleans Theses and Dissertations (3)
- Mathematics, Physics, and Computer Science Faculty Articles and Research (3)
- OES Faculty Publications (3)
- Research Collection School Of Computing and Information Systems (3)
- Western Libraries Presentations (3)
- Publication Type
- File Type
Articles 331 - 360 of 418
Full-Text Articles in Data Science
The Pandemic’S Effects On The Use Of Personal Listening Devices And Prevalence Of Hearing Damage In College Students, Morgan Fink
The Pandemic’S Effects On The Use Of Personal Listening Devices And Prevalence Of Hearing Damage In College Students, Morgan Fink
Senior Honors Projects
Personal listening devices (PLDs), such as earbuds and headphones, are prevalent in today’s society, and overuse of these devices can cause hearing damage. Since the pandemic caused lockdowns and online classes, college students have presumably had more time to be indoors and to use PLDs, leading to a higher risk or developing hearing damage. Previous studies have explored the PLD use and the prevalence of hearing damage in college students, but this study investigates whether the coronavirus pandemic has affected college students’ PLD listening habits and whether these changes are related to the students’ disclosure of suspected symptoms of hearing …
Understanding The Enumerated World: Making Sense Of Data As An Information Source, Kristi Thompson, Elizabeth Hill, Alexandra Cooper
Understanding The Enumerated World: Making Sense Of Data As An Information Source, Kristi Thompson, Elizabeth Hill, Alexandra Cooper
Western Libraries Publications
Chapter in ACRL publication The Data Literacy Cookbook.
This recipe is a guide to preparing an instructional session aimed at postsecondary students in the social or health sciences or related disciplines on locating, evaluating, and using secondary data sources as information resources. Who collects data? Where can you access them? Why are data available on some topics and not others? Why are some statistics available at a detailed level of geography and others only nationally? What are some key limitations of official statistics, and where can information be found to fill in the gaps? This recipe uses these questions to …
Estimating Efforts For Various Activities In Agile Software Development: An Empirical Study, Lan Cao
Estimating Efforts For Various Activities In Agile Software Development: An Empirical Study, Lan Cao
Information Technology & Decision Sciences Faculty Publications
Effort estimation is an important practice in agile software development. The agile community believes that developers’ estimates get more accurate over time due to the cumulative effect of learning from short and frequent feedback. However, there is no empirical evidence of an improvement in estimation accuracy over time, nor have prior studies examined effort estimation in different development activities, which are associated with substantial costs. This study fills the knowledge gap in the field of software estimation in agile software development by investigating estimations across time and different development activities based on data collected from a large agile project. This …
Predicting Outcomes Of El Clásico Using Random Forests And Extreme Gradient Boosting, Emanuel Jarquin
Predicting Outcomes Of El Clásico Using Random Forests And Extreme Gradient Boosting, Emanuel Jarquin
CMC Senior Theses
In the modern era, sports betting is becoming increasingly popular. This is especially true in the realm of soccer (or ‘football’ as it is known outside the United States). As a result, the concept of attempting to predict the outcomes of soccer matches using machine learning has garnered much attention in recent years. In this thesis, I utilize well-known machine learning techniques to predict the outcomes of El Clásico matchups and compare the predictive performance of these techniques. The predictive methods employed for this thesis are random forests using the party package in R and extreme gradient boosting using the …
Integrated Gradients Is A Nonlinear Generalization Of The Industry Standard Approach To Variable Attribution For Credit Risk Models, Jonathan Boardman, Md Shafiul Alam, Xiao Huang, Ying Xie
Integrated Gradients Is A Nonlinear Generalization Of The Industry Standard Approach To Variable Attribution For Credit Risk Models, Jonathan Boardman, Md Shafiul Alam, Xiao Huang, Ying Xie
Published and Grey Literature from PhD Candidates
In modern society, epistemic uncertainty limits trust in financial relationships, necessitating transparency and accountability mechanisms for both consumers and lenders. One upshot is that credit risk assessments must be explainable to the consumer. In the United States regulatory milieu, this entails both the identification of key factors in a decision and the provision of consistent actions that would improve standing. The traditionally accepted approach to explainable credit risk modeling involves generating scores with Generalized Linear Models (GLMs) - usually logistic regression, calculating the contribution of each predictor to the total points lost from the theoretical maximum, and generating reason codes …
Using Neural Networks To Model Guitar Distortion, Caleb Koch, Scott Hawley, Andrew Fyfe
Using Neural Networks To Model Guitar Distortion, Caleb Koch, Scott Hawley, Andrew Fyfe
Science University Research Symposium (SURS)
Guitar players have been modifying their guitar tone with audio effects ever since the mid-20th century. Traditionally, these effects have been achieved by passing a guitar signal through a series of electronic circuits which modify the signal to produce the desired audio effect. With advances in computer technology, audio “plugins” have been created to produce audio effects digitally through programming algorithms. More recently, machine learning researchers have been exploring the use of neural networks to produce audio effects that yield strikingly similar results to their analog counterparts. Recurrent Neural Networks and Temporal Convolutional Networks have proven to be exceptional at …
Towards A Burden-Free Implicit Authentication For Wearable Device Users, Bryan Lee, Sudip Vhaduri
Towards A Burden-Free Implicit Authentication For Wearable Device Users, Bryan Lee, Sudip Vhaduri
Discovery Undergraduate Interdisciplinary Research Internship
The state of current knowledge-based wearable authentication systems requires users to physically interact with a device to initiate and validate their presence, thereby imposing a burden on the user. However, with the recent advancements of sensor technologies in consumer smart wearables (e.g., Fitbit and Apple watches), we were able to utilize vectors of statistical features extracted from the continuous stream of data from these IoT devices to implicitly validate a user's activities and its spatiotemporal context via the use of machine learning techniques. To improve the performance of our models, additional soft biometric data (i.e., respiratory sounds) was collected, and …
Non-Inferiority Testing: Kernel Estimation And Overlap Measure, Larie C. Ward
Non-Inferiority Testing: Kernel Estimation And Overlap Measure, Larie C. Ward
College of Graduate Studies: Theses & Dissertations
In non-inferiority testing, the decision of whether a proposed treatment is non-inferior to a reference treatment depends on model assumptions and choices of acceptable tolerance limits. Here, we consider a method that employs kernels to estimate the probability density functions of both the experimental and reference populations from two independent samples. Based on these densities, we introduce a quantity called the overlap coefficient or overlap measure. A bootstrap technique is helpful in exploring the distribution and variance empirically. We derive the distribution of this measure and define a hypothesis test that can be applied to the non-inferiority setting under some …
Eeg Signals Classification Using Lstm-Based Models And Majority Logic, James A. Orgeron
Eeg Signals Classification Using Lstm-Based Models And Majority Logic, James A. Orgeron
College of Graduate Studies: Theses & Dissertations
The study of elecroencephalograms (EEGs) has gained enormous interest in the last decade with the increase of computational power and availability of EEG signals collected from various human activities or produced during medical tests. The applicability of analyzing EEG signals ranges from helping impaired people communicate or move (using appropriate medical equipment) to understanding people's feelings and detecting diseases.
We proposed new methodology and models for analyzing and classifying EEG signals collected from individuals observing visual stimuli. Our models rely on powerful Long-Short Term Memory (LSTM) Neural Network models, which are currently the state of the art models for performing …
A Study On Developing Novel Methods For Relation Extraction, Darshini Mahendran
A Study On Developing Novel Methods For Relation Extraction, Darshini Mahendran
Theses and Dissertations
Relation Extraction (RE) is a task of Natural Language Processing (NLP) to detect and classify the relations between two entities. Relation extraction in the biomedical and scientific literature domain is challenging as text can contain multiple pairs of entities in the same instance. During the course of this research, we developed an RE framework (RelEx), which consists of five main RE paradigms: rule-based, machine learning-based, Convolutional Neural Network (CNN)-based, Bidirectional Encoder Representations from Transformers (BERT)-based, and Graph Convolutional Networks (GCNs)-based approaches. RelEx's rule-based approach uses co-location information of the entities to determine whether a relation exists between a selected entity …
Universal Design In Bci: Deep Learning Approaches For Adaptive Speech Brain-Computer Interfaces, Srdjan Lesaja
Universal Design In Bci: Deep Learning Approaches For Adaptive Speech Brain-Computer Interfaces, Srdjan Lesaja
Theses and Dissertations
In the last two decades, there have been many breakthrough advancements in non-invasive and invasive brain-computer interface (BCI) systems. However, the majority of BCI model designs still follow a paradigm whereby neural signals are preprocessed and task-related features extracted using static, and generally customized, data-independent designs. Such BCI designs commonly optimize narrow task performance over generalizability, adaptability, and robustness, which is not well suited to meeting individual user needs. If one day BCIs are to be capable of decoding our higher-order cognitive commands and conceptual maps, their designs will need to be adaptive architectures that will evolve and grow in …
A Unified Health Information System Framework For Connecting Data, People, Devices, And Systems, Wu He, Justin Zuopeng Zhang, Huanmei Wu, Wenzhuo Li, Sachin Shetty
A Unified Health Information System Framework For Connecting Data, People, Devices, And Systems, Wu He, Justin Zuopeng Zhang, Huanmei Wu, Wenzhuo Li, Sachin Shetty
Information Technology & Decision Sciences Faculty Publications
The COVID-19 pandemic has heightened the necessity for pervasive data and system interoperability to manage healthcare information and knowledge. There is an urgent need to better understand the role of interoperability in improving the societal responses to the pandemic. This paper explores data and system interoperability, a very specific area that could contribute to fighting COVID-19. Specifically, the authors propose a unified health information system framework to connect data, systems, and devices to increase interoperability and manage healthcare information and knowledge. A blockchain-based solution is also provided as a recommendation for improving the data and system interoperability in healthcare.
Finding The Best Predictors For Foot Traffic In Us Seafood Restaurants, Isabel Paige Beaulieu
Finding The Best Predictors For Foot Traffic In Us Seafood Restaurants, Isabel Paige Beaulieu
Honors Theses and Capstones
COVID-19 caused state and nation-wide lockdowns, which altered human foot traffic, especially in restaurants. The seafood sector in particular suffered greatly as there was an increase in illegal fishing, it is made up of perishable goods, it is seasonal in some places, and imports and exports were slowed. Foot traffic data is useful for business owners to have to know how much to order, how many employees to schedule, etc. One issue is that the data is very expensive, hard to get, and not available until months after it is recorded. Our goal is to not only find covariates that …
Towards Exchanging Wearable-Pghd With Ehrs: Developing A Standardized Information Model For Wearable-Based Patient Generated Health Data, Abdullahi Abubakar Kawu, Dympna O'Sullivan, Lucy Hederman
Towards Exchanging Wearable-Pghd With Ehrs: Developing A Standardized Information Model For Wearable-Based Patient Generated Health Data, Abdullahi Abubakar Kawu, Dympna O'Sullivan, Lucy Hederman
Articles
Wearables have become commonplace for tracking and making sense of patient lifestyle, wellbeing and health data. Most of this tracking is done by individuals outside of clinical settings, however some data from wearables may be useful in a clinical context. As such, wearables may be considered a prominent source of Patient Generated Health Data (PGHD). Studies have attempted to maximize the use of the data from wearables including integrating with Electronic Health Records (EHRs). However, usually a limited number of wearables are considered for integration and, in many cases, only one brand is investigated. In addition, we find limited studies …
Image-Data-Driven Deep Learning For Slope Stability Analysis, Behnam Azmoon
Image-Data-Driven Deep Learning For Slope Stability Analysis, Behnam Azmoon
Dissertations, Master's Theses and Master's Reports
Landslides cause major infrastructural issues, damage the environment, and cause socio-economic disruptions. Therefore, various slope stability analysis methods have been developed to evaluate the stability of slopes and the probability of their failure. This dissertation attempts to take advantage of the recent advancements in remote sensing and computer technology to implement a deep-learning-based landslide prediction method.
Considering the novelty of this approach, this dissertation leads with proof-of-concept studies to evaluate and establish the suitability of deep learning models for slope stability analysis. To achieve this, a simulated 2D dataset of slope images was created with different geometries and soil properties. …
Advanced Full-Text Search Based On Synonyms In Postgres, Joey Bodoia
Advanced Full-Text Search Based On Synonyms In Postgres, Joey Bodoia
CMC Senior Theses
This paper discusses the advanced full-text search queries based on synonyms that are supported in Chajda, which is a postgres extension and corresponding python library for highly multi-lingual full-text search in postgres. This discussion will include the motivations for using advanced queries based on synonyms, examples of how to use these advanced queries in Chajda, current limitiations of the advanced queries, and performance testing of the advanced queries.
Towards A Better Understanding Of The Development Of Collaborative Consumption Pricing, Funda Sarican
Towards A Better Understanding Of The Development Of Collaborative Consumption Pricing, Funda Sarican
2022
With the rise of social platforms, consumers started creating value by sharing human and physical resources. As a result, there has been an increase in collaborative consumption. My papers focus on the peer-to-peer travel accommodation service Airbnb, a popular online market for short-term housing rentals.
In paper one, an integrated framework for collaborative consumption pricing is developed by adopting the hedonic demand theory and leveraging the multilevel modeling method that accounts for various factors from nested data. This paper contributes to the research literature by extending the usage of multilevel modeling across cities to consider city effects and broadening the …
The Performance Optimization Of Asp Solving Based On Encoding Rewriting And Encoding Selection, Liu Liu
The Performance Optimization Of Asp Solving Based On Encoding Rewriting And Encoding Selection, Liu Liu
Theses and Dissertations--Computer Science
Answer set programming (ASP) has long been used for modeling and solving hard search problems. These problems are modeled in ASP as encodings, a collection of rules that declaratively describe the logic of the problem without explicitly listing how to solve it. It is common that the same problem has several different but equivalent encodings in ASP. Experience shows that the performance of these ASP encodings may vary greatly from instance to instance when processed by current state-of-the-art ASP grounder/solver systems. In particular, it is rarely the case that one encoding outperforms all others. Moreover, running an ASP system on …
Batch Normalization Preconditioning For Neural Network Training, Susanna Luisa Gertrude Lange
Batch Normalization Preconditioning For Neural Network Training, Susanna Luisa Gertrude Lange
Theses and Dissertations--Mathematics
Batch normalization (BN) is a popular and ubiquitous method in deep learning that has been shown to decrease training time and improve generalization performance of neural networks. Despite its success, BN is not theoretically well understood. It is not suitable for use with very small mini-batch sizes or online learning. In this work, we propose a new method called Batch Normalization Preconditioning (BNP). Instead of applying normalization explicitly through a batch normalization layer as is done in BN, BNP applies normalization by conditioning the parameter gradients directly during training. This is designed to improve the Hessian matrix of the loss …
Statistical Theory For Specialized Linear Regression Adjustment Methods Compared To Multiple Linear Regression In The Presence And Absence Of Interaction Effects, Leon Su
Theses and Dissertations--Statistics
When building models to investigate outcomes and variables of interest, researchers often want to adjust for other variables. There is a variety of ways that these adjustments are performed. In this work, we will consider four approaches to adjustment utilized by researchers in various fields. We will compare the efficacy of these methods to what we call the ”true model method”, fitting a multiple linear regression model in which adjustment variables are model covariates. Our goal is to show that these adjustment methods have inferior performance to the true model method by comparing model parameter estimates, power, type I error, …
Addressing Ascertainment Bias In The Study Of Cardiovascular Disease Burden In Opioid Use Disorders - Application Of Natural Language Processing Of Electronic Health Records, Jade Huang Singleton
Addressing Ascertainment Bias In The Study Of Cardiovascular Disease Burden In Opioid Use Disorders - Application Of Natural Language Processing Of Electronic Health Records, Jade Huang Singleton
Theses and Dissertations--Epidemiology and Biostatistics
In the United States, the prevalence of long-term exposure to opioid drugs, for both medically and nonmedically indicated purposes, has increased considerably since the mid-1990’s. Concerns have emerged about the potential health effects of opioid use. There is also growing interest in other possible connections with opioid use including cardiovascular disease. Electronic health records (EHR) contain information about patient care in the form of structured codes and unstructured notes. Natural language processing (NLP) provides a tool for processing unstructured textual data in EHR clinical notes and extracts useful information for research with structured formats. The purpose of this dissertation was …
Exploration Of Lignin-Based Superabsorbent Polymers (Hydrogels) For Soil Water Management And As A Carrier For Delivering Rhizobium Spp., Toby Adjuik
Theses and Dissertations--Biosystems and Agricultural Engineering
Superabsorbent polymers (hydrogels) as soil amendments may improve soil hydraulic properties and act as carrier materials beneficial to soil microorganisms. Researchers have mostly explored synthetic hydrogels which may not be environmentally sustainable. This dissertation focused on the development and application of lignin-based hydrogels as sustainable soil amendments. This dissertation also explores the development of pedotransfer transfer functions (PTFs) for predicting saturated hydraulic conductivity using statistical and machine learning methods with a publicly available large data set. A lignin-based hydrogel was synthesized, and its impact on soil water retention was determined in silt loam and loamy fine sand soils. Hydrogel treatment …
Largemouth Bass In The Upper Mississippi River: An Evaluation Of Management Strategies And Understanding Potential Factors Influencing Dynamic Rate Functions, Kylie Beth Sterling
Largemouth Bass In The Upper Mississippi River: An Evaluation Of Management Strategies And Understanding Potential Factors Influencing Dynamic Rate Functions, Kylie Beth Sterling
Graduate Theses/Dissertations
The Upper Mississippi River (UMR) supports ecologically and economically important commercial and recreational fisheries. One recreational fishery in the UMR is the Largemouth Bass fishery. Recreational fisheries can be effectively managed using information on population dynamics, though little is known about Largemouth Bass population dynamics in large river ecosystems. Therefore, the objectives of this study were to 1) evaluate recruitment, growth, and mortality of three Largemouth Bass populations in the UMR, specifically within Pools 4, 8, and 13, and 2) to use those estimates of recruitment, growth and mortality to inform exploitation models to evaluate best management practices for each …
Introduction To Analytics And Data Science, Leila Halawi, Amal Clarke, Kelly George
Introduction To Analytics And Data Science, Leila Halawi, Amal Clarke, Kelly George
Publications
Learning Objectives
-
Identify the different data types and measurements
-
Identify the need for data and data sources
-
Explain data partitioning and honest assessment
-
Identify the necessity of data preparation and curation
-
Identify the data preparation process
-
Explore SAS VIYA platform and import data
Automated Identification Of Missing Is-A Relations In The Human Phenotype Ontology, Maryamsadat Mohtashamian, Ran Hu, Rashmie Abeysinghe, Xubing Hao, Hua Xu, Licong Cui
Automated Identification Of Missing Is-A Relations In The Human Phenotype Ontology, Maryamsadat Mohtashamian, Ran Hu, Rashmie Abeysinghe, Xubing Hao, Hua Xu, Licong Cui
Faculty, Staff and Student Publications
Auditing the Human Phenotype Ontology (HPO) is necessary to provide accurate terminology for its use in clinical research. We investigate an approach leveraging the lexical features of concepts in HPO to identify missing IS-A relations among HPO concepts. We first model the names of HPO concepts as sets of words in lower case. Then, we generate two types of concept-pairs which have at least a single common word: (1) Linked concept-pairs generated from concept-pairs having an IS-A relation; (2) Unlinked concept-pairs generated from concept-pairs without an IS- A relation. Concept-pairs generate Derived Term Pairs (DTPs) emphasizing unique lexical information of …
Cost Of Care For Asylum Seekers And Refugees Entering The United States: The Case Of Volunteer Medical Providers In El Paso, Texas, Rigoberto I Delgado, Manuel De La Rosa, Marlon A Picado, Lisa Ayoub-Rodriguez, Celia E Gonzalez, Leopold Gemoets
Cost Of Care For Asylum Seekers And Refugees Entering The United States: The Case Of Volunteer Medical Providers In El Paso, Texas, Rigoberto I Delgado, Manuel De La Rosa, Marlon A Picado, Lisa Ayoub-Rodriguez, Celia E Gonzalez, Leopold Gemoets
Faculty, Staff and Student Publications
BACKGROUND: Between October 2018, and February 2020, the United States saw an unprecedented increase in the number of asylum seekers and refugees arriving unexpectedly at international crossings along the US-Mexico Border. Many of these migrants needed proper medical attention, and consequently created significant pressure on local health systems. In El Paso, Texas, volunteer clinicians, collaborating closely with religious organizations and non-governmental organizations, provided outpatient medical care for the new arrivals; the county hospital provided in-patient care at local tax payers' expense. The objective of this study was to estimate costs of healthcare services offered by these volunteers in order to …
Dat3m: A Data Tracker For Multi-Faceted Management Of Multi-Site Clinical Research Data Submission, Curation, Master Inventorying, And Sharing, Shiqiang Tao, Licong Cui, Wei-Chun Chou, Samden Lhatoo, Guo-Qiang Zhang
Dat3m: A Data Tracker For Multi-Faceted Management Of Multi-Site Clinical Research Data Submission, Curation, Master Inventorying, And Sharing, Shiqiang Tao, Licong Cui, Wei-Chun Chou, Samden Lhatoo, Guo-Qiang Zhang
Faculty, Staff and Student Publications
Managing research data is an important and challenging aspect of clinical studies, especially for multi-site collaboratives. To address this challenge, we designed, developed and deployed a multi-faceted, multi-level interactive data tracker (DaT3M) for multi-site clinical research data submission, curation, master inventorying, and sharing. Components of DaT3M include data overview, data portal, data status panel, data query engine, and data downloader. DaT3M managed clinical research data for the Center for SUDEP Research (CSR). The CSR instance of DaT3M includes 2,743 subjects from seven data contributing institutions, 7 data modalities and 10,678 data components: 3,398 Epilepsy Monitoring Unit reports, 3,440 electroencephalography recordings, …
Temporal Cohort Logic, Guo-Qiang Zhang, Xiaojin Li, Yan Huang, Licong Cui
Temporal Cohort Logic, Guo-Qiang Zhang, Xiaojin Li, Yan Huang, Licong Cui
Faculty, Staff and Student Publications
We introduce a new logic, called Temporal Cohort Logic (TCL), for cohort specification and discovery in clinical and population health research. TCL is created to fill a conceptual gap in formalizing temporal reasoning in biomedicine, in a similar role that temporal logics play for computer science and its applications. We provide formal syntax and semantics for TCL and illustrate the various logical constructs using examples related to human health. Relationships and distinctions with existing temporal logical frameworks are discussed. Applications in electronic health record (EHR) and in neurophysiological data resource are provided. Our approach differs from existing temporal logics, in …
Causal Inference Of Genetic Variants And Genes In Amyotrophic Lateral Sclerosis, Siyu Pan, Xinxuan Liu, Tianzi Liu, Zhongming Zhao, Yulin Dai, Yin-Ying Wang, Peilin Jia, Fan Liu
Causal Inference Of Genetic Variants And Genes In Amyotrophic Lateral Sclerosis, Siyu Pan, Xinxuan Liu, Tianzi Liu, Zhongming Zhao, Yulin Dai, Yin-Ying Wang, Peilin Jia, Fan Liu
Faculty, Staff and Student Publications
Amyotrophic lateral sclerosis (ALS) is a fatal progressive multisystem disorder with limited therapeutic options. Although genome-wide association studies (GWASs) have revealed multiple ALS susceptibility loci, the exact identities of causal variants, genes, cell types, tissues, and their functional roles in the development of ALS remain largely unknown. Here, we reported a comprehensive post-GWAS analysis of the recent large ALS GWAS (n = 80,610), including functional mapping and annotation (FUMA), transcriptome-wide association study (TWAS), colocalization (COLOC), and summary data-based Mendelian randomization analyses (SMR) in extensive multi-omics datasets. Gene property analysis highlighted inhibitory neuron 6, oligodendrocytes, and GABAergic neurons (Gad1/Gad2) as …
Deep Graph Convolutional Network For Us Birth Data Harmonization, Lishan Yu, Hamisu M Salihu, Deepa Dongarwar, Luyao Chen, Xiaoqian Jiang
Deep Graph Convolutional Network For Us Birth Data Harmonization, Lishan Yu, Hamisu M Salihu, Deepa Dongarwar, Luyao Chen, Xiaoqian Jiang
Faculty, Staff and Student Publications
In this paper, we developed a feasible and efficient deep-learning-based framework to combine the United States (US) natality data for the last five decades, with changing variables and factors, into a consistent database. We constructed a graph based on the property and elements of databases, including variables, and conducted a graph convolutional network (GCN) to learn the embeddings of variables on the constructed graph, where the learned embeddings implied the similarity of variables. Specifically, we devised a loss function with a slack margin and a banlist mechanism (for a random walk) to learn the desired structure (two nodes sharing more …