Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Computer Sciences (19)
- Artificial Intelligence and Robotics (13)
- Statistics and Probability (9)
- Medicine and Health Sciences (7)
- Life Sciences (6)
-
- Applied Statistics (5)
- Statistical Methodology (5)
- Theory and Algorithms (4)
- Biostatistics (3)
- Engineering (3)
- Public Health (3)
- Statistical Models (3)
- Applied Mathematics (2)
- Bioinformatics (2)
- Chemical Engineering (2)
- Databases and Information Systems (2)
- Epidemiology (2)
- Health Information Technology (2)
- Interprofessional Education (2)
- Mathematics (2)
- Medical Education (2)
- Numerical Analysis and Computation (2)
- Telemedicine (2)
- Analytical, Diagnostic and Therapeutic Techniques and Equipment (1)
- Biochemistry, Biophysics, and Structural Biology (1)
- Bioelectrical and Neuroengineering (1)
- Biomedical Engineering and Bioengineering (1)
- Categorical Data Analysis (1)
- Keyword
-
- Machine learning (10)
- Information Extraction (3)
- Natural Language Processing (3)
- Natural language processing (3)
- Concept Drift (2)
-
- Data science (2)
- Deep Learning (2)
- Deep learning (2)
- Dynamic heat map (2)
- Epidemiology (2)
- Interprofessional collaboration (2)
- Large Language Models (2)
- Machine Learning (2)
- Named Entity Recognition (2)
- Opioid (2)
- Overdose epidemic (2)
- Predictive model (2)
- Real-time surveillance (2)
- Unsupervised learning (2)
- Arctic shipping (1)
- Artificial Intelligence (1)
- B-spline (1)
- BERT (1)
- Big data (1)
- Biomedical entity linking (1)
- Causality (1)
- Chemical prediction (1)
- Classification Models (1)
- Clinical notes (1)
- Clinical research (1)
- Publication
- Publication Type
Articles 1 - 30 of 35
Full-Text Articles in Data Science
Remote Sensing To Detect Crop Damage By An Invasive Species, Cole Butler
Remote Sensing To Detect Crop Damage By An Invasive Species, Cole Butler
Biology and Medicine Through Mathematics Conference
No abstract provided.
An Integrated Data-Driven Framework For Arctic Shipping: Analyzing Vessel Speed, Environmental And Ecological Factors Through Innovative Statistical Spatio-Temporal Methods, Inverse Optimization And Machine Learning, Mauli Pant
Theses and Dissertations
This dissertation develops an integrated data-driven framework to analyze vessel navigation and ecological risk in the United States Arctic from 2010 to 2019. As environmental change and maritime activity increase in the region, understanding how vessels respond to dynamic conditions and how those responses interact with marine ecosystems has become increasingly important. A central theme of this dissertation is the treatment of vessel speed as both an observed outcome and a decision variable reflecting trade- offs among operational, environmental, and ecological factors. The first chapter develops a predictive framework for vessel speed over ground (SOG) using Gaussian Process Boosting (GPBoost), …
Leveraging Distributed Semantics From Deep Learning Architectures For Literature-Based Discovery, Clint A. Cuffy
Leveraging Distributed Semantics From Deep Learning Architectures For Literature-Based Discovery, Clint A. Cuffy
Theses and Dissertations
Literature-based discovery (LBD) is a scientific process that introduces methods to automatically identify novel insights between non-interacting sets of literature. To date, numerous statistical and machine learning-based methods have been applied in the biomedical domain to find treatments for diseases such as Raynaud's disease, Parkinson's disease, and Multiple Sclerosis. However, the lack of standardized practices and creation of bespoke methodologies produces a scenario where the adoption of LBD remains challenging in real-world systems. Our work addresses these concerns through the improvement of five critical areas: 1) error propagation within LBD's a priori dependent tasks, 2) exploring the integration of modern …
Beyond Homogeneity: Exploring Causal Heterogeneity In Psychopathology, Philip B. Vinh
Beyond Homogeneity: Exploring Causal Heterogeneity In Psychopathology, Philip B. Vinh
Theses and Dissertations
Traditional models in psychiatric research often impose assumptions of causal homogeneity, treating population-level associations as reflective of uniform underlying mechanisms. This dissertation challenges that assumption by introducing statistical and machine learning frameworks designed to detect and model causal heterogeneity in the development of psychopathology. Central to this approach is the advancement of finite mixture structural equation modeling (FM-SEM) to identify latent subgroups characterized by distinct, and sometimes opposing, causal pathways.
The dissertation comprises three integrated empirical studies. The first introduces mixDoC, a finite mixture extension of the classical Direction of Causation (DoC) model applied to twin data, enabling the detection …
Multitask Learning For Named Entity Recognition And Relationship Extraction, Adrienne D. Hembrick
Multitask Learning For Named Entity Recognition And Relationship Extraction, Adrienne D. Hembrick
Theses and Dissertations
Information Extraction (IE) is a fundamental task in Natural Language Processing (NLP), involving the identification of structured information from unstructured text. Two core components of IE—Named Entity Recognition (NER) and Relation Extraction (RE)—are widely used to extract key concepts and the relationships between them across various domains. However, the sequential dependency of RE on the output of NER makes it vulnerable to error propagation: inaccuracies in entity recognition can negatively affect downstream relation extraction.
To mitigate this issue, Multitask Learning (MTL) has been proposed as an approach that jointly models NER and RE, aiming to improve overall performance and reduce …
Optimal Data Splitting Methods, Sujay Mudalgi
Optimal Data Splitting Methods, Sujay Mudalgi
Theses and Dissertations
In predictive modeling, effective data splitting is crucial for creating statistically representative training and validation sets. The state-of-the-art data splitting methods are based on minimizing the energy distance between the split subsets. However, there are a number of limitations in the existing methods, which this dissertation aims to address. First, the existing methods were computationally inefficient. Thus, Chapter 2 proposes a method to scale up these approaches for big data. Here, we introduce scalable Twinning (s-Twinning), which significantly improves the execution speed of data splitting without sacrificing accuracy. Second, the existing methods did not consider the predictive relationship in the …
Machine Learning Models Leveraging Patient-Similarity And Clinical Temporality For Disease Prognoses, Ahmad F. Al Musawi
Machine Learning Models Leveraging Patient-Similarity And Clinical Temporality For Disease Prognoses, Ahmad F. Al Musawi
Theses and Dissertations
Electronic Health Records (EHRs) constitute a comprehensive and high-dimensional repository of clinical data, encompassing a wide array of patient-level information such as diagnoses, procedures, medications, laboratory results, and unstructured clinical narratives. These data hold immense potential for advancing predictive modeling in healthcare, including tasks such as disease progression modeling, hospital readmission prediction, and length of stay (LoS) estimation. However, the intrinsic complexity of EHR data—manifested in its heterogeneity, sparsity, and temporal dynamics—poses significant analytical challenges that limit the generalizability and interpretability of conventional machine learning models. Recent methodological advancements in deep learning and graph-based learning, particularly Graph Neural Networks (GNNs), …
Learning From Non-Stationary Data Streams, Gabriel Jonas Aguiar
Learning From Non-Stationary Data Streams, Gabriel Jonas Aguiar
Theses and Dissertations
The rapid growth of data from sources such as mobile applications, sensors, and network monitoring has increased the need for machine learning algorithms capable of handling non-stationary data streams. However, learning from such streams presents significant challenges due to their evolving nature and the presence of concept drift. One of the most complex issues is learning from imbalanced data streams, where shifting data distributions, combined with feature space drifts, complicate continuous adaptation. These challenges become even more pronounced in multi-class scenarios, which are common in real-world applications. Detecting concept drift in such contexts is particularly demanding, as it requires tracking …
Dual-Domain Clustering Of Spatiotemporal Infectious Disease Data, Samuel R. Thornton, Erin C.S. Acquesta, Patrick D. Finley, Mansoor A. Haider
Dual-Domain Clustering Of Spatiotemporal Infectious Disease Data, Samuel R. Thornton, Erin C.S. Acquesta, Patrick D. Finley, Mansoor A. Haider
Biology and Medicine Through Mathematics Conference
No abstract provided.
Adaptive Multi-Label Classification On Drifting Data Streams, Martha Roseberry
Adaptive Multi-Label Classification On Drifting Data Streams, Martha Roseberry
Theses and Dissertations
Drifting data streams and multi-label data are both challenging problems. When multi-label data arrives as a stream, the challenges of both problems must be addressed along with additional challenges unique to the combined problem. Algorithms must be fast and flexible, able to match both the speed and evolving nature of the stream. We propose four methods for learning from multi-label drifting data streams. First, a multi-label k Nearest Neighbors with Self Adjusting Memory (ML-SAM-kNN) exploits short- and long-term memories to predict the current and evolving states of the data stream. Second, a punitive k nearest neighbors algorithm with a self-adjusting …
Developing Machine Learning And Time-Series Analysis Methods With Applications In Diverse Fields, Muhammed Aljifri
Developing Machine Learning And Time-Series Analysis Methods With Applications In Diverse Fields, Muhammed Aljifri
Theses and Dissertations
This dissertation introduces methodologies that combine machine learning models with time-series analysis to tackle data analysis challenges in varied fields. The first study enhances the traditional cumulative sum control charts with machine learning models to leverage their predictive power for better detection of process shifts, applying this advanced control chart to monitor hospital readmission rates. The second project develops multi-layer models for predicting chemical concentrations from ultraviolet-visible spectroscopy data, specifically addressing the challenge of analyzing chemicals with a wide range of concentrations. The third study presents a new method for detecting multiple changepoints in autocorrelated ordinal time series, using the …
Large Language Models, Prompting, And Synthetic Data Generation For Continual Named Entity Recognition, Charles I. Cutler
Large Language Models, Prompting, And Synthetic Data Generation For Continual Named Entity Recognition, Charles I. Cutler
Theses and Dissertations
With the ever-growing amount of textual data, the task of Named Entity Recognition (NER) is vital to Natural Language Processing (NLP), a field which focuses on enabling computers to understand and manipulate human language. NER enables the extraction of information from unstructured text. Accurate information extraction is crucial for applications ranging from information retrieval to systems for question-answering. To ensure that NER models are robust to changes in data distributions and capable of recognizing new entity types, one may consider expanding the capabilities of an existing model. Continual learning is a paradigm within machine learning. It studies the objective of …
Penalized Interpolating B-Splines And Their Applications, Kylee L. Hartman-Caballero
Penalized Interpolating B-Splines And Their Applications, Kylee L. Hartman-Caballero
Theses and Dissertations
One of the most studied data analysis techniques in Numerical Analysis is interpolation. Interpolation is used in a variety of fields, namely computer graphic design and biomedical research. Among interpolation techniques, cubic splines have been viewed as the standard since at least the 1960s, due to their ease of computation, numerical stability, and the relative smoothness of the interpolating curve. However, cubic splines have notable drawbacks, such as their lack of local control and necessary knowledge of boundary conditions. Arguably a more versatile interpolation technique is the use of B-splines. B-splines, a relative of Bézier curves, allow local control through …
Title: I: L1-Norm Matrix Completion For Recommender Systems Ii: Conjecturing-Based Classification, Fatemeh Valizadeh Gamchi
Title: I: L1-Norm Matrix Completion For Recommender Systems Ii: Conjecturing-Based Classification, Fatemeh Valizadeh Gamchi
Theses and Dissertations
Recommendation systems are essential for providing personalized user experiences, but their performance can be affected by outliers especially in traditional collaborative filtering methods that use the L2-norm. To address this challenge, we developed two new algorithms, SharpEl1rs and SharpEl1rs-Impute, based on the L1-norm to improve resistance against extreme values and effectively handle missing data. Our experimental setting was designed to compare these proposed methods with existing techniques. Then our algorithms are applied to real datasets to assess their performance, with findings indicating that our proposed models offer improved accuracy in some cases and solid performance in others for industrial-scale recommendation …
Analytical Approach For Monitoring The Behavior Of Patients With Pancreatic Adenocarcinoma At Different Stages As A Function Of Time, Aditya Chakaborty Dr, Chris P. Tsokos Dr
Analytical Approach For Monitoring The Behavior Of Patients With Pancreatic Adenocarcinoma At Different Stages As A Function Of Time, Aditya Chakaborty Dr, Chris P. Tsokos Dr
Biology and Medicine Through Mathematics Conference
No abstract provided.
Face Anti-Spoofing And Deep Learning Based Unsupervised Image Recognition Systems, Enoch Solomon
Face Anti-Spoofing And Deep Learning Based Unsupervised Image Recognition Systems, Enoch Solomon
Theses and Dissertations
One of the main problems of a supervised deep learning approach is that it requires large amounts of labeled training data, which are not always easily available. This PhD dissertation addresses the above-mentioned problem by using a novel unsupervised deep learning face verification system called UFace, that does not require labeled training data as it automatically, in an unsupervised way, generates training data from even a relatively small size of data. The method starts by selecting, in unsupervised way, k-most similar and k-most dissimilar images for a given face image. Moreover, this PhD dissertation proposes a new loss function to …
Inferring Dynamics Of Biological Systems, Tracey G. Oellerich
Inferring Dynamics Of Biological Systems, Tracey G. Oellerich
Biology and Medicine Through Mathematics Conference
No abstract provided.
A Study On Developing Novel Methods For Relation Extraction, Darshini Mahendran
A Study On Developing Novel Methods For Relation Extraction, Darshini Mahendran
Theses and Dissertations
Relation Extraction (RE) is a task of Natural Language Processing (NLP) to detect and classify the relations between two entities. Relation extraction in the biomedical and scientific literature domain is challenging as text can contain multiple pairs of entities in the same instance. During the course of this research, we developed an RE framework (RelEx), which consists of five main RE paradigms: rule-based, machine learning-based, Convolutional Neural Network (CNN)-based, Bidirectional Encoder Representations from Transformers (BERT)-based, and Graph Convolutional Networks (GCNs)-based approaches. RelEx's rule-based approach uses co-location information of the entities to determine whether a relation exists between a selected entity …
Universal Design In Bci: Deep Learning Approaches For Adaptive Speech Brain-Computer Interfaces, Srdjan Lesaja
Universal Design In Bci: Deep Learning Approaches For Adaptive Speech Brain-Computer Interfaces, Srdjan Lesaja
Theses and Dissertations
In the last two decades, there have been many breakthrough advancements in non-invasive and invasive brain-computer interface (BCI) systems. However, the majority of BCI model designs still follow a paradigm whereby neural signals are preprocessed and task-related features extracted using static, and generally customized, data-independent designs. Such BCI designs commonly optimize narrow task performance over generalizability, adaptability, and robustness, which is not well suited to meeting individual user needs. If one day BCIs are to be capable of decoding our higher-order cognitive commands and conceptual maps, their designs will need to be adaptive architectures that will evolve and grow in …
Temporal Disambiguation Of Relative Temporal Expressions In Clinical Texts Using Temporally Fine-Tuned Contextual Word Embeddings., Amy L. Olex
Theses and Dissertations
Temporal reasoning is the ability to extract and assimilate temporal information to reconstruct a series of events such that they can be reasoned over to answer questions involving time. Temporal reasoning in the clinical domain is challenging due to specialized medical terms and nomenclature, shorthand notation, fragmented text, a variety of writing styles used by different medical units, redundancy of information that has to be reconciled, and an increased number of temporal references as compared to general domain texts. Work in the area of clinical temporal reasoning has progressed, but the current state-of-the-art still has a ways to go before …
Estimating Weighted Panel Sizes For Primary Care Providers: An Assessment Of Clustering And Novel Methods Of Panel Size Estimation On Electronic Medical Records, Martin A. Lavallee
Estimating Weighted Panel Sizes For Primary Care Providers: An Assessment Of Clustering And Novel Methods Of Panel Size Estimation On Electronic Medical Records, Martin A. Lavallee
Theses and Dissertations
Primary Care is on the frontlines of healthcare, thus they see the most diverse set of patients. In order to achieve high functioning primary care, a practice must establish empanelment, the pairing of patients to providers. Enumeration of empanelment, or estimating panel sizes, helps ensure that the demands of the patients demand the supply of providers and optimize the balance of primary care resources to improve quality of care. Further we can adjust panel sizes by using patient-level data on healthcare utilization and complexity extracted from the electronic medial record to determine the amount of care or burden of work …
Multi-Modality Automatic Lung Tumor Segmentation Method Using Deep Learning And Radiomics, Siqiu Wang
Multi-Modality Automatic Lung Tumor Segmentation Method Using Deep Learning And Radiomics, Siqiu Wang
Theses and Dissertations
Delineation of the tumor volume is the initial and fundamental step in the radiotherapy planning process. The current clinical practice of manual delineation is time-consuming and suffers from observer variability. This work seeks to develop an effective automatic framework to produce clinically usable lung tumor segmentations. First, to facilitate the development and validation of our methodology, an expansive database of planning CTs, diagnostic PETs, and manual tumor segmentations was curated, and an image registration and preprocessing pipeline was established. Then a deep learning neural network was constructed and optimized to utilize dual-modality PET and CT images for lung tumor segmentation. …
Incorporating Ontological Information In Biomedical Entity Linking Of Phrases In Clinical Text, Evan French
Incorporating Ontological Information In Biomedical Entity Linking Of Phrases In Clinical Text, Evan French
Theses and Dissertations
Biomedical Entity Linking (BEL) is the task of mapping spans of text within biomedical documents to normalized, unique identifiers within an ontology. Translational application of BEL on clinical notes has enormous potential for augmenting discretely captured data in electronic health records, but the existing paradigm for evaluating BEL systems developed in academia is not well aligned with real-world use cases. In this work, we demonstrate a proof of concept for incorporating ontological similarity into the training and evaluation of BEL systems to begin to rectify this misalignment. This thesis has two primary components: 1) a comprehensive literature review and 2) …
Computational Analysis Of Drug Targets And Prediction Of Protein-Compound Interactions, Sina Ghadermarzi
Computational Analysis Of Drug Targets And Prediction Of Protein-Compound Interactions, Sina Ghadermarzi
Theses and Dissertations
Computational prediction of compound-protein interactions generated a substantial amount of interest in the recent years owing to the importance of the knowledge of these interaction for drug discovery and drug repurposing efforts. Research suggests that the currently known drug targets constitute only a fraction of a complete set of drug targets, limiting our ability to identify suitable targets to develop new drugs or to repurpose current drugs for new diseases. These efforts are further thwarted by our limited knowledge of protein-drug (and more generally protein-compound) interactions, where only a subset of drug targets is typically known for the currently used …
Statistical Approaches For Estimation And Comparison Of Brain Functional Connectivity, Jifang Zhao
Statistical Approaches For Estimation And Comparison Of Brain Functional Connectivity, Jifang Zhao
Theses and Dissertations
Drug addiction can lead to many health-related problems and social concerns. Functional connectivity obtained from functional magnetic resonance imaging (fMRI) data promotes a variety of fundamental understandings in such association. Due to its complex correlation structure and large dimensionality, the modeling and analysis of the functional connectivity from neuroimage are challenging. By proposing a spatio-temporal model for multi-subject neuroimage data, we incorporate voxel-level spatio-temporal dependencies of whole-brain measurements to improve the accuracy of statistical inference. To tackle large-scale spatio-temporal neuroimage data, we develop a computationally efficient algorithm to estimate the parameters. Our method is used to identify functional connectivity and …
Information Architecture For A Chemical Modeling Knowledge Graph, Adam R. Luxon
Information Architecture For A Chemical Modeling Knowledge Graph, Adam R. Luxon
Theses and Dissertations
Machine learning models for chemical property predictions are high dimension design challenges spanning multiple disciplines. Free and open-source software libraries have streamlined the model implementation process, but the design complexity remains. In order better navigate and understand the machine learning design space, model information needs to be organized and contextualized. In this work, instances of chemical property models and their associated parameters were stored in a Neo4j property graph database. Machine learning model instances were created with permutations of dataset, learning algorithm, molecular featurization, data scaling, data splitting, hyperparameters, and hyperparameter optimization techniques. The resulting graph contains over 83,000 nodes …
Reliable And Interpretable Machine Learning For Modeling Physical And Cyber Systems, Daniel L. Marino Lizarazo
Reliable And Interpretable Machine Learning For Modeling Physical And Cyber Systems, Daniel L. Marino Lizarazo
Theses and Dissertations
Over the past decade, Machine Learning (ML) research has predominantly focused on building extremely complex models in order to improve predictive performance. The idea was that performance can be improved by adding complexity to the models. This approach proved to be successful in creating models that can approximate highly complex relationships while taking advantage of large datasets. However, this approach led to extremely complex black-box models that lack reliability and are difficult to interpret. By lack of reliability, we specifically refer to the lack of consistent (unpredictable) behavior in situations outside the training data. Lack of interpretability refers to the …
Improving Space Efficiency Of Deep Neural Networks, Aliakbar Panahi
Improving Space Efficiency Of Deep Neural Networks, Aliakbar Panahi
Theses and Dissertations
Language models employ a very large number of trainable parameters. Despite being highly overparameterized, these networks often achieve good out-of-sample test performance on the original task and easily fine-tune to related tasks. Recent observations involving, for example, intrinsic dimension of the objective landscape and the lottery ticket hypothesis, indicate that often training actively involves only a small fraction of the parameter space. Thus, a question remains how large a parameter space needs to be in the first place — the evidence from recent work on model compression, parameter sharing, factorized representations, and knowledge distillation increasingly shows that models can be …
Learning From Multi-Class Imbalanced Big Data With Apache Spark, William C. Sleeman Iv
Learning From Multi-Class Imbalanced Big Data With Apache Spark, William C. Sleeman Iv
Theses and Dissertations
With data becoming a new form of currency, its analysis has become a top priority in both academia and industry, furthering advancements in high-performance computing and machine learning. However, these large, real-world datasets come with additional complications such as noise and class overlap. Problems are magnified when with multi-class data is presented, especially since many of the popular algorithms were originally designed for binary data. Another challenge arises when the number of examples are not evenly distributed across all classes in a dataset. This often causes classifiers to favor the majority class over the minority classes, leading to undesirable results …
K-Nearest Neighbors Density-Based Clustering, Avory C. Bryant
K-Nearest Neighbors Density-Based Clustering, Avory C. Bryant
Theses and Dissertations
Traditional density-based clustering approaches rely on a distance-based parameter to define data connectivity and density. However, an appropriate value of this parameter can be difficult to determine as it is highly dependent on the underlying distribution of the data. In particular, distribution parameters affect the scale of inter-group distances (e.g., variance); this dependence leads to a well-known inability to simultaneously detect clusters at varying levels of density. In this work, connectivity and density are defined according to the rank-order induced by the distance metric (i.e., invariant to the expected scale of the distances). Connectivity by k-nearest neighbors and density by …