Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

Algorithms

Discipline
Institution
Publication Year
Publication
Publication Type

Articles 1 - 30 of 38

Full-Text Articles in Data Science

Machine Learning-Enabled Chemical Ecology For Integrated Pest Management: From Volatiles To Field Applications, Steve B.S. Baleba, Victor O. Omondi, Pascal Aigbedion-Atalor, Emmanuel Peter, Souleymane Diallo, Komi Mensah Agboka Sep 2026

Machine Learning-Enabled Chemical Ecology For Integrated Pest Management: From Volatiles To Field Applications, Steve B.S. Baleba, Victor O. Omondi, Pascal Aigbedion-Atalor, Emmanuel Peter, Souleymane Diallo, Komi Mensah Agboka

All Peer-Reviewed Publications

Machine learning is transforming chemical ecology by accelerating the discovery and deployment of semiochemical-based tools for precision pest management. These advances are particularly important in the face of climate change, pesticide resistance, and the growing need for sustainable agricultural intensification. This review synthesizes how machine learning can be applied across the semiochemical discovery and implementation pipeline, from chemical signal detection to field deployment and decision support for integrated pest management. We review major machine learning approaches and demonstrate how they extract biologically relevant information from high-dimensional chemical, electrophysiological, behavioral, sensor, and field datasets. These methods accelerate semiochemical discovery, prioritize candidate …


Artificial Intelligence In Medicine: Barriers, Solutions, And Strategies, Anil Harrison, Melissa Stradley Moreno, Caroline E. Williams, Munevver Mine Subasi, Ersoy Subasi Apr 2026

Artificial Intelligence In Medicine: Barriers, Solutions, And Strategies, Anil Harrison, Melissa Stradley Moreno, Caroline E. Williams, Munevver Mine Subasi, Ersoy Subasi

HCA Healthcare Journal of Medicine

The integration of artificial intelligence (AI) and machine learning (ML) into health care holds the potential to revolutionize patient care by enhancing clinical decision-making, improving diagnostic accuracy, and reducing costs. Despite this promise, adoption remains limited due to a range of technical, regulatory, educational, and cultural barriers. This paper examines these challenges and proposes strategies to support safe and effective implementation of AI in clinical practice.

Key barriers include the lack of model interpretability, often referred to as the "black box" problem, which undermines clinician trust and accountability in clinical settings, evolving regulatory frameworks and unresolved questions surrounding liability, and …


Precision Phenotyping For Curating Research Cohorts Of Patients With Unexplained Post-Acute Sequelae Of Covid-19, Alaleh Azhir, Jonas Hügel, Jiazi Tian, Jingya Cheng, Ingrid V Bassett, Douglas S Bell, Elmer V Bernstam, Maha R Farhat, Darren W Henderson, Emily S Lau, Michele Morris, Yevgeniy R Semenov, Virginia A Triant, Shyam Visweswaran, Zachary H Strasser, Jeffrey G Klann, Shawn N Murphy, Hossein Estiri Mar 2025

Precision Phenotyping For Curating Research Cohorts Of Patients With Unexplained Post-Acute Sequelae Of Covid-19, Alaleh Azhir, Jonas Hügel, Jiazi Tian, Jingya Cheng, Ingrid V Bassett, Douglas S Bell, Elmer V Bernstam, Maha R Farhat, Darren W Henderson, Emily S Lau, Michele Morris, Yevgeniy R Semenov, Virginia A Triant, Shyam Visweswaran, Zachary H Strasser, Jeffrey G Klann, Shawn N Murphy, Hossein Estiri

Faculty, Staff and Student Publications

BACKGROUND: Scalable identification of patients with post-acute sequelae of COVID-19 (PASC) is challenging due to a lack of reproducible precision phenotyping algorithms, which has led to suboptimal accuracy, demographic biases, and underestimation of the PASC.

METHODS: In a retrospective case-control study, we developed a precision phenotyping algorithm for identifying cohorts of patients with PASC. We used longitudinal electronic health records data from over 295,000 patients from 14 hospitals and 20 community health centers in Massachusetts. The algorithm employs an attention mechanism to simultaneously exclude sequelae that prior conditions can explain and include infection-associated chronic conditions. We performed independent chart reviews …


Proxy Panels Enable Privacy-Aware Outsourcing Of Genotype Imputation, Degui Zhi, Xiaoqian Jiang, Arif Harmanci Feb 2025

Proxy Panels Enable Privacy-Aware Outsourcing Of Genotype Imputation, Degui Zhi, Xiaoqian Jiang, Arif Harmanci

Faculty, Staff and Student Publications

One of the major challenges in genomic data sharing is protecting participants' privacy in collaborative studies and in cases when genomic data are outsourced to perform analysis tasks, for example, genotype imputation services and federated collaborations genomic analysis. Although numerous cryptographic methods have been developed, these methods may not yet be practical for population-scale tasks in terms of computational requirements, rely on high-level expertise in security, and require each algorithm to be implemented from scratch. In this study, we focus on outsourcing of genotype imputation, a fundamental task that utilizes population-level reference panels, and develop protocols that rely on using …


Crisprofft: Comprehensive Database Of Crispr/Cas Off-Targets, Grant Wang, Xiaona Liu, Aoqi Wang, Jianguo Wen, Pora Kim, Qianqian Song, Xiaona Liu, Xiaobo Zhou Jan 2025

Crisprofft: Comprehensive Database Of Crispr/Cas Off-Targets, Grant Wang, Xiaona Liu, Aoqi Wang, Jianguo Wen, Pora Kim, Qianqian Song, Xiaona Liu, Xiaobo Zhou

Faculty, Staff and Student Publications

The CRISPR (clustered regularly interspaced short palindromic repeats)/Cas (CRISPR-associated protein) programmable nuclease system continues to evolve, with in vivo therapeutic gene editing increasingly applied in clinical settings. However, off-target effects remain a significant challenge, hindering its broader clinical application. To enhance the development of gene-editing therapies and the accuracy of prediction algorithms, we developed CRISPRoffT (https://ccsm.uth.edu/CRISPRoffT/). Users can access a comprehensive repository of off-target regions predicted and validated by a diverse range of technologies across various cell lines, Cas enzyme variants, engineered sgRNAs (single guide RNAs) and CRISPR editing systems. CRISPRoffT integrates results of off-target analysis from 74 studies, encompassing …


A Bayesian Deep Segmentation Framework For Glioblastoma Tumor Segmentation Using Follow-Up Mris, Tanjida Kabir, Kang-Lin Hsieh, Luis Nunez, Yu-Chun Hsu, Juan C Rodriguez Quintero, Octavio Arevalo, Kangyi Zhao, Jay-Jiguang Zhu, Roy F Riascos, Mahboubeh Madadi, Xiaoqian Jiang, Shayan Shams Jan 2025

A Bayesian Deep Segmentation Framework For Glioblastoma Tumor Segmentation Using Follow-Up Mris, Tanjida Kabir, Kang-Lin Hsieh, Luis Nunez, Yu-Chun Hsu, Juan C Rodriguez Quintero, Octavio Arevalo, Kangyi Zhao, Jay-Jiguang Zhu, Roy F Riascos, Mahboubeh Madadi, Xiaoqian Jiang, Shayan Shams

Faculty, Staff and Student Publications

Background: Glioblastoma (GBM) is the most common malignant brain tumor with an abysmal prognosis. Since complete tumor cell removal is impossible due to the infiltrative nature of GBM, accurate measurement is paramount for GBM assessment. Preoperative magnetic resonance images (MRIs) are crucial for initial diagnosis and surgical planning, while follow-up MRIs are vital for evaluating treatment response. The structural changes in the brain caused by surgical and therapeutic measures create significant differences between preoperative and follow-up MRIs. In clinical research, advanced deep learning models trained on preoperative MRIs are often applied to assess follow-up scans, but their effectiveness in this …


De-Identification Is Not Enough: A Comparison Between De-Identified And Synthetic Clinical Notes, Atiquer Rahman Sarkar, Yao-Shun Chuang, Noman Mohammed, Xiaoqian Jiang Nov 2024

De-Identification Is Not Enough: A Comparison Between De-Identified And Synthetic Clinical Notes, Atiquer Rahman Sarkar, Yao-Shun Chuang, Noman Mohammed, Xiaoqian Jiang

Faculty, Staff and Student Publications

For sharing privacy-sensitive data, de-identification is commonly regarded as adequate for safeguarding privacy. Synthetic data is also being considered as a privacy-preserving alternative. Recent successes with numerical and tabular data generative models and the breakthroughs in large generative language models raise the question of whether synthetically generated clinical notes could be a viable alternative to real notes for research purposes. In this work, we demonstrated that (i) de-identification of real clinical notes does not protect records against a membership inference attack, (ii) proposed a novel approach to generate synthetic clinical notes using the current state-of-the-art large language models, (iii) evaluated …


Generalizing Parkinson’S Disease Detection Using Keystroke Dynamics: A Self-Supervised Approach, Shikha Tripathi, Alejandro Acien, Ashley A Holmes, Teresa Arroyo-Gallego, Luca Giancardo May 2024

Generalizing Parkinson’S Disease Detection Using Keystroke Dynamics: A Self-Supervised Approach, Shikha Tripathi, Alejandro Acien, Ashley A Holmes, Teresa Arroyo-Gallego, Luca Giancardo

Faculty, Staff and Student Publications

Objective: Passive monitoring of touchscreen interactions generates keystroke dynamic signals that can be used to detect and track neurological conditions such as Parkinson's disease (PD) and psychomotor impairment with minimal burden on the user. However, this typically requires datasets with clinically confirmed labels collected in standardized environments, which is challenging, especially for a large subject pool. This study validates the efficacy of a self-supervised learning method in reducing the reliance on labels and evaluates its generalizability.

Materials and methods: We propose a new type of self-supervised loss combining Barlow Twins loss, which attempts to create similar feature representations with reduced …


Development And Validation Of A Rule-Based Algorithm To Identify Periodontal Diagnosis Using Structured Electronic Health Record Data, Bunmi Tokede, Ryan Brandon, Chun-Teh Lee, Guo-Hao Lin, Joel White, Alfa Yansane, Xiaoqian Jiang, Elsbeth Kalenderian, Muhammad Walji May 2024

Development And Validation Of A Rule-Based Algorithm To Identify Periodontal Diagnosis Using Structured Electronic Health Record Data, Bunmi Tokede, Ryan Brandon, Chun-Teh Lee, Guo-Hao Lin, Joel White, Alfa Yansane, Xiaoqian Jiang, Elsbeth Kalenderian, Muhammad Walji

Faculty, Staff and Student Publications

AIM: To develop and validate an automated electronic health record (EHR)-based algorithm to suggest a periodontal diagnosis based on the 2017 World Workshop on the Classification of Periodontal Diseases and Conditions.

MATERIALS AND METHODS: Using material published from the 2017 World Workshop, a tool was iteratively developed to suggest a periodontal diagnosis based on clinical data within the EHR. Pertinent clinical data included clinical attachment level (CAL), gingival margin to cemento-enamel junction distance, probing depth, furcation involvement (if present) and mobility. Chart reviews were conducted to confirm the algorithm's ability to accurately extract clinical data from the EHR, and then …


Generalizable Pipeline For Constructing Hiv Risk Prediction Models Across Electronic Health Record Systems, Sarah B May, Thomas P Giordano, Assaf Gottlieb Feb 2024

Generalizable Pipeline For Constructing Hiv Risk Prediction Models Across Electronic Health Record Systems, Sarah B May, Thomas P Giordano, Assaf Gottlieb

Faculty, Staff and Student Publications

OBJECTIVE: The HIV epidemic remains a significant public health issue in the United States. HIV risk prediction models could be beneficial for reducing HIV transmission by helping clinicians identify patients at high risk for infection and refer them for testing. This would facilitate initiation on treatment for those unaware of their status and pre-exposure prophylaxis for those uninfected but at high risk. Existing HIV risk prediction algorithms rely on manual construction of features and are limited in their application across diverse electronic health record systems. Furthermore, the accuracy of these models in predicting HIV in females has thus far been …


An Open Natural Language Processing (Nlp) Framework For Ehr-Based Clinical Research: A Case Demonstration Using The National Covid Cohort Collaborative (N3c), Sijia Liu, Andrew Wen, Liwei Wang, Huan He, Sunyang Fu, Robert Miller, Andrew Williams, Daniel Harris, Ramakanth Kavuluru, Mei Liu, Noor Abu-El-Rub, Dalton Schutte, Rui Zhang, Masoud Rouhizadeh, John D Osborne, Yongqun He, Umit Topaloglu, Stephanie S Hong, Joel H Saltz, Thomas Schaffter, Emily Pfaff, Christopher G Chute, Tim Duong, Melissa A Haendel, Rafael Fuentes, Peter Szolovits, Hua Xu, Hongfang Liu Nov 2023

An Open Natural Language Processing (Nlp) Framework For Ehr-Based Clinical Research: A Case Demonstration Using The National Covid Cohort Collaborative (N3c), Sijia Liu, Andrew Wen, Liwei Wang, Huan He, Sunyang Fu, Robert Miller, Andrew Williams, Daniel Harris, Ramakanth Kavuluru, Mei Liu, Noor Abu-El-Rub, Dalton Schutte, Rui Zhang, Masoud Rouhizadeh, John D Osborne, Yongqun He, Umit Topaloglu, Stephanie S Hong, Joel H Saltz, Thomas Schaffter, Emily Pfaff, Christopher G Chute, Tim Duong, Melissa A Haendel, Rafael Fuentes, Peter Szolovits, Hua Xu, Hongfang Liu

Faculty, Staff and Student Publications

Despite recent methodology advancements in clinical natural language processing (NLP), the adoption of clinical NLP models within the translational research community remains hindered by process heterogeneity and human factor variations. Concurrently, these factors also dramatically increase the difficulty in developing NLP models in multi-site settings, which is necessary for algorithm robustness and generalizability. Here, we reported on our experience developing an NLP solution for Coronavirus Disease 2019 (COVID-19) signs and symptom extraction in an open NLP framework from a subset of sites participating in the National COVID Cohort (N3C). We then empirically highlight the benefits of multi-site data for both …


Minimal Positional Substring Cover Is A Haplotype Threading Alternative To Li And Stephens Model, Ahsan Sanaullah, Degui Zhi, Shaojie Zhang Jul 2023

Minimal Positional Substring Cover Is A Haplotype Threading Alternative To Li And Stephens Model, Ahsan Sanaullah, Degui Zhi, Shaojie Zhang

Faculty, Staff and Student Publications

The Li and Stephens (LS) hidden Markov model (HMM) models the process of reconstructing a haplotype as a mosaic copy of haplotypes in a reference panel. For small panels, the probabilistic parameterization of LS enables modeling the uncertainties of such mosaics. However, LS becomes inefficient when sample size is large, because of its linear time complexity. Recently the PBWT, an efficient data structure capturing the local haplotype matching among haplotypes, was proposed to offer a fast method for giving some optimal solution (Viterbi) to the LS HMM. Previously, we introduced the minimal positional substring cover (MPSC) problem as an alternative …


Rapid-Query For Fast Identity By Descent Search And Genealogical Analysis, Yuan Wei, Ardalan Naseri, Degui Zhi, Shaojie Zhang Jun 2023

Rapid-Query For Fast Identity By Descent Search And Genealogical Analysis, Yuan Wei, Ardalan Naseri, Degui Zhi, Shaojie Zhang

Faculty, Staff and Student Publications

MOTIVATION: Due to the rapid growth of the genetic database size, genealogical search, a process of inferring familial relatedness by identifying DNA matches, has become a viable approach to help individuals finding missing family members or law enforcement agencies locating suspects. A fast and accurate method is needed to search an out-of-database individual against millions of individuals. Most existing approaches only offer all-versus-all within panel match. Some prototype algorithms offer one-versus-all query from out-of-panel individual, but they do not tolerate errors.

RESULTS: A new method, random projection-based identity-by-descent (IBD) detection (RaPID) query, is introduced to make fast genealogical search possible. …


Algorithmic Bias: Causes And Effects On Marginalized Communities, Katrina M. Baha May 2023

Algorithmic Bias: Causes And Effects On Marginalized Communities, Katrina M. Baha

Undergraduate Honors Theses

Individuals from marginalized backgrounds face different healthcare outcomes due to algorithmic bias in the technological healthcare industry. Algorithmic biases, which are the biases that arise from the set of steps used to solve or analyze a problem, are evident when people from marginalized communities use healthcare technology. For example, many pulse oximeters, which are the medical devices used to measure oxygen saturation in the blood, are not able to accurately read people who have darker skin tones. Thus, people with darker skin tones are not able to receive proper health care due to their pulse oximetry data being inaccurate. This …


Why Is Biomedical Informatics Hard? A Fundamental Framework, Todd R Johnson, Elmer V Bernstam Apr 2023

Why Is Biomedical Informatics Hard? A Fundamental Framework, Todd R Johnson, Elmer V Bernstam

Faculty, Staff and Student Publications

Building on previous work to define the scientific discipline of biomedical informatics, we present a framework that categorizes fundamental challenges into groups based on data, information, and knowledge, along with the transitions between these levels. We define each level and argue that the framework provides a basis for separating informatics problems from non-informatics problems, identifying fundamental challenges in biomedical informatics, and provides guidance regarding the search for general, reusable solutions to informatics problems. We distinguish between processing data (symbols) and processing meaning. Computational systems, that are the basis for modern information technology (IT), process data. In contrast, many important challenges …


Automated Contouring And Planning In Radiation Therapy: What Is 'Clinically Acceptable'?, Hana Baroudi, Kristy K Brock, Wenhua Cao, Xinru Chen, Caroline Chung, Laurence E Court, Mohammad D El Basha, Maguy Farhat, Skylar Gay, Mary P Gronberg, Aashish Chandra Gupta, Soleil Hernandez, Kai Huang, David A Jaffray, Rebecca Lim, Barbara Marquez, Kelly Nealon, Tucker J Netherton, Callistus M Nguyen, Brandon Reber, Dong Joo Rhee, Ramon M Salazar, Mihir D Shanker, Carlos Sjogreen, Mckell Woodland, Jinzhong Yang, Cenji Yu, Yao Zhao Feb 2023

Automated Contouring And Planning In Radiation Therapy: What Is 'Clinically Acceptable'?, Hana Baroudi, Kristy K Brock, Wenhua Cao, Xinru Chen, Caroline Chung, Laurence E Court, Mohammad D El Basha, Maguy Farhat, Skylar Gay, Mary P Gronberg, Aashish Chandra Gupta, Soleil Hernandez, Kai Huang, David A Jaffray, Rebecca Lim, Barbara Marquez, Kelly Nealon, Tucker J Netherton, Callistus M Nguyen, Brandon Reber, Dong Joo Rhee, Ramon M Salazar, Mihir D Shanker, Carlos Sjogreen, Mckell Woodland, Jinzhong Yang, Cenji Yu, Yao Zhao

Faculty, Staff and Student Publications

Developers and users of artificial-intelligence-based tools for automatic contouring and treatment planning in radiotherapy are expected to assess clinical acceptability of these tools. However, what is 'clinical acceptability'? Quantitative and qualitative approaches have been used to assess this ill-defined concept, all of which have advantages and disadvantages or limitations. The approach chosen may depend on the goal of the study as well as on available resources. In this paper, we discuss various aspects of 'clinical acceptability' and how they can move us toward a standard for defining clinical acceptability of new autocontouring and planning tools.


Syllable-Pbwt For Space-Efficient Haplotype Long-Match Query, Victor Wang, Ardalan Naseri, Shaojie Zhang, Degui Zhi Jan 2023

Syllable-Pbwt For Space-Efficient Haplotype Long-Match Query, Victor Wang, Ardalan Naseri, Shaojie Zhang, Degui Zhi

Faculty, Staff and Student Publications

MOTIVATION: The positional Burrows-Wheeler transform (PBWT) has led to tremendous strides in haplotype matching on biobank-scale data. For genetic genealogical search, PBWT-based methods have optimized the asymptotic runtime of finding long matches between a query haplotype and a predefined panel of haplotypes. However, to enable fast query searches, the full-sized panel and PBWT data structures must be kept in memory, preventing existing algorithms from scaling up to modern biobank panels consisting of millions of haplotypes. In this work, we propose a space-efficient variation of PBWT named Syllable-PBWT, which divides every haplotype into syllables, builds the PBWT positional prefix arrays on …


Sensitive Data Detection With High-Throughput Machine Learning Models In Electrical Health Records, Kai Zhang, Xiaoqian Jiang Jan 2023

Sensitive Data Detection With High-Throughput Machine Learning Models In Electrical Health Records, Kai Zhang, Xiaoqian Jiang

Faculty, Staff and Student Publications

In the era of big data, there is an increasing need for healthcare providers, communities, and researchers to share data and collaborate to improve health outcomes, generate valuable insights, and advance research. The Health Insurance Portability and Accountability Act of 1996 (HIPAA) is a federal law designed to protect sensitive health information by defining regulations for protected health information (PHI). However, it does not provide efficient tools for detecting or removing PHI before data sharing. One of the challenges in this area of research is the heterogeneous nature of PHI fields in data across different parties. This variability makes rule-based …


Comparative Analysis Of Fullstack Development Technologies: Frontend, Backend And Database, Qozeem Odeniran Jan 2023

Comparative Analysis Of Fullstack Development Technologies: Frontend, Backend And Database, Qozeem Odeniran

College of Graduate Studies: Theses & Dissertations

Accessing websites with various devices has brought changes in the field of application development. The choice of cross-platform, reusable frameworks is very crucial in this era. This thesis embarks in the evaluation of front-end, back-end, and database technologies to address the status quo. Study-a explores front-end development, focusing on angular.js and react.js. Using these frameworks, comparative web applications were created and evaluated locally. Important insights were obtained through benchmark tests, lighthouse metrics, and architectural evaluations. React.js proves to be a performance leader in spite of the possible influence of a virtual machine, opening the door for additional research. Study b …


Federated Learning Algorithms For Generalized Mixed-Effects Model (Glmm) On Horizontally Partitioned Data From Distributed Sources, Wentao Li, Jiayi Tong, Md Monowar Anjum, Noman Mohammed, Yong Chen, Xiaoqian Jiang Oct 2022

Federated Learning Algorithms For Generalized Mixed-Effects Model (Glmm) On Horizontally Partitioned Data From Distributed Sources, Wentao Li, Jiayi Tong, Md Monowar Anjum, Noman Mohammed, Yong Chen, Xiaoqian Jiang

Faculty, Staff and Student Publications

OBJECTIVES: This paper developed federated solutions based on two approximation algorithms to achieve federated generalized linear mixed effect models (GLMM). The paper also proposed a solution for numerical errors and singularity issues. And showed the two proposed methods can perform well in revealing the significance of parameter in distributed datasets, comparing to a centralized GLMM algorithm from R package ('lme4') as the baseline model.

METHODS: The log-likelihood function of GLMM is approximated by two numerical methods (Laplace approximation and Gaussian Hermite approximation, abbreviated as LA and GH), which supports federated decomposition of GLMM to bring computation to data. To solve …


Cov-Inception: Covid-19 Detection Tool Using Chest X-Ray, Aswini Thota, Ololade Awodipe, Rashmi Patel Sep 2022

Cov-Inception: Covid-19 Detection Tool Using Chest X-Ray, Aswini Thota, Ololade Awodipe, Rashmi Patel

SMU Data Science Review

Since the pandemic started, researchers have been trying to find a way to detect COVID-19 which is a cost-effective, fast, and reliable way to keep the economy viable and running. This research details how chest X-ray radiography can be utilized to detect the infection. This can be for implementation in Airports, Schools, and places of business. Currently, Chest imaging is not a first-line test for COVID-19 due to low diagnostic accuracy and confounding with other viral pneumonia. Different pre-trained algorithms were fine-tuned and applied to the images to train the model and the best model obtained was fine-tuned InceptionV3 model …


Secure Human Action Recognition By Encrypted Neural Network Inference, Miran Kim, Xiaoqian Jiang, Kristin Lauter, Elkhan Ismayilzada, Shayan Shams Aug 2022

Secure Human Action Recognition By Encrypted Neural Network Inference, Miran Kim, Xiaoqian Jiang, Kristin Lauter, Elkhan Ismayilzada, Shayan Shams

Faculty, Staff and Student Publications

Advanced computer vision technology can provide near real-time home monitoring to support "aging in place" by detecting falls and symptoms related to seizures and stroke. Affordable webcams, together with cloud computing services (to run machine learning algorithms), can potentially bring significant social benefits. However, it has not been deployed in practice because of privacy concerns. In this paper, we propose a strategy that uses homomorphic encryption to resolve this dilemma, which guarantees information confidentiality while retaining action detection. Our protocol for secure inference can distinguish falls from activities of daily living with 86.21% sensitivity and 99.14% specificity, with an average …


Real-World Matching Performance Of Deidentified Record-Linking Tokens, Elmer V Bernstam, Reuben Joseph Applegate, Alvin Yu, Deepa Chaudhari, Tian Liu, Alex Coda, Jonah Leshin Aug 2022

Real-World Matching Performance Of Deidentified Record-Linking Tokens, Elmer V Bernstam, Reuben Joseph Applegate, Alvin Yu, Deepa Chaudhari, Tian Liu, Alex Coda, Jonah Leshin

Faculty, Staff and Student Publications

OBJECTIVE: Our objective was to evaluate tokens commonly used by clinical research consortia to aggregate clinical data across institutions.

METHODS: This study compares tokens alone and token-based matching algorithms against manual annotation for 20,002 record pairs extracted from the University of Texas Houston's clinical data warehouse (CDW) in terms of entity resolution.

RESULTS: The highest precision achieved was 99.9% with a token derived from the first name, last name, gender, and date-of-birth. The highest recall achieved was 95.5% with an algorithm involving tokens that reflected combinations of first name, last name, gender, date-of-birth, and social security number.

DISCUSSION: To protect …


External Validation Of A Laboratory Prediction Algorithm For The Reduction Of Unnecessary Labs In The Critical Care Setting, Linda T Li, Tongtong Huang, Elmer V Bernstam, Xiaoqian Jiang Jun 2022

External Validation Of A Laboratory Prediction Algorithm For The Reduction Of Unnecessary Labs In The Critical Care Setting, Linda T Li, Tongtong Huang, Elmer V Bernstam, Xiaoqian Jiang

Faculty, Staff and Student Publications

BACKGROUND: Unnecessary laboratory tests contribute to iatrogenic harm and are a major source of waste in the health care system. We previously developed a machine learning algorithm to help clinicians identify unnecessary laboratory tests, but it has not been externally validated. In this study, we externally validate our machine learning algorithm.

METHODS: To externally validate the machine learning algorithm that was originally trained on the Medical Information Mart for Intensive Care (MIMIC) III database, we tested the algorithm in a separate institution. We identified and abstracted data for all patients older than 18 years admitted to the intensive care unit …


Data Ethics: An Investigation Of Data, Algorithms, And Practice, Gabrialla S. Cockerell May 2022

Data Ethics: An Investigation Of Data, Algorithms, And Practice, Gabrialla S. Cockerell

Honors Projects

This paper encompasses an examination of defective data collection, algorithms, and practices that continue to be cycled through society under the illusion that all information is processed uniformly, and technological innovation consistently parallels societal betterment. However, vulnerable communities, typically the impoverished and racially discriminated, get ensnared in these harmful cycles due to their disadvantages. Their hindrances are reflected in their information due to the interconnectedness of data, such as race being highly correlated to wealth, education, and location. However, their information continues to be analyzed with the same measures as populations who are not significantly affected by racial bias. Not …


Data And Algorithmic Modeling Approaches To Count Data, Andraya Hack May 2022

Data And Algorithmic Modeling Approaches To Count Data, Andraya Hack

Honors College Theses

Various techniques are used to create predictions based on count data. This type of data takes the form of a non-negative integers such as the number of claims an insurance policy holder may make. These predictions can allow people to prepare for likely outcomes. Thus, it is important to know how accurate the predictions are. Traditional statistical approaches for predicting count data include Poisson regression as well as negative binomial regression. Both methods also have a zero-inflated version that can be used when the data has an overabundance of zeros. Another procedure is to use computer algorithms, also known as …


Privacy-Preserving Logistic Regression With Secret Sharing, Ali Reza Ghavamipour, Fatih Turkmen, Xiaoqian Jiang Apr 2022

Privacy-Preserving Logistic Regression With Secret Sharing, Ali Reza Ghavamipour, Fatih Turkmen, Xiaoqian Jiang

Faculty, Staff and Student Publications

BACKGROUND: Logistic regression (LR) is a widely used classification method for modeling binary outcomes in many medical data classification tasks. Researchers that collect and combine datasets from various data custodians and jurisdictions can greatly benefit from the increased statistical power to support their analysis goals. However, combining data from different sources creates serious privacy concerns that need to be addressed.

METHODS: In this paper, we propose two privacy-preserving protocols for performing logistic regression with the Newton-Raphson method in the estimation of parameters. Our proposals are based on secure Multi-Party Computation (MPC) and tailored to the honest majority and dishonest majority …


Particle Identification And Tracking In Real Time Using Machine Learning On Fpga, F. Barbosa, L. Belfore, C. Dickover, C. Fanelli, S. Furletov, Y. Furletova, L. Jokhovets, D. Lawrence, D. Romanov Jan 2022

Particle Identification And Tracking In Real Time Using Machine Learning On Fpga, F. Barbosa, L. Belfore, C. Dickover, C. Fanelli, S. Furletov, Y. Furletova, L. Jokhovets, D. Lawrence, D. Romanov

Electrical & Computer Engineering Faculty Publications

This project is a multi-disciplinary endeavour between Physics, Electrical Engineering, and Computer Engineering. The purpose is to develop and implement an FPGA(*) based Machine Learning algorithm for real-time particle identification, filtering, and data reduction. This is important research that can be applied to streaming readout systems being developed now at JLab and other facilities. Real-time data processing is a frontier field in experimental physics, especially in HEP. The application of FPGAs at the trigger level is used by many current and planned experiments (CMS, LHCb, Belle2, PANDA). Usually they use conventional processing algorithms. LHCb has implemented ML elements for real-time …


An Introduction To Calling Bullshit: Learning To Think Outside The Black Box, Jevin D. West, Carl T. Bergstrom Aug 2021

An Introduction To Calling Bullshit: Learning To Think Outside The Black Box, Jevin D. West, Carl T. Bergstrom

Numeracy

Bergstrom, Carl T. and Jevin D. West. 2020. Calling Bullshit: The Art of Skepticism in a Data-Driven World. (New York: Random House) 336 pp. ISBN 978-0525509202.

While statistical methods receive greater attention, the art of critically evaluating information in everyday life more commonly depends on thinking outside the black box of the algorithm. In this piece we introduce readers to our book and associated online teaching materials—for readers who want to more capably call “bullshit” or to teach their students to do the same.


Awegnn: Auto-Parametrized Weighted Element-Specific Graph Neural Networks For Molecules., Timothy Szocinski, Duc Duy Nguyen, Guo-Wei Wei Jul 2021

Awegnn: Auto-Parametrized Weighted Element-Specific Graph Neural Networks For Molecules., Timothy Szocinski, Duc Duy Nguyen, Guo-Wei Wei

Mathematics Faculty Publications

While automated feature extraction has had tremendous success in many deep learning algorithms for image analysis and natural language processing, it does not work well for data involving complex internal structures, such as molecules. Data representations via advanced mathematics, including algebraic topology, differential geometry, and graph theory, have demonstrated superiority in a variety of biomolecular applications, however, their performance is often dependent on manual parametrization. This work introduces the auto-parametrized weighted element-specific graph neural network, dubbed AweGNN, to overcome the obstacle of this tedious parametrization process while also being a suitable technique for automated feature extraction on these internally complex …