Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- The Texas Medical Center Library (510)
- Department of Primary Industries and Regional Development, Western Australia (23)
- Chinese Academy of Sciences (20)
- Old Dominion University (13)
- City University of New York (CUNY) (11)
-
- University of Kentucky (11)
- Purdue University (9)
- University of Nebraska - Lincoln (9)
- University of South Alabama (9)
- Clemson University (7)
- Illinois State University (7)
- University of Central Florida (6)
- Virginia Commonwealth University (6)
- Kennesaw State University (5)
- SIT Graduate Institute/SIT Study Abroad (5)
- Southern Methodist University (5)
- West Virginia University (5)
- Central Washington University (4)
- Chapman University (4)
- Claremont Colleges (4)
- Kutztown University (4)
- Michigan Technological University (4)
- University of Texas at Arlington (4)
- Dartmouth College (3)
- International Centre of Insect Physiology and Ecology (3)
- Nova Southeastern University (3)
- Bowling Green State University (2)
- California Polytechnic State University, San Luis Obispo (2)
- Case Western Reserve University (2)
- Gonzaga University (2)
- Keyword
-
- Humans (268)
- Female (51)
- Male (49)
- Animals (41)
- Electronic Health Records (38)
-
- Machine learning (38)
- Mice (29)
- Adult (28)
- Machine Learning (27)
- Middle Aged (27)
- Algorithms (26)
- COVID-19 (25)
- Aged (21)
- Genome-Wide Association Study (20)
- Natural Language Processing (20)
- Neoplasms (20)
- Deep Learning (19)
- Deep learning (19)
- Genomics (19)
- Software (18)
- United States (18)
- Neural Networks, Computer (17)
- Brain (16)
- Retrospective Studies (16)
- Western Australia (16)
- Artificial Intelligence (15)
- Gene Expression Profiling (15)
- Genetic (15)
- Privacy (15)
- Adolescent (14)
- Publication Year
- Publication
-
- Faculty, Staff and Student Publications (506)
- Bulletin of Chinese Academy of Sciences (Chinese Version) (20)
- University Faculty and Staff Publications (9)
- All Dissertations (7)
- Fisheries Research Reports (7)
-
- Theses and Dissertations (7)
- Dissertations, Theses, and Capstone Projects (6)
- Annual Symposium on Biomathematics and Ecology Education and Research (5)
- Data Science and Data Mining (5)
- Fisheries Research Articles (5)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (5)
- Independent Study Project (ISP) Collection (5)
- Computer Science and Information Technology Faculty (4)
- Dissertations, Master's Theses and Master's Reports (4)
- Electronic Theses and Dissertations (4)
- Graduate Industrial Research Symposium (4)
- Symposium of Student Scholars (4)
- All Peer-Reviewed Publications (3)
- Biology and Medicine Through Mathematics Conference (3)
- Computer Science Faculty Publications (3)
- Dissertations (3)
- Master's Theses (3)
- SMU Data Science Review (3)
- All Faculty Scholarship for the College of the Sciences (2)
- All HCAS Student Capstones, Theses, and Dissertations (2)
- Biological Sciences Faculty Publications (2)
- Computational and Data Sciences (PhD) Dissertations (2)
- Computer Science Faculty Scholarship (2)
- Dartmouth College Ph.D Dissertations (2)
- Dissertations and Doctoral Documents, University of Nebraska-Lincoln, 2023– (2)
- Publication Type
- File Type
Articles 31 - 60 of 765
Full-Text Articles in Data Science
Ibi-Dt: A Novel Approach Combining Individualized Bayesian Inference And Decision Tree For Identifying Cancer Drivers And Their Interactions, Md Asad Rahman, Gregory F. Cooper, Jinying Zhao, Xinghua Lu, Jinling Liu
Ibi-Dt: A Novel Approach Combining Individualized Bayesian Inference And Decision Tree For Identifying Cancer Drivers And Their Interactions, Md Asad Rahman, Gregory F. Cooper, Jinying Zhao, Xinghua Lu, Jinling Liu
Engineering Management and Systems Engineering Faculty Research & Creative Works
Cancer is mainly caused by a relatively small portion of somatic genome alterations (SGAs), called cancer drivers. Despite success in identifying a good number of cancer drivers, many more remain to be discovered to explain various cancers. Moreover, limited tools are available to identify potential interactions among cancer drivers for a better understanding of oncogenesis. To tackle these challenges, we have developed a novel approach called individualized Bayesian inference using a decision tree (IBI-DT). IBI-DT recognizes the genetic heterogeneity among cancer patients, where different individuals or patient subgroups of distinct genomic makeup may have different drivers. IBI-DT works by constructing …
Performance Analysis Of Computational Methods For Predicting Protein Function In Rare Diseases, Aichetou Mohamed Sidiya, Hanin Alzaher, Razan Almahdi, Tayeb Brahimi
Performance Analysis Of Computational Methods For Predicting Protein Function In Rare Diseases, Aichetou Mohamed Sidiya, Hanin Alzaher, Razan Almahdi, Tayeb Brahimi
Effat Undergraduate Research Journal
Protein function prediction is crucial for understanding the underlying mechanisms of rare diseases. With the increasing availability of computational methods including machine learning-based approaches, network-based methods, and sequence-based methods, predicting protein functions has become more accessible. However, it is not clear which of these methods performs better or how they compare to each other in terms of accuracy, efficiency, and scalability. In this study, we evaluate several computational methods for predicting protein functions in rare diseases using key performance indicators (KPIs). We analyze the strengths and weaknesses of each method and provide recommendations for researchers and clinicians interested in using …
Crop Yield Prediction At Multiple Spatial Scales With Statistical Machine Learning, Vaibhav Charan, Pratishtha Poudel
Crop Yield Prediction At Multiple Spatial Scales With Statistical Machine Learning, Vaibhav Charan, Pratishtha Poudel
Discovery Undergraduate Interdisciplinary Research Internship
Understanding and accurately predicting crop yield is becoming increasingly important today in the face of global food security challenges, and thus, the availability of standardized data and scalable models is the need of the hour. To support this, researchers have developed CY-Bench (Crop Yield Benchmark), a comprehensive dataset that helps forecast maize and wheat yields on a global scale. This research project primarily involved working with the CY-Bench dataset aiming to improve crop yield prediction through machine learning. Initially, papers explaining the CY-Bench dataset and other papers for agriculture modeling were studied and analyzed in detail. The research then progressed …
Farming With Data: Tracing Critical Tensions Using Data Science For Food Justice, Marc Sager, Maximilan Sherard, Anthony Petrosino
Farming With Data: Tracing Critical Tensions Using Data Science For Food Justice, Marc Sager, Maximilan Sherard, Anthony Petrosino
Publications
In this manuscript, we explore the intersection of artificial intelligence (AI) and equitable learning in higher education, focusing on data science as a subset of AI and social justice as the core theme of equity. Our investigation sheds light on the nuanced tensions inherent in employing data science for social justice. Rooted in situated perspectives of learning and consequential learning, our study employs an instrumental case-study methodology and analysis techniques from interaction and conversation analysis. Collaborating with three undergraduate students and an urban farm, the students used data science practices to highlight inequities surrounding food justice and access to food. …
Application Of Artificial Neural Network Algorithms For Irrigation Scheduling, Lisa Umutoni
Application Of Artificial Neural Network Algorithms For Irrigation Scheduling, Lisa Umutoni
All Dissertations
Neural networks have been extensively used in predicting soil water tension for improved irrigation scheduling and management. However, their lack of interpretability constrains their efficacy in grasping the nuanced patterns prevalent in soil water tension time series data. The first goal of this research was to develop interpretable deep neural network models for soil water tension prediction across multiple soil depths (0.15m, 0.3m, 0.46m and 0.6m) and prediction horizons (1h, 6h, and 12h). The Neural Hierarchical Interpolation for Time Series (N-HiTS) and Neural Basis Expansion Analysis Time Series (N-BEATS) models were used in this research. Historical soil water tension data …
Authentication And Message Integrity Verification For Emerging Wireless Networks, Ebuka Philip Oguchi
Authentication And Message Integrity Verification For Emerging Wireless Networks, Ebuka Philip Oguchi
Dissertations and Doctoral Documents, University of Nebraska-Lincoln, 2023–
This dissertation presents a comprehensive body of research on authentication and message integrity verification for emerging wireless networks, focusing on secret-free and physical layer security techniques across diverse, challenging, and unconventional environments.
It comprises four first-author contributions that span underground wireless systems, over-the-air (OTA) channels, vehicular communications, and nanoscale molecular networks.
The first contribution, Soil-Assisted Trust Establishment for Underground Wireless Networks (STUN), introduces a physical-layer trust bootstrapping protocol that achieves authentication and message integrity without pre-shared secrets. Leveraging underground-to-air propagation laws and trusted relay nodes, STUN resists active signal injection attacks and demonstrates security comparable to the unbalanced oil and …
A Cancer Education Needs Assessment: Informing Middle-Aged Female Patients About The Relationships Between Obesity And Women’S Health Concerns In The Reproductive System, Breast, And Endometrial Health, Batul Mirza
MUSC Theses and Dissertations
Obesity significantly impacts women’s health, particularly among middle-aged women, by increasing the risk of hormone-sensitive cancers such as breast, endometrial, and reproductive system cancers. This study examines the educational needs of this demographic group regarding obesity-related cancer risks and explores effective intervention strategies. Obesity-induced mechanisms – hormonal imbalances, chronic inflammation, and insulin resistance – drive cancer susceptibility, emphasizing the need for targeted health education. The study employs a qualitative design, which includes interviews with subject matter experts (SMEs) and surveys of middle-aged women. The goal is to assess awareness, perceived barriers, and preferred learning methods. Findings suggest that with many …
Use Of High-Throughput Phenomics With And Without Fungicide As A Wheat Breeding Tool, Gerardo Ivan Rivera Collazo
Use Of High-Throughput Phenomics With And Without Fungicide As A Wheat Breeding Tool, Gerardo Ivan Rivera Collazo
Dissertations and Doctoral Documents, University of Nebraska-Lincoln, 2023–
Wheat (Triticum aestivum) is an important food staple for many countries around the world and current consumption and production data demonstrate a production deficit. Breeders are challenged to help close this gap in production by selecting better cultivars with improved yields. Several methods aim to accelerate the breeding process to reduce the time required for releasing an elite variety. One of the main breeding bottlenecks is the lack of fast and reliable phenotypic data acquisition that could dissect physiological and morphological plant data. High-throughput remote sensing could have the potential to reduce this bottleneck by streamlining data acquisition, …
Detecting Physical Activity Using Wearable Sensor Data, Dipok Deb
Detecting Physical Activity Using Wearable Sensor Data, Dipok Deb
Data Science and Data Mining
This study focuses on detecting physical activity using wearable sensor data, specifically distinguishing between walking and running. A dataset comprising accelerometer and gyroscope readings is used to train and evaluate various machine learning models, including logistic regression, random forest, k-nearest neighbors, naïve Bayes, and XGBoost. Extensive preprocessing, such as creating lag features and rolling statistics, is performed to enhance temporal data representation. The models are evaluated using metrics like accuracy, precision, recall, and F1 score. Incorporating lag and rolling features significantly improves model performance, with logistic regression achieving perfect scores across all metrics. These findings demonstrate the effectiveness of enhanced …
Topodino: Self-Supervised Topological Representation Learning For Neuronal Morphologies, Yasser Binbisher
Topodino: Self-Supervised Topological Representation Learning For Neuronal Morphologies, Yasser Binbisher
Master's Theses
Neuronal cell types are categorized by transcriptomic identity, yet their morphological heterogeneity defies this classification. In response, researchers have adopted unsupervised graph representation learning as a tool to reveal morphological variation within single-class transcriptomic types. However, the complex geometry of neuronal morphology—especially long axons and dense dendrites—challenges graph neural networks, which struggle with message propagation across extended structures. To mitigate this, current approaches enforce sub-sampling on neuronal graphs and omit axons entirely, sacrificing critical biological features for computational efficiency. To overcome this trade-off, this thesis introduces TopoDINO, a self-supervised, topology-aware representation learning model designed to preserve the full hierarchical organization …
Computerized Diagnostic Decision Support Systems-Isabel Pro Versus Chatgpt-4 Part Ii, Joe M Bridges, Xiaoqian Jiang, Michael Ige, Oluwatoniloba Toyobo
Computerized Diagnostic Decision Support Systems-Isabel Pro Versus Chatgpt-4 Part Ii, Joe M Bridges, Xiaoqian Jiang, Michael Ige, Oluwatoniloba Toyobo
Faculty, Staff and Student Publications
Objective: Does a Tree-of-Thought prompt and reconsideration of Isabel Pro's differential improve ChatGPT-4's accuracy; does increasing expert panel size improve ChatGPT-4's accuracy; does ChatGPT-4 produce consistent outputs in sequential requests; what is the frequency of fabricated references?
Materials and methods: Isabel Pro, a computerized diagnostic decision support system, and ChatGPT-4, a large language model. Using 201 cases from the New England Journal of Medicine, each system produced a differential diagnosis ranked by likelihood. Statistics were Mean Reciprocal Rank, Recall at Rank, Average Rank, Number of Correct Diagnoses, and Rank Improvement. For reproducibility, the study compared the initial expert panel run …
Protein Word Detection Using Text Segmentation Techniques, Ashish V. Tendulkar, Sutanu Chakraborti, Ganesh Devi
Protein Word Detection Using Text Segmentation Techniques, Ashish V. Tendulkar, Sutanu Chakraborti, Ganesh Devi
Journal of Global Awareness
Literature in Molecular Biology is abundant with linguistic metaphors. There has been works in the past that attempt to draw parallels between linguistics and biology, driven by the fundamental premise that proteins have a language of their own. Since word detection is crucial to the decipherment of any unknown language, we attempt to establish a problem mapping from natural language text to protein sequences at the level of words. Towards this end, we explore the use of an unsupervised text segmentation algorithm for the task of extracting "biological words” from protein sequences. We demonstrate the effectiveness of using domain knowledge …
A Complete Transfer Learning-Based Pipeline For Discriminating Between Select Pathogenic Yeasts From Microscopy Photographs, Ryan A. Parker, Danielle S. Hannagan, Jan H. Strydom, Christopher J. Boon, Jessica Fussell, Chelbie A. Mitchell, Katie L. Moerschel, Aura G. Valter-Franco, Christopher Cornelison
A Complete Transfer Learning-Based Pipeline For Discriminating Between Select Pathogenic Yeasts From Microscopy Photographs, Ryan A. Parker, Danielle S. Hannagan, Jan H. Strydom, Christopher J. Boon, Jessica Fussell, Chelbie A. Mitchell, Katie L. Moerschel, Aura G. Valter-Franco, Christopher Cornelison
Faculty Articles
Pathogenic yeasts are an increasing concern in healthcare, with species like Candida auris often displaying drug resistance and causing high mortality in immunocompromised patients. The need for rapid and accessible diagnostic methods for accurate yeast identification is critical, especially in resource-limited settings. This study presents a convolutional neural network (CNN)-based approach for classifying pathogenic yeast species from microscopy images. Using transfer learning, we trained the model to identify six yeast species from simple micrographs, achieving high classification accuracy (93.91% at the patch level, 99.09% at the whole image level) and low misclassification rates across species, with the best performing model. …
Histone Methyltransferase Ash1l Primes Metastases And Metabolic Reprogramming Of Macrophages In The Bone Niche, Chenling Meng, Kevin Lin, Wei Shi, Hongqi Teng, Xinhai Wan, Anna Debruine, Yin Wang, Xin Liang, Javier Leo, Feiyu Chen, Qianlin Gu, Jie Zhang, Vivien Van, Kiersten L Maldonado, Boyi Gan, Li Ma, Yue Lu, Di Zhao
Histone Methyltransferase Ash1l Primes Metastases And Metabolic Reprogramming Of Macrophages In The Bone Niche, Chenling Meng, Kevin Lin, Wei Shi, Hongqi Teng, Xinhai Wan, Anna Debruine, Yin Wang, Xin Liang, Javier Leo, Feiyu Chen, Qianlin Gu, Jie Zhang, Vivien Van, Kiersten L Maldonado, Boyi Gan, Li Ma, Yue Lu, Di Zhao
Faculty, Staff and Student Publications
Bone metastasis is a major cause of cancer death; however, the epigenetic determinants driving this process remain elusive. Here, we report that histone methyltransferase ASH1L is genetically amplified and is required for bone metastasis in men with prostate cancer. ASH1L rewires histone methylations and cooperates with HIF-1α to induce pro-metastatic transcriptome in invading cancer cells, resulting in monocyte differentiation into lipid-associated macrophage (LA-TAM) and enhancing their pro-tumoral phenotype in the metastatic bone niche. We identified IGF-2 as a direct target of ASH1L/HIF-1α and mediates LA-TAMs' differentiation and phenotypic changes by reprogramming oxidative phosphorylation. Pharmacologic inhibition of the ASH1L-HIF-1α-macrophages axis elicits …
Evaluating Predictive Models For Predicting Total Score Of Beef Carcasses, Emmanuel Forson
Evaluating Predictive Models For Predicting Total Score Of Beef Carcasses, Emmanuel Forson
Electronic Theses and Dissertations
The beef industry plays a vital role in global agriculture, with carcass quality and consumer preference being key determinants of market success. This thesis examines predictive modeling techniques for estimating the Total Score of beef carcasses, a composite measure representing yield and quality, primarily used by the Nebraska Cattlemen Association. Using data from the Nebraska Cattlemen’s Foundation Retail Value Steer Challenge (2000–2023), the study compares the performance of First Order Multiple Linear Regression (MLR) with three machine learning techniques: K-Nearest Neighbors (KNN), Random Forest, and Gradient Boosting Machine (GBM).
The analysis focuses on six key predictors: Hot Carcass Weight, Back …
Performance Of Lasso And Ridge Regression For Variable Selection In Genome-Wide Association Studies Of Maize Flowering Time, Dipok Deb
Data Science and Data Mining
Genome-Wide Association Studies (GWAS) are instrumental in identifying genetic variants linked to complex traits, providing valuable insights into trait heritability and biological mechanisms. This study applies GWAS to investigate flowering time in maize, a critical adaptive trait, using a diverse dataset of 5,000 recombinant inbred lines across eight environments. Traditional GWAS methods often encounter challenges in high-dimensional datasets due to the presence of multiple small-effect genetic loci. To address this, we compared two penalized regression methods—LASSO and Ridge regression—to perform variable selection and regression analysis within a GWAS framework. LASSO effectively reduced the number of predictors by selecting the most …
Toward The Application Of Natural Language Processing In Electronic Health Record Analysis For Taxonomy Development, Latoya Mcdonald
Toward The Application Of Natural Language Processing In Electronic Health Record Analysis For Taxonomy Development, Latoya Mcdonald
All Dissertations
Electronic health records (EHRs) are pivotal resources for nurse practice because they increase the timeliness and reliability of patient information at the point of care and support access by multiple healthcare providers and the individual patients themselves. However, it is widely recognized that data extraction from EHRs is challenging due to the variability in the language used in clinical care notes and the lack of standardized terminology across healthcare systems. The broad objective of this dissertation is to develop taxonomy-based classification models for nursing care by applying feature engineering approaches to EHRs that include nursing care of ostomy patients following …
A Comparison Of The Bacterial Communities Of The Yamuna River (India) And Mississippi River (Usa), Jacob Gareis, Osvaldo Martinez, Silas Bergen
A Comparison Of The Bacterial Communities Of The Yamuna River (India) And Mississippi River (Usa), Jacob Gareis, Osvaldo Martinez, Silas Bergen
Research & Creative Achievement Day
This study compared the bacteria in the Yamuna River in India and the Mississippi River in the USA. Water samples were taken from 11 sites (2 in the Mississippi River and 9 in the Yamuna River). The bacteria were identified using 16S rRNA gene sequencing. Some phyla such as Proteobacteria and Bacteroidetes were seen in abundance regardless of country. However, principal-component analysis showed three distinct groupings: Mississippi/Ganga/Tons and other Yamuna River locations (7 sites), Yamuna River at Delhi (2 sites), Yamuna River below Glacier (2 sites) Additionally, the Yamuna River below Glacier had the highest bacterial diversity, with a Shannon …
Leveraging Attention Mechanism To Unlock Gene And Protein Attributes, Ala Jararweh
Leveraging Attention Mechanism To Unlock Gene And Protein Attributes, Ala Jararweh
Computer Science ETDs
Advancing personalized medicine depends on effectively integrating and interpreting the vast, heterogeneous landscape of biological data, from genomic sequences and transcriptomics to the insights embedded in scientific literature. Current machine learning models often focus on single data modalities, limiting their capacity to capture the multifaceted nature of biological systems. We address this gap by developing three attention-based machine-learning models integrating diverse data modalities. Firstly, DeepVul is a multi-task model that leverages cancer transcriptome data to predict genes critical for cancer survival and their corresponding drugs. Subsequently, LitGene refines gene representations by integrating textual information from the scientific literature. Finally, Protein2Text …
Resurrecting The Past: Engineering Ancestral Enzymes For Modern Bacterial Applications, Gregory M. Gladkowski Iii, Masa Watanabe
Resurrecting The Past: Engineering Ancestral Enzymes For Modern Bacterial Applications, Gregory M. Gladkowski Iii, Masa Watanabe
SACAD: Scholarly Activities
Ancestral Sequence Reconstruction (ASR) is a computational technique that infers and resurrects ancient protein sequences to explore molecular evolution and enable protein engineering. By integrating phylogenetics and statistical modeling, ASR produces stable, mutation-tolerant enzymes ideal for directed evolution. This poster presents a workflow for reconstructing enzymes from pathogenic bacteria and highlights ASR’s value in uncovering novel functions and advancing applications in drug discovery and synthetic biology.
Enhancing Remote Sensing Imagery Temporal Resolution Using Starfm Data Fusion Approach For Improved Land Surface Monitoring, Ahmadreza Pourghodrat
Enhancing Remote Sensing Imagery Temporal Resolution Using Starfm Data Fusion Approach For Improved Land Surface Monitoring, Ahmadreza Pourghodrat
School of Computing: Dissertations, Theses, and Student Research
High-resolution remote sensing imagery plays a critical role in various domains, such as farm-level agricultural operations, environmental monitoring, and natural resource management. However, data with high spatial resolution typically have low temporal resolution, and those with high temporal resolution often lack spatial detail. For example, Landsat 8 and 9 satellites deliver high spatial resolution images with a 30-meter pixel size but suffer from low temporal resolution, with a 16-day revisit cycle. In contrast, satellites like MODIS and VIIRS provide daily images but with a much coarser spatial resolution (375 meters or more), reducing spatial details. Additionally, there is a lack …
Student Expectations And Outcomes In Virtual Vs In-Person Interprofessional Simulations: A Qualitative Analysis, Padmavathy Ramaswamy, Abbey M Bachmann, Tiffany Champagne-Langabeer, Chasisty L Gilder, Samuel E Neher, Jennifer L Swails
Student Expectations And Outcomes In Virtual Vs In-Person Interprofessional Simulations: A Qualitative Analysis, Padmavathy Ramaswamy, Abbey M Bachmann, Tiffany Champagne-Langabeer, Chasisty L Gilder, Samuel E Neher, Jennifer L Swails
Faculty, Staff and Student Publications
Background: Health-related programs frequently integrate interprofessional education (IPE) into their training. The COVID-19 pandemic transitioned many IPE programs online, making it essential to assess student expectations and perceived learning outcomes across virtual simulations and in-person settings.
Methods: This qualitative study compared student expectations and self-reported outcomes across in-person and virtual case scenarios at a Texas health science center. Responses to open-ended questions from two data collection periods were analyzed using inductive coding and thematic analysis.
Results: Students from nursing, medicine, dentistry, public health, and informatics participated in each group. Three major themes emerged from this study: communication, teamwork, and …
Linking Water Quality And Climate Change To Long-Term Trends In Species Abundance In Norwalk Harbor, Viktoria Savatorova, Aidan Kieft, Nicole C. Spiller, Kasey Burns
Linking Water Quality And Climate Change To Long-Term Trends In Species Abundance In Norwalk Harbor, Viktoria Savatorova, Aidan Kieft, Nicole C. Spiller, Kasey Burns
Spora: A Journal of Biomathematics
This study examines the effects of environmental changes on fish populations in Norwalk Harbor, focusing on winter flounder (Pseudopleuronectes americanus), cunner (Tautogolabrus adspersus), northern pipefish (Syngnathus fuscus), and naked goby (Gobiosoma bosci) as examples of species responding to climate-related shifts. We analyze how water temperature, salinity, and dissolved oxygen correlate with fish abundance. To assess statistically significant differences in catch per unit effort (CPUE) across harbor regions, we applied the Kruskal-Wallis test followed by Dunn's post-hoc test. Seasonal variations in CPUE were examined by comparing monthly catch data for each species. K-means …
Precision Phenotyping For Curating Research Cohorts Of Patients With Unexplained Post-Acute Sequelae Of Covid-19, Alaleh Azhir, Jonas Hügel, Jiazi Tian, Jingya Cheng, Ingrid V Bassett, Douglas S Bell, Elmer V Bernstam, Maha R Farhat, Darren W Henderson, Emily S Lau, Michele Morris, Yevgeniy R Semenov, Virginia A Triant, Shyam Visweswaran, Zachary H Strasser, Jeffrey G Klann, Shawn N Murphy, Hossein Estiri
Precision Phenotyping For Curating Research Cohorts Of Patients With Unexplained Post-Acute Sequelae Of Covid-19, Alaleh Azhir, Jonas Hügel, Jiazi Tian, Jingya Cheng, Ingrid V Bassett, Douglas S Bell, Elmer V Bernstam, Maha R Farhat, Darren W Henderson, Emily S Lau, Michele Morris, Yevgeniy R Semenov, Virginia A Triant, Shyam Visweswaran, Zachary H Strasser, Jeffrey G Klann, Shawn N Murphy, Hossein Estiri
Faculty, Staff and Student Publications
BACKGROUND: Scalable identification of patients with post-acute sequelae of COVID-19 (PASC) is challenging due to a lack of reproducible precision phenotyping algorithms, which has led to suboptimal accuracy, demographic biases, and underestimation of the PASC.
METHODS: In a retrospective case-control study, we developed a precision phenotyping algorithm for identifying cohorts of patients with PASC. We used longitudinal electronic health records data from over 295,000 patients from 14 hospitals and 20 community health centers in Massachusetts. The algorithm employs an attention mechanism to simultaneously exclude sequelae that prior conditions can explain and include infection-associated chronic conditions. We performed independent chart reviews …
Evaluating The Meditation Practices And Barriers To Adopting Mindful Medicine Among Physicians, Tiffany Champagne-Langabeer, Chelsea G Ratcliff, Christine Bakos-Block, Francine Vega, Marylou Cardenas-Turanzas, Aila Malik, Radha Korupolu
Evaluating The Meditation Practices And Barriers To Adopting Mindful Medicine Among Physicians, Tiffany Champagne-Langabeer, Chelsea G Ratcliff, Christine Bakos-Block, Francine Vega, Marylou Cardenas-Turanzas, Aila Malik, Radha Korupolu
Faculty, Staff and Student Publications
Background: Chronic pain affects over 25% of U.S. adults and is a leading cause of disability. Mindfulness meditation (MM) is a nonpharmacologic approach to manage pain and improve well-being. Despite mounting evidence supporting its efficacy, MM remains underutilized in medical practice. Understanding physicians' engagement with MM and the barriers they face can inform strategies for integration into clinical care. This study assessed physicians' attitudes toward MM, including barriers to practice and their likelihood of recommending it to patients.
Methods: A cross-sectional survey of U.S. physicians was conducted from April to July 2024. Participants provided information on demographics, health struggles, and …
Leveraging Large Language Models For Knowledge-Free Weak Supervision In Clinical Natural Language Processing, Enshuo Hsu, Kirk Roberts
Leveraging Large Language Models For Knowledge-Free Weak Supervision In Clinical Natural Language Processing, Enshuo Hsu, Kirk Roberts
Faculty, Staff and Student Publications
The performance of deep learning-based natural language processing systems is based on large amounts of labeled training data which, in the clinical domain, are not easily available or affordable. Weak supervision and in-context learning offer partial solutions to this issue, particularly using large language models (LLMs), but their performance still trails traditional supervised methods with moderate amounts of gold-standard data. In particular, inferencing with LLMs is computationally heavy. We propose an approach leveraging fine-tuning LLMs and weak supervision with virtually no domain knowledge that still achieves consistently dominant performance. Using a prompt-based approach, the LLM is used to generate weakly-labeled …
Proxy Panels Enable Privacy-Aware Outsourcing Of Genotype Imputation, Degui Zhi, Xiaoqian Jiang, Arif Harmanci
Proxy Panels Enable Privacy-Aware Outsourcing Of Genotype Imputation, Degui Zhi, Xiaoqian Jiang, Arif Harmanci
Faculty, Staff and Student Publications
One of the major challenges in genomic data sharing is protecting participants' privacy in collaborative studies and in cases when genomic data are outsourced to perform analysis tasks, for example, genotype imputation services and federated collaborations genomic analysis. Although numerous cryptographic methods have been developed, these methods may not yet be practical for population-scale tasks in terms of computational requirements, rely on high-level expertise in security, and require each algorithm to be implemented from scratch. In this study, we focus on outsourcing of genotype imputation, a fundamental task that utilizes population-level reference panels, and develop protocols that rely on using …
Modeling The Relationship Between Calories And Activity Metrics: A Regression Analysis With Variable Selection, Felix Yeboah
Modeling The Relationship Between Calories And Activity Metrics: A Regression Analysis With Variable Selection, Felix Yeboah
Data Science and Data Mining
Physical activity monitors have become integral to daily routines, with wearable devices such as the Apple Watch and Fitbit offering continuous data on users’ physical activity. This study compares the measurement accuracy of these devices by examining how they record parameters relevant to fitness and health. Employing multiple linear regression, we modeled the relationship between calories expended and a set of explanatory variables, including heart rate, steps, distance, age, activity level, weight, and device type. Evaluation of all possible variable combinations identified heart rate, steps, distance, weight, and watch type as the most effective predictors of calorie expenditure. Although the …
Classification Of Schizophrenia, Bipolar Disorder And Major Depressive Disorder With Comorbid Traits And Deep Learning Algorithms, Xiangning Chen, Yimei Lu, Joan Manuel Cue, Mira V Han, Vishwajit L Nimgaonkar, Daniel R Weinberger, Shizhong Han, Zhongming Zhao, Jingchun Chen
Classification Of Schizophrenia, Bipolar Disorder And Major Depressive Disorder With Comorbid Traits And Deep Learning Algorithms, Xiangning Chen, Yimei Lu, Joan Manuel Cue, Mira V Han, Vishwajit L Nimgaonkar, Daniel R Weinberger, Shizhong Han, Zhongming Zhao, Jingchun Chen
Faculty, Staff and Student Publications
Many psychiatric disorders share genetic liabilities, but whether these shared liabilities can be utilized to classify and differentiate psychiatric disorders remains unclear. In this study, we use polygenic risk scores (PRSs) of 42 traits comorbid with schizophrenia (SCZ), bipolar disorder (BIP), and major depressive disorder (MDD) to evaluate their utilities. We found that combining target specific PRS with PRSs of comorbid traits can improve the classification of the target disorders. Importantly, without inclusion of PRSs from targeted disorders, we can still classify SCZ (accuracy 0.710 ± 0.008, AUC 0.789 ± 0.011), BIP (accuracy 0.782 ± 0.006, AUC 0.852 ± 0.004), …
Quantified Lives: Data, Bias, And The Cost Of Categorization, Colin F. Geraghty
Quantified Lives: Data, Bias, And The Cost Of Categorization, Colin F. Geraghty
Dissertations, Theses, and Capstone Projects
This project critically examines how statistical methods and data visualization have historically been used to categorize, rank, and control human populations. By exploring the works of key figures like Francis Galton, Karl Pearson, and R.A. Fisher, it traces the evolution of categorization frameworks from their eugenic roots to modern applications in artificial intelligence and global metrics like the World Happiness Report. Through interactive visualizations and historical analysis, the project invites users to question the legitimacy of these frameworks and reflect on their impact on society. At its core, this work emphasizes the need to move beyond rigid classification systems, challenging …