Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

3,235 Full-Text Articles 9,310 Authors 1,358,113 Downloads 221 Institutions

All Articles in Data Science

Faceted Search

3,235 full-text articles. Page 42 of 155.

A Framework For Human Evaluation Of Large Language Models In Healthcare Derived From Literature Review, Thomas Yu Chow Tam, Sonish Sivarajkumar, Sumit Kapoor, Alisa V Stolyar, Katelyn Polanska, Karleigh R McCarthy, Hunter Osterhoudt, Xizhi Wu, Shyam Visweswaran, Sunyang Fu, Piyush Mathur, Giovanni E Cacciamani, Cong Sun, Yifan Peng, Yanshan Wang 2024 The Texas Medical Center Library

A Framework For Human Evaluation Of Large Language Models In Healthcare Derived From Literature Review, Thomas Yu Chow Tam, Sonish Sivarajkumar, Sumit Kapoor, Alisa V Stolyar, Katelyn Polanska, Karleigh R Mccarthy, Hunter Osterhoudt, Xizhi Wu, Shyam Visweswaran, Sunyang Fu, Piyush Mathur, Giovanni E Cacciamani, Cong Sun, Yifan Peng, Yanshan Wang

Faculty, Staff and Student Publications

With generative artificial intelligence (GenAI), particularly large language models (LLMs), continuing to make inroads in healthcare, assessing LLMs with human evaluations is essential to assuring safety and effectiveness. This study reviews existing literature on human evaluation methodologies for LLMs in healthcare across various medical specialties and addresses factors such as evaluation dimensions, sample types and sizes, selection, and recruitment of evaluators, frameworks and metrics, evaluation process, and statistical analysis type. Our literature review of 142 studies shows gaps in reliability, generalizability, and applicability of current human evaluation practices. To overcome such significant obstacles to healthcare LLM developments and deployments, we …


Exploring How Uncertain Labels From Non-Consensus Panels Affect Machine Learning, Amal Almansour 2024 DePaul University

Exploring How Uncertain Labels From Non-Consensus Panels Affect Machine Learning, Amal Almansour

College of Computing and Digital Media Dissertations

A dataset becomes meaningful for analysis when it contains more representative features. Machine and deep learning models rely on annotated instances for training. The annotation process is usually done either by humans (experts or crowdsourcing) or by models. In many cases, the variability between humans (the inter-observer variability) in evaluation leads to uncertainty in the learning process. Due to the lack of reliable labels in large datasets, the inter-observer variability can be quantified with different methods to estimate the ground truth label (i.e., referenced standard label) for model learning.

In health care, with the rise of artificial intelligence in clinical …


The Aimag Project: Using Machine Learning To Predict Crustal Magnetic Anomaly Values, Xavier Gobble, Marlie Mollett, Dr. Dawn King, Dr. Cory Reed, Erin Knese 2024 University of Missouri-St. Louis

The Aimag Project: Using Machine Learning To Predict Crustal Magnetic Anomaly Values, Xavier Gobble, Marlie Mollett, Dr. Dawn King, Dr. Cory Reed, Erin Knese

Undergraduate Research Symposium

A detailed model of the Earth’s total magnetic field is important for acquiring the means for GPS-alternative, magnetic anomaly-based navigation. The Earth’s total magnetic field is an amalgam of 5 mechanisms: the geodynamo generated by the rotation of the Earth’s molten iron core, the fields induced by the flows of electric current in the atmosphere and oceans, the disturbance of the ionosphere by solar wind, and local anomalies attributable to ferromagnetic minerals present in the crust; the lattermost compose the crustal magnetic field. The EMAG2v3 dataset comprises a compilation of satellite, shipborne, and airborne magnetic measurements differenced from the Comprehensive …


Trends From 20 Years Of Artificial Intelligence In Financial Services In Africa, Nthabiseng Moela, Lerato Matlala, Jackie Ma, Dipuo Maphutha, Hossana Twinomurinzi 2024 University of Johannesburg

Trends From 20 Years Of Artificial Intelligence In Financial Services In Africa, Nthabiseng Moela, Lerato Matlala, Jackie Ma, Dipuo Maphutha, Hossana Twinomurinzi

African Conference on Information Systems and Technology

The need for financial inclusion in Africa, particularly for marginalised groups like women and small businesses, highlights the importance of leveraging Artificial Intelligence (AI). This study provides a bibliometric analysis of AI's integration into African financial services from 2003 to 2023. The key results show a significant increase in AI use, particularly in fraud detection, credit risk prediction, and stock market volatility forecasting, with 49% of the research coming from South Africa, Nigeria, and Tunisia. However, areas like financial development management, inflation control, and gender disparities in loan access remain underexplored. The emphasis has been on the technical implementation of …


Review Of Data Bias In Healthcare Applications, Atharva Prakash Parate, Aditya Ajay Iyer, Kanav Gupta, Harsh Porwal, P. C. Kishoreraja, R. Sivakumar, Rahul Soangra 2024 Vellore Institute of Technology

Review Of Data Bias In Healthcare Applications, Atharva Prakash Parate, Aditya Ajay Iyer, Kanav Gupta, Harsh Porwal, P. C. Kishoreraja, R. Sivakumar, Rahul Soangra

Physical Therapy Faculty Articles and Research

In the area of medical artificial intelligence (AI), data bias is a major difficulty that affects several phases of data collection, processing, and model building. The many forms of data bias that are common in AI in healthcare are thoroughly examined in this review study, encompassing biases related to socioeconomic status, race, and ethnicity as well as biases in machine learning models and datasets. We examine how data bias affects the provision of healthcare, emphasizing how it might worsen health inequalities and jeopardize the accuracy of AI-driven clinical tools. We address methods for reducing data bias in AI and focus …


Institutional Data Repositories Are Vital, Jen Darragh, Mikala R. Narlock, Halle Burns, Peter A. Cerda, Wind Cowles, Leslie M. Delserone, Seth Erickson, Joel Herndon, Heidi Imker, Lisa R. Johnston, Sherry Lake, Michael Lenard, Alicia Hofelich Mohr, Jennifer Moore, Jonathan Petters, Brandie Pullen, Shawna Taylor, Briana Wham 2024 Duke University

Institutional Data Repositories Are Vital, Jen Darragh, Mikala R. Narlock, Halle Burns, Peter A. Cerda, Wind Cowles, Leslie M. Delserone, Seth Erickson, Joel Herndon, Heidi Imker, Lisa R. Johnston, Sherry Lake, Michael Lenard, Alicia Hofelich Mohr, Jennifer Moore, Jonathan Petters, Brandie Pullen, Shawna Taylor, Briana Wham

University of Nebraska-Lincoln Libraries: Faculty Publications

As funding agencies and publishers reiterate research data sharing expectations (1), many higher-education institutions have demonstrated their commitment to the long-term stewardship of research data by connecting researchers to local infrastructure, with dedicated staffing, that eases the burden of data sharing. Institutional repositories are an example of this investment (2). They provide support for researchers in sharing data that might otherwise be lost: data without a disciplinary repository, data from projects with limited funding, or data that are too large to sustainably store elsewhere. The staffing and technical infrastructure provided by institutional repositories ensures responsible access to information while considering …


20 Years Of Repo Interest Rate Determination Using Ai: Global Trends And Africa, Takalani Rasalanavho, Henry Hondo, Kevin Julius, Marius Alembong, Sikelela Madonsela, Hossana Twinomurinzi 2024 University of Johannesburg

20 Years Of Repo Interest Rate Determination Using Ai: Global Trends And Africa, Takalani Rasalanavho, Henry Hondo, Kevin Julius, Marius Alembong, Sikelela Madonsela, Hossana Twinomurinzi

African Conference on Information Systems and Technology

This study investigated the application of artificial intelligence (AI) in determining repo interest rates, which play a vital role in guiding monetary policy, controlling inflation, and ensuring economic stability. Through a bibliometric review of research from 2004 to 2024, the findings highlight AI's transformative impact, particularly in forecasting, optimising repo rate decisions, and improving risk assessment for more effective monetary policy. However, the study also identifies a significant gap in AI usage for repo rate determination in African countries, with contributions largely limited to Ghana, Nigeria, Egypt, and South Africa. This underrepresentation poses a risk of Africa falling behind in …


Revolutionizing Medical Education: Harnessing Ai To Cultivate Critical Thinking Skills, Alex Zuo, Anthony L. Alanis 2024 The University of Texas Rio Grande Valley

Revolutionizing Medical Education: Harnessing Ai To Cultivate Critical Thinking Skills, Alex Zuo, Anthony L. Alanis

Research Colloquium

Our study intends to integrate AI-enabled tools in medical education, including but not limited to adaptive learning systems, virtual patients, and AI-enhanced assessment methods, to develop and foster critical thinking and problem-solving skills with the engagement of medical students. It is expected that this provides for a good environment for both teachers and students through workshops, online resources, and collaborative academic projects. We also consider AI-generated images and open educational resources that could augment curricula and personalize the experience for learners. Medical educators use storytelling, including AI for data storytelling, to package complex clinical data in approachable yet revealing ways …


Data-Driven Mathematical Modeling And Simulation Of Migration Dynamics During The Russian-Ukrainian War, Danielle Sitalo, Alonso Ogueda-Oliva, Padmanabhan Seshaiyer 2024 Texas A&M University

Data-Driven Mathematical Modeling And Simulation Of Migration Dynamics During The Russian-Ukrainian War, Danielle Sitalo, Alonso Ogueda-Oliva, Padmanabhan Seshaiyer

Spora: A Journal of Biomathematics

In this work, we employ a governing system of ordinary differential equations (ODEs) to create a mathematical model for getting insights into the dynamics of migration of Ukrainians evacuating due to war. A suitable assumption on coefficients of this model results in the well-known logistic growth. Additionally, stability analyses of equilibrium solutions for these ODEs are performed, and we employ parameter estimation techniques to identify coefficients using online datasets via both a least-squares approach as well as a physics informed neural network approach. Our findings indicate that over time, the daily influx of Ukrainian refugees to Poland stabilizes at a …


Big Geospatiotemporal Data Approaches To Monitoring And Mitigating Environmental Impacts In Agriculture, Olatunde D. Akanbi, Vibha S. Mandayam, Erika I. Barcelos, Arafath Nihar, Yinghui Wu, Jeffrey Yarus, Roger H. French 2024 Case Western Reserve University

Big Geospatiotemporal Data Approaches To Monitoring And Mitigating Environmental Impacts In Agriculture, Olatunde D. Akanbi, Vibha S. Mandayam, Erika I. Barcelos, Arafath Nihar, Yinghui Wu, Jeffrey Yarus, Roger H. French

Student Scholarship

This research explores the application of geospatial techniques for global agricultural monitoring, integrating satellite imagery and soil data to assess crop health and soil conditions. Our approach provides actionable insights to improve agricultural productivity and sustainability, addressing food security challenges through advanced machine learning models.


Supplementary Files For "Impact Of Snow Accumulation On Structural Integrity: Present And Future Perspectives", Kenneth Pomeyie, Brennan Bean 2024 Utah State University

Supplementary Files For "Impact Of Snow Accumulation On Structural Integrity: Present And Future Perspectives", Kenneth Pomeyie, Brennan Bean

Browse all Datasets

Evaluating the impact of weight exerted by settled snow (i.e., snow load) on structures poses numerous statistical challenges, including missing data, biased distribution parameters, and the influence of climate change. This dissertation aims to address challenges related to the use both direct and indirect measurements of snow load (or equivalently, snow water equivalent), as well as the anticipated impact of climate change on future extreme snow loads. The first paper within this dissertation investigates short-term snow loads by comparing various techniques for estimating extreme values of short-term snow accumulations. Additionally, the first paper includes a comparative analysis of short-term and …


Design And Implementation Of An Opioid Scorecard For Hospital System-Wide Peer Comparison Of Opioid Prescribing Habits: Observational Study, Benjamin Slovis, Soonyip Huang, Melanie McArthur, Cara Martino, Tasia Beers, Meghan Labella, Jeffrey Riggio, Edmund Pribitkin 2024 Thomas Jefferson University

Design And Implementation Of An Opioid Scorecard For Hospital System-Wide Peer Comparison Of Opioid Prescribing Habits: Observational Study, Benjamin Slovis, Soonyip Huang, Melanie Mcarthur, Cara Martino, Tasia Beers, Meghan Labella, Jeffrey Riggio, Edmund Pribitkin

Jefferson Hospital Staff Papers and Presentations

BACKGROUND: Reductions in opioid prescribing by health care providers can lead to a decreased risk of opioid dependence in patients. Peer comparison has been demonstrated to impact providers' prescribing habits, though its effect on opioid prescribing has predominantly been studied in the emergency department setting.

OBJECTIVE: The purpose of this study is to describe the development of an enterprise-wide opioid scorecard, the architecture of its implementation, and plans for future research on its effects.

METHODS: Using data generated by the author's enterprise vendor-based electronic health record, the enterprise analytics software, and expertise from a dedicated group of informaticists, physicians, and …


A Case Demonstration Of The Open Health Natural Language Processing Toolkit From The National Covid-19 Cohort Collaborative And The Researching Covid To Enhance Recovery Programs For A Natural Language Processing System For Covid-19 Or Postacute Sequelae Of Sars Cov-2 Infection: Algorithm Development And Validation, Andrew Wen, Liwei Wang, Huan He, Sunyang Fu, Sijia Liu, David A Hanauer, Daniel R Harris, Ramakanth Kavuluru, Rui Zhang, Karthik Natarajan, Nishanth P Pavinkurve, Janos Hajagos, Sritha Rajupet, Veena Lingam, Mary Saltz, Corey Elowsky, Richard A Moffitt, Farrukh M Koraishy, Matvey B Palchuk, Jordan Donovan, Lora Lingrey, Garo Stone-DerHagopian, Robert T Miller, Andrew E Williams, Peter J Leese, Paul I Kovach, Emily R Pfaff, Mikhail Zemmel, Robert D Pates, Nick Guthe, Melissa A Haendel, Christopher G Chute, Hongfang Liu, National COVID Cohort Collaborative, RECOVER Initiative 2024 The Texas Medical Center Library

A Case Demonstration Of The Open Health Natural Language Processing Toolkit From The National Covid-19 Cohort Collaborative And The Researching Covid To Enhance Recovery Programs For A Natural Language Processing System For Covid-19 Or Postacute Sequelae Of Sars Cov-2 Infection: Algorithm Development And Validation, Andrew Wen, Liwei Wang, Huan He, Sunyang Fu, Sijia Liu, David A Hanauer, Daniel R Harris, Ramakanth Kavuluru, Rui Zhang, Karthik Natarajan, Nishanth P Pavinkurve, Janos Hajagos, Sritha Rajupet, Veena Lingam, Mary Saltz, Corey Elowsky, Richard A Moffitt, Farrukh M Koraishy, Matvey B Palchuk, Jordan Donovan, Lora Lingrey, Garo Stone-Derhagopian, Robert T Miller, Andrew E Williams, Peter J Leese, Paul I Kovach, Emily R Pfaff, Mikhail Zemmel, Robert D Pates, Nick Guthe, Melissa A Haendel, Christopher G Chute, Hongfang Liu, National Covid Cohort Collaborative, Recover Initiative

Faculty, Staff and Student Publications

BACKGROUND: A wealth of clinically relevant information is only obtainable within unstructured clinical narratives, leading to great interest in clinical natural language processing (NLP). While a multitude of approaches to NLP exist, current algorithm development approaches have limitations that can slow the development process. These limitations are exacerbated when the task is emergent, as is the case currently for NLP extraction of signs and symptoms of COVID-19 and postacute sequelae of SARS-CoV-2 infection (PASC).

OBJECTIVE: This study aims to highlight the current limitations of existing NLP algorithm development approaches that are exacerbated by NLP tasks surrounding emergent clinical concepts and …


Data Analysis On Predicting The Top 12 Fantasy Football Players By Position, Alan Abadzic, Jacquelyn Cheun, Milan Patel 2024 Southern Methodist University

Data Analysis On Predicting The Top 12 Fantasy Football Players By Position, Alan Abadzic, Jacquelyn Cheun, Milan Patel

SMU Data Science Review

Fantasy football enthusiasts rely on rankings populated by their platform of choice to draft winning teams and make strategic roster decisions. This study presents a comprehensive analysis of player performance data to forecast the top 12 fantasy points performers per position for the upcoming season. Leveraging machine learning techniques and historical data, our model identifies key performance indicators and trends to inform player evaluations. Insights gleaned from positional trends, breakout candidates, risk assessment, and matchup analysis offer a competitive edge. By addressing limitations, ethical considerations, and avenues for future research, this study contributes to the advancement of fantasy sports analysis …


Enhancing Imputation Accuracy: A Multi-Faceted Approach For Missing Data In Chicago Arrest Records, Steve Bramhall, Jae Chung, Nicholas Mueller 2024 Southern Methodist University

Enhancing Imputation Accuracy: A Multi-Faceted Approach For Missing Data In Chicago Arrest Records, Steve Bramhall, Jae Chung, Nicholas Mueller

SMU Data Science Review

This paper introduces a novel approach to enhance the imputation process for missing data, utilizing crime records from Chicago with arrests as the target feature. Robust imputation techniques are crucial in the era of burgeoning datasets for generating reliable insights. Our core objective is to present an innovative method that improves imputation techniques, augmenting model performance and bolstering the reliability of analytical outcomes. Leveraging numeric crime data, we establish a Gradient Boosting (GBM) baseline model, then introduce ensemble methods including Random Forest and Decision Trees for further refinement. By systematically exploring multiple imputation processes, we establish a baseline for comparative …


Geospatial Temporal Crime Prediction Using Convolution And Lstm Neural Networks: Enhancing The Las Vegas Cardiff Model, Corey D. Holmes, Christian Orji, Chris Papesh 2024 Southern Methodist University

Geospatial Temporal Crime Prediction Using Convolution And Lstm Neural Networks: Enhancing The Las Vegas Cardiff Model, Corey D. Holmes, Christian Orji, Chris Papesh

SMU Data Science Review

According to the Department of Justice, more than half of violent crimes go unreported to law enforcement in the United States (Kollar et al., 2018). This data gap reduces the opportunity to implement proven solutions in the areas with the greatest need. In 1996, Dr. Shepherd developed the Cardiff Model with the aim of bringing together hospitals, law enforcement, and community leaders through the sharing of data. We partnered with ongoing efforts to implement the Cardiff Model in Las Vegas, Nevada. Our goal was to provide a geospatial temporal model that can predict the next 30 days of crime. By …


Rethinking Retrieval Augmented Fine-Tuning In An Evolving Llm Landscape, Nicholas Sager, Timothy Cabaza, Matthew Cusack, Ryan Bass, Joaquin Dominguez 2024 Southern Methodist University

Rethinking Retrieval Augmented Fine-Tuning In An Evolving Llm Landscape, Nicholas Sager, Timothy Cabaza, Matthew Cusack, Ryan Bass, Joaquin Dominguez

SMU Data Science Review

This study explores the utilization of Retrieval Augmented Fine-Tuning (RAFT) to enhance the performance of Large Language Models (LLMs) in domain-specific Retrieval Augmented Generation (RAG) tasks. By integrating domain-specific information during the retrieval process, RAG aims to reduce hallucination and improve the accuracy of LLM outputs. We investigate the use of RAFT, an approach that enhances LLMs by incorporating domain-specific knowledge and effectively handling distractor documents. This paper validates previous work, which found that RAFT can considerably improve the performance of Llama2-7B in specific domains. We also expand upon previous work into new state-of-the-art open-source models and other datasets with …


Enhancing Shap With Multi-Core Parallelization And Distributed Computation, matthew david, William Jones, Hayley Horn 2024 Southern Methodist University

Enhancing Shap With Multi-Core Parallelization And Distributed Computation, Matthew David, William Jones, Hayley Horn

SMU Data Science Review

In recent years, the adoption of complex machine learning algorithms, often perceived as “black box” models, has grown exponentially across various disciplines. However, the lack of understanding regarding how these models come to their predictions often fosters skepticism and mistrust. In response to the demand for transparency and interpretability, Explainable AI techniques, such as SHapley Additive exPlanations (SHAP), have emerged as powerful tools for comprehending and trusting these algorithms. However, SHAP has an exponential computational demand O( x2 ), where x is the number of features. This becomes increasingly problematic with the larger datasets standard in most industries. Many frameworks …


Ensemble Pretrained Language Models To Extract Biomedical Knowledge From Literature, Zhao Li, Qiang Wei, Liang-Chin Huang, Jianfu Li, Yan Hu, Yao-Shun Chuang, Jianping He, Avisha Das, Vipina Kuttichi Keloth, Yuntao Yang, Chiamaka S Diala, Kirk E Roberts, Cui Tao, Xiaoqian Jiang, W Jim Zheng, Hua Xu 2024 The Texas Medical Center Library

Ensemble Pretrained Language Models To Extract Biomedical Knowledge From Literature, Zhao Li, Qiang Wei, Liang-Chin Huang, Jianfu Li, Yan Hu, Yao-Shun Chuang, Jianping He, Avisha Das, Vipina Kuttichi Keloth, Yuntao Yang, Chiamaka S Diala, Kirk E Roberts, Cui Tao, Xiaoqian Jiang, W Jim Zheng, Hua Xu

Faculty, Staff and Student Publications

OBJECTIVES: The rapid expansion of biomedical literature necessitates automated techniques to discern relationships between biomedical concepts from extensive free text. Such techniques facilitate the development of detailed knowledge bases and highlight research deficiencies. The LitCoin Natural Language Processing (NLP) challenge, organized by the National Center for Advancing Translational Science, aims to evaluate such potential and provides a manually annotated corpus for methodology development and benchmarking.

MATERIALS AND METHODS: For the named entity recognition (NER) task, we utilized ensemble learning to merge predictions from three domain-specific models, namely BioBERT, PubMedBERT, and BioM-ELECTRA, devised a rule-driven detection method for cell line and …


Improving Large Language Models For Clinical Named Entity Recognition Via Prompt Engineering, Yan Hu, Qingyu Chen, Jingcheng Du, Xueqing Peng, Vipina Kuttichi Keloth, Xu Zuo, Yujia Zhou, Zehan Li, Xiaoqian Jiang, Zhiyong Lu, Kirk Roberts, Hua Xu 2024 The Texas Medical Center Library

Improving Large Language Models For Clinical Named Entity Recognition Via Prompt Engineering, Yan Hu, Qingyu Chen, Jingcheng Du, Xueqing Peng, Vipina Kuttichi Keloth, Xu Zuo, Yujia Zhou, Zehan Li, Xiaoqian Jiang, Zhiyong Lu, Kirk Roberts, Hua Xu

Faculty, Staff and Student Publications

IMPORTANCE: The study highlights the potential of large language models, specifically GPT-3.5 and GPT-4, in processing complex clinical data and extracting meaningful information with minimal training data. By developing and refining prompt-based strategies, we can significantly enhance the models' performance, making them viable tools for clinical NER tasks and possibly reducing the reliance on extensive annotated datasets.

OBJECTIVES: This study quantifies the capabilities of GPT-3.5 and GPT-4 for clinical named entity recognition (NER) tasks and proposes task-specific prompts to improve their performance.

MATERIALS AND METHODS: We evaluated these models on 2 clinical NER tasks: (1) to extract medical problems, treatments, …


Digital Commons powered by bepress