Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

Machine Learning

Discipline
Institution
Publication Year
Publication
Publication Type

Articles 31 - 60 of 215

Full-Text Articles in Data Science

Toward Robust Semantic Segmentation In Levee Infrastructure Monitoring: Enhancing Accuracy With High-Fidelity Synthetic Data And Ensemble Learning, Padam Jung Thapa May 2025

Toward Robust Semantic Segmentation In Levee Infrastructure Monitoring: Enhancing Accuracy With High-Fidelity Synthetic Data And Ensemble Learning, Padam Jung Thapa

LSU New Orleans Theses and Dissertations

Abstract: Levees serve as critical flood protection structures, but failures due to inadequate maintenance and extreme water pressures have led to devastating events such as Hurricane Katrina. Manual inspections are slow, labor-intensive, and prone to human error, necessitating the development of automated solutions. This study proposes an AI-driven framework for levee inspection utilizing deep learning-based semantic segmentation to detect rutting and enhance the identification of sand boils. To address dataset limitations, high-fidelity synthetic images are generated using DreamBooth for fine-tuning, while ControlNet adds structural constraints to enhance realism and consistency. A semi-automatic convex hull annotation technique enhances labeling efficiency, and …


Car Price Prediction Using Machine Learning: Analyzing The Dvm-Car Dataset, Yaman Abu Ghareebaih May 2025

Car Price Prediction Using Machine Learning: Analyzing The Dvm-Car Dataset, Yaman Abu Ghareebaih

Electronic Theses and Dissertations

The objective of this study is to predict car prices using machine learning models and the DVM-CAR dataset, which includes over 1.4 million images and car specifi- cations from 899 car models. Key factors such as mileage, engine power, and year of registration were analyzed for their correlation with car prices. Extensive data cleaning was performed, including filling missing values, identifying outliers, and normalizing numerical variables. Discrete variables like car make and body type were encoded using one-hot encoding. Linear relationships were analyzed with Multiple Logistic Regression, and Random Forest models were used for nonlinear patterns. Model performance was evaluated …


Ml Playground: Data Modification/Preprocessing And Model Simulation Tool, Marco D. Cerrato May 2025

Ml Playground: Data Modification/Preprocessing And Model Simulation Tool, Marco D. Cerrato

Electronic Theses, Projects, and Dissertations

There is a heavy reliance on programming when it comes to learning machine learning (ML). This often creates barriers for students and newcomers unfamiliar with coding. While the lessons you learn in the classroom provide essential foundational understanding, some technical or practical aspects of ML—such as data preprocessing, feature engineering, and model tuning—are best learned through hands-on interaction. ML Playground was developed to act as a proof-of-concept application to address this gap by offering a browser-based, graphical user interface that lets users engage with core ML workflows without writing code. Designed with educational accessibility in mind, the application allows users …


Explainable Ai (Xai) For A Machine Learning Heart Disease Prediction Model, Sai Abhishek Sanchula May 2025

Explainable Ai (Xai) For A Machine Learning Heart Disease Prediction Model, Sai Abhishek Sanchula

Electronic Theses, Projects, and Dissertations

Cardiovascular diseases (CVDs) remain the leading cause of mortality worldwide, necessitating the development of accurate and interpretable machine learning (ML) models for early diagnosis and risk assessment (World Health Organization, 2021). While ML algorithms such as logistic regression, decision trees, support vector machines (SVM) (Cortes & Vapnik, 1995), and deep learning models (LeCun et al., 2015) have demonstrated high predictive accuracy, their adoption in clinical practice is hindered by their black-box nature (Rudin, 2019). Explainable AI (XAI) techniques, including SHapley Additive Explanations (SHAP) (Lundberg & Lee, 2017), Local Interpretable Model-agnostic Explanations (LIME) (Ribeiro et al., 2016), and feature importance analysis …


Mortgage Default Classification Modeling For Variable Analysis, Brendan R. Goggins Apr 2025

Mortgage Default Classification Modeling For Variable Analysis, Brendan R. Goggins

Honors College Theses

The financial crisis of the early 2000’s is a prime example of the severe consequences that mortgage default and borrower insolvency can have on economies at large. Mortgage default specifically is a prime case with the popularization of mortgage backed securities and the commonality of this loan structure. Multiple hypotheses and models have been formed to understand the reasons, causes, and consequences of mortgage default. This paper uses both machine learning and statistical classification models to inform an understanding of the variables most significant and impactful to the default outcome of mortgages. Consideration is given to both loan-level microeconomic variables …


Applying Software Engineering Black-Box Methods For Testing Machine Learning Models, Timothy Elvira Apr 2025

Applying Software Engineering Black-Box Methods For Testing Machine Learning Models, Timothy Elvira

Doctoral Dissertations and Master's Theses

This dissertation proposes researching an approach to incorporate and align Software black-box testing methods into Machine Learning (ML) applications, specifically in the context of computer vision models. Typically, testing methods within Software Engineering (SE) encompass a range of test types that assess levels of a software system, such as Unit, Integration, Functional, and System testing [1]. The testing spectrum offers two perspectives on the system: black-box, where the system’s code is hidden, and white-box, where the system's code is exposed for testing. Software Quality pairs testing with requirements, in a many-to-one relationship, to ensure proper validation of the software system. …


Credit Card Fraud Detection Via Model Retraining And Fine-Tuning, Anamol Khadka Jan 2025

Credit Card Fraud Detection Via Model Retraining And Fine-Tuning, Anamol Khadka

Computer Science and Engineering Student Research - Archive

Credit card fraud detection is a critical task in financial systems, especially given the rarity and evolving nature of the fraudulent behavior. The highly imbalanced class levels of the fraudulent and non-fraudulent transactions make it a challenging classification problem to solve. This study investigates the effectiveness of machine learning models: Logistic Regression, XGBoost, and Multi-Layer Perceptron (Neural Network), evaluated under temporal retraining and fine-tuning scenarios using a publicly available, highly imbalanced dataset of European credit card transactions. The dataset includes 284,807 transactions, of which only 492 (0.172%) are labeled as fraudulent, making it a well-known example of an imbalanced classification …


Demand Forecasting And Inventory Optimization In Mid-Sized Grocery Retail Using Machine Learning: A Data-Driven Approach To Minimizing Stock-Outs And Waste., Dragos Andrei Ungureanu Jan 2025

Demand Forecasting And Inventory Optimization In Mid-Sized Grocery Retail Using Machine Learning: A Data-Driven Approach To Minimizing Stock-Outs And Waste., Dragos Andrei Ungureanu

ICT

Mid-sized grocery retailers face a persistent challenge in balancing on-shelf availability with minimizing spoilage of perishable goods. This dissertation addresses this issue by developing a data-driven forecasting and inventory simulation framework within Microsoft Fabric, leveraging scalable data ingestion, Spark-based processing, and advanced machine learning. Using multi-year transactional data enriched with holiday schedules, promotions, and macroeconomic indicators, the study compares classical ARIMA models with XGBoost to capture complex demand patterns. Rigorous hyperparameter tuning in a distributed environment demonstrates that XGBoost outperforms baseline models in terms of MAE and MAPE, particularly during promotion-driven spikes. Inventory simulations based on these forecasts reduce stock-outs …


Temporal Machine Learning For Predicting Accidents And Violations In The Mining Industry, Nathan T. Kelley Jan 2025

Temporal Machine Learning For Predicting Accidents And Violations In The Mining Industry, Nathan T. Kelley

Theses and Dissertations--Mining Engineering

This thesis examines the predictive capability of a temporal machine learning model for forecasting future accidents and violations at individual mines, based on historical data. Mine accidents were categorized by accident classification and violations were categorized by the Part Section. The primary datasets utilized were the mine safety and health administration’s (MSHA’s) Accident Injuries and Violations datasets. The available datasets were cleaned and organized by mine type and commodity, then divided into separate subsets for training, validating, and testing. Different models, cutoff metrics, learning rates, number of hidden layers, data processing methods, data processing divisions, number of points observed …


A Comparative Study Of Machine Learning Models For Javanese Wuku Classification: Exploring Svm, Naïve Bayes, And Cnn For Cultural Texts, Danang Arbian Sulistyo, Aji Prasetya Wibawa, Didik Dwi Prasetya, Fadhli Almu'iini Ahda, Agung Bella Putra Utama Jan 2025

A Comparative Study Of Machine Learning Models For Javanese Wuku Classification: Exploring Svm, Naïve Bayes, And Cnn For Cultural Texts, Danang Arbian Sulistyo, Aji Prasetya Wibawa, Didik Dwi Prasetya, Fadhli Almu'iini Ahda, Agung Bella Putra Utama

Knowledge Engineering and Data Science

This study rigorously evaluates machine learning models for classifying culturally significant Javanese Wuku texts from the “Keagamaan atau Spiritual” category, a domain challenged by unique linguistic nuances and limited digitized resources. We compared Support Vector Machine (SVM), Naïve Bayes, and Convolutional Neural Network (CNN) on texts from five pivotal Wuku types (Sinta, Galungan, Kuningan, Sungsang, Warigalit) sourced from sastra.org, aiming to identify the most effective computational approach. The dataset comprises N = 1419 documents (T = 751.290 tokens), with per-class document counts reported for all five Wuku types. Our evaluation uses accuracy, precision, recall, F1-score, and …


Comparative Performance Of Vgg16 And Efficientnetb0-Based Transfer Learning For Brain Tumor Classification, Huzain Azis, Rizqi Ananda Jalil, Abdul Rachman Manga' Jan 2025

Comparative Performance Of Vgg16 And Efficientnetb0-Based Transfer Learning For Brain Tumor Classification, Huzain Azis, Rizqi Ananda Jalil, Abdul Rachman Manga'

Knowledge Engineering and Data Science

The classification of brain tumors using Magnetic Resonance Imaging (MRI) images is essential for early diagnosis but remains challenging due to tumor diversity. This study evaluates the effectiveness of two distinct architectural approaches for feature extraction: VGG16, representing a classic sequential design, and EfficientNetB0, a modern architecture optimized for parameter efficiency through compound scaling. Using a dataset of 2,870 MRI images categorized into four classes, we implemented a static transfer learning strategy by freezing all pre-trained ImageNet weights to act as fixed feature extractors. Features were extracted from specific layers, the final pooling layer for VGG16 and the Global Average …


Machine Learning Methods For Intrusion Detection And Response In Network Security, Ayomide Oyemaja Jan 2025

Machine Learning Methods For Intrusion Detection And Response In Network Security, Ayomide Oyemaja

College of Graduate Studies: Theses & Dissertations

Intrusion Detection Systems (IDS) play a crucial role in computer network security by identifying malicious activities and potential cyberattacks. This thesis combines machine learning and cybersecurity by applying Reinforcement Learning (RL) in intrusion detection and response using the NSL-KDD dataset.

We designed and implemented a Q-learning framework where an agent learns to classify network traffic over time by interacting with the environment and receiving rewards based on detection accuracy. We also look at the importance of feature selection and classification techniques and how effective they are in improving model performance, reducing the complexity of computation, and producing more desirable results. …


Predicting Heart Disease Using Machine Learning Models, Zeynep Cetin Jan 2025

Predicting Heart Disease Using Machine Learning Models, Zeynep Cetin

Williams Honors College, Honors Research Projects

Heart disease remains the leading cause of death in the United States, particularly among the elderly population. The growing availability of large-scale health data and the advancement of machine learning tools present an opportunity to create more accurate and individualized predictive models. This study utilizes a subset of the 2020 Behavioral Risk Factor Surveillance System (BRFSS) dataset, focusing on individuals aged 70 and above, to explore predictive modeling using logistic regression, random forests, and XGBoost. The models were evaluated using key performance metrics, including sensitivity, specificity, accuracy, and the area under the ROC curve (AUC). The findings suggest that while …


Application Of Machine Learning And Large Language Models In Healthcare For Data Prediction And Summarization, Chiazam Chisom Izuchukwu Jan 2025

Application Of Machine Learning And Large Language Models In Healthcare For Data Prediction And Summarization, Chiazam Chisom Izuchukwu

College of Graduate Studies: Theses & Dissertations

This study aims to examine the use of machine learning (ML) and large language models (LLMs) in healthcare to enhance disease prediction, clinical decision-making, and information management. Five supervised ML models—Logistic Regression (LR), Support Vector Machine (SVM), Random Forest (RF), Decision Trees (DT), and Naïve Bayes (NB)—on three different computing platforms—Google Colab, Databricks, and Snowflake—were employed for disease classification. Data preprocessing included treating missing values, encoding categorical variables utilizing one-hot-encoding, feature scaling when needed, and tackling class imbalance with Synthetic Minority Over-sampling Technique (SMOTE) before an 80-20 train-test separation. Models were created with Scikit-learn (Google Collab), Spark MLlib (Databricks), and …


Utilizing Deep Learning Audio Models For Blind And Low Vision Crosswalk Assistance, Wayne Lam Jan 2025

Utilizing Deep Learning Audio Models For Blind And Low Vision Crosswalk Assistance, Wayne Lam

Dissertations and Theses

Navigating urban environments poses significant challenges for blind and low vision (BLV) individuals, particularly at street intersections where determining when it is safe to cross can be life-threatening. In New York City, where pedestrian fatalities are on the rise and only 2% of intersections are equipped with Accessible Pedestrian Signals (APS), alternative solutions are urgently needed. This thesis proposes an audio-based deep learning approach to support BLV individuals at crosswalks by detecting traffic movement direction and idling states using spatial sound. With 4-channel audio capturing capabilities of wearables, such as Meta Project Aria glasses, we explore state-of-the-art sound event localization …


Learning From Non-Stationary Data Streams, Gabriel Jonas Aguiar Jan 2025

Learning From Non-Stationary Data Streams, Gabriel Jonas Aguiar

Theses and Dissertations

The rapid growth of data from sources such as mobile applications, sensors, and network monitoring has increased the need for machine learning algorithms capable of handling non-stationary data streams. However, learning from such streams presents significant challenges due to their evolving nature and the presence of concept drift. One of the most complex issues is learning from imbalanced data streams, where shifting data distributions, combined with feature space drifts, complicate continuous adaptation. These challenges become even more pronounced in multi-class scenarios, which are common in real-world applications. Detecting concept drift in such contexts is particularly demanding, as it requires tracking …


Unlocking The Power Of Data: Enhancing Public Policy Through Advanced Data Infrastructure And Language Model Analysis, Zahid Asghar Dec 2024

Unlocking The Power Of Data: Enhancing Public Policy Through Advanced Data Infrastructure And Language Model Analysis, Zahid Asghar

CBER Conference

Data is the fundamental building block for advancements in artificial intelligence (AI), general AI (GAI), machine learning (ML), and large language models (LLMs). This study emphasizes the critical need for robust data infrastructure, arguing that without it, countries cannot fully benefit from technological advancements in various economic sectors. Governments possess vast repositories of both structured and unstructured data across multiple domains such as the judiciary, parliaments, and civil bureaucracy. However, these potential goldmines remain untapped due to inadequate data management capabilities and a lack of appreciation for the necessity of high-quality data. The research identifies key issues in public data …


Graph Neural Networks Powered Scientific Paper Recommendation, Junhao Shen Dec 2024

Graph Neural Networks Powered Scientific Paper Recommendation, Junhao Shen

Computer Science and Engineering Theses and Dissertations

Scientific paper recommendation systems aim to help researchers discover relevant papers amidst the vast and ever-growing body of literature. With the exponential yearly increase in scientific publications, the demand for effective paper recommendation solutions has become both critical and increasingly challenging. In recent years, deep learning techniques have revolutionized recommender systems, and scientific paper recommendations have naturally integrated these advancements. In this dissertation, we address these challenges through three progressive contributions.

First, we enhance traditional content-based methods using Graph Neural Networks (GNNs) by introducing a Graph Convolutional Network-strengthened Topic Modeling (GCN-TM) approach. This method improves upon conventional topic modeling techniques …


Enhancing Password Security And Memorability Using Machine Learning And Linguistic Patterns, Jared Wise Dec 2024

Enhancing Password Security And Memorability Using Machine Learning And Linguistic Patterns, Jared Wise

LSU New Orleans Theses and Dissertations

In the digital age, text-based passwords remain a primary method for securing online accounts. Yet, users frequently face a dilemma between creating passwords that are easy to remember and sufficiently secure against cyberattacks. This research introduces an approach to password generation that bridges this gap by utilizing linguistic patterns, particularly song lyrics, to develop highly secure and naturally memorable passwords. Using large lyric datasets gained from web scrapes from popular song lyric websites (AZ Lyrics, Genius), features are extracted from a corpus of over 5 million lyrics using sentence structure and natural language processing in a novel way. In using …


Computational Representation, Analysis And Verification Of Requirements In Engineering Design And Systems Engineering, Chandan Kumar Sahu Dec 2024

Computational Representation, Analysis And Verification Of Requirements In Engineering Design And Systems Engineering, Chandan Kumar Sahu

All Dissertations

Systems are developed to satisfy a set of requirements derived from stakeholders’ needs, defining the problem space for which the system is created as a feasible solution. The system design process begins with eliciting these requirements and concludes with validating whether the created system meets them. Requirements engineering (RE) encompasses elicitation, representation, analysis, documentation, verification, and validation. However, challenges in RE, such as imprecision in natural language (NL), proprietary restrictions, and a lack of standardized quality metrics, hinder the creation of well-formed and comprehensive requirements. These challenges complicate formalization and analysis of requirements.

This dissertation addresses these challenges by proposing …


Competitive Conquest: Charting The Climb To Pokémon Supremacy, Robert Dilworth Nov 2024

Competitive Conquest: Charting The Climb To Pokémon Supremacy, Robert Dilworth

BCoE Publications

This manuscript presents a comprehensive exploration of optimizing Pokémon gameplay through data-driven methodologies, aimed at enhancing competitive performance in high-stakes environments. In the first section, we introduce a robust Pokémon teambuilding algorithm that leverages statistical analysis of championship-winning compositions. By employing multiple linear regression techniques, we predict team performance based on critical factors such as Base Stat Totals (BSTs) and various coverage types. This integration of data science principles into Pokémon strategy underscores the importance of offensive capabilities over defensive considerations, ultimately contributing to advancements in teambuilding strategies. Our proficiency in R programming facilitated the development of an efficient codebase …


Characterizing The Progression From Mild Cognitive Impairment To Dementia: A Network Analysis Of Longitudinal Clinical Visits, Muskan Garg, Sara Hejazi, Sunyang Fu, Maria Vassilaki, Ronald C Petersen, Jennifer St Sauver, Sunghwan Sohn Oct 2024

Characterizing The Progression From Mild Cognitive Impairment To Dementia: A Network Analysis Of Longitudinal Clinical Visits, Muskan Garg, Sara Hejazi, Sunyang Fu, Maria Vassilaki, Ronald C Petersen, Jennifer St Sauver, Sunghwan Sohn

Faculty, Staff and Student Publications

Background: With the recent surge in the utilization of electronic health records for cognitive decline, the research community has turned its attention to conducting fine-grained analyses of dementia onset using advanced techniques. Previous works have mostly focused on machine learning-based prediction of dementia, lacking the analysis of dementia progression and its associations with risk factors over time. The black box nature of machine learning models has also raised concerns regarding their uncertainty and safety in decision making, particularly in sensitive domains like healthcare.

Objective: We aimed to characterize the progression of health conditions, such as chronic diseases and neuropsychiatric symptoms, …


Exploring Machine Learning, Feature Engineering, And Explainability To Constrain Spica’S Apsidal Constant Through Mesa Simulations, Hannah C. Woodruff Oct 2024

Exploring Machine Learning, Feature Engineering, And Explainability To Constrain Spica’S Apsidal Constant Through Mesa Simulations, Hannah C. Woodruff

Doctoral Dissertations and Master's Theses

Spica (α-Virginis) is a notable binary star system located in the constellation of Virgo and offers valuable insights into stellar interiors and dynamics. Within binary systems, gravitational forces between the two stars cause minor distortions that alter their orbital motion. The steady rate of this alteration is known as the apsidal constant, which provides key information about a star’s internal structure and its evolutionary state. Traditionally, stellar environments like Spica are studied using simulations, such as MESA (Modules for Experiments in Stellar Astrophysics). These simulations allow researchers to explore various aspects of stellar behavior through the entire evolution of the …


Rethinking Retrieval Augmented Fine-Tuning In An Evolving Llm Landscape, Nicholas Sager, Timothy Cabaza, Matthew Cusack, Ryan Bass, Joaquin Dominguez Sep 2024

Rethinking Retrieval Augmented Fine-Tuning In An Evolving Llm Landscape, Nicholas Sager, Timothy Cabaza, Matthew Cusack, Ryan Bass, Joaquin Dominguez

SMU Data Science Review

This study explores the utilization of Retrieval Augmented Fine-Tuning (RAFT) to enhance the performance of Large Language Models (LLMs) in domain-specific Retrieval Augmented Generation (RAG) tasks. By integrating domain-specific information during the retrieval process, RAG aims to reduce hallucination and improve the accuracy of LLM outputs. We investigate the use of RAFT, an approach that enhances LLMs by incorporating domain-specific knowledge and effectively handling distractor documents. This paper validates previous work, which found that RAFT can considerably improve the performance of Llama2-7B in specific domains. We also expand upon previous work into new state-of-the-art open-source models and other datasets with …


Ensemble Pretrained Language Models To Extract Biomedical Knowledge From Literature, Zhao Li, Qiang Wei, Liang-Chin Huang, Jianfu Li, Yan Hu, Yao-Shun Chuang, Jianping He, Avisha Das, Vipina Kuttichi Keloth, Yuntao Yang, Chiamaka S Diala, Kirk E Roberts, Cui Tao, Xiaoqian Jiang, W Jim Zheng, Hua Xu Sep 2024

Ensemble Pretrained Language Models To Extract Biomedical Knowledge From Literature, Zhao Li, Qiang Wei, Liang-Chin Huang, Jianfu Li, Yan Hu, Yao-Shun Chuang, Jianping He, Avisha Das, Vipina Kuttichi Keloth, Yuntao Yang, Chiamaka S Diala, Kirk E Roberts, Cui Tao, Xiaoqian Jiang, W Jim Zheng, Hua Xu

Faculty, Staff and Student Publications

OBJECTIVES: The rapid expansion of biomedical literature necessitates automated techniques to discern relationships between biomedical concepts from extensive free text. Such techniques facilitate the development of detailed knowledge bases and highlight research deficiencies. The LitCoin Natural Language Processing (NLP) challenge, organized by the National Center for Advancing Translational Science, aims to evaluate such potential and provides a manually annotated corpus for methodology development and benchmarking.

MATERIALS AND METHODS: For the named entity recognition (NER) task, we utilized ensemble learning to merge predictions from three domain-specific models, namely BioBERT, PubMedBERT, and BioM-ELECTRA, devised a rule-driven detection method for cell line and …


Enhancing Fundraising Strategies In Higher Education Through Machine Learning, Laith Alatwah Aug 2024

Enhancing Fundraising Strategies In Higher Education Through Machine Learning, Laith Alatwah

Electrical Engineering Theses

This thesis presents a comprehensive application of machine learning techniques, namely Fine Gaussian SVM and RUS Boosted Trees, to enhance fundraising strategies in higher education institutions. Analyzing a rich dataset from Blackbaud Raiser's Edge NXT, spanning 2012 to 2022, the study focuses on donor profiles, including demographics, donation history, and engagement patterns. Key demographic insights include the increasing engagement of younger donors (20-29 age group) and significant contributions from older donors (70-99 age group). Geographical trends are also examined, revealing distinct patterns based on donors' city, state, and ZIP code. The Fine Gaussian SVM model demonstrates moderate discriminatory power, with …


Automatic Uncovering Of Patient Primary Concerns In Portal Messages Using A Fusion Framework Of Pretrained Language Modelsautomatic Uncovering Of Patient Primary Concerns In Portal Messages Using A Fusion Framework Of Pretrained Language Models, Yang Ren, Yuqi Wu, Jungwei W Fan, Aditya Khurana, Sunyang Fu, Dezhi Wu, Hongfang Liu, Ming Huang Aug 2024

Automatic Uncovering Of Patient Primary Concerns In Portal Messages Using A Fusion Framework Of Pretrained Language Modelsautomatic Uncovering Of Patient Primary Concerns In Portal Messages Using A Fusion Framework Of Pretrained Language Models, Yang Ren, Yuqi Wu, Jungwei W Fan, Aditya Khurana, Sunyang Fu, Dezhi Wu, Hongfang Liu, Ming Huang

Faculty, Staff and Student Publications

OBJECTIVES: The surge in patient portal messages (PPMs) with increasing needs and workloads for efficient PPM triage in healthcare settings has spurred the exploration of AI-driven solutions to streamline the healthcare workflow processes, ensuring timely responses to patients to satisfy their healthcare needs. However, there has been less focus on isolating and understanding patient primary concerns in PPMs-a practice which holds the potential to yield more nuanced insights and enhances the quality of healthcare delivery and patient-centered care.

MATERIALS AND METHODS: We propose a fusion framework to leverage pretrained language models (LMs) with different language advantages via a Convolution Neural …


Physics-Informed Machine Learning Methods For Inverse Design Of Multi-Phase Materials With Targeted Mechanical Properties, Yunpeng Wu Aug 2024

Physics-Informed Machine Learning Methods For Inverse Design Of Multi-Phase Materials With Targeted Mechanical Properties, Yunpeng Wu

All Dissertations

Advances in machine learning algorithms and applications have significantly enhanced engineering inverse design capabilities. This work focuses on the machine learning-based inverse design of material microstructures with targeted linear and nonlinear mechanical properties. It involves developing and applying predictive and generative physics-informed neural networks for both 2D and 3D multiphase materials.

The first investigation aims to develop a machine learning method for the inverse design of 2D multiphase materials, particularly porous materials. We first develop machine learning methods to understand the implicit relationship between a material's microstructure and its mechanical behavior. Specifically, we use ResNet-based models to predict the elastic …


Genomic Data Science Approaches For Understanding Human Diseases, Snehal Shah Aug 2024

Genomic Data Science Approaches For Understanding Human Diseases, Snehal Shah

All Dissertations

The intricate interplay of genetic predisposition, environmental influences, and lifestyle acts as the multifactorial landscape of diseases. Understanding this complexity presents a significant challenge. Molecular insights into disease mechanisms, particularly the interactions of DNA, RNA, and proteins with environmental and lifestyle factors, have revolutionized disease diagnosis, prognosis, and treatment. High-throughput technologies, such as next-generation sequencing, generate large amounts of molecular data, holding a wealth of knowledge. These datasets unveil the roles of genes and their interactions with various factors through analysis, shedding light on previously unknown molecular mechanisms underlying disease pathogenesis. Furthermore, they facilitate the discovery of biomarkers crucial for …


Study Of Prognostic Splicing Factors In Cancer Using Machine Learning Approaches, Mengyuan Yang, Jiajia Liu, Pora Kim, Xiaobo Zhou Jun 2024

Study Of Prognostic Splicing Factors In Cancer Using Machine Learning Approaches, Mengyuan Yang, Jiajia Liu, Pora Kim, Xiaobo Zhou

Faculty, Staff and Student Publications

Splicing factors (SFs) are the major RNA-binding proteins (RBPs) and key molecules that regulate the splicing of mRNA molecules through binding to mRNAs. The expression of splicing factors is frequently deregulated in different cancer types, causing the generation of oncogenic proteins involved in cancer hallmarks. In this study, we investigated the genes that encode RNA-binding proteins and identified potential splicing factors that contribute to the aberrant splicing applying a random forest classification model. The result suggested 56 splicing factors were related to the prognosis of 13 cancers, two SF complexes in liver hepatocellular carcinoma, and one SF complex in esophageal …