Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 1231 - 1260 of 3233

Full-Text Articles in Data Science

Using Unsupervised Learning Methods In Extracting Features For Classifying Rice Varieties From Rice Grains Images., Kevin Anthony Martinez Jan 2024

Using Unsupervised Learning Methods In Extracting Features For Classifying Rice Varieties From Rice Grains Images., Kevin Anthony Martinez

ICT

Rice, a staple food for nearly half of the global population, requires accurate classification of its varieties to ensure food quality, support agricultural trade, and enhance yield optimisation. Traditional manual classification methods are time-intensive and error-prone, prompting this study's exploration of unsupervised learning for feature extraction from rice grain images. The research tested classifiers on 75,000 rice samples across five classes, with 15,000 samples per class.

The study's DCGAN-CNN model achieved the highest classification accuracy of 99.67%. However, the PCA-CNN model underperformed, with only 20% accuracy, due to implementation errors. Recommendations for improvement include optimising model parameters such as learning …


Judging Our New Judges: Why We Must Remove Artificial Intelligence From Our Courtrooms Now, Kieran Duffy Newcomb Jan 2024

Judging Our New Judges: Why We Must Remove Artificial Intelligence From Our Courtrooms Now, Kieran Duffy Newcomb

Honors Theses and Capstones

In this paper, I explore some of the ways in which artificial intelligence might enhance the sentencing process through recidivism prediction technology. Notably, this technology can increase the accuracy of risk predictions and the speed with which sentencing decisions are reached. I then show, however, that the recidivism prediction technology is likely to turn into what data scientist Cathy O’Neil calls a Weapon of Math Destruction. The potential harmfulness of this technology is due not to the inherent nature of the technology, but the symbiotic relationship it will have with our already harmful criminal justice system. I argue that the …


A Deep Learning Model For Early Diagnosis Of Systemic Lupus Erythematosus From Facial Images, Shourav Bikash Dey Jan 2024

A Deep Learning Model For Early Diagnosis Of Systemic Lupus Erythematosus From Facial Images, Shourav Bikash Dey

All Graduate Theses, Dissertations, and Other Capstone Projects

Systemic Lupus Erythematosus (SLE) poses significant challenges due to its complex and varied symptoms making diagnosis extremely challenging and time consuming. Symptoms of SLE often mimics other autoimmune or physical conditions and around 5 million people worldwide suffers from this condition, as reported by the Lupus Foundation of America during their study in 2019. However, diagnosis is much more difficult in developing countries with backdated clinical technology and setup therefore, making it virtually unknown the exact number of SLE patient count worldwide. Among all the heterogeneous symptoms presented by SLE, Butterfly Malar Rash (BMR) is one of the symptoms that …


Comparative Analysis Of Data Augmentation On Sentiment Analysis In Three Distinct Languages, Hyesu Lee Jan 2024

Comparative Analysis Of Data Augmentation On Sentiment Analysis In Three Distinct Languages, Hyesu Lee

All Graduate Theses, Dissertations, and Other Capstone Projects

Machine learning in natural language processing analyzes datasets to make future predictions for various filed in the real world. By training machine algorithms on the datasets of text, the model can learn patterns and structure of the text in many different languages. Then the model enables to perform the text classification, sentiment analysis, and other tasks. A large and balanced dataset is required to develop an accurate machine learning model. However, the collection of a reliable, large, and equally distributed dataset is a challenging and requires significant resources and time. As a solution to this challenge, a data augmentation technique …


Applying Neural Networks To Predict Factors Affecting Harmful Algal Blooms For Timely Alerting And Implementing Preventive Measures In Ireland's Marine Ecosystem., Nikolai Potapov Jan 2024

Applying Neural Networks To Predict Factors Affecting Harmful Algal Blooms For Timely Alerting And Implementing Preventive Measures In Ireland's Marine Ecosystem., Nikolai Potapov

ICT

This study applies neural networks to predict harmful algal blooms (HABs) along the Irish coast, addressing ecological, health, and economic risks. Using primary interviews and secondary data on HAB species like Alexandrium and Karenia mikimotoi, the research incorporated Exploratory Data Analysis and tested three neural models: LSTM, Ensemble Stacking LSTM, and CNN-LSTM. Key factors influencing HABs, such as sea surface temperature and euphotic zone depth, were identified.

Results demonstrate the potential of neural networks to improve HAB prediction and monitoring, despite limitations. Future work aims to enhance model accuracy and integrate them into HAB warning systems.


Assessment Of The Impact Of Various Feature Extraction Techniques On The Effectiveness Of Music Genre Classification In Neural Network Models., Sabhdh Grace Jan 2024

Assessment Of The Impact Of Various Feature Extraction Techniques On The Effectiveness Of Music Genre Classification In Neural Network Models., Sabhdh Grace

ICT

This research focuses on Music Genre Classification (MGC) using Convolutional Neural Networks (CNNs) and various datasets, including raw audio files (WAV) and extracted features such as Mel Spectrograms (MS), Mel-Frequency Cepstral Coefficients (MFCC), and Chroma Features (CF). The study employs Explanatory Sequential Mixed Methods (ESMM), combining qualitative research and experimental analysis to explore different model inputs and their performance. Several CNN-based models, including 2D CNN, 2D CNN-LSTM, 1D CNN, and 1D CNN-LSTM, were tested. However, the models generally underperformed, with most achieving accuracy of 10% or lower, and the best model (raw audio 1D CNN) reaching only 20%. The research …


Deep Learning Model Compression For Resource-Constrained Environments., Stephen Burke Jan 2024

Deep Learning Model Compression For Resource-Constrained Environments., Stephen Burke

ICT

This study examines the effects of three Deep Neural Network compression techniques—Quantisation, Pruning, and Weight Sharing/Clustering—on CNN and ANN models trained for image classification tasks. The models were tested on the CIFAR-10 dataset for multiclass classification and a binary classification task using a dataset derived from COCO. The best validation accuracy achieved was 74.7% with a CNN on CIFAR-10 and 53% with the best ANN. On the COCO dataset, a modified CIFAR-10 CNN model achieved 75%. The models were compressed using the three techniques and benchmarked on a ThinkPad laptop and Raspberry Pi 3B+ based on metrics relevant for resource-constrained …


Data Analysis Of Twitter’S Nasdaq100 Sentiments And Topics As Indicators For News Articles Retrieval: Fine-Tuning Roberta And Rag., Kagan Timur Jan 2024

Data Analysis Of Twitter’S Nasdaq100 Sentiments And Topics As Indicators For News Articles Retrieval: Fine-Tuning Roberta And Rag., Kagan Timur

ICT

This study investigates the combination of sentiment analysis using the VADER lexicon and semantic analysis through Latent Dirichlet Allocation (LDA) to identify real-life events, focusing on Twitter datasets. The research shows that while sentiment analysis alone may be insufficient, combining it with semantic analysis improves the process, particularly for identifying relevant news articles and understanding brand perception on social media. The study also fine-tunes the RoBERTa model for question-answering tasks, yielding significant improvements in the SQuAD evaluation metric. The exact match (EM) score rose dramatically from 2.06% to 62%, and the F1 score improved from 9.41% to 65%. A retrieval …


Development And Optimisation Of Convolutional Neural Networks (Cnns) To Predict The Nutrition And Sustainability Scores Of Foods From Crowd Sourced Images., Cormac Mcelhinney Jan 2024

Development And Optimisation Of Convolutional Neural Networks (Cnns) To Predict The Nutrition And Sustainability Scores Of Foods From Crowd Sourced Images., Cormac Mcelhinney

ICT

This research explores the use of Convolutional Neural Networks (CNNs) for the automated classification and profiling of food products based on publicly sourced data. With the vast array of food products available worldwide and the complexity of labelling regulations, food business operators face challenges in ensuring compliance, while regulators struggle to verify adherence. This study addresses the need for efficient and accurate methods for food classification and eco/nutritional profiling. It begins with a comprehensive literature review on the application of CNNs in food product classification, followed by the collection of a large-scale dataset from Open Food Facts. A CNN architecture …


Evaluating The Performance Of Different Long Short-Term Memory Networks (Lstm’S) On Financial Timeseries Data Using Mean Squared Error In Order To Identify The Optimum Lstm Variant For Regression Performance On Financial Timeseries Data., Patrick O’ Connor Jan 2024

Evaluating The Performance Of Different Long Short-Term Memory Networks (Lstm’S) On Financial Timeseries Data Using Mean Squared Error In Order To Identify The Optimum Lstm Variant For Regression Performance On Financial Timeseries Data., Patrick O’ Connor

ICT

This study explores the use of Long Short Term Memory (LSTM) networks, a variant of Recurrent Neural Networks (RNNs), in the context of financial forecasting, specifically oil price prediction. The research follows the CRISP-DM (Cross-Industry Standard Process for Data Mining) methodology and tests six different LSTM variants. The models are evaluated based on Mean Squared Error (MSE), aiming to determine the optimal parameter settings for each LSTM type. Among the variants tested, the Gated Recurrent Unit (GRU) emerged as the highest performer, achieving an MSE of 0.100. This was surprising, as simpler variants outperformed more complex ones, suggesting that simpler …


Ml Predictive Model For Earthquakes Integrating Mass, Distance, Gravity, And Magnitude., Aadarsh Kushwaha Jan 2024

Ml Predictive Model For Earthquakes Integrating Mass, Distance, Gravity, And Magnitude., Aadarsh Kushwaha

ICT

This research investigates the application of machine learning regression models to improve earthquake prediction by integrating geophysical and astronomical factors such as Earth-Moon gravitational forces, their varying distances, and localized gravity fluctuations. Using data from 2011 to 2024, sourced from the US Geological Survey (USGS) and web scraping, the study tested models across four dataset proportions (25%, 50%, 80%, and 100%) with a 70:30 train-test split. The XGBRegressor model emerged as the best performer, achieving an R² score of 0.8706 on training data and 0.8632 on test data, along with a Mean Squared Error (MSE) of 0.1114 and Mean Absolute …


Maize Crop Pests And Diseases Classification Using Hybrid Models., Diana Flora Namaemba Jan 2024

Maize Crop Pests And Diseases Classification Using Hybrid Models., Diana Flora Namaemba

ICT

This research focuses on improving the detection and classification of maize crop pests and diseases to enhance agricultural yield and food security. A dataset comprising 5389 images of maize conditions (healthy, pest-affected, and disease-affected) across seven classes was used. The images underwent preprocessing, including resizing to 299x299, class balancing using augmentation techniques, and noise reduction with Gaussian filtering.

Feature extraction utilised EfficientNetB0 and InceptionV3 architectures, with PCA employed for feature selection. Classification was conducted using a Support Vector Machine (SVM) with a One-vs-One strategy, alongside a baseline 2D CNN model. Data engineering included label encoding, standardisation, and an 80:10:10 train-test-validation …


The Hazard Prediction Problem, Mary E. Helander, Brendan Smith, Sylvia Charchut, Erika Swiatowy, Calvin Nau, Gregory Cavaretta, Timothy Schuler, Adam Schunk, Héctor Ortiz-Peña Jan 2024

The Hazard Prediction Problem, Mary E. Helander, Brendan Smith, Sylvia Charchut, Erika Swiatowy, Calvin Nau, Gregory Cavaretta, Timothy Schuler, Adam Schunk, Héctor Ortiz-Peña

Social Science - All Scholarship

This work formulates the hazard prediction problem while addressing the research question: Can machine learning create a model to automatically recognize patterns that correspond to hazard state conditions during a mission-critical operation? Supervised learning models were trained and tested on data observed from mission simulators, which allowed for safe observation of dynamic system states and undesirable casualty events. The prediction task was formulated as a binary classification problem, producing the probability of being in a hazard state at time t and providing situational awareness of a possible imminent loss. Several modeling architectures were investigated: neural networks, logistic regression, a support …


Machine Learning Approaches For Cyberbullying Detection, Roland Fiagbe Jan 2024

Machine Learning Approaches For Cyberbullying Detection, Roland Fiagbe

Data Science and Data Mining

Cyberbullying refers to the act of bullying using electronic means and the internet. In recent years, this act has been identifed to be a major problem among young people and even adults. It can negatively impact one’s emotions and lead to adverse outcomes like depression, anxiety, harassment, and suicide, among others. This has led to the need to employ machine learning techniques to automatically detect cyberbullying and prevent them on various social media platforms. In this study, we want to analyze the combination of some Natural Language Processing (NLP) algorithms (such as Bag-of-Words and TFIDF) with some popular machine learning …


Data Science In Finance: Challenges And Opportunities, Xianrong Zheng, Elizabeth Gildea, Sheng Chai, Tongxiao Zhang, Shuxi Wang Jan 2024

Data Science In Finance: Challenges And Opportunities, Xianrong Zheng, Elizabeth Gildea, Sheng Chai, Tongxiao Zhang, Shuxi Wang

Information Technology & Decision Sciences Faculty Publications

Data science has become increasingly popular due to emerging technologies, including generative AI, big data, deep learning, etc. It can provide insights from data that are hard to determine from a human perspective. Data science in finance helps to provide more personal and safer experiences for customers and develop cutting-edge solutions for a company. This paper surveys the challenges and opportunities in applying data science to finance. It provides a state-of-the-art review of financial technologies, algorithmic trading, and fraud detection. Also, the paper identifies two research topics. One is how to use generative AI in algorithmic trading. The other is …


A Novel K-Nearest Neighbors Method Based On Generalized Feature Optimization For Precipitation Forecasting, Sean Guidry Stanteen Jan 2024

A Novel K-Nearest Neighbors Method Based On Generalized Feature Optimization For Precipitation Forecasting, Sean Guidry Stanteen

Mathematics Dissertations - Archive

This study introduces a novel k-nearest neighbors (kNN) method of forecasting precipitation at weather-observing stations. The method identifies numerous monthly temporal patterns to produce precipitation forecasts for a specific month. Compared to climatological forecasts, which average the observed precipitation over the prior thirty years, and other existing contemporary iterations of kNN, the proposed novel kNN method produces more accurate forecasts on a consistent basis. Specifically, the novel kNN method produces improved root mean square errors (RMSE), mean relative errors, and Nash-Sutcliffe coefficients when compared to climatological and other kNN forecasts at five weather …


Integrating Machine Learning With Cure Models And Associated Inference, Wisdom Aselisewine Jan 2024

Integrating Machine Learning With Cure Models And Associated Inference, Wisdom Aselisewine

Mathematics Dissertations - Archive

Recent advancements in medical treatments have significantly enhanced the rates of recovery for numerous chronic illnesses. This progress has sparked growing interest in developing suitable statistical models capable of handling survival data that includes substantial cure fractions. The mixture cure model finds extensive application in analyzing survival data when there exists a cured subgroup. Standard logistic regression-based approaches for modeling the incidence part of the mixture cure model may suffer from poor predictive accuracy, especially in the presence of high dimensional covariates and/or non-linear covariate effects. To overcome this limitation, we propose the integration of distinct machine learning algorithms with …


Language Models For Rare Disease Information Extraction: Empirical Insights And Model Comparisons, Shashank Gupta Jan 2024

Language Models For Rare Disease Information Extraction: Empirical Insights And Model Comparisons, Shashank Gupta

Theses and Dissertations--Computer Science

End-to-end relation extraction (E2ERE) is a crucial task in natural language processing (NLP) that involves identifying and classifying semantic relationships between entities in text. This thesis compares three paradigms for end-to-end relation extraction (E2ERE) in biomedicine, focusing on rare diseases with discontinuous and nested entities. We evaluate Named Entity Recognition (NER) to Relation Extraction (RE) pipelines, sequence-to-sequence models, and generative pre-trained transformer (GPT) models using the RareDis information extraction dataset. Our findings indicate that pipeline models are the most effective, followed closely by sequence-to-sequence models. GPT models, despite having eight times as many parameters, perform worse than sequence-to-sequence models and …


Three Essays On Energy Related To State Policies And Low Carbon Transitions, Pinky Thomas Jan 2024

Three Essays On Energy Related To State Policies And Low Carbon Transitions, Pinky Thomas

Graduate Theses, Dissertations, and Problem Reports (ETD)

This dissertation consists of three essays on energy-related state policies and energy transition. Each paragraph below refers to the three abstracts for the three chapters in this dissertation, respectively.

The first essay is entitled: “Impacts of State Tax and Resource Ownership Policies on Extraction: Evidence from U.S. Natural Gas Production”. The innovation of combined use of horizontal drilling and hydraulic fracturing technologies during the 2000s has allowed natural gas producers in the United States to extract natural gas and liquids from deep shale formations in a cost-efficient manner. This essay evaluates whether unconventional gas production responds to tax changes, and …


Investigation Of Space Charge Effects On Co2 Electrocatalytic Reduction On Gd-Doped Ceria Via Scanning Kelvin Probe And Model-Based Bayesian Analysis, Alejandro Mejia Jan 2024

Investigation Of Space Charge Effects On Co2 Electrocatalytic Reduction On Gd-Doped Ceria Via Scanning Kelvin Probe And Model-Based Bayesian Analysis, Alejandro Mejia

Graduate Theses, Dissertations, and Problem Reports (ETD)

In studying novel energy conversion and storage systems, such as high-temperature electrolysis, numerous underlying fundamental physical processes remain unclear or inadequately understood. Among these, the modeling and comprehension of surface reaction mechanisms, coupled with the intricate effects of space‑charge interfaces, remains an unclear and challenging area of research.

The work of this dissertation involves the development of a 2D finite element analysis model, leveraging the robust MOOSE framework from INL. This model, featuring inhomogeneous defect thermodynamics for near-surface chemistry, formulated through Poisson‑Cahn variational theory, has been exploited for studying the electrocatalytic reduction of CO2 on gadolinia doped ceria. The …


Efficient Classification Of Very High Resolution Images, Mohammad I. Nouyed Jan 2024

Efficient Classification Of Very High Resolution Images, Mohammad I. Nouyed

Graduate Theses, Dissertations, and Problem Reports (ETD)

In recent decades, deep learning approaches have shown significant improvement in various image understanding tasks. However, analysis of high-resolution images remains a major challenge. In this work, we address the challenge of very high-resolution histopathological image (VHRHI) classification using a new information-theoretic discriminative patch selection approach. We show results on a high-resolution image dataset, namely, gigapixel whole slide tissue images for cancer tumors. Then we address how to efficiently classify challenging histopathology images, such as gigapixel whole-slide images for cancer diagnostics with image-level annotation. These ``weak labels'' are applied throughout the image but describe tumor regions of variable sizes and …


A Bayesian Inversion For Emissions And Export Productivity Across The End-Cretaceous Boundary, Alexander A. Cox Jan 2024

A Bayesian Inversion For Emissions And Export Productivity Across The End-Cretaceous Boundary, Alexander A. Cox

Dartmouth College Master’s Theses

The end-Cretaceous mass extinction was marked by both the Chicxulub impact and the ongoing emplacement of the Deccan Traps flood basalt province. Both of these events perturbed the environment by the emission of climate-active volatiles, primarily CO2 and SO2. To understand the mechanism of extinction, we must disentangle the timing, duration, and intensity of volcanic and meteoritic environmental forcings. In this thesis, we used a parallel Markov chain Monte Carlo approach to invert for the aforementioned volatile emissions, export productivity, and remineralization from 67 to 65 million years ago using the LOSCAR (Long-term Ocean-atmosphere-Sediment CArbon cycle Reservoir) model. The parallel …


Erratum: Toward Standardization, Harmonization, And Integration Of Social Determinants Of Health Data: A Texas Clinical And Translational Science Award Institutions Collaboration - Corrigendum, Catherine K Craven, Linda Highfield, Mujeeb Basit, Elmer V Bernstam, Byeong Yeob Choi, Robert L Ferrer, Jonathan A Gelfond, Sandi L Pruitt, Vaishnavi Kannan, Paula K Shireman, Heidi Spratt, Kayla J Torres Morales, Chen-Pin Wang, Zhan Wang, Meredith N Zozus, Edward C Sankary, Susanne Schmidt Jan 2024

Erratum: Toward Standardization, Harmonization, And Integration Of Social Determinants Of Health Data: A Texas Clinical And Translational Science Award Institutions Collaboration - Corrigendum, Catherine K Craven, Linda Highfield, Mujeeb Basit, Elmer V Bernstam, Byeong Yeob Choi, Robert L Ferrer, Jonathan A Gelfond, Sandi L Pruitt, Vaishnavi Kannan, Paula K Shireman, Heidi Spratt, Kayla J Torres Morales, Chen-Pin Wang, Zhan Wang, Meredith N Zozus, Edward C Sankary, Susanne Schmidt

Faculty, Staff and Student Publications

This corrects the article "Toward standardization, harmonization, and integration of social determinants of health data: A Texas Clinical and Translational Science Award institutions collaboration" in volume 8, e17.


Predicting Endothelium-Dependent Diastolic Function (Fmd)And Its Correlation With The Degree Of Coronary Artery Disease (Cad) And Plaque Vulnerability For Cardiovascular Events, Guangming Zhang, Jing Yang, Hanghang Xing, Hongning Yin, Guoqing Gu Jan 2024

Predicting Endothelium-Dependent Diastolic Function (Fmd)And Its Correlation With The Degree Of Coronary Artery Disease (Cad) And Plaque Vulnerability For Cardiovascular Events, Guangming Zhang, Jing Yang, Hanghang Xing, Hongning Yin, Guoqing Gu

Faculty, Staff and Student Publications

OBJECTIVE: This study aims to investigate the correlation between vascular endothelium-dependent diastolic function (FMD) and the degree of coronary artery disease (CAD), plaque vulnerability, and its predictive value for cardiovascular events.

METHODS: Initially, patients (n=100) who were admitted from January 2020 to January 2021 and intended to undergo percutaneous coronary intervention (PCI) were selected. Further, FMD in all patients was determined before the procedure and divided into a high-FMD group (≥4.2%) and a low-FMD group (

RESULTS: No significant differences were observed concerning general information, number of coronary arteries-associated branches, lesion type, involvement of the left main stem (LM), the …


A Holistic Approach To Performance Prediction In Collegiate Athletics: Player, Team, And Conference Perspectives, Christopher Taber, S. Sharma, Mehul S. Raval, Samah Senbel, Allison Keefe, Jui Shah, Emma Patterson, Julie K. Nolan, N.S. Artan, Tolga Kaya Jan 2024

A Holistic Approach To Performance Prediction In Collegiate Athletics: Player, Team, And Conference Perspectives, Christopher Taber, S. Sharma, Mehul S. Raval, Samah Senbel, Allison Keefe, Jui Shah, Emma Patterson, Julie K. Nolan, N.S. Artan, Tolga Kaya

Exercise Science Faculty Publications

Predictive sports data analytics can be revolutionary for sports performance. Existing literature discusses players' or teams' performance, independently or in tandem. Using Machine Learning (ML), this paper aims to holistically evaluate player-, team-, and conference (season)-level performances in Division-1 Women's basketball. The players were monitored and tested through a full competitive year. The performance was quantified at the player level using the reactive strength index modified (RSImod), at the team level by the game score (GS) metric, and finally at the conference level through Player Efficiency Rating (PER). The data includes parameters from training, subjective stress, sleep, and recovery (WHOOP …


How Can A Cloud Computing It Framework Be Created And Applied Effectively In The Online Printing Industry?, Stefan Meissner Jan 2024

How Can A Cloud Computing It Framework Be Created And Applied Effectively In The Online Printing Industry?, Stefan Meissner

Dissertations

This research aims to design a cloud computing IT framework for the online printing industry based on a detailed literature review, the development of proof of concepts (PoC), and the conduction of a focus group. The framework can be adopted by the online printing industry or by vendors of print-specific applications to optimize their products for the online printing industry. The author has been working in the online printing process optimization and automation since 2007. During this time, he got deep insight into many industry-specific applications, their architectural design, and their challenges being used in the context of online printing. …


Building A Human Digital Twin (Hdtwin) Using Large Language Models For Cognitive Diagnosis: Algorithm Development And Validation, Gina Sprint, Maureen Schmitter-Edgecombe, Diane Cook Jan 2024

Building A Human Digital Twin (Hdtwin) Using Large Language Models For Cognitive Diagnosis: Algorithm Development And Validation, Gina Sprint, Maureen Schmitter-Edgecombe, Diane Cook

Computer Science Faculty Scholarship

Background: Human digital twins have the potential to change the practice of personalizing cognitive health diagnosis because these systems can integrate multiple sources of health information and influence into a unified model. Cognitive health is multifaceted, yet researchers and clinical professionals struggle to align diverse sources of information into a single model. Objective: This study aims to introduce a method called HDTwin, for unifying heterogeneous data using large language models. HDTwin is designed to predict cognitive diagnoses and offer explanations for its inferences. Methods: HDTwin integrates cognitive health data from multiple sources, including demographic, behavioral, ecological momentary assessment, n-back test, …


Social Networks And Large Language Models For Division I Basketball Game Winner Prediction, Gina Sprint Jan 2024

Social Networks And Large Language Models For Division I Basketball Game Winner Prediction, Gina Sprint

Computer Science Faculty Scholarship

Sporting event outcome prediction is a well-established and actively researched domain, with a particular focus on college basketball’s March Madness tournament. Researchers, fans, and gamblers alike seek accurate game-level predictions using features such as tournament seeds, season performance, and expert opinions. While machine learning algorithms have been harnessed to build prediction models, no perfect model or human-created bracket has emerged. This paper explores a novel approach to basketball game outcome prediction by utilizing the power of social networks and large language models (LLMs). LLMs are trained to understand and generate text, often eliminating the need for a feature engineering step. …


Optimal Network Analysis Through Vertex Order Coloring Of Intuitionistic Fuzzy Graph Operations, A. Meenakshi, S. Dhanushiya, Hong Qin, Maniyandy Elangovan Jan 2024

Optimal Network Analysis Through Vertex Order Coloring Of Intuitionistic Fuzzy Graph Operations, A. Meenakshi, S. Dhanushiya, Hong Qin, Maniyandy Elangovan

Data Science Faculty Publications

Intuitionistic fuzzy graphs IFGs are a powerful tool for modeling uncertainty and complex relationships. They offer versatile frameworks for addressing real-world challenges. In this research, we have introduced intuitionistic fuzzy vertex order coloring IFVOC and analyzed the alpha-strong (alpha str), beta-strong (beta str), and gamma-strong (gamma str) vertices through their degree. We explored important theorems based on the types of strong vertices, broadening the scope of our study. We analyzed multiple IFG products to determine the most optimal network based on some important metrics, including the weight and total number of alpha str vertices, the chromatic number, and the weight …


Examining Information Systems Use To Facilitate The Workplace Accommodation Process, Shiya Cao Jan 2024

Examining Information Systems Use To Facilitate The Workplace Accommodation Process, Shiya Cao

Statistical and Data Sciences: Faculty Publications

BACKGROUND: The workplace accommodation process is often affected by ineffective and inefficient communications and information exchanges among disabled employees and other stakeholders. Information systems (IS) can play a key role in facilitating a more effective and efficient accommodation process since IS has been shown to facilitate business processes and effect positive organizational changes.

OBJECTIVE: Since there is little to no research that exists on IS use to facilitate the workplace accommodation process, this paper, as a critical first step, examines how IS have been used in the accommodation process.

METHODS: Thirty-six interviews were conducted with disabled employees from various organizations. …