Open Access. Powered by Scholars. Published by Universities.®

Data Storage Systems Commons

Open Access. Powered by Scholars. Published by Universities.®

Machine Learning

Discipline
Institution
Publication Year
Publication
Publication Type

Articles 1 - 17 of 17

Full-Text Articles in Data Storage Systems

Clear Skies, Avery C. Munn Jan 2026

Clear Skies, Avery C. Munn

Williams Honors College, Honors Research Projects

Air quality impacts public health, environmental sustainability, and quality of life. However, accurate and easily accessible short-term air quality forecasting is challenging to find. This project, Clear Skies, presents a machine learning–based system for forecasting next-day Air Quality Index (AQI) levels across regions in Ohio. By using historical pollutant data with variables such as temperature, humidity, wind speed, and atmospheric pressure, the system finds relationships that traditional statistical models often don’t show.

Machine learning models are evaluated alongside AI techniques to find the environmental factors that influence AQI predictions. This helps reduce the “black box” nature of many AI systems …


Csc36000 - Modern Distributed Computing Assignment, Saptarashmi Bandyopadhyay Oct 2025

Csc36000 - Modern Distributed Computing Assignment, Saptarashmi Bandyopadhyay

Open Educational Resources

This assignment covers standard performance metrics for Distributed Systems and the basics of Multiprocessing for CSC36000 - Modern Distributed Computing at the City College of New York CUNY. It is an interactive coding assignment intended to be executed in a Python notebook.


Computer Vision In Soccer: Yolov11 Analytics Engine For Quantifying Game Strategy, Connor S. Maurer May 2025

Computer Vision In Soccer: Yolov11 Analytics Engine For Quantifying Game Strategy, Connor S. Maurer

Data Science Undergraduate Honors Theses

Single-shot object detection capabilities significantly reduce computational overhead for real-time computer vision in sports analytics at 60 FPS. YOLO11’s lightweight CNN gives promising accuracy while meeting the low-latency demand of dynamic soccer matches. As data-driven approaches take over the sport of soccer, efficient player tracking systems become critical for informing coach’s strategies. I prototype the ETL (Extract, Transform, Load) process of data collected from a single- shot detection program and evaluate its viability for estimating player fatigue. YOLO11 detects players, the ball, and other characteristics, with the output transformed by homography to estimate the positions in the real world. These …


Virtual Makeup And Technology Integration, Vishwa Bhatt May 2025

Virtual Makeup And Technology Integration, Vishwa Bhatt

Electronic Theses, Projects, and Dissertations

The Virtual Makeup Streamlit application presents an advanced approach to digital cosmetic try-on by allowing users to apply makeup to their facial images in real time. This project uses computer vision and web technologies to create an interactive and user-friendly platform that capitalizes on the increasing popularity of virtual try-on solutions in the cosmetics industry.

At its core, the system uses effective facial detection and semantic segmentation techniques to recognize and separate facial areas such as lips and hair. Techniques such as U-Net and Resnet, and Midepipe are used to create accurate segmentation masks, which are essential for accurate makeup …


Predicting Crises On The African Frontier Stock Markets With Investor Sentiment Indicators: A Machine Learning Approach, David Korsah, Lord Mensah Jan 2025

Predicting Crises On The African Frontier Stock Markets With Investor Sentiment Indicators: A Machine Learning Approach, David Korsah, Lord Mensah

Journal of International Technology and Information Management

This study examined the predictive ability of machine learning algorithms in identifying crises within African stock markets. The study employed seven distinct machine-learning models, analyzing historical stock prices from eight stock markets, three major sentiment indicators, and the exchange rates of local currencies against the US dollar, with each data spanning from May 1, 2007, to April 1, 2023. Extreme Gradient Boosting (XGBoost) emerged as the most effective algorithm for predicting crises. Historical stock prices and exchange rates were identified as the most critical features for prediction. On the sentiment side, investors’ perceptions of potential volatility on the S&P 500, …


Predicting Global Healthcare Supply Chain Delays: A Machine Learning Approach Leveraging Country-Level Logistics Metrics, Jeevan Sai Gali, Nima Molavi, Sepideh Alavi Jan 2025

Predicting Global Healthcare Supply Chain Delays: A Machine Learning Approach Leveraging Country-Level Logistics Metrics, Jeevan Sai Gali, Nima Molavi, Sepideh Alavi

Journal of International Technology and Information Management

In global healthcare logistics, ensuring the timely delivery of medical commodities is critical, particularly in low- and middle-income countries characterized by infrastructural limitations and operational uncertainties. This research introduces an advanced, data-driven predictive framework designed to forecast delivery delays by synthesizing granular, internal shipment-level data from the USAID Global Health Supply Chain Program (GHSC-PSM) with external country-level logistics capabilities indicators derived from the World Bank’s Logistics Performance Index (LPI). Rather than relying on retrospective trend analyses, this study employs machine learning algorithms such as Random Forest, XGBoost, Support Vector Machines (SVM), and Multi-Layer Perceptron (MLP) to detect …


Application Of Lossy And Lossless Compression To Dicom Files, Yizhe Yang Dec 2024

Application Of Lossy And Lossless Compression To Dicom Files, Yizhe Yang

All Theses

The Digital Imaging and Communications in Medicine (DICOM) standard is widely utilized for the management, storage, and transfer of medical images. However, the substantial file sizes associated with DICOM data present challenges in terms of storage and data transmission. Data reduction techniques help address these challenges by minimizing the size of the data while preserving its integrity. This thesis examines various compression methods aimed at reducing the size of DICOM files. We evaluate five lossless compressors and four lossy compressors on DICOM data to compare and assess their performance. Through an analysis of each compressor’s compression efficiency and resulting image …


Analyzing Information Cascades Through Machine Learning And Data Analytics, Betul Agirman May 2024

Analyzing Information Cascades Through Machine Learning And Data Analytics, Betul Agirman

Honors Scholar Theses

In today's digital age, social media platforms have become pivotal in influencing public opinion and behavior, with information spreading being both beneficial and detrimental. This rapid spread is typically called an information cascade, and they are important in further understanding social influence, managing misinformation, and even predicting potential trends of public responses. With social media, people are connected so easily to one another like a network, wherein it becomes possible for them to influence each other’s behavior and decisions. Utilizing a dataset from Weibo that spans critical periods of the COVID-19 outbreak, this study integrates machine learning and data analytics …


From Leanstore To Learnedstore: Using A Learned Index To Improve Database Index Search, Sujit Maharjan Dec 2023

From Leanstore To Learnedstore: Using A Learned Index To Improve Database Index Search, Sujit Maharjan

Computer Science and Engineering Faculty Publications - Archive

In the realm of database systems, optimizing B+-tree index performance is of paramount importance to overall database performance. LeanStore, a high-performance OLTP storage engine, has extensively optimized its in-memory B+-tree component as well as its B+-tree -indexed database on the disk. However, B+-tree's lookup time increases linearly with the tree height. This is especially problematic when all or part of its lookup path is on the disk. Recently proposed learned index technique has the potential to significantly improve the performance of the B+-tree -based index by predicting location of the search key, instead of the level-by-Ievel path walk. However, this …


Comparison Of Machine Learning Algorithms For Species Family Classification Using Dna Barcode, Lala Septem Riza, M Ammar Fadhlur Rahman, Yudi Prasetyo, Muhammad Iqbal Zain, Herbert Siregar, Topik Hidayat, Khyrina Airin Fariza Abu Samah, Miftahurrahma Rosyda Dec 2023

Comparison Of Machine Learning Algorithms For Species Family Classification Using Dna Barcode, Lala Septem Riza, M Ammar Fadhlur Rahman, Yudi Prasetyo, Muhammad Iqbal Zain, Herbert Siregar, Topik Hidayat, Khyrina Airin Fariza Abu Samah, Miftahurrahma Rosyda

Knowledge Engineering and Data Science

Classifying plant species within the Liliaceae and Amaryllidaceae families presents inherent challenges due to the complex genetic diversity and overlapping morphological traits among species. This study explores the difficulties in accurate classification by comparing 11 supervised learning algorithms applied to DNA barcode data, aiming to enhance the precision of species family classification in these taxonomically intricate plant families. The ribulose-1,5-bisphosphate carboxylase-oxygenase large sub-unit (rbcL) gene, selected as a DNA barcode locus for plants, is used to represent species within the Amaryllidaceae and Liliaceae families. The experimental results demonstrate that nearly all tested models achieve accurate species classification into the appropriate …


Chatgpt As Metamorphosis Designer For The Future Of Artificial Intelligence (Ai): A Conceptual Investigation, Amarjit Kumar Singh (Library Assistant), Dr. Pankaj Mathur (Deputy Librarian) Mar 2023

Chatgpt As Metamorphosis Designer For The Future Of Artificial Intelligence (Ai): A Conceptual Investigation, Amarjit Kumar Singh (Library Assistant), Dr. Pankaj Mathur (Deputy Librarian)

Library Philosophy and Practice (e-journal)

Abstract

Purpose: The purpose of this research paper is to explore ChatGPT’s potential as an innovative designer tool for the future development of artificial intelligence. Specifically, this conceptual investigation aims to analyze ChatGPT’s capabilities as a tool for designing and developing near about human intelligent systems for futuristic used and developed in the field of Artificial Intelligence (AI). Also with the helps of this paper, researchers are analyzed the strengths and weaknesses of ChatGPT as a tool, and identify possible areas for improvement in its development and implementation. This investigation focused on the various features and functions of ChatGPT that …


A Comparison Of Machine Learning Models To Prioritise Emailsusing Emotion Analysis For Customer Service Excellence, Mohammad Yasser Chuttur, Yashinee Parianen Jul 2022

A Comparison Of Machine Learning Models To Prioritise Emailsusing Emotion Analysis For Customer Service Excellence, Mohammad Yasser Chuttur, Yashinee Parianen

Knowledge Engineering and Data Science

There has been little research on machine learning for email prioritization for customer service excellence. To fill this gap, we propose and assess the efficacy of various machine learning techniques for classifying emails into three degrees of priority: high, low, and neutral, based on the emotions inherent in the email content. It is predicted that after emails are classified into those three categories, recipients will be able to respond to emails more efficiently and provide better customer service. We use the NRC Emotion Lexicon to construct a labeled email dataset of 517,401 messages for our proposal. Following that, we train …


Total Sky Imager Project, Ryan D. Maier, Benjamin Jack Forest, Kyle X. Mcgrath Jun 2022

Total Sky Imager Project, Ryan D. Maier, Benjamin Jack Forest, Kyle X. Mcgrath

Mechanical Engineering

Solar farms like the Gold Tree Solar Farm at Cal Poly San Luis Obispo have difficulty delivering a consistent level of power output. Cloudy days can trigger a significant drop in the utility of a farm’s solar panels, and an unexpected loss of power from the farm could potentially unbalance the electrical grid. Being able to predict these power output drops in advance could provide valuable time to prepare a grid and keep it stable. Furthermore, with modern data analysis methods such as machine learning, these predictions are becoming more and more accurate – given a sufficient data set. The …


How The Power Of Machine – Machine Learning, Data Science And Nlp Can Be Used To Prevent Spoofing And Reduce Financial Risks, Sasibhushan Rao Chanthati Mar 2021

How The Power Of Machine – Machine Learning, Data Science And Nlp Can Be Used To Prevent Spoofing And Reduce Financial Risks, Sasibhushan Rao Chanthati

Harrisburg University Other Works

This paper discusses the potential of machine learning, data science, and natural language processing (NLP) in mitigating the incidence of spoofing and financial risks hinged on cyber threats. Another one is spoofing; it is the act of impersonating legitimate entities to gain unauthorized information and it is indeed a threat to the public and companies to some extent. The research introduces two primary methodologies to combat spoofing: an email filtering system using a machine learning algorithm and an encryption and decryption system using a Caesar Cipher and Python programming language. It distinguishes between approved domains and unapproved domains by using …


Improving Vix Futures Forecasts Using Machine Learning Methods, James Hosker, Slobodan Djurdjevic, Hieu Nguyen, Robert Slater Jan 2019

Improving Vix Futures Forecasts Using Machine Learning Methods, James Hosker, Slobodan Djurdjevic, Hieu Nguyen, Robert Slater

SMU Data Science Review

The problem of forecasting market volatility is a difficult task for most fund managers. Volatility forecasts are used for risk management, alpha (risk) trading, and the reduction of trading friction. Improving the forecasts of future market volatility assists fund managers in adding or reducing risk in their portfolios as well as in increasing hedges to protect their portfolios in anticipation of a market sell-off event. Our analysis compares three existing financial models that forecast future market volatility using the Chicago Board Options Exchange Volatility Index (VIX) to six machine/deep learning supervised regression methods. This analysis determines which models provide best …


Crude Oil Prices Forecasting: Time Series Vs. Svr Models, Xin James He Dec 2018

Crude Oil Prices Forecasting: Time Series Vs. Svr Models, Xin James He

Journal of International Technology and Information Management

This research explores the weekly crude oil price data from U.S. Energy Information Administration over the time period 2009 - 2017 to test the forecasting accuracy by comparing time series models such as simple exponential smoothing (SES), moving average (MA), and autoregressive integrated moving average (ARIMA) against machine learning support vector regression (SVR) models. The main purpose of this research is to determine which model provides the best forecasting results for crude oil prices in light of the importance of crude oil price forecasting and its implications to the economy. While SVR is often considered the best forecasting model in …


Semantic Visualization For Short Texts With Word Embeddings, Van Minh Tuan Le, Hady W. Lauw Aug 2017

Semantic Visualization For Short Texts With Word Embeddings, Van Minh Tuan Le, Hady W. Lauw

Research Collection School Of Computing and Information Systems

Semantic visualization integrates topic modeling and visualization, such that every document is associated with a topic distribution as well as visualization coordinates on a low-dimensional Euclidean space. We address the problem of semantic visualization for short texts. Such documents are increasingly common, including tweets, search snippets, news headlines, or status updates. Due to their short lengths, it is difficult to model semantics as the word co-occurrences in such a corpus are very sparse. Our approach is to incorporate auxiliary information, such as word embeddings from a larger corpus, to supplement the lack of co-occurrences. This requires the development of a …