Open Access. Powered by Scholars. Published by Universities.®

Databases and Information Systems Commons™

Open Access. Powered by Scholars. Published by Universities.®

2024

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 301 - 312 of 312

Full-Text Articles in Databases and Information Systems

Optimizing Sports Outcome Prediction Through Feature Engineering And Machine Learning, Vitor S. Freitas Jan 2024

Optimizing Sports Outcome Prediction Through Feature Engineering And Machine Learning, Vitor S. Freitas

Graduate Theses/Dissertations

The challenge of predicting the outcome of a team game lies in the high complexity and dynamics of the sports data. This thesis focuses on the aspect of using feature engineering and the genetic algorithm to predict the winner and the score of various sports events. Generally, it deals with how machine learning algorithms are combined with state-of-the-art feature engineering techniques in sports datasets derived from various sports disciplines. In this thesis, five different machine learning models have been applied, classification and regression trees (CART), random forest (RF), stochastic gradient boosting (SGB), eXtreme gradient boosting (XGBoost), and extreme learning machine …


Predicting Viral Rumors And Vulnerable Users With Graph-Based Neural Multi-Task Learning For Infodemic Surveillance, Xuan Zhang, Wei Gao Jan 2024

Predicting Viral Rumors And Vulnerable Users With Graph-Based Neural Multi-Task Learning For Infodemic Surveillance, Xuan Zhang, Wei Gao

Research Collection School Of Computing and Information Systems

In the age of the infodemic, it is crucial to have tools for effectively monitoring the spread of rampant rumors that can quickly go viral, as well as identifying vulnerable users who may be more susceptible to spreading such misinformation. This proactive approach allows for timely preventive measures to be taken, mitigating the negative impact of false information on society. We propose a novel approach to predict viral rumors and vulnerable users using a unified graph neural network model. We pre-train network-based user embeddings and leverage a cross-attention mechanism between users and posts, together with a community-enhanced vulnerability propagation (CVP) …


Privobfnet: A Weakly Supervised Semantic Segmentation Model For Data Protection, Chiat Pin Tay, Vigneshwaran Subbaraju, Thivya Kandappu Jan 2024

Privobfnet: A Weakly Supervised Semantic Segmentation Model For Data Protection, Chiat Pin Tay, Vigneshwaran Subbaraju, Thivya Kandappu

Research Collection School Of Computing and Information Systems

The use of social media has made it easy to communicate and share information over the internet. However, it also brings issues such as data privacy leakage, which can be exploited by recipients with malicious intentions to harm the sender. In this paper, we propose a deep neural network that analyzes user’s image for privacy sensitive content and automatically locates sensitive regions for obfuscation. Our approach relies solely on image level annotations and learns to (a) predict an overall privacy score, (b) detect sensitive attributes and (c) demarcate the sensitive regions for obfuscation, in a given input image. We validated …


Railroad Condition Monitoring Using Distributed Acoustic Sensing And Deep Learning Techniques, Md Arifur Rahman Jan 2024

Railroad Condition Monitoring Using Distributed Acoustic Sensing And Deep Learning Techniques, Md Arifur Rahman

College of Graduate Studies: Theses & Dissertations

Proper condition monitoring has been a major issue among railroad administrations since it might cause catastrophic dilemmas that lead to fatalities or damage to the infrastructure. Although various aspects of train safety have been conducted by scholars, in-motion monitoring detection of defect occurrence, cause, and severity is still a big concern. Hence extensive studies are still required to enhance the accuracy of inspection methods for railroad condition monitoring (CM). Distributed acoustic sensing (DAS) has been recognized as a promising method because of its sensing capabilities over long distances and for massive structures. As DAS produces large datasets, algorithms for precise …


Short: Can Citations Tell Us About A Paper's Reproducibility? A Case Study Of Machine Learning Papers, Rochana R. Obadage, Sarah M. Rajtmajer, Jian Wu Jan 2024

Short: Can Citations Tell Us About A Paper's Reproducibility? A Case Study Of Machine Learning Papers, Rochana R. Obadage, Sarah M. Rajtmajer, Jian Wu

Computer Science Faculty Publications

The iterative character of work in machine learning (ML) and artificial intelligence (AI) and reliance on comparisons against benchmark datasets emphasize the importance of reproducibility in that literature. Yet, resource constraints and inadequate documentation can make running replications particularly challenging. Our work explores the potential of using downstream citation contexts as a signal of reproducibility. We introduce a sentiment analysis framework applied to citation contexts from papers involved in Machine Learning Reproducibility Challenges in order to interpret the positive or negative outcomes of reproduction attempts. Our contributions include training classifiers for reproducibility-related contexts and sentiment analysis, and exploring correlations between …


Retrogressive Document Manipulation Of Us Federal Environmental Websites, Lesley Frew, Michael L. Nelson, Michele C. Weigle Jan 2024

Retrogressive Document Manipulation Of Us Federal Environmental Websites, Lesley Frew, Michael L. Nelson, Michele C. Weigle

Computer Science Faculty Publications

Changes made to webpages can affect their retrievability. Often this is done with the intention of increasing the page's search engine ranking to improve overall access to information on the page. The Environmental Data and Governance Initiative (EDGI) created a dataset that describes changes on US federal environmental webpages between 2016 and 2020. EDGI noted that many environmental terms were deleted from the pages, but without user data, claims that page retrievability and public information access were lowered are only anecdotal. The Open Resource for Click Analysis in Search (ORCAS) dataset was created during the same time frame, from 2017 …


Causal Disentangled Recommendation Against User Preference Shifts, Wenjie Wang, Xinyu Lin, Liuhui Wang, Fuli Feng, Yunshan Ma, Tat‑Seng Chua Jan 2024

Causal Disentangled Recommendation Against User Preference Shifts, Wenjie Wang, Xinyu Lin, Liuhui Wang, Fuli Feng, Yunshan Ma, Tat‑Seng Chua

Research Collection School Of Computing and Information Systems

Recommender systems easily face the issue of user preference shifts. User representations will become outof-date and lead to inappropriate recommendations if user preference has shifted over time. To solve theissue, existing work focuses on learning robust representations or predicting the shifting pattern. Therelacks a comprehensive view to discover the underlying reasons for user preference shifts. To understand thepreference shift, we abstract a causal graph to describe the generation procedure of user interaction sequences.Assuming user preference is stable within a short period, we abstract the interaction sequence as a set ofchronological environments. From the causal graph, we find that the changes …


Quantifying The Competitiveness Of A Dataset In Relation To General Preferences, Kyriakos Mouratidis, Keming Li, Bo Tang Jan 2024

Quantifying The Competitiveness Of A Dataset In Relation To General Preferences, Kyriakos Mouratidis, Keming Li, Bo Tang

Research Collection School Of Computing and Information Systems

Typically, a specific market (e.g., of hotels, restaurants, laptops, etc.) is represented as a multi-attribute dataset of the available products. The topic of identifying and shortlisting the products of most interest to a user has been well-explored. In contrast, in this work we focus on the dataset, and aim to assess its competitiveness with regard to different possible preferences. We define measures of competitiveness, and represent them in the form of a heat-map in the domain of preferences. Our work finds application in market analysis and in business development. These applications are further enhanced when the competitiveness heat-map is used …


A Unified Framework For Contextual And Factoid Question Generation, Chenhe Dong, Ying Shen, Shiyang Lin, Zhenzhou Lin, Yang Deng Jan 2024

A Unified Framework For Contextual And Factoid Question Generation, Chenhe Dong, Ying Shen, Shiyang Lin, Zhenzhou Lin, Yang Deng

Research Collection School Of Computing and Information Systems

Question generation (QG) aims to automatically generate fluent and relevant questions, where the two most mainstream directions are generating questions from unstructured contextual texts (CQG), such as news articles, and generating questions from structured factoid texts (FQG), such as knowledge graphs or tables. Existing methods for these two tasks mainly face challenges of limited internal structural information as well as scarce background information, while these two tasks can benefit each other for alleviating these issues. For example, when meeting the entity mention “United Kingdom” in CQG, it can be inferred that it is a country in European continent based on …


Societal Impacts Of Artificial Intelligence: Ethics, Legal, And Governance Issues, Yuzhou Qian, Keng Siau, Fiona Fui-Hoon Nah Jan 2024

Societal Impacts Of Artificial Intelligence: Ethics, Legal, And Governance Issues, Yuzhou Qian, Keng Siau, Fiona Fui-Hoon Nah

Research Collection School Of Computing and Information Systems

Artificial intelligence (AI) is quickly changing the way we work and the way we live. The emergence of ChatGPT has thrust AI, especially Generative AI, into the spotlight. The societal impact of AI is on most people's minds. This article presents several research projects on how AI impacts work and society. Three research works are discussed in this article. The first study develops a theoretical framework structuring the legal and ethical objectives that are needed and the means to achieve them. The second study concentrates on bias and discrimination issues embedded in AI applications. It focuses on enhancing the collaboration …


Dynamic Memory Management For Key-Value Store, Yuchen Wang Jan 2024

Dynamic Memory Management For Key-Value Store, Yuchen Wang

Dissertations, Master's Theses and Master's Reports

To minimize the latency of accessing back-end servers, modern web services often use in-memory key-value (k-v) stores at the front end to cache frequently accessed objects. Due to the limited memory capacity, these stores must be configured with a fixed amount of memory. Consequently, cache replacement is required when the footprint of the accessed objects exceeds the cache size.

This thesis presents a comprehensive exploration of advanced dynamic memory management techniques for k-v stores. The first study conducts a detailed analysis of K-LRU, a random sampling-based replacement policy, proposing a dynamic K configuration scheme to exploit the potential miss ratio …


Model Guided Memory Optimization For Key-Value Caches, Daniel Byrne Jan 2024

Model Guided Memory Optimization For Key-Value Caches, Daniel Byrne

Dissertations, Master's Theses and Master's Reports

Modern web services deploy key-value caches to store popular requests to backend systems. As such, how the cache stores data impacts both the cache miss ratio and throughput. Therefore, in this thesis, we introduce and apply cache modeling techniques to optimize the memory organization of a key-value cache to improve overall cache performance.

Specifically, we begin with a single-level key-value cache and use miss ratio curves to adjust the memory assigned to the residing applications dynamically. This leads to an improvement in miss ratio up to 25% over state-of-the-art techniques and an 8.8% improvement in cache throughput. We then consider …