Open Access. Powered by Scholars. Published by Universities.®

Databases and Information Systems Commons™

Open Access. Powered by Scholars. Published by Universities.®

2025

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 181 - 199 of 199

Full-Text Articles in Databases and Information Systems

Learning To Rank Aspects And Opinions For Comparative Explanations, Trung Hoang Le, Hady Wirawan Lauw Jan 2025

Learning To Rank Aspects And Opinions For Comparative Explanations, Trung Hoang Le, Hady Wirawan Lauw

Research Collection School Of Computing and Information Systems

Comparative recommendation explanations help to make sense of recommendations by comparing a recommended item along some aspects of interest with one or many items being considered. This work extends the notion of comparative explanations, by going beyond merely better/worse statements, to further incorporate aspect-level opinions for more informative comparisons. To enhance the quality of both the personalized recommendation and the explanation, we incorporate optimization objectives that preserve relative rankings of aspects and opinions, in addition to the classical rankings of overall preferences for items. We integrate the multiple ranking objectives and multi-tensor factorization together. Experiments on datasets of different domains …


Performance Evaluation Of Newsql Databases In A Distributed Architecture, Zhiyao Zhang, Alan @ Ali Madjelisi Megargel, Lingxiao Jiang Jan 2025

Performance Evaluation Of Newsql Databases In A Distributed Architecture, Zhiyao Zhang, Alan @ Ali Madjelisi Megargel, Lingxiao Jiang

Research Collection School Of Computing and Information Systems

In the last decade, application architectures have evolved drastically, moving from monolithic architectures to distributed architectures where deployment has shifted from dedicated on-premises servers to the cloud. Distributed architectures and cloud computing has enabled businesses to scale their application components across different geographical locations. While it is easy to scale the application layer, scaling its database layer that relies on traditional SQL databases is challenging and often is a common source of bottlenecks when it comes to application performance. This paper evaluates the performance characteristics between two NewSQL databases solutions, MySQL NDB Cluster vs. TIBCO ActiveSpaces IMDG. Serving as an …


Risk Spillover Effect Of China-Asean Supply Chains: Insights Of Industrial Transfer, Zeyang Bian, Yuning Zhang, Keng Siau, Yaqian Zhang, Jianjia He Jan 2025

Risk Spillover Effect Of China-Asean Supply Chains: Insights Of Industrial Transfer, Zeyang Bian, Yuning Zhang, Keng Siau, Yaqian Zhang, Jianjia He

Research Collection School Of Computing and Information Systems

As labour costs in China increase, labour-intensive industries are migrating to ASEAN countries, attracted by lower labour costs and market potential. This shift not only affects the economies of China and ASEAN but also reshapes the global manufacturing landscape. This paper investigates the correlation and spillover of supply chain risks using production exposure indicators derived from inter-country input-output data and the R-Vine Copula model. We assess the risk spillover of each country within the global supply chain. Our findings indicate that industrial relocation can significantly alter supply chain structures, thereby affecting the concentration and direction of risks. While China's role …


Label Correlated Contrastive Learning For Medical Report Generation, Xinyao Liu, Junchang Xin, Bing Tian Dai, Qi Shen, Zhihong Huang, Zhiqiong Wang Jan 2025

Label Correlated Contrastive Learning For Medical Report Generation, Xinyao Liu, Junchang Xin, Bing Tian Dai, Qi Shen, Zhihong Huang, Zhiqiong Wang

Research Collection School Of Computing and Information Systems

Background and Objective: Automatic generation of medical reports reduces both the burden on radiologists and the possibility of errors due to the inexperience of radiologists. The model that utilizes attention mechanism and contrastive learning can generate medical reports by capturing both general and specific semantics. However, existing contrastive learning methods ignore the specificity of medical data, that is, a patient may suffer from multiple diseases at the same time. This means that the lack of fine-grained relationships for contrastive learning will lead to the problem of insufficient specificity. Methods: To address the above problem, a label correlated contrastive learning method …


Flowing Together Or Alone: Impact Of Collaboration In The Metaverse, Fiona Fui-Hoon Nah, Brenda Eschenbrenner, Langtao Chen Jan 2025

Flowing Together Or Alone: Impact Of Collaboration In The Metaverse, Fiona Fui-Hoon Nah, Brenda Eschenbrenner, Langtao Chen

Research Collection School Of Computing and Information Systems

The metaverse is the next-generation Internet (Web3) that facilitates social connections and collaborations in a virtual world environment. Given the potential of the metaverse to provide more satisfying and effective means of remote collaborations, exploring the possibility of leveraging the metaverse for these endeavors is warranted. Therefore, an important question to address is whether greater engagement occurs when tasks are completed collaboratively versus individually in the metaverse. We address this question by drawing on flow and transportation theories to hypothesize the effect of carrying out a creative task in the metaverse collaboratively versus alone on one's cognitive absorption, a contextually …


Your Cursor Reveals: On Analyzing Workers’ Browsing Behavior And Annotation Quality In Crowdsourcing Tasks, Pei-Chi Lo, Ee-Peng Lim Jan 2025

Your Cursor Reveals: On Analyzing Workers’ Browsing Behavior And Annotation Quality In Crowdsourcing Tasks, Pei-Chi Lo, Ee-Peng Lim

Research Collection School Of Computing and Information Systems

In this work, we investigate the connection between browsing behavior and task quality of crowdsourcing workers performing annotation tasks that require information judgements. Such information judgements are often required to derive ground truth answers to information retrieval queries. We explore the use of workers’ browsing behavior to directly determine their annotation result quality. We hypothesize user attention to be the main factor contributing to a worker’s annotation quality. To predict annotation quality at the task level, we model two aspects of task-specific user attention, also known as general and semantic user attentions . Both aspects of user attention can be …


Llms-Based Augmentation For Domain Adaptation In Long-Tailed Food Datasets, Qing Wang, Chong-Wah Ngo, Ee-Peng Lim, Qianru Sun Jan 2025

Llms-Based Augmentation For Domain Adaptation In Long-Tailed Food Datasets, Qing Wang, Chong-Wah Ngo, Ee-Peng Lim, Qianru Sun

Research Collection School Of Computing and Information Systems

Training a model for food recognition is challenging because the training samples, which are typically crawled from the Internet, are visually different from the pictures captured by users in the free-living environment. In addition to this domain-shift problem, the real-world food datasets tend to be long-tailed distributed and some dishes of different categories exhibit subtle variations that are difficult to distinguish visually. In this paper, we present a framework empowered with large language models (LLMs) to address these challenges in food recognition. We first leverage LLMs to parse food images to generate food titles and ingredients. Then, we project the …


Empowering Crisis Information Extraction Through Actionability Event Schemata And Domain-Adaptive Pre-Training, Yuhao Zhang, Siaw Ling Lo, Phyo Yi Win Myint Jan 2025

Empowering Crisis Information Extraction Through Actionability Event Schemata And Domain-Adaptive Pre-Training, Yuhao Zhang, Siaw Ling Lo, Phyo Yi Win Myint

Research Collection School Of Computing and Information Systems

One of the persistent challenges in crisis detection is inferring actionable information to support emergency response. Existing methods focus on situational awareness but often lack actionable insights. This study proposes a holistic approach to implementing an actionability extraction system on social media, including requirement gathering, selection of machine learning tasks, data preparation, and integration with existing resources, providing guidance for governments, civil services, emergency workers, and researchers on supplementing existing channels with actionable information from social media. Our solution leverages an actionability schema and domain-adaptive pre-training, improving upon the state-of-the-art model by 5.5% and 10.1% in micro and macro F1 …


A Review On Knowledge And Information Extraction From Pdf Documents And Storage Approaches, Salvador D. Atagong, Henri Tonnang, Kennedy Senagi, Mark Wamalwa, Komi M. Agboka, John Odindi Jan 2025

A Review On Knowledge And Information Extraction From Pdf Documents And Storage Approaches, Salvador D. Atagong, Henri Tonnang, Kennedy Senagi, Mark Wamalwa, Komi M. Agboka, John Odindi

All Peer-Reviewed Publications

Introduction: Automating the extraction of information from Portable Document Format (PDF) documents represents a major advancement in information extraction, with applications in various domains such as healthcare, law, or biochemistry. However, existing solutions face challenges related to accuracy, domain adaptability, and implementation complexity. Methods: A systematic review of the literature was conducted using the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) methodology to examine approaches and trends in PDF information extraction and storage approaches. Results: The review revealed three dominant methodological categories: rule-based systems, statistical learning models, and neural network-based approaches. Key limitations include the rigidity of rule-based …


Coming Back Differently: An Exploratory Case Study Of Near Death Experiences Of Webpages, Lesley Frew, Michael L. Nelson, Michele Weigle Jan 2025

Coming Back Differently: An Exploratory Case Study Of Near Death Experiences Of Webpages, Lesley Frew, Michael L. Nelson, Michele Weigle

Computer Science Faculty Publications

In this case study, we use web archives to analyze 8,824 webpages that were taken offline and subsequently put back online, thus experiencing a “near death experience.” We enumerate the stages of a webpage’s near death experience, including the change from a successful HTTP status code to non-successful and back, the intermediate stage with markers such as an under construction banner, and an analysis of how the pages came back differently.


Defending Federated Recommender Systems Against Untargeted Attacks: A Contribution-Aware Robust Aggregation Scheme, Ruicheng Liang, Yuanchun Jiang, Feida Zhu, Ling Cheng, Huiwen Liu Jan 2025

Defending Federated Recommender Systems Against Untargeted Attacks: A Contribution-Aware Robust Aggregation Scheme, Ruicheng Liang, Yuanchun Jiang, Feida Zhu, Ling Cheng, Huiwen Liu

Research Collection School Of Computing and Information Systems

Federated recommender systems (FedRSs) effectively tackle the tradeoff between recommendation accuracy and privacy preservation. However, recent studies have revealed severe vulnerabilities in FedRSs, particularly against untargeted attacks seeking to undermine their overall performance. Defense methods employed in traditional recommender systems are not applicable to FedRSs, and existing robust aggregation schemes for other federated learning-based applications have proven ineffective in FedRSs. Building on the observation that malicious clients contribute negatively to the training process, we design a novel contribution-aware robust aggregation scheme to defend FedRSs against untargeted attacks, named contribution-aware Bayesian knowledge distillation aggregation (ConDA), comprising two key components for the …


A Survey Of Multilingual Large Language Models, Libo Qin, Qiguang Chen, Yuhang Zhou, Zhi Chen, Yinghui Li, Lizi Liao, Min Li, Wanxiang Che, Philip S. Yu Jan 2025

A Survey Of Multilingual Large Language Models, Libo Qin, Qiguang Chen, Yuhang Zhou, Zhi Chen, Yinghui Li, Lizi Liao, Min Li, Wanxiang Che, Philip S. Yu

Research Collection School Of Computing and Information Systems

Multilingual large language models (MLLMs) leverage advanced large language models to process and respond to queries across multiple languages, achieving significant success in polyglot tasks. Despite these breakthroughs, a comprehensive survey summarizing existing approaches and recent developments remains absent. To this end, this paper presents a unified and thorough review of the field, highlighting recent progress and emerging trends in MLLM research. The contributions of this paper are as follows. (1) Extensive survey: to our knowledge, this is the pioneering thorough review of multilingual alignment in MLLMs. (2) Unified taxonomy: we provide a unified framework to summarize the current progress …


Fedart: A Neural Model Integrating Federated Learning And Adaptive Resonance Theory, Shubham Pateria, Budhitama Subagdja, Ah-Hwee Tan Jan 2025

Fedart: A Neural Model Integrating Federated Learning And Adaptive Resonance Theory, Shubham Pateria, Budhitama Subagdja, Ah-Hwee Tan

Research Collection School Of Computing and Information Systems

Federated Learning (FL) has emerged as a promising paradigm for collaborative model training across distributed clients while preserving data privacy. However, prevailing FL approaches aggregate the clients’ local models into a global model through multi-round iterative parameter averaging. This leads to the undesirable bias of the aggregated model towards certain clients in the presence of heterogeneous data distributions among the clients. Moreover, such approaches are restricted to supervised classification tasks and do not support unsupervised clustering. To address these limitations, we propose a novel one-shot FL approach called Federated Adaptive Resonance Theory (FedART) which leverages self-organizing Adaptive Resonance Theory (ART) …


Marrying Top-K With Skyline Queries: Operators With Relaxed Preference Input And Controllable Output Size, Kyriakos Mouratidis, Keming Li, Bo Tang Jan 2025

Marrying Top-K With Skyline Queries: Operators With Relaxed Preference Input And Controllable Output Size, Kyriakos Mouratidis, Keming Li, Bo Tang

Research Collection School Of Computing and Information Systems

The two most common paradigms to identify records of preference in a multi-objective setting rely either on dominance (e.g., the skyline operator) or on a utility function defined over the records' attributes (typically, using a top-k query). Despite their proliferation, each of them has its own palpable drawbacks. Motivated by these drawbacks, we identify three hard requirements for practical decision support, namely, personalization, controllable output size, and flexibility in preference specification. With these requirements as a guide, we combine elements from both paradigms and propose two new operators, ORD and ORU. We perform a qualitative study to demonstrate how they …


Impact Of Achievement-Oriented Gamification In Erp Systems: Examining Subjective And Objective User Outcomes, E. Adeborna, Fiona Fui-Hoon Nah, L. Motiwalla Jan 2025

Impact Of Achievement-Oriented Gamification In Erp Systems: Examining Subjective And Objective User Outcomes, E. Adeborna, Fiona Fui-Hoon Nah, L. Motiwalla

Research Collection School Of Computing and Information Systems

This research explores the effect of gamification using achievement-oriented affordances in Enterprise Resource Planning (ERP) systems on subjective (behavioral intention) and objective (performance) outcomes. Drawing on the cognitive-affectiveconative (CAC) framework, a research model was developed to explain behavioral intention and tested in a pilot experiment with 63 participants. These participants completed a post-study questionnaire for assessing the impact of gamification on users’ behavioral intention that is mediated by CAC constructs: focused immersion, enjoyment, and selfrewarding experience. Preliminary results show that gamification enhances enjoyment and self-rewarding experience, which in turn positively influence and fully mediate behavioral intention. Objective performance outcomes were …


Interactive Video Search With Multi-Modal Llm Video Captioning, Yu-Tong Cheng, Jiaxin Wu, Zhixin Ma, Jiangshan He, Xiao-Yong Wei, Chong-Wah Ngo Jan 2025

Interactive Video Search With Multi-Modal Llm Video Captioning, Yu-Tong Cheng, Jiaxin Wu, Zhixin Ma, Jiangshan He, Xiao-Yong Wei, Chong-Wah Ngo

Research Collection School Of Computing and Information Systems

Cross-modal representation learning is essential for interactive text-to-video search tasks. However, the representation learning is limited by the size and quality of video-caption pairs. To improve the search accuracy, we propose to enlarge the size of available video-caption pairs by leveraging multi-model LLM on video captioning. Specifically, we use LLM to generate video captions for a large video collection (i.e., WebVid dataset) and use the generated video-caption pairs to pre-train a text-to-video search model. Additionally, we use LLM to generate fine-grained captions for test video collections to enable text-to-caption retrieval. Furthermore, we build a semantic overview of the retrieved rank …


Interpreting Topic Models In Byte-Pair Encoding Space, Jia Peng Lim, Hady Wirawan Lauw Jan 2025

Interpreting Topic Models In Byte-Pair Encoding Space, Jia Peng Lim, Hady Wirawan Lauw

Research Collection School Of Computing and Information Systems

Byte-pair encoding (BPE) is pivotal for processing text into chunksize tokens, particularly in Large Language Model (LLM). From a topic modeling perspective, as these chunksize tokens might be mere parts of valid words, evaluating and interpreting these tokens for coherence is challenging. Most, if not all, of coherence evaluation measures are incompatible as they benchmark using valid words. We propose to interpret the recovery of valid words from these tokens as a ranking problem and present a model-agnostic and training-free recovery approach from the topic-token distribution onto a selected vocabulary space, following which we could apply existing evaluation measures. Results …


Dims: Distributed Index For Similarity Search In Metric Spaces, Yifan Zhu, Chengyang Luo, Tang Qian, Lu Chen, Yunjun Gao, Baihua Zheng Jan 2025

Dims: Distributed Index For Similarity Search In Metric Spaces, Yifan Zhu, Chengyang Luo, Tang Qian, Lu Chen, Yunjun Gao, Baihua Zheng

Research Collection School Of Computing and Information Systems

Similarity search finds objects that are similar to a given query object based on a similarity metric. As the amount and variety of data continue to grow, similarity search in metric spaces has gained significant attention. Metric spaces can accommodate any type of data and support flexible distanc e metrics, making similarity search in metric spaces beneficial for many real-world applications, such as multimedia retrieval, personalized recommendation, trajectory analytics, data mining, decision planning, and distributed servers. However, existing studies mostly focus on indexing metric spaces on a single machine, which faces efficiency and scalability limitations with increasing data volume and …


Development And Application Of Computational Tools For Data-Driven Materials Science., Logan L. Lang Jan 2025

Development And Application Of Computational Tools For Data-Driven Materials Science., Logan L. Lang

Graduate Theses, Dissertations, and Problem Reports (ETD)

Modern materials science generates vast amounts of data from computational simulations and experiments, creating significant challenges for data processing and analysis. This thesis addresses these challenges through the development and application of computational tools within the framework of Material Data Science (MDS). Contributions span the four pillars of MDS: Material/Molecular Data, Algorithms, Databases, and High-Throughput Processes—with a primary focus on the Algorithm, Data, Database pillars.

For the Algorithm pillar, two Python libraries were developed to streamline common analysis tasks. PyProcar simplifies the post-processing and visualization of electronic structure data (band structures, density of states, Fermi surfaces) obtained from various Density …