Open Access. Powered by Scholars. Published by Universities.®
Databases and Information Systems Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Social and Behavioral Sciences (14)
- Communication (11)
- Social Media (11)
- Engineering (7)
- Artificial Intelligence and Robotics (6)
-
- Computer Engineering (6)
- Data Storage Systems (6)
- Numerical Analysis and Scientific Computing (6)
- Business (5)
- Software Engineering (5)
- Graphics and Human Computer Interfaces (3)
- Information Security (3)
- Systems Architecture (2)
- Theory and Algorithms (2)
- Accounting (1)
- Arts and Humanities (1)
- Asian Studies (1)
- Business Administration, Management, and Operations (1)
- E-Commerce (1)
- East Asian Languages and Societies (1)
- Environmental Policy (1)
- International and Area Studies (1)
- Mining Engineering (1)
- OS and Networks (1)
- Programming Languages and Compilers (1)
- Public Affairs, Public Policy and Public Administration (1)
- Real Estate (1)
- Keyword
-
- Social media (5)
- Natural language processing (4)
- Machine learning (3)
- Recommender Systems (3)
- Text mining (3)
-
- Classification (2)
- Correlation (2)
- Data analytics (2)
- Deep learning (2)
- Neural networks (2)
- Online learning (2)
- Opinion mining (2)
- Preference Learning (2)
- Social network (2)
- Topic model (2)
- API recommendation (1)
- Adaptive Large Neighborhood Search (1)
- Adversarial Learning (1)
- Adversarial Robustness (1)
- Airport (1)
- Algorithm (1)
- Analytics (1)
- Analyzing Subjectivity (1)
- And dialogue state tracking (1)
- Anomalous behavior (1)
- Anomaly detection (1)
- Antagonistic Community (1)
- Application protection (1)
- Aspect-Level Sentiment (1)
- Authentication misuse (1)
Articles 1 - 30 of 58
Full-Text Articles in Databases and Information Systems
Addressing Sparsity For Knowledge Graph Completion: Data And Model Perspectives, Ran Liu
Addressing Sparsity For Knowledge Graph Completion: Data And Model Perspectives, Ran Liu
Dissertations and Theses Collection (Open Access)
Knowledge graphs (KGs) are powerful tools for structuring factual knowledge into relational triples, yet their practical utility is often adversely affected by data sparsity. Many entities and relations are associated with only a few observations, which limits the quality of learned embeddings and weakens generalization in downstream tasks. The problem of sparsity led to two interrelated challenges. Firstly, it restricts the informativeness of training samples: positive examples are scarce, and conventional negative sampling often produces trivial or redundant negatives that resulting in limited guidance. Secondly, in few-shot relation learning scenarios, sparsity worsens distribution shifts between training and test relations, as …
A Data-Driven Framework For Optimal Retail Store Location, Ming Hui Tan
A Data-Driven Framework For Optimal Retail Store Location, Ming Hui Tan
Dissertations and Theses Collection (Open Access)
This study develops a data-driven framework for optimal retail store location planning that integrates road network analysis, mobility data and optimization techniques. By addressing the limitations of traditional approaches that rely on outdated census data and manual site selection, this research offers a scalable and adaptable solution for retail expansion in diverse urban environments. Chapters 1 and 2 establish the foundational context and theoretical underpinnings of this research. Chapter 1 introduces the research problem and motivation, highlighting the limitations of existing approaches and defining three key research objectives: automating candidate site identification, improving footfall estimation, and developing a scalable multi-site …
Memory-Efficient Graph Processing On Gpus: Reducing Intermediate Data Structure Overhead, Chang Ye
Memory-Efficient Graph Processing On Gpus: Reducing Intermediate Data Structure Overhead, Chang Ye
Dissertations and Theses Collection (Open Access)
The increasing scale of real-world graphs in domains such as fraud detection, community detection, and biological analysis demands high-throughput, memory-efficient graph processing solutions. GPUs offer massive parallelism for accelerating such workloads, and numerous frameworks have been developed to leverage their computational power. These frameworks primarily focus on optimizing scheduling to better align graph processing with GPU architectures. It performs well for algorithms with low memory demands, such as BFS, SSSP, and PageRank. However, for algorithms that require substantial memory, such as label propagation, and subgraph counting, the limited memory capacity of GPUs often becomes a significant bottleneck.
This dissertation addresses …
Rich Models And Methods For On-Demand Same Day Deliveries, Zhiqin Zhang
Rich Models And Methods For On-Demand Same Day Deliveries, Zhiqin Zhang
Dissertations and Theses Collection (Open Access)
Same-day delivery has brought numerous conveniences to people’s lives, but it has also presented challenges in terms of service management. To effectively optimize on-demand same-day delivery operations within urban logistics, intelligent decision-making strategies capable of adapting to rapidly changing circumstances are essential. Employing effective decisionmaking strategies that account for order allocation, route planning, courier scheduling, and other relevant factors, is pivotal in advancing logistics operations, enhancing efficiency, customer satisfaction, and resource utilization in the context of dynamic same-day delivery problems.
The focus of this thesis revolves around different emerging challenges presented by on-demand same-day delivery problems, with a particular emphasis …
Towards Reliable Ml: Data Attribution And Adversarial Robustness, Xiaosen Zheng
Towards Reliable Ml: Data Attribution And Adversarial Robustness, Xiaosen Zheng
Dissertations and Theses Collection (Open Access)
Modern machine learning (ML) models achieve remarkable success, but face critical reliability challenges. This thesis advances two pillars of reliable ML systems: interpretability through data attribution and robustness against adversarial threats.
In the first part, we develop novel data attribution methods to elucidate the data-model relationship. We establish the critical role of memorization in model generalization through token-level influence analysis, extend sample-level attribution to diffusion models with effective approximation techniques, and introduce REGMIX, a group-level approach that predicts data mixture performance using small-scale experiments. These contributions provide practitioners with scalable tools to audit training data impacts across modalities.
The second …
Correlation And Causation Analysis For Cross-Sectional And Panel Data, Barry Nuqoba
Correlation And Causation Analysis For Cross-Sectional And Panel Data, Barry Nuqoba
Dissertations and Theses Collection (Open Access)
This dissertation investigates how data, algorithms, and expert knowledge can be harnessed to better understand human behavior and enhance well-being. It emphasizes the critical importance of interdisciplinary collaboration to bridge knowledge gaps and foster insights that support preventive care, causal theory advancement, and policy development.
The first study, part of the SHINESeniors project, shed light on the potential usefulness of passive, unobtrusive sensors for detecting nocturia and poor sleep quality, symptoms commonly observed in chronic diseases, thereby enabling live-alone older adults to age in place. Utilizing machine learning techniques on sensor-derived features, the study can identify nocturia and poor sleep …
Towards Trustworthy Recommendation Systems: Beyond Collaborative Filtering, Zhongzhou Liu, Zhongzhou
Towards Trustworthy Recommendation Systems: Beyond Collaborative Filtering, Zhongzhou Liu, Zhongzhou
Dissertations and Theses Collection (Open Access)
Recommendation systems have been widely deployed in various scenarios and applications, such as e-commerce, social media, and streaming services. Recommendation systems have significantly influenced how we interact with various items in a wide range of platforms. They help users discover their preferred items and provide efficient and enjoyable experiences. They also help item providers and platforms to quickly find their potential customers, thus increasing the total revenue and user engagement.
The majority of existing recommendation systems merely focus on the matching between users and items, aiming for higher recommendation accuracy. Collaborative filtering is regarded as one of the most successful …
Food Computing: Domain Adaptation And Causal Inference, Qing Wang
Food Computing: Domain Adaptation And Causal Inference, Qing Wang
Dissertations and Theses Collection (Open Access)
This dissertation addresses two challenges in food computing: food recognition and food image-to-recipe retrieval. The main research ideas are: (1) leveraging Large Language Models (LLMs) to augment food image representations to mitigate the combined challenges of domain gaps and data imbalance in fine-grained food recognition; (2) proposing a causal-theory inspired cross-modal representation learning formulation for reducing the bias caused by the emphasis on certain ingredients for cross-modal recipe retrieval; and (3) extending the framework to incorporate multiple confounding factors, particularly ingredients and cooking actions, allows for more comprehensive modeling of the food image-torecipe retrieval problem.
We first explore the challenges …
Sequential Decision Learning For Social Good And Fairness, Dexun Li
Sequential Decision Learning For Social Good And Fairness, Dexun Li
Dissertations and Theses Collection (Open Access)
Sequential decision learning is one of the key research areas in artificial intelligence. Typically, a sequence of events is observed through a transformation that introduces uncertainty into the observations and based on these observations, the recognition process produces a hypothesis of the underlying events. This learning process is characterized by maximizing the sum of the reward signals. However, many real-life problems are inherently constrained by limited resources. Besides, when the learning algorithms are used to inform decisions involving human beings (e.g., Security and justice, health intervention, etc), they may inherit the potential, pre-existing bias in the dataset and exhibit similar …
Decentralized Consensus And Governance For Collaborative Intelligence, Huiwen Liu
Decentralized Consensus And Governance For Collaborative Intelligence, Huiwen Liu
Dissertations and Theses Collection (Open Access)
The data economy today is becoming increasingly collaborative in nature. Take business intelligence, for example. To unleash the full potential of big data, it is essential to integrate multi-source data depicting entities from a multi-faceted and multi-modal perspective, which, not surprisingly, is not achievable by any company alone. In collaborative intelligence, there are two core issues, namely "trust" and "incentive". The core mechanisms to solve these two problems are consensus and tokenization separately.
To solve the trust problem more effectively, we propose a systematic consensus evaluation framework to investigate whether existing consensus algorithms can do so. After a lot of …
Improving The Performance Of Wi-Fi Indoor Localization In Both Dense And Unknown Environments, Quang Truong Hai
Improving The Performance Of Wi-Fi Indoor Localization In Both Dense And Unknown Environments, Quang Truong Hai
Dissertations and Theses Collection (Open Access)
Indoor localization is important for various pervasive applications, garnering considerable research attention over recent decades. Despite numerous proposed solutions, the practical application of these methods in real-world environments with high applicability remains challenging. One compelling use case for building owners is the ability to track individuals as they navigate through the building, whether for security, customer analytics, space utilization planning, or other management purposes. However, this task becomes exceedingly difficult in environments with hundreds or thousands of people in motion. Conversely, the need to track oneself’s location is also meaningful from the perspective of individuals traversing in crowded spaces. These …
The Effect Of Internet Firms’ Data Analytics Capability On Their Innovation Speed And Innovation Quality: A Dynamic Capability Perspective, Yeyu Hua
Dissertations and Theses Collection (Open Access)
With the advent of big data era, data plays a pivotal role in sustainingfirms’ competitive advantages. Although a few studies have shown that data analytics capability contributes to firms’ innovative performance, these studies either focus on general innovative performance or specific types of innovation, such as incremental innovation, radical innovation, and supply chaininnovation. In this thesis, I enrich this stream of literature by conducting twostudies to further examine the relationship between data analytics capabilityand innovation speed as well as innovation quality. This thesis consists of twostudies. Study 1 is a survey study, in which I investigate the relationshipbetween data analytics …
Data-Driven Optimization Approaches For Dynamic Urban Logistics Operational Problems, Jingfeng Yang
Data-Driven Optimization Approaches For Dynamic Urban Logistics Operational Problems, Jingfeng Yang
Dissertations and Theses Collection (Open Access)
Given the rapid pace of urbanization, there is a pressing need to optimize urban logistics delivery operations for enhanced capacity and efficiency. Over recent decades, a multitude of optimization approaches have been put forth to address urban logistics challenges, encompassing routing and scheduling within both static and dynamic contexts. In light of the rising computational capabilities and the widespread adoption of machine learning in recent times, there is a growing body of research aimed at elucidating the seamless integration of data and machine learning within conventional urban logistics optimization models. Additionally, the ubiquitous utilization of smartphones and internet innovations presents …
Analyzing Taxi Drivers’ Decision-Making And Recommending Strategies For Enhanced Performance: A Data-Driven Approach, Mengyu Ji
Dissertations and Theses Collection (Open Access)
This thesis focuses on analyzing the decision-making process of taxi drivers and providing data-driven strategies to enhance their performance. By examin- ing comprehensive historical data encompassing passenger demand patterns, drivers’ spatial dynamics, and fare structures, valuable insights are gained into drivers’ choices regarding optimal routes, timing, and areas with high demand. Integrating real-time information sources, such as GPS data and passenger updates, allows drivers to adapt their strategies dynamically to changing traffic conditions and emerging demand patterns. Predictive analytics models, includ- ing ARIMA, XGBoost, and Linear Regression, are utilized to forecast demand flow at key locations, enabling proactive decision-making and …
Connecting The Dots For Contextual Information Retrieval, Pei-Chi Lo
Connecting The Dots For Contextual Information Retrieval, Pei-Chi Lo
Dissertations and Theses Collection (Open Access)
There are many information retrieval tasks that depend on knowledge graphs to return contextually relevant result of the query. We call them Knowledgeenriched Contextual Information Retrieval (KCIR) tasks and these tasks come in many different forms including query-based document retrieval, query answering and others. These KCIR tasks often require the input query to contextualized by additional facts from a knowledge graph, and using the context representation to perform document or knowledge graph retrieval and prediction. In this dissertation, we present a meta-framework that identifies Contextual Representation Learning (CRL) and Contextual Information Retrieval (CIR) to be the two key components in …
A Study Of The Impact Of Data Intelligence On Software Delivery Performance, Yongdong Dong
A Study Of The Impact Of Data Intelligence On Software Delivery Performance, Yongdong Dong
Dissertations and Theses Collection (Open Access)
With the rise of big data and artificial intelligence, data intelligence has gradually become the focus of academia and industry. Data intelligence has two obvious characteristics: big data drive and application scene drive. More and more enterprises extract valuable patterns contained in data with prediction and decision analysis methods and technologies such as large-scale data mining, machine learning and deep learning and use them to improve the management and decision in complex practice, so as to promote changes of new business modes, organizational structures and even business strategies, and improve the operational efficiency of organizations. However, there are few studies …
Mining Product Textual Data For Recommendation Explanations, Le Trung Hoang
Mining Product Textual Data For Recommendation Explanations, Le Trung Hoang
Dissertations and Theses Collection (Open Access)
Recommendation explanations help to make sense of recommendations, increasing the likelihood of adoption. Here, we are interested in mining product textual data, an unstructured data type, coming from manufacturers, sellers, or consumers, appearing in many places including title, summary, description, review, question and answers, etc., can be a rich source of information to explain the recommendation. As the explanation task could be decoupled from that of recommendation objective, we can categorize recommendation explanation into integrated approach, that uses a single interpretable model to produce both recommendation and explanation, or pipeline approach, that uses a post-hoc explanation model to produce explanation …
Robustness And Cross-Lingual Transfer: An Exploration Of Out-Of-Distribution Scenario In Natural Language Processing, Yu, Sicheng
Robustness And Cross-Lingual Transfer: An Exploration Of Out-Of-Distribution Scenario In Natural Language Processing, Yu, Sicheng
Dissertations and Theses Collection (Open Access)
Most traditional machine learning or deep learning methods are based on the premise that training data and test data are independent and identical distributed, i.e., IID. However, it is just an ideal situation. In real-world applications, test set and training data often follow different distributions, which we refer to as the out of distribution, i.e., OOD, setting. As a result, models trained with traditional methods always suffer from an undesirable performance drop on the OOD test set. It's necessary to develop techniques to solve this problem for real applications. In this dissertation, we present four pieces of work in the …
Finding Top-M Leading Records In Temporal Data, Yiyi Wang
Finding Top-M Leading Records In Temporal Data, Yiyi Wang
Dissertations and Theses Collection (Open Access)
A traditional top-k query retrieves the records that stand out at a certain point in time. On the other hand, a durable top-k query considers how long the records retain their supremacy, i.e., it reports those records that are consistently among the top-k in a given time interval. In this thesis, we introduce a new query to the family of durable top-k formulations. It finds the top-m leading records, i.e., those that rank among the top-k for the longest duration within the query interval. Practically, this query assesses the records based on how long …
Chinese Idiom Understanding With Transformer-Based Pretrained Language Models, Minghuan Tan
Chinese Idiom Understanding With Transformer-Based Pretrained Language Models, Minghuan Tan
Dissertations and Theses Collection (Open Access)
In this dissertation, I study the understanding of Chinese idioms using transformer-based pretrained language models. By ``understanding", I confine the topics to word embeddings learning, contextualized word representations learning, multiple-choice cloze-test reading comprehension and conditional text generation. Chinese idioms are fixed phrases that have special meanings usually derived from an ancient story. The meanings of these idioms are oftentimes not directly related to their component characters, which makes it hard to model them compared with standard phrases whose meanings are compositional. We initiate the work with studying idiom representations derived from pretrained language models, in particular, BERT. We adopt probing-based …
Modeling Sentiments And Preferences From Multimodal Data, Quoc Tuan Truong
Modeling Sentiments And Preferences From Multimodal Data, Quoc Tuan Truong
Dissertations and Theses Collection (Open Access)
Online reviews are prevalent in many modern Web applications, such as e-commerce, crowd-sourced location and check-in platforms. Fueled by the rise of mobile phones that are often the only cameras on hand, reviews are increasingly multimodal, with photos in addition to textual content. In this thesis, we focus on modeling the subjectivity carried in this form of data, with two research objectives.
In the first part, we tackle the problem of detecting sentiment expressed by a review. This is a key unlocking many applications, e.g., analyzing opinions, monitoring consumer satisfaction, assessing product quality.
Traditionally, the task of sentiment analysis primarily …
The Effects Of Recommender System On Sales Promotion Of High-Value Products: Evidence From A Field Experiment In The Real Estate Industry, Lian Liu
Dissertations and Theses Collection (Open Access)
Real estate sales industry in China has long suffered the problem of inefficient matching of customers to projects. Inspired by the design of recommender systems, which have been widely used in the online retail industry, and are shown to facility customer-product matching and improve sales, we apply this system to the real estate sales industry using a novel approach. Instead of recommending products to customers, we suggest the best potential customers to salespeople with whom they will conduct sales with. Using city-wide sales data from the largest real estate sales company in China, we first develop a recommend system based …
Deep Learning For Video-Grounded Dialogue Systems, Hung Le
Deep Learning For Video-Grounded Dialogue Systems, Hung Le
Dissertations and Theses Collection (Open Access)
In recent years, we have witnessed significant progress in building systems with artificial intelligence. However, despite advancements in machine learning and deep learning, we are still far from achieving autonomous agents that can perceive multi-dimensional information from the surrounding world and converse with humans in natural language. Towards this goal, this thesis is dedicated to building intelligent systems in the task of video-grounded dialogues. Specifically, in a video-grounded dialogue, a system is required to hold a multi-turn conversation with humans about the content of a video. Given an input video, a dialogue history, and a question about the video, the …
Can We Make It Better? Assessing And Improving Quality Of Github Repositories, Gede Artha Azriadi Prana
Can We Make It Better? Assessing And Improving Quality Of Github Repositories, Gede Artha Azriadi Prana
Dissertations and Theses Collection (Open Access)
The code hosting platform GitHub has gained immense popularity worldwide in recent years, with over 200 million repositories hosted as of June 2021. Due to its popularity, it has great potential to facilitate widespread improvements across many software projects. Naturally, GitHub has attracted much research attention, and the source code in the various repositories it hosts also provide opportunity to apply techniques and tools developed by software engineering researchers over the years. However, much of existing body of research applicable to GitHub focuses on code quality of the software projects and ways to improve them. Fewer work focus on potential …
Novel Techniques In Recovering, Embedding, And Enforcing Policies For Control-Flow Integrity, Yan Lin
Novel Techniques In Recovering, Embedding, And Enforcing Policies For Control-Flow Integrity, Yan Lin
Dissertations and Theses Collection (Open Access)
Control-Flow Integrity (CFI) is an attractive security property with which most injected and code-reuse attacks can be defeated, including advanced attacking techniques like Return-Oriented Programming (ROP). CFI extracts a control-flow graph (CFG) for a given program and instruments the program to respect the CFG. Specifically, checks are inserted before indirect branch instructions. Before these instructions are executed during runtime, the checks consult the CFG to ensure that the indirect branch is allowed to reach the intended target. Hence, any sort of controlflow hijacking would be prevented. There are three fundamental components in CFI enforcement. The first component is accurately recovering …
Vision-Based Analytics For Improved Ai-Driven Iot Applications, Amit Sharma
Vision-Based Analytics For Improved Ai-Driven Iot Applications, Amit Sharma
Dissertations and Theses Collection (Open Access)
Proliferation of Internet of Things (IoT) sensor systems, primarily driven by cheaper embedded hardware platforms and wide availability of light-weight software platforms, has opened up doors for large-scale data collection opportunities. The availability of massive amount of data has in-turn given way to rapidly growing machine learning models e.g. You Only Look Once (YOLO), Single-Shot-Detectors (SSD) and so on. There has been a growing trend of applying machine learning techniques, e.g., object detection, image classification, face detection etc., on data collected from camera sensors and therefore enabling plethora of vision-sensing applications namely self-driving cars, automatic crowd monitoring, traffic-flow analysis, occupancy …
Deep Learning For Real-World Object Detection, Xiongwei Wu
Deep Learning For Real-World Object Detection, Xiongwei Wu
Dissertations and Theses Collection (Open Access)
Despite achieving significant progresses, most existing detectors are designed to detect objects in academic contexts but consider little in real-world scenarios. In real-world applications, the scale variance of objects can be significantly higher than objects in academic contexts; In addition, existing methods are designed for achieving localization with relatively low precision, however more precise localization is demanded in real-world scenarios; Existing methods are optimized with huge amount of annotated data, but in certain real-world scenarios, only a few samples are available. In this dissertation, we aim to explore novel techniques to address these research challenges to make object detection algorithms …
Using Knowledge Bases For Question Answering, Yunshi Lan
Using Knowledge Bases For Question Answering, Yunshi Lan
Dissertations and Theses Collection (Open Access)
A knowledge base (KB) is a well-structured database, which contains many of entities and their relations. With the fast development of large-scale knowledge bases such as Freebase, DBpedia and YAGO, knowledge bases have become an important resource, which can serve many applications, such as dialogue system, textual entailment, question answering and so on. These applications play significant roles in real-world industry.
In this dissertation, we try to explore the entailment information and more general entity-relation information from the KBs. Recognizing textual entailment (RTE) is a task to infer the entailment relations between sentences. We need to decide whether a hypothesis …
Question Answering With Textual Sequence Matching, Shuohang Wang
Question Answering With Textual Sequence Matching, Shuohang Wang
Dissertations and Theses Collection (Open Access)
Question answering (QA) is one of the most important applications in natural language processing. With the explosive text data from the Internet, intelligently getting answers of questions will help humans more efficiently collect useful information. My research in this thesis mainly focuses on solving question answering problem with textual sequence matching model which is to build vectorized representations for pairs of text sequences to enable better reasoning. And our thesis consists of three major parts.
In Part I, we propose two general models for building vectorized representations over a pair of sentences, which can be directly used to solve the …
Modeling Sequential And Basket-Oriented Associations For Top-K Recommendation, Duc-Trong Le Duc Trong
Modeling Sequential And Basket-Oriented Associations For Top-K Recommendation, Duc-Trong Le Duc Trong
Dissertations and Theses Collection (Open Access)
Top-K recommendation is a typical task in Recommender Systems. In traditional approaches, it mainly relies on the modeling of user-item associations, which emphasizes the user-specific factor or personalization. Here, we investigate another direction that models item-item associations, especially with the notions of sequence-aware and basket-level adoptions . Sequences are created by sorting item adoptions chronologically. The associations between items along sequences, referred to as “sequential associations”, indicate the influence of the preceding adoptions on the following adoptions. Considering a basket of items consumed at the same time step (e.g., a session, a day), “basket-oriented associations” imply correlative dependencies among these …