Open Access. Powered by Scholars. Published by Universities.®
Databases and Information Systems Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Numerical Analysis and Scientific Computing (671)
- Social and Behavioral Sciences (374)
- Artificial Intelligence and Robotics (356)
- Graphics and Human Computer Interfaces (313)
- Business (253)
-
- Software Engineering (239)
- Communication (234)
- Social Media (202)
- Engineering (199)
- Theory and Algorithms (178)
- Computer Engineering (171)
- Information Security (148)
- OS and Networks (116)
- Programming Languages and Compilers (98)
- E-Commerce (84)
- Data Storage Systems (75)
- Medicine and Health Sciences (70)
- Public Affairs, Public Policy and Public Administration (60)
- Education (55)
- Management Information Systems (53)
- International and Area Studies (51)
- Asian Studies (50)
- Health Information Technology (48)
- Transportation (47)
- Finance and Financial Management (43)
- Digital Communications and Networking (32)
- Technology and Innovation (32)
- Keyword
-
- Social media (59)
- Machine learning (56)
- Online learning (46)
- Deep learning (43)
- Data mining (42)
-
- Artificial intelligence (36)
- Twitter (30)
- Query processing (29)
- Classification (26)
- Neural networks (25)
- Reinforcement learning (25)
- Deep Learning (24)
- Algorithms (23)
- Clustering (21)
- Social network (21)
- Algorithm (20)
- Graph neural networks (20)
- Machine Learning (20)
- Natural language processing (20)
- Recommender systems (20)
- Semantics (20)
- Task analysis (20)
- Anomaly detection (19)
- Cloud computing (19)
- Visualization (19)
- Image retrieval (18)
- Performance (18)
- Sentiment analysis (18)
- Singapore (18)
- Social networks (17)
- Publication Year
- Publication
-
- Research Collection School Of Computing and Information Systems (3436)
- Dissertations and Theses Collection (Open Access) (58)
- Research Collection Lee Kong Chian School Of Business (11)
- Asian Management Insights (8)
- Research Collection School Of Accountancy (7)
-
- Dissertations and Theses Collection (5)
- PhD Student’s Publications Collection (5)
- Research Collection College of Integrative Studies (5)
- Research Collection Yong Pung How School Of Law (5)
- MITB Thought Leadership Series (3)
- LARC Research Publications (2)
- Perspectives@SMU (2)
- Research Collection School of Computing and Information Systems (2)
- 2024 AI for Research Week (1)
- CCX Research (1)
- Research Collection School Of Economics (1)
- Research Collection School of Accountancy (1)
- Research Collection School of Social Sciences (1)
- Research@SMU Infographics (1)
- Publication Type
Articles 61 - 90 of 3555
Full-Text Articles in Databases and Information Systems
Nondeterministic Polynomial-Time Problem Challenge: An Ever-Scaling Reasoning Benchmark For Llms, Chang Yang, Ruiyu Wang, Junzhe Jiang, Qi Jiang, Qinggang Zhang, Yanchen Deng, Shuxin Li, Shuyue Hu, Bo Li, Florian T. Pokorny, Xiao Huang, Xinrun Wang
Nondeterministic Polynomial-Time Problem Challenge: An Ever-Scaling Reasoning Benchmark For Llms, Chang Yang, Ruiyu Wang, Junzhe Jiang, Qi Jiang, Qinggang Zhang, Yanchen Deng, Shuxin Li, Shuyue Hu, Bo Li, Florian T. Pokorny, Xiao Huang, Xinrun Wang
Research Collection School Of Computing and Information Systems
Reasoning is the fundamental capability of large language models (LLMs). Due to the rapid progress of LLMs, there are two main issues of current benchmarks: i) these benchmarks can be crushed in a short time (less than 1 year), and ii) these benchmarks may be easily hacked. To handle these issues, we propose the ever-scalingness for building the benchmarks which are scaling over complexity against crushing, instance against hacking and exploitation, oversight for easy verification, and coverage for real-world relevance. This paper presents Nondeterministic Polynomial-time Problem Challenge (NPPC), an ever-scaling reasoning benchmark for LLMs. Specifically, the NPPC has three main …
Reinforce Trustworthiness In Multimodal Emotional Support System, Huy M. Le, Dat Tien Nguyen, Ngan T. T. Vo, Tuan D. Q. Nguyen, Nguyen Le Binh, Duy Minh Ho Nguyen, Daniel Sonntag, Lizi Liao, Binh T. Nguyen
Reinforce Trustworthiness In Multimodal Emotional Support System, Huy M. Le, Dat Tien Nguyen, Ngan T. T. Vo, Tuan D. Q. Nguyen, Nguyen Le Binh, Duy Minh Ho Nguyen, Daniel Sonntag, Lizi Liao, Binh T. Nguyen
Research Collection School Of Computing and Information Systems
In today's world, emotional support is increasingly essential, yet it remains challenging for both those seeking help and those offering it. Multimodal approaches to emotional support show great promise by integrating diverse data sources to provide empathetic, contextually relevant responses, fostering more effective interactions. However, current methods have notable limitations, often relying solely on text or converting other data types into text, or providing emotion recognition only, thus overlooking the full potential of multimodal inputs. Moreover, many studies prioritize response generation without accurately identifying critical emotional support elements or ensuring the reliability of outputs. To overcome these issues, we introduce …
Leveraging Large Language Models For Career Mobility Analysis: A Study Of Gender, Race, And Job Change Using Us Online Resume Profiles, Palakorn Achananuparp, Ye Xu, Yao Lu, Xavier Jayaraj Siddarth Ashok, Ee-Peng Lim
Leveraging Large Language Models For Career Mobility Analysis: A Study Of Gender, Race, And Job Change Using Us Online Resume Profiles, Palakorn Achananuparp, Ye Xu, Yao Lu, Xavier Jayaraj Siddarth Ashok, Ee-Peng Lim
Research Collection School Of Computing and Information Systems
We present a large-scale analysis of career mobility of college-educated U.S. workers using online resume profiles to investigate how gender, race, and job change options are associated with upward mobility. This study addresses key research questions of how the job changes affect their upward career mobility, and how the outcomes of upward career mobility differ by gender and race. We address data challenges – such as missing demographic attributes, missing wage data, and noisy occupation labels – through various data processing and Artificial Intelligence (AI) methods. In particular, we develop a large language models (LLMs) based occupation classification method known …
Towards Inclusive Digital Futures Of Cultural Heritage: Insights From A Critical Discourse Analysis Of Unesco Dialogues, Shiqing Huang, Keng Siau, Xiaoting Chen
Towards Inclusive Digital Futures Of Cultural Heritage: Insights From A Critical Discourse Analysis Of Unesco Dialogues, Shiqing Huang, Keng Siau, Xiaoting Chen
Research Collection School Of Computing and Information Systems
Digital technologies are shaping many aspects of cultural heritage, but very little research has examined the implications of digital transformation. Drawing on concepts from Fairclough’s three-dimensional critical discourse analysis, this research examines the discourse using seven online dialogues (available on the UNESCO website) between 18 professionals who have different backgrounds and cultures to identify social practices related to the digital transformation of cultural heritage. We identify four digital transformation discourse types in professional dialogues: documentation, management, interpretation, and interaction. We also identify seven main groups: memory institutions including libraries, archives, and museums (LAMs), governments, international organizations, art and creative supporters, …
Reliable-Data-Split (Rds): Maximizing Model Potential With Reinforced Selection Strategy, Hoang D. Nguyen, Xuan-Son Vu, Quoc Tuan Truong, Duc-Trong Le
Reliable-Data-Split (Rds): Maximizing Model Potential With Reinforced Selection Strategy, Hoang D. Nguyen, Xuan-Son Vu, Quoc Tuan Truong, Duc-Trong Le
Research Collection School Of Computing and Information Systems
The nexus between data characteristics and parametric models is fundamental for developing effective and reliable artificial intelligence (AI) systems. Mismatches in data properties for model development may lead to deleterious effects on AI model performance in machine learning practice. This paper proposes a Reliable Data Split (RDS) procedure to learn how to select data points that will generalise the target domain adequately by employing prior knowledge of the data generative process. We introduce a reinforced selection strategy using deep reinforcement learning with diverse black box predictors in maximising ensemble rewards as the proxy of model performance potential while maintaining an …
Digital Communications Between Firms And Investors: Impact Of Explanatory Responses On Investor Engagement In Online Financial Q&A, Runyu Wang, Zili Zhang, Keng Siau, Ziqiong Zhang
Digital Communications Between Firms And Investors: Impact Of Explanatory Responses On Investor Engagement In Online Financial Q&A, Runyu Wang, Zili Zhang, Keng Siau, Ziqiong Zhang
Research Collection School Of Computing and Information Systems
The emerging trend of digital communications between firms and investors through online question-and-answer (Q&A) platforms is recognized as a vital strategy for managing investor relations, contributing to enhanced market efficiency and information transparency through increased information exchange. Potential investors can seek responses from firm managers to address their information needs, thereby mitigating market uncertainties. To provide foundational insights, we conduct a survey of investors to assess their awareness, usage, and perceptions of firm-investor Q&A platforms. In the subsequent empirical study, we specifically focus on the substance of managers’ responses, which are primarily aimed at clarifying firm events or information. In …
Real-Time Estimated Sequential Organ Failure Assessment (Sofa) Score With Intervals: Improved Risk Monitoring With Estimated Uncertainty In Health Condition For Patients In Intensive Care Units, Yan He, Qian Luo, Hai Wang, Zhichao Zheng, Haidong Luo, Oon Cheong Ooi
Real-Time Estimated Sequential Organ Failure Assessment (Sofa) Score With Intervals: Improved Risk Monitoring With Estimated Uncertainty In Health Condition For Patients In Intensive Care Units, Yan He, Qian Luo, Hai Wang, Zhichao Zheng, Haidong Luo, Oon Cheong Ooi
Research Collection Lee Kong Chian School Of Business
Purpose: Real-time risk monitoring is critical but challenging in intensive care units (ICUs) due to the lack of real-time updates for most clinical variables. Although real-time predictions have been integrated into various risk-scoring systems to aid monitoring, existing systems do not address uncertainties in risk assessments. We developed an enhanced risk monitoring framework based on commonly used systems like the Sequential Organ Failure Assessment (SOFA) score by incorporating uncertainties to improve the effectiveness of real-time risk monitoring in ICUs.Methods: This study included 5,351 patients admitted to the Cardiothoracic ICU in the National University Hospital in Singapore. We developed machine learning …
Scalable Graph Indexing Using Gpus For Approximate Nearest Neighbor Search, Zhonggen Li, Xiangyu Ke, Yifan Zhu, Bocheng Yu, Baihua Zheng, Yunjun Gao
Scalable Graph Indexing Using Gpus For Approximate Nearest Neighbor Search, Zhonggen Li, Xiangyu Ke, Yifan Zhu, Bocheng Yu, Baihua Zheng, Yunjun Gao
Research Collection School Of Computing and Information Systems
Approximate nearest neighbor search (ANNS) in high-dimensional vector spaces has a wide range of real-world applications. Numerous methods have been proposed to handle ANNS efficiently, while graph-based indexes have gained prominence due to their high accuracy and efficiency. However, the indexing overhead of graph-based indexes remains substantial. With exponential growth in data volume and increasing demands for dynamic index adjustments, this overhead continues to escalate, posing a critical challenge.In this paper, we introduce Tagore, a fasT library accelerated by GPUs for graph indexing, which has powerful capabilities of constructing refinement-based graph indexes such as NSG and Vamana. We first introduce …
Pilot-C: Physics-Informed Low-Distortion Optimal Trajectory Compression, Kefei Wu, Baihua Zheng, Weiwei Sun
Pilot-C: Physics-Informed Low-Distortion Optimal Trajectory Compression, Kefei Wu, Baihua Zheng, Weiwei Sun
Research Collection School Of Computing and Information Systems
Location-aware devices continuously generate massive volumes of trajectory data, creating demand for efficient compression. Line simplification is a common solution but typically assumes 2D trajectories and ignores time synchronization and motion continuity. We propose PILOT-C, a novel trajectory compression framework that integrates frequency-domain physics modeling with error-bounded optimization. Unlike existing line simplification methods, PILOT-C supports trajectories in arbitrary dimensions, including 3D, by compressing each spatial axis independently. Evaluated on four real-world datasets, PILOT-C achieves superior performance across multiple dimensions. In terms of compression ratio, PILOT-C outperforms CISED-W, the current state-of-the-art SED-based line simplification algorithm, by an average of 19.2%. For …
The Rise Of Parameter Specialization For Knowledge Storage In Large Language Models, Yihuai Hong, Yiran Zhao, Wei Tang, Yang Deng, Yu Rong, Wenxuan Zhang
The Rise Of Parameter Specialization For Knowledge Storage In Large Language Models, Yihuai Hong, Yiran Zhao, Wei Tang, Yang Deng, Yu Rong, Wenxuan Zhang
Research Collection School Of Computing and Information Systems
Over time, a growing wave of large language models from various series has been introduced to the community. Researchers are striving to maximize the performance of language models with constrained parameter sizes. However, from a microscopic perspective, there has been limited research on how to better store knowledge in model parameters, particularly within MLPs, to enable more effective utilization of this knowledge by the model. In this work, we analyze twenty publicly available open-source large language models to investigate the relationship between their strong performance and the way knowledge is stored in their corresponding MLP parameters. Our findings reveal that …
A Partition Cover Approach To Tokenization, Jia Peng Lim, Shawn Tan, Davin Choo, Hady Wirawan Lauw
A Partition Cover Approach To Tokenization, Jia Peng Lim, Shawn Tan, Davin Choo, Hady Wirawan Lauw
Research Collection School Of Computing and Information Systems
Tokenization is the process of encoding strings into tokens of a fixed vocabulary size, and is widely utilized in Natural Language Processing applications. The leading tokenization algorithm today is Byte Pair Encoding (BPE), which formulates the tokenization problem as a compression problem and tackles it by performing sequences of merges. In this work, we formulate tokenization as an optimization objective, show that it is NP-hard via a simple reduction from vertex cover, and propose a polynomial-time greedy algorithm GreedTok. Our formulation naturally relaxes to the well-studied weighted maximum coverage problem which has a simple -approximation algorithm GreedWMC. Through empirical evaluations …
Design Principles For Customer‑Engaging Digital Service Systems: An Action Research Study, Keng Siau, Xiaofeng Chen, Xin Tan
Design Principles For Customer‑Engaging Digital Service Systems: An Action Research Study, Keng Siau, Xiaofeng Chen, Xin Tan
Research Collection School Of Computing and Information Systems
Digital services represent a business approach employed by organizations to operate in the digital environment. However, systematic development guidelines for developing quality digital service systems are lacking in the literature. The authors identified four general challenges for developing and implementing customer-engaging digital service systems (CEDSS). By employing the method of canonical action research in a digital service system project, they derived 10 design principles for developing high-quality CEDSS. They empirically evaluated the design principles in the development project and through follow-up focus group sessions. The design principles provide applicable and actionable guidelines for the development of CEDSS.
Island-Based Evolutionary Computation With Diverse Surrogates And Adaptive Knowledge Transfer For High-Dimensional Data-Driven Optimization, Xianrong Zhang, Yuejiao Gong, Zhiguang Cao, Jun Zhang
Island-Based Evolutionary Computation With Diverse Surrogates And Adaptive Knowledge Transfer For High-Dimensional Data-Driven Optimization, Xianrong Zhang, Yuejiao Gong, Zhiguang Cao, Jun Zhang
Research Collection School Of Computing and Information Systems
In recent years, there has been a growing interest in data-driven evolutionary algorithms (DDEAs) employing surrogate models to approximate the objective functions with limited data. However, current DDEAs are primarily designed for lower-dimensional problems and their performance drops significantly when applied to large-scale optimization problems (LSOPs). To address the challenge, this paper proposes an offline DDEA named DSKT-DDEA. DSKT-DDEA leverages multiple islands that utilize different data to establish diverse surrogate models, fostering diverse subpopulations and mitigating the risk of premature convergence. In the intra-island optimization phase, a semi-supervised learning method is devised to fine-tune the surrogates. It not only facilitates …
Addressing Sparsity For Knowledge Graph Completion: Data And Model Perspectives, Ran Liu
Addressing Sparsity For Knowledge Graph Completion: Data And Model Perspectives, Ran Liu
Dissertations and Theses Collection (Open Access)
Knowledge graphs (KGs) are powerful tools for structuring factual knowledge into relational triples, yet their practical utility is often adversely affected by data sparsity. Many entities and relations are associated with only a few observations, which limits the quality of learned embeddings and weakens generalization in downstream tasks. The problem of sparsity led to two interrelated challenges. Firstly, it restricts the informativeness of training samples: positive examples are scarce, and conventional negative sampling often produces trivial or redundant negatives that resulting in limited guidance. Secondly, in few-shot relation learning scenarios, sparsity worsens distribution shifts between training and test relations, as …
Usefulness And Diminishing Returns: Evaluating Social Information In Recommender Systems, Qing Meng, Huiyu Min, Ming Shan Hee, Roy Ka-Wei Lee, Bing Tian Dai, Shuai Xu
Usefulness And Diminishing Returns: Evaluating Social Information In Recommender Systems, Qing Meng, Huiyu Min, Ming Shan Hee, Roy Ka-Wei Lee, Bing Tian Dai, Shuai Xu
Research Collection School Of Computing and Information Systems
Social recommendation, which leverages users’ social information to predict users’ preferences, is a popular branch of recommender systems. Many existing studies have attempted to advance the performance of collaborative filtering methods by leveraging the user-user matrix to enhance user embedding learning with user’s social connections. While the existing social recommender systems have demonstrated good performance in various recommendation tasks, the extent of social information usefulness in recommender systems remains unclear. This paper addresses the research gap by designing experiments to answer three research questions: (i) How useful is social information in varying user-item data sparsity? (ii) How much social information …
International Workshop On Multimodal Generative Search And Recommendation (Mmgensr@Cikm 2025), Yi Bin, Haoxuan Li, Haokai Ma, Yang Zhang, Wenjie Wang, Yunshan Ma, Yang Yang, Tat‑Seng Chua
International Workshop On Multimodal Generative Search And Recommendation (Mmgensr@Cikm 2025), Yi Bin, Haoxuan Li, Haokai Ma, Yang Zhang, Wenjie Wang, Yunshan Ma, Yang Yang, Tat‑Seng Chua
Research Collection School Of Computing and Information Systems
Recent breakthroughs in generative Artificial Intelligence (AI) have ignited a revolutionary wave across information retrieval and recommender systems. This workshop serves as a premier interdisciplinary platform to explore how generative models, particularly Large Language Models (LLMs) and Large Multimodal Models (LMMs), are transforming multimodal search and recommendation paradigms [3, 6, 9, 10, 12-14]. We aim to convene researchers and practitioners to discuss innovative architectures, methodologies, and evaluation strategies spanning generative document retrieval [5, 8] generative image retrieval [ 7, 16], grounded answer generation [17], generative recommendation [2, 4, 11], and related tasks involving multiple modalities [1,15]. The workshop will facilitate …
Damslnet: Dual-Attention Multi-Scale Lightweight Network For Plant Disease Classification, Linfan Deng, Juan Qin, Kun Li, Jinhua Zhu, Zhaoxia Wang
Damslnet: Dual-Attention Multi-Scale Lightweight Network For Plant Disease Classification, Linfan Deng, Juan Qin, Kun Li, Jinhua Zhu, Zhaoxia Wang
Research Collection School Of Computing and Information Systems
Accurately identifying crop diseases plays a crucial role in advancing intelligent and modern agricultural production. Deep learning techniques have performed robust performance in classifying plant disease images. However, current studies face the challenge that many plant disease datasets are generated in controlled environments, leading to reduced model performance in real-world agricultural settings. This paper aims to provide a lightweight model that can accurately classify plant diseases in natural environments. Specifically, this paper investigates the Dual-Attention Multi-Scale Lightweight Network (DAMSLNet), which combines dual-attention-based multi-scale feature extraction and deep information fusion, to classify plant diseases. At the front end, the model employs …
Predict Social Economic Outcomes By Transferred Knowledge With Satellite Imagery, Yang Tang, Shih-Fen Cheng, Yunqiang Zhu, Yichen Yang, Zhiqiang Zou
Predict Social Economic Outcomes By Transferred Knowledge With Satellite Imagery, Yang Tang, Shih-Fen Cheng, Yunqiang Zhu, Yichen Yang, Zhiqiang Zou
Research Collection School Of Computing and Information Systems
Traditional deep learning methods and econometric models have played a crucial role in the field of data mining, particularly in the prediction of socioeconomic outcomes. However, socio-economic information is unable to be directly extracted from remote sensing data. So, in this paper, we propose a method to leverage transfer learning to predict socioeconomic indicators (outcomes) through satellite imagery. Specifically, we use road network types as a proxy for socioeconomic factors, which is more effective and stable than using nightlight. We have extracted eleven distinct road topological features to generate reasonable road network types. Given the unique characteristics of road networks, …
When Deep Learning Meets Information Retrieval-Based Bug Localization: A Survey, Feifei Niu, Chuanyi Li, Kui Liu, Xin Xia, David Lo
When Deep Learning Meets Information Retrieval-Based Bug Localization: A Survey, Feifei Niu, Chuanyi Li, Kui Liu, Xin Xia, David Lo
Research Collection School Of Computing and Information Systems
Bug localization is a crucial aspect of software maintenance, running through the entire software lifecycle. Information retrieval-based bug localization (IRBL) identifies buggy code based on bug reports, expediting the bug resolution process for developers. Recent years have witnessed significant achievements in IRBL, propelled by the widespread adoption of deep learning (DL). To provide a comprehensive overview of the current state of the art and delve into key issues, we conduct a survey encompassing 61 IRBL studies leveraging DL. We summarize best practices in each phase of the IRBL workflow, undertake a meta-analysis of prior studies, and suggest future research directions. …
Contrastrepair: Enhancing Conversation-Based Automated Program Repair Via Contrastive Test Case Pairs, Jiaolong Kong, Xiaofei Xie, Mingfei Cheng, Shangqing Liu, Xiaoning Du, Qi Guo
Contrastrepair: Enhancing Conversation-Based Automated Program Repair Via Contrastive Test Case Pairs, Jiaolong Kong, Xiaofei Xie, Mingfei Cheng, Shangqing Liu, Xiaoning Du, Qi Guo
Research Collection School Of Computing and Information Systems
Automated Program Repair (APR) aims to automatically generate patches for rectifying software bugs. Recentstrides in Large Language Models (LLM), such as ChatGPT, have yielded encouraging outcomes in APR,especially within the conversation-driven APR framework. Nevertheless, the efficacy of conversation-drivenAPR is contingent on the quality of the feedback information. In this article, we propose ContrastRepair, anovel conversation-based APR approach that augments conversation-driven APR by providing LLMs withcontrastive test pairs. A test pair consists of a failing test and a passing test, which offer contrastive feedback tothe LLM. Our key insight is to minimize the difference between the generated passing test and the …
A Data-Driven Framework For Optimal Retail Store Location, Ming Hui Tan
A Data-Driven Framework For Optimal Retail Store Location, Ming Hui Tan
Dissertations and Theses Collection (Open Access)
This study develops a data-driven framework for optimal retail store location planning that integrates road network analysis, mobility data and optimization techniques. By addressing the limitations of traditional approaches that rely on outdated census data and manual site selection, this research offers a scalable and adaptable solution for retail expansion in diverse urban environments. Chapters 1 and 2 establish the foundational context and theoretical underpinnings of this research. Chapter 1 introduces the research problem and motivation, highlighting the limitations of existing approaches and defining three key research objectives: automating candidate site identification, improving footfall estimation, and developing a scalable multi-site …
Impact Of Original Versus Reposted Social Endorsements On Content Consumption: The Moderating Role Of Endorsers’ Network Characteristics, Anqi Zhao, Qian Tang
Impact Of Original Versus Reposted Social Endorsements On Content Consumption: The Moderating Role Of Endorsers’ Network Characteristics, Anqi Zhao, Qian Tang
Research Collection School Of Computing and Information Systems
Social endorsements broadcast endorsers’ positive attitudes toward content or products, especially to their social ties. Original endorsements created by endorsers can be propagated further as reposted endorsements. Both are important marketing tools to increase content consumption, yet their differences are unclear. This study compares the impacts of original and reposted endorsements on content consumption and their contingencies on the endorsers’ network characteristics. Using data on social endorsements of YouTube videos on Twitter, we find that original endorsements (i.e., original tweets) significantly boost content consumption, and the effect is positively moderated by the endorsers’ network size but not their tie strength. …
Filterfl: Knowledge Filtering-Based Data-Free Backdoor Defense For Federated Learning, Yanxin Yang, Ming Hu, Xiaofei Xie, Yue Cao, Pengyu Zhang, Yihao Huang, Mingsong Chen
Filterfl: Knowledge Filtering-Based Data-Free Backdoor Defense For Federated Learning, Yanxin Yang, Ming Hu, Xiaofei Xie, Yue Cao, Pengyu Zhang, Yihao Huang, Mingsong Chen
Research Collection School Of Computing and Information Systems
As a distributed machine learning paradigm, Federated Learning (FL) enables large-scale clients to collaboratively train a model without sharing their raw data. However, due to the lack of data auditing for untrusted clients, FL is vulnerable to poisoning attacks, especially backdoor attacks. By using poisoned data for local training or directly changing the model parameters, attackers can easily inject backdoors into the model, which can trigger the model to make misclassification of targeted patterns in images. To address these issues, we propose a novel data-free trigger-generation-based defense approach based on the two characteristics of backdoor attacks: i) triggers are learned …
Website Owner Identification Through Multi-Level Contrastive Representation Learning, Cheng Tu, Yunshan Ma, Yang Li, Min Zhang, Miao Hu, Fan Shi, Xiang Wang
Website Owner Identification Through Multi-Level Contrastive Representation Learning, Cheng Tu, Yunshan Ma, Yang Li, Min Zhang, Miao Hu, Fan Shi, Xiang Wang
Research Collection School Of Computing and Information Systems
Website owner identification aims to recognize the organization or individual who owns a given website that is served on the web. It is a crucial step for cyberspace surveying and mapping, playing a significant role in cyberspace administration and governance. Existing widely employed solutions for website owner identification mainly fall into two paradigms: (1) querying the public information databases such as WHOIS, which store the Internet resource’s registered users or assignees; and (2) directly extracting the organization or individual name of the website owner from the webpage using the technique of named entity recognition. However, the former is less reliable …
Deep Learning For Hate Speech Detection: A Comparative Study, Jitendra Singh Malik, Hezhe Qiao, Guansong Pang, Anton Van Den Hengel
Deep Learning For Hate Speech Detection: A Comparative Study, Jitendra Singh Malik, Hezhe Qiao, Guansong Pang, Anton Van Den Hengel
Research Collection School Of Computing and Information Systems
Automated hate speech detection is an important tool in combating the spread of hate speech, particularly in social media. Numerous methods have been developed for the task, including a recent proliferation of deep-learning based approaches. A variety of datasets have also been developed, exemplifying various manifestations of the hate-speech detection problem. We present here a largescale empirical comparison of deep and shallow hate-speech detection methods, mediated through the three most commonly used datasets. Our goal is to illuminate progress in the area, and identify strengths and weaknesses in the current state-of-the-art. We particularly focus our analysis on measures of practical …
Probabilistic Modeling, Learnability And Uncertainty Estimation For Interaction Prediction In Movie Rating Datasets, Jennifer Poernomo, Nicole Gabrielle Lee Tan, Rodrigo Alves, Antoine Ledent
Probabilistic Modeling, Learnability And Uncertainty Estimation For Interaction Prediction In Movie Rating Datasets, Jennifer Poernomo, Nicole Gabrielle Lee Tan, Rodrigo Alves, Antoine Ledent
Research Collection School Of Computing and Information Systems
In this paper, we examine the hypothesis that the interactions recorded in many Recommendation Systems datasets are distributed according to a low-rank distribution, i.e. a mixture of factorizable distributions. Surprisingly, we find that on several popular datasets, a simple non-negative matrix factorization method equals or outperforms more modern methods such as LightGCN, which indicates that the sampling distribution over interactions is indeed low-rank. Furthermore, we mathematically prove that low-rank distributions are learnable with a sparse number of observations (where m/n and r refer to the number of users/items and the non-negative rank respectively) both in terms of the total variation …
Managing Rumors On Electronic Interaction Platforms: How Management Responses Affect Investor Reaction, Runyu Wang, Zili Zhang, Keng Siau, Ziqiong Zhang
Managing Rumors On Electronic Interaction Platforms: How Management Responses Affect Investor Reaction, Runyu Wang, Zili Zhang, Keng Siau, Ziqiong Zhang
Research Collection School Of Computing and Information Systems
This study investigates how listed firms respond to investors’ rumor-related inquiries and examines the impact of these responses on investor reactions, as indicated by subsequent daily abnormal stock returns (ARs). Using a unique dataset of question-and-answer (Q&A) interactions from China’s major e-interaction platforms, established by the stock exchanges, our study provides insights into regulated firm-investor communications in a structured Q&A setting. Unlike informal social media channels, these platforms enable official responses from firm representatives, typically board secretaries, under direct regulatory oversight. By analyzing rumor-related Q&A pairs with regression models and several robustness checks, we find that firms can benefit from …
Memory-Efficient Graph Processing On Gpus: Reducing Intermediate Data Structure Overhead, Chang Ye
Memory-Efficient Graph Processing On Gpus: Reducing Intermediate Data Structure Overhead, Chang Ye
Dissertations and Theses Collection (Open Access)
The increasing scale of real-world graphs in domains such as fraud detection, community detection, and biological analysis demands high-throughput, memory-efficient graph processing solutions. GPUs offer massive parallelism for accelerating such workloads, and numerous frameworks have been developed to leverage their computational power. These frameworks primarily focus on optimizing scheduling to better align graph processing with GPU architectures. It performs well for algorithms with low memory demands, such as BFS, SSSP, and PageRank. However, for algorithms that require substantial memory, such as label propagation, and subgraph counting, the limited memory capacity of GPUs often becomes a significant bottleneck.
This dissertation addresses …
Recurrent Autoregressive Linear Model For Next-Basket Recommendation, Tereza Zmeskalova, Antoine Ledent, Martin Spisak, Pavel Kordik, Rodrigo Alves
Recurrent Autoregressive Linear Model For Next-Basket Recommendation, Tereza Zmeskalova, Antoine Ledent, Martin Spisak, Pavel Kordik, Rodrigo Alves
Research Collection School Of Computing and Information Systems
Next-basket recommendation aims to predict the (sets of) items that a user is most likely to purchase during their next visit, capturing both short-term sequential patterns and long-term user preferences. However, effectively modeling these dynamics remains a challenge for traditional methods, which often struggle with interpretability and computational efficiency, particularly when dealing with intricate temporal dependencies and inter-item relationships. In this paper, we propose ReALM, a Recurrent Autoregressive Linear Model that explicitly captures temporal item-to-item dependencies across multiple time steps. By leveraging a recurrent loss function and a closed-form optimization solution, our approach offers both interpretability and scalability while maintaining …
An Efficient Security-Enhanced Accountable Access Control For Named Data Networking, Jianfei Sun, Yuxian Li, Xuehuan Yang, Guomin Yang, Robert H. Deng
An Efficient Security-Enhanced Accountable Access Control For Named Data Networking, Jianfei Sun, Yuxian Li, Xuehuan Yang, Guomin Yang, Robert H. Deng
Research Collection School Of Computing and Information Systems
Named Data Networking (NDN) is embraced as the crucial implementation of Information-Centric Networking (ICN), enhancing content distribution and caching efficiency through edge routers. However, existing NDN architectures face significant security and privacy challenges, including: (a) a lack of secure and efficient access control; (b) inadequate support for flexible and selective content management by content publishers; (c) insufficient implementation of accountability and privilege revocation mechanisms. To handle these challenges, we propose ESAS, the first-ever Efficient Security-enhanced Accountable Access Control Scheme for NDN. Specifically, our ESAS incorporates anonymous authentication using group signatures at network routers to prevent unauthorized access, employs key-aggregation-based access …