Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences Commons™

Open Access. Powered by Scholars. Published by Universities.®

Research Collection School Of Computing and Information Systems

Discipline
Keyword
Publication Year
File Type

Articles 901 - 930 of 8458

Full-Text Articles in Computer Sciences

On The Lossiness Of 2k-Th Power And The Instantiability Of Rabin-Oaep, Haiyang Xue, Bao Li, Xianhui Lu, Kunpeng Wang, Yamin Liu Oct 2024

On The Lossiness Of 2k-Th Power And The Instantiability Of Rabin-Oaep, Haiyang Xue, Bao Li, Xianhui Lu, Kunpeng Wang, Yamin Liu

Research Collection School Of Computing and Information Systems

Seurin PKC 2014 proposed the 2-ï /4-hiding assumption which asserts the indistinguishability of Blum Numbers from pseudo Blum Numbers. In this paper, we investigate the lossiness of 2 k -th power based on the 2 k -ï /4-hiding assumption, which is an extension of the 2-ï /4-hiding assumption. And we prove that 2 k -th power function is a lossy trapdoor permutation over Quadratic Residuosity group. This new lossy trapdoor function has 2 k -bits lossiness for k -bits exponent, while the RSA lossy trapdoor function given by Kiltz et al. Crypto 2010 has k -bits lossiness for k -bits …


Foss: Towards Fine-Grained Unknown Class Detection Against The Open-Set Attack Spectrum With Variable Legitimate Traffic, Ziming Zhao, Zhaoxuan Li, Xiaofei Xie, Jiongchi Yu, Fan Zhang, Rui Zhang, Binbin Chen, Xiangyang Luo, Ming Hu, Wenrui Ma Oct 2024

Foss: Towards Fine-Grained Unknown Class Detection Against The Open-Set Attack Spectrum With Variable Legitimate Traffic, Ziming Zhao, Zhaoxuan Li, Xiaofei Xie, Jiongchi Yu, Fan Zhang, Rui Zhang, Binbin Chen, Xiangyang Luo, Ming Hu, Wenrui Ma

Research Collection School Of Computing and Information Systems

Anomaly-based network intrusion detection systems (NIDSs) are essential for ensuring cybersecurity. However, the security communities realize some limitations when they put most existing proposals into practice. The challenges are mainly concerned with (i) fine-grained unknown attack detection and (ii) ever-changing legitimate traffic adaptation. To tackle these problem, we present three key design norms. The core idea is to construct a model to split the data distribution hyperplane and leverage the concept of isolation, as well as advance the incremental model update. We utilize the isolation tree as the backbone to design our model, named FOSS, to echo back three norms. …


Large-Scale Graph Label Propagation On Gpus, Chang Ye, Yuchen Li, Bingsheng He, Zhao Li, Jianling Sun Oct 2024

Large-Scale Graph Label Propagation On Gpus, Chang Ye, Yuchen Li, Bingsheng He, Zhao Li, Jianling Sun

Research Collection School Of Computing and Information Systems

Graph label propagation (LP) is a core component in many downstream applications such as fraud detection, recommendation and image segmentation. In this paper, we propose GLP, a GPU-based framework to enable efficient LP processing on large-scale graphs. By investigating the data processing pipeline in a large e-commerce platform, we have identified two key challenges on integrating GPU-accelerated LP processing to the pipeline: (1) programmability for evolving application logics; (2) demand for real-time performance. Motivated by these challenges, we offer a set of expressive APIs that data engineers can customize and deploy efficient LP algorithms on GPUs with ease. To achieve …


Calibrated One-Class Classification For Unsupervised Time Series Anomaly Detection, Hongzuo Xu, Yijie Wang, Songlei Jian, Qing Liao, Yongjun Wang, Guansong Pang Oct 2024

Calibrated One-Class Classification For Unsupervised Time Series Anomaly Detection, Hongzuo Xu, Yijie Wang, Songlei Jian, Qing Liao, Yongjun Wang, Guansong Pang

Research Collection School Of Computing and Information Systems

Time series anomaly detection is instrumental in maintaining system availability in various domains. Current work in this research line mainly focuses on learning data normality deeply and comprehensively by devising advanced neural network structures and new reconstruction/prediction learning objectives. However, their one-class learning process can be misled by latent anomalies in training data (i.e., anomaly contamination) under the unsupervised paradigm. Their learning process also lacks knowledge about the anomalies. Consequently, they often learn a biased, inaccurate normality boundary. To tackle these problems, this paper proposes calibrated one-class classification for anomaly detection, realizing contamination-tolerant, anomaly-informed learning of data normality via uncertainty …


Nigerian Software Engineer Or American Data Scientist? Github Profile Recruitment Bias In Large Language Models, Takashi Nakano, Kazumasa Shimari, Raula Gaikovina Kula, Christoph Treude, Marc Cheong, Kenichi Matsumoto Oct 2024

Nigerian Software Engineer Or American Data Scientist? Github Profile Recruitment Bias In Large Language Models, Takashi Nakano, Kazumasa Shimari, Raula Gaikovina Kula, Christoph Treude, Marc Cheong, Kenichi Matsumoto

Research Collection School Of Computing and Information Systems

Large Language Models (LLMs) have taken the world by storm, demonstrating their ability not only to automate tedious tasks, but also to show some degree of proficiency in completing software engineering tasks. A key concern with LLMs is their “black-box” nature, which obscures their internal workings and could lead to societal biases in their outputs. In the software engineering context, in this early results paper, we empirically explore how well LLMs can automate recruitment tasks for a geographically diverse software team. We use OpenAI's ChatGPT to conduct an initial set of experiments using GitHub User Profiles from four regions to …


Retrofitting A Legacy Cutlery Washing Machine Using Computer Vision, Hua Leong Fwa Oct 2024

Retrofitting A Legacy Cutlery Washing Machine Using Computer Vision, Hua Leong Fwa

Research Collection School Of Computing and Information Systems

Industry 4.0, the digitalization of manufacturing promises to lead to lowered cost, efficient processes and even discovery of new business models. However, many of the enterprises have huge investments in legacy machines which are not 'smart'. In this study, we thus designed a cost-efficient solution to retrofit a legacy conveyor belt-based cutlery washing machine with a commodity web camera. We then applied computer vision (using both traditional image processing and deep learning techniques) to infer the speed and utilization of the machine. We detailed the algorithms that we designed for computing both speed andutilization. With the existing operational constraints of …


An Empirical Study To Evaluate Aigc Detectors On Code Content, Jian Wang, Shangqing Liu, Xiaofei Xie, Yi Li Oct 2024

An Empirical Study To Evaluate Aigc Detectors On Code Content, Jian Wang, Shangqing Liu, Xiaofei Xie, Yi Li

Research Collection School Of Computing and Information Systems

Artificial Intelligence Generated Content (AIGC) has garnered considerable attention for its impressive performance, with Large Language Models (LLMs), like ChatGPT, emerging as a leading AIGC model that produces high-quality responses across various applications, including software development and maintenance. Despite its potential, the misuse of LLMs, especially in security and safetycritical domains, such as academic integrity and answering questions on Stack Overflow, poses significant concerns. Numerous AIGC detectors have been developed and evaluated on natural language data. However, their performance on code-related content generated by LLMs remains unexplored. To fill this gap, in this paper, we present an empirical study evaluating …


Audio Description Customization, Rosiana Natalie, Ruei-Che Chang, Sheshadri Smitha, Anhong Guo, Kotaro Hara Oct 2024

Audio Description Customization, Rosiana Natalie, Ruei-Che Chang, Sheshadri Smitha, Anhong Guo, Kotaro Hara

Research Collection School Of Computing and Information Systems

Blind and low-vision (BLV) people use audio descriptions (ADs) to access videos. However, current ADs are unalterable by end users, thus are incapable of supporting BLV individuals’ potentially diverse needs and preferences. This research investigates if customizing AD could improve how BLV individuals consume videos. We conducted an interview study (Study 1) with fifteen BLV participants, which revealed desires for customizing properties like length, emphasis, speed, voice, format, tone, and language. At the same time, concerns like interruptions and increased interaction load due to customization emerged. To examine AD customization’s effectiveness and tradeoffs, we designed CustomAD, a prototype that enables …


Direct Range Proofs For Paillier Cryptosystem And Their Applications, Zhikang Xie, Mengling Liu, Haiyang Xue, Man Ho Au, Robert H. Deng, Siu-Ming Yiu Oct 2024

Direct Range Proofs For Paillier Cryptosystem And Their Applications, Zhikang Xie, Mengling Liu, Haiyang Xue, Man Ho Au, Robert H. Deng, Siu-Ming Yiu

Research Collection School Of Computing and Information Systems

The Paillier cryptosystem is renowned for its applications in electronic voting, threshold ECDSA, multi-party computation, and more, largely due to its additive homomorphism. In these applications, range proofs for the Paillier cryptosystem are crucial for maintaining security, because of the mismatch between the message space in the Paillier system and the operation space in application scenarios. In this paper, we present novel range proofs for the Paillier cryptosystem, specifically aimed at optimizing those for both Paillier plaintext and affine operation. We interpret encryptions and affine operations as commitments over integers, as opposed to solely over ZN. Consequently, we propose direct …


A Survey Of Protocol Fuzzing, Xiaohan Zhang, Cen Zhang, Xinghua Li, Zhengjie Du, Bing Mao, Yeting Li, Pan Li Oct 2024

A Survey Of Protocol Fuzzing, Xiaohan Zhang, Cen Zhang, Xinghua Li, Zhengjie Du, Bing Mao, Yeting Li, Pan Li

Research Collection School Of Computing and Information Systems

Communication protocols form the bedrock of our interconnected world, yet vulnerabilities within their implementations pose significant security threats. Recent developments have seen a surge in fuzzing-based research dedicated to uncovering these vulnerabilities within protocol implementations. However, there still lacks a systematic overview of protocol fuzzing for answering the essential questions such as what the unique challenges are, how existing works solve them, and so on. To bridge this gap, we conducted a comprehensive investigation of related works from both academia and industry. Our study includes a detailed summary of the specific challenges in protocol fuzzing and provides a systematic categorization …


Self-Supervised Learning For Time Series Analysis : Taxonomy, Progress, And Prospects, Zhang Kexin, Qingsong Wen, Chaoli Zhang, Rongyao Cai, Ming Jin, Yong Liu, James Y. Zhang, Guansong Pang, Guansong Pang, Pan Shirui Oct 2024

Self-Supervised Learning For Time Series Analysis : Taxonomy, Progress, And Prospects, Zhang Kexin, Qingsong Wen, Chaoli Zhang, Rongyao Cai, Ming Jin, Yong Liu, James Y. Zhang, Guansong Pang, Guansong Pang, Pan Shirui

Research Collection School Of Computing and Information Systems

Self-supervised learning (SSL) has recently achieved impressive performance on various time series tasks. The most prominent advantage of SSL is that it reduces the dependence on labeled data. Based on the pre-training and fine-tuning strategy, even a small amount of labeled data can achieve high performance. Compared with many published self-supervised surveys on computer vision and natural language processing, a comprehensive survey for time series SSL is still missing. To fill this gap, we review current state-of-the-art SSL methods for time series data in this article. To this end, we first comprehensively review existing surveys related to SSL and time …


Demystifying And Extracting Fault-Indicating Information From Logs For Failure Diagnosis, Junjie Huang, Zhihan Jiang, Jinyang Liu, Yintong Huo, Jiazhen Gu, Zhuangbin Chen, Cong Feng, Hui Dong, Zengyin Yang, Michael R. Lyu Oct 2024

Demystifying And Extracting Fault-Indicating Information From Logs For Failure Diagnosis, Junjie Huang, Zhihan Jiang, Jinyang Liu, Yintong Huo, Jiazhen Gu, Zhuangbin Chen, Cong Feng, Hui Dong, Zengyin Yang, Michael R. Lyu

Research Collection School Of Computing and Information Systems

Logs are imperative in the maintenance of online service systems, which often encompass important information for effective failure mitigation. While existing anomaly detection methodologies facilitate the identification of anomalous logs within extensive runtime data, manual investigation of log messages by engineers remains essential to comprehend faults, which is labor-intensive and error-prone. Upon examining the log-based troubleshooting practices at CloudA 1, we find that engineers typically prioritize two categories of log information for diagnosis. These include fault-indicating descriptions, which record abnormal system events, and fault-indicating parameters, which specify the associated entities. Motivated by this finding, we propose an approach to automatically …


An End-To-End Bi-Objective Approach To Deep Graph Partitioning, Pengcheng Wei, Yuan Fang, Zhihao Wen, Zheng Xiao, Binbin Chen Oct 2024

An End-To-End Bi-Objective Approach To Deep Graph Partitioning, Pengcheng Wei, Yuan Fang, Zhihao Wen, Zheng Xiao, Binbin Chen

Research Collection School Of Computing and Information Systems

Graphs are ubiquitous in real-world applications, such as computation graphs and social networks. Partitioning large graphs into smaller, balanced partitions is often essential, with the biobjective graph partitioning problem aiming to minimize both the“cut” across partitions and the imbalance in partition sizes. However, existing heuristic methods face scalability challenges or overlook partition balance, leading to suboptimal results. Recent deep learning approaches, while promising, typically focus only on node-level features and lack a truly end-to-end framework, resulting in limited performance. In this paper, we introduce a novel method based on graph neural networks (GNNs) that leverages multilevel graph features and addresses …


Kpiroot: Efficient Monitoring Metric-Based Root Cause Localization In Large-Scale Cloud Systems, Wenwei Gu, Xinying Sun, Jinyang Liu, Yintong Huo, Zhuangbin Chen, Jianping Zhang, Jiazhen Gu, Yongqiang Yang, Michael R. Lyu Oct 2024

Kpiroot: Efficient Monitoring Metric-Based Root Cause Localization In Large-Scale Cloud Systems, Wenwei Gu, Xinying Sun, Jinyang Liu, Yintong Huo, Zhuangbin Chen, Jianping Zhang, Jiazhen Gu, Yongqiang Yang, Michael R. Lyu

Research Collection School Of Computing and Information Systems

To ensure the reliability of cloud systems, their run-time status reflecting the service quality is periodically monitored with monitoring metrics, i.e., KPIs (key performance indicators). When performance issues happen, root cause localization pinpoints the specific KPIs that are responsible for the degradation of overall service quality, facilitating prompt problem diagnosis and resolution. To this end, existing methods generally locate root-cause KPIs by identifying the KPIs that exhibit a similar anomalous trend to the overall service performance. While straightforward, solely relying on the similarity calculation may be ineffective when dealing with cloud systems with complicated interdependent services. Recent deep learning-based methods …


Transformer-Based Joint Learning Approach For Text Normalization In Vietnamese Automatic Speech Recognition Systems, The Viet Bui, Tho Chi Luong, Oanh Thi Tran Oct 2024

Transformer-Based Joint Learning Approach For Text Normalization In Vietnamese Automatic Speech Recognition Systems, The Viet Bui, Tho Chi Luong, Oanh Thi Tran

Research Collection School Of Computing and Information Systems

In this article, we investigate the task of normalizing transcribed texts in Vietnamese Automatic Speech Recognition (ASR) systems in order to improve user readability and the performance of downstream tasks. This task usually consists of two main sub-tasks: predicting and inserting punctuation (i.e., period, comma); and detecting and standardizing named entities (i.e., numbers, person names) from spoken forms to their appropriate written forms. To achieve these goals, we introduce a complete corpus including of 87,700 sentences and investigate conditional joint learning approaches which globally optimize two sub-tasks simultaneously. The experimental results are quite promising. Overall, the proposed architecture outperformed the …


Data Provenance Via Differential Auditing, Xin Mu, Ming Pang, Feida Zhu Oct 2024

Data Provenance Via Differential Auditing, Xin Mu, Ming Pang, Feida Zhu

Research Collection School Of Computing and Information Systems

With the rising awareness of data assets, data governance, which is to understand where data comes from, how it is collected, and how it is used, has been assuming evergrowing importance. One critical component of data governance gaining increasing attention is auditing machine learning models to determine if specific data has been used for training. Existing auditing techniques, like shadow auditing methods, have shown feasibility under specific conditions such as having access to label information and knowledge of training protocols. However, these conditions are often not met in most real-world applications. In this paper, we introduce a practical framework for …


Hisoma: A Hierarchical Multi-Agent Model Integrating Self-Organizing Neural Networks With Multi-Agent Deep Reinforcement Learning, Minghong Geng, Shubham Pateria, Budhitama Subagdja, Ah-Hwee Tan Oct 2024

Hisoma: A Hierarchical Multi-Agent Model Integrating Self-Organizing Neural Networks With Multi-Agent Deep Reinforcement Learning, Minghong Geng, Shubham Pateria, Budhitama Subagdja, Ah-Hwee Tan

Research Collection School Of Computing and Information Systems

Multi-agent deep reinforcement learning (MADRL) has shown remarkable advancements in the past decade. However, most current MADRL models focus on task-specific short-horizon problems involving a small number of agents, limiting their applicability to long-horizon planning in complex environments. Hierarchical multi-agent models offer a promising solution by organizing agents into different levels, effectively addressing tasks with varying planning horizons. However, these models often face constraints related to the number of agents or levels of hierarchies. This paper introduces HiSOMA, a novel hierarchical multi-agent model designed to handle long-horizon, multi-agent, multi-task decision-making problems. The top-level controller, FALCON, is modeled as a class …


Does Ceo Agreeableness Personality Mitigate Real Earnings Management?, Shan Liu, Xingying Wu, Nan Hu Oct 2024

Does Ceo Agreeableness Personality Mitigate Real Earnings Management?, Shan Liu, Xingying Wu, Nan Hu

Research Collection School Of Computing and Information Systems

Despite efforts to mitigate aggressive financial reporting, earnings management remains challenging to parties interested in inhibiting its dysfunctional effects. Using linguistic algorithms to assess CEO agreeableness personality from their unscripted texts in conference calls, we find that it is a determinant that mitigates a firm's real earnings management. Furthermore, such an effect is more pronounced when firms confront intensive market competition and financial distress and have weaker managerial entrenchment or when CEOs face stronger internal governance. Our findings persist even after we utilize several alternative real earnings management metrics and control other confounding personalities in prior earnings management studies. The …


D2sr: Decentralized Detection, De-Synchronization, And Recovery Of Lidar Interference, Darshana Rathnayake, Hemanth Sabbella, Meera Radhakrishnan, Archan Misra Oct 2024

D2sr: Decentralized Detection, De-Synchronization, And Recovery Of Lidar Interference, Darshana Rathnayake, Hemanth Sabbella, Meera Radhakrishnan, Archan Misra

Research Collection School Of Computing and Information Systems

We address the challenge of multi-LiDAR interference, an issue of growing importance as LiDAR sensors are embedded in a growing set of pervasive devices. We introduce a novel approach named D2SR, enabling decentralized interference detection, mitigation, and recovery without explicit coordination among nearby LiDAR devices. D2SR comprises three stages: (a) Detection, which identifies interfered frames, (b) Mitigation, which performs time-shifting of a LiDAR’s active period to reduce interference, and (c) Recovery, which corrects or reconstructs the depth values in interfered regions of a depth frame. Key contributions include a lightweight interference detection algorithm achieving an F1-score of 92%, a simple …


Motif Graph Neural Network, Xuexin Chen, Ruicui Cai, Yuan Fang, Min Wu, Zijian Li, Zhifeng Hao Oct 2024

Motif Graph Neural Network, Xuexin Chen, Ruicui Cai, Yuan Fang, Min Wu, Zijian Li, Zhifeng Hao

Research Collection School Of Computing and Information Systems

Graphs can model complicated interactions between entities, which naturally emerge in many important applications. These applications can often be cast into standard graph learning tasks, in which a crucial step is to learn low-dimensional graph representations. Graph neural networks (GNNs) are currently the most popular model in graph embedding approaches. However, standard GNNs in the neighborhood aggregation paradigm suffer from limited discriminative power in distinguishing high-order graph structures as opposed to low-order structures. To capture high-order structures, researchers have resorted to motifs and developed motif-based GNNs. However, the existing motif-based GNNs still often suffer from less discriminative power on high-order …


Predicting The Limits: Tailoring Unnoticeable Hand Redirection Offsets In Virtual Reality To Individuals' Perceptual Boundaries, Martin Feick, Kora Persephone Regitz, Lukas Gehrke, André Zenner, Anthony Tang, Tobias Patrick Jungbluth, Maurice Rekrut, Antonio Krüger Oct 2024

Predicting The Limits: Tailoring Unnoticeable Hand Redirection Offsets In Virtual Reality To Individuals' Perceptual Boundaries, Martin Feick, Kora Persephone Regitz, Lukas Gehrke, André Zenner, Anthony Tang, Tobias Patrick Jungbluth, Maurice Rekrut, Antonio Krüger

Research Collection School Of Computing and Information Systems

Many illusion and interaction techniques in Virtual Reality (VR) rely on Hand Redirection (HR), which has proved to be effective as long as the introduced offsets between the position of the real and virtual hand do not noticeably disturb the user experience. Yet calibrating HR offsets is a tedious and time-consuming process involving psychophysical experimentation, and the resulting thresholds are known to be affected by many variables—limiting HR’s practical utility. As a result, there is a clear need for alternative methods that allow tailoring HR to the perceptual boundaries of individual users. We conducted an experiment with 18 participants combining …


Joint Weakly Supervised Image Emotion Analysis Based On Interclass Discrimination And Intraclass Correlation, Xinyue Zhang, Zhaoxia Wang, Guitao Cao, Seng-Beng Ho Oct 2024

Joint Weakly Supervised Image Emotion Analysis Based On Interclass Discrimination And Intraclass Correlation, Xinyue Zhang, Zhaoxia Wang, Guitao Cao, Seng-Beng Ho

Research Collection School Of Computing and Information Systems

Regional information-based image emotion analysis has recently garnered significant attention. However, existing methods often focus on identifying region proposals through layered steps or merely rely on visual saliency. These approaches may lead to an underestimation of emotional categories and a lack of comprehensive interclass discrimination perception and emotional intraclass contextual mining. To address these limitations, we propose a novel approach named InterIntraIEA, which combines interclass discrimination and intraclass correlation joint learning capabilities for image emotion analysis. The proposed method not only employs category-specific dictionary learning for class adaptation, but also models intraclass contextual relationships and perceives correlations at the channel …


Exploring Conversations Between A Practitioner And A Person With Dementia, Kotaro Hara, Rosiana Natalie, Wei Soon Cheong, Jingjing Gu, Qianli Xu Oct 2024

Exploring Conversations Between A Practitioner And A Person With Dementia, Kotaro Hara, Rosiana Natalie, Wei Soon Cheong, Jingjing Gu, Qianli Xu

Research Collection School Of Computing and Information Systems

In social service centers, practitioners engage in conversations with clients with dementia to facilitate their daily activities and provide support when they are distressed. However, the nature of the care demands the practitioner’s active engagement, which becomes difficult to deliver as the number of people who need care expands. Researchers have been investigating the efficacy of developing agents that assume conversational tasks to alleviate this work. To contribute to the future design of agents for caregiving, we collected and analyzed ten conversations between clients with mild dementia and practitioners who provide care. Our analyses of turn-taking dynamics and dialogue acts …


Temporal Relational Graph Convolutional Network Approach To Financial Performance Prediction, Jeyaraman Brindha Priyadarshini, Bing Tian Dai, Yuan Fang Oct 2024

Temporal Relational Graph Convolutional Network Approach To Financial Performance Prediction, Jeyaraman Brindha Priyadarshini, Bing Tian Dai, Yuan Fang

Research Collection School Of Computing and Information Systems

Accurately predicting financial entity performance remains a challenge due to the dynamic nature of financial markets and vast unstructured textual data. Financial knowledge graphs (FKGs) offer a structured representation for tackling this problem by representing complex financial relationships and concepts. However, constructing a comprehensive and accurate financial knowledge graph that captures the temporal dynamics of financial entities is non-trivial. We introduce FintechKG, a comprehensive financial knowledge graph developed through a three-dimensional information extraction process that incorporates commercial entities and temporal dimensions and uses a financial concept taxonomy that ensures financial domain entity and relationship extraction. We propose a temporal and …


Efficient Cascaded Multiscale Adaptive Network For Image Restoration, Yichen Zhou, Pan Zhou, Teck Khim Ng Oct 2024

Efficient Cascaded Multiscale Adaptive Network For Image Restoration, Yichen Zhou, Pan Zhou, Teck Khim Ng

Research Collection School Of Computing and Information Systems

Image restoration, encompassing tasks such as deblurring, denoising, and super-resolution, remains a pivotal area in computer vision. However, efficiently addressing the spatially varying artifacts of various low-quality images with local adaptiveness and handling their degradations at different scales poses significant challenges. To efficiently tackle these issues, we propose the novel Efficient Cascaded Multiscale Adaptive (ECMA) Network. ECMA employs Local Adaptive Module, LAM, which dynamically adjusts convolution kernels across local image regions to efficiently handle varying artifacts. Thus, LAM addresses the local adaptiveness challenge more efficiently than costlier mechanisms like self-attention, due to its less computationally intensive convolutions. To construct a …


Self-Adaptive Fine-Grained Multi-Modal Data Augmentation For Semi-Supervised Muti-Modal Coreference Resolution, Li Zheng, Boyu Chen, Hao Fei, Fei Li, Shengqiong Wu, Lizi Liao, Donghong Ji Oct 2024

Self-Adaptive Fine-Grained Multi-Modal Data Augmentation For Semi-Supervised Muti-Modal Coreference Resolution, Li Zheng, Boyu Chen, Hao Fei, Fei Li, Shengqiong Wu, Lizi Liao, Donghong Ji

Research Collection School Of Computing and Information Systems

Coreference resolution, an essential task in natural language processing, is particularly challenging in multi-modal scenarios where data comes in various forms and modalities. Despite advancements, limitations due to scarce labeled data and underleveraged unlabeled data persist. We address these issues with a self-adaptive fine-grained multi-modal data augmentation framework for semi-supervised MCR, focusing on enriching training data from labeled datasets and tapping into the untapped potential of unlabeled data. Regarding the former issue, we first leverage text coreference resolution datasets and diffusion models,to perform fine-grained text-to-image generation with aligned text entities and image bounding boxes. We then introduce a self-adaptive selection …


Enhancing Recipe Retrieval With Foundation Models: A Data Augmentation Perspective, Fangzhou Song, Bin Zhu, Yanbin Hao, Shuo Wang Oct 2024

Enhancing Recipe Retrieval With Foundation Models: A Data Augmentation Perspective, Fangzhou Song, Bin Zhu, Yanbin Hao, Shuo Wang

Research Collection School Of Computing and Information Systems

Learning recipe and food image representation in common embedding space is non-trivial but crucial for cross-modal recipe retrieval. In this paper, we propose a new perspective for this problem by utilizing foundation models for data augmentation. Leveraging on the remarkable capabilities of foundation models (i.e., Llama2 and SAM), we propose to augment recipe and food image by extracting alignable information related to the counterpart. Specifically, Llama2 is employed to generate a textual description from the recipe, aiming to capture the visual cues of a food image, and SAM is used to produce image segments that correspond to key ingredients in …


Risurconv : Rotation Invariant Surface Attention-Augmented Convolutions For 3d Point Cloud Classification And Segmentation, Zhiyuan Zhang, Licheng Yang, Xiang Zhiyu Oct 2024

Risurconv : Rotation Invariant Surface Attention-Augmented Convolutions For 3d Point Cloud Classification And Segmentation, Zhiyuan Zhang, Licheng Yang, Xiang Zhiyu

Research Collection School Of Computing and Information Systems

Despite the progress on 3D point cloud deep learning, most prior works focus on learning features that are invariant to translation and point permutation, and very limited efforts have been devoted for rotation invariant property. Several recent studies achieve rotation invariance at the cost of lower accuracies. In this work, we close this gap by proposing a novel yet effective rotation invariant architecture for 3D point cloud classification and segmentation. Instead of traditional pointwise operations, we construct local triangle surfaces to capture more detailed surface structure, based on which we can extract highly expressive rotation invariant surface properties which are …


Desk2desk : Optimization-Based Mixed Reality Workspace Integration For Remote Side-By-Side Collaboration, Ludwig Sidenmark, Tianyu Zhang, Leen Al Lababidi, Jiannan Li, Tovi Grossman Oct 2024

Desk2desk : Optimization-Based Mixed Reality Workspace Integration For Remote Side-By-Side Collaboration, Ludwig Sidenmark, Tianyu Zhang, Leen Al Lababidi, Jiannan Li, Tovi Grossman

Research Collection School Of Computing and Information Systems

Mixed Reality enables hybrid workspaces where physical and virtual monitors are adaptively created and moved to suit the current environment and needs. However, in shared settings, individual users’ workspaces are rarely aligned and can vary significantly in the number of monitors, available physical space, and workspace layout, creating inconsistencies between workspaces which may cause confusion and reduce collaboration. We present Desk2Desk, an optimization-based approach for remote collaboration in which the hybrid workspaces of two collaborators are fully integrated to enable immersive side-by-side collaboration. The optimization adjusts each user’s workspace in layout and number of shared monitors and creates a mapping …


Improving Out-Of-Distribution Detection With Disentangled Foreground And Background Features, Choubo Ding, Guansong Pang Oct 2024

Improving Out-Of-Distribution Detection With Disentangled Foreground And Background Features, Choubo Ding, Guansong Pang

Research Collection School Of Computing and Information Systems

Detecting out-of-distribution (OOD) inputs is a principal task for ensuring the safety of deploying deep-neural-network classifiers in open-set scenarios. OOD samples can be drawn from arbitrary distributions and exhibit deviations from in-distribution (ID) data in various dimensions, such as foreground features (e.g., objects in CIFAR100 images vs. those in CIFAR10 images) and background features (e.g., textural images vs. objects in CIFAR10). Existing methods can confound foreground and background features in training, failing to utilize the background features for OOD detection. This paper considers the importance of feature disentanglement in out-of-distribution detection and proposes the simultaneous exploitation of both foreground and …