Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences Commons™

Open Access. Powered by Scholars. Published by Universities.®

Artificial Intelligence and Robotics

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 2821 - 2850 of 11188

Full-Text Articles in Computer Sciences

Checklist For Reproducibility Of Deep Learning In Medical Imaging, Mana Moassefi, Yashbir Singh, Gian Marco Conte, Bardia Khosravi, Pouria Rouzrokh, Sanaz Vahdati, Nabile Safdar, Linda Moy, Felipe Kitamura, Amilcare Gentili, Paras Lakhani, Nina Kottler, Safwan Halabi, Joseph Yacoub, Yuankai Hou, Khaled Younis, Bradley Erickson, Elizabeth Krupinski, Shahriar Faghani Aug 2024

Checklist For Reproducibility Of Deep Learning In Medical Imaging, Mana Moassefi, Yashbir Singh, Gian Marco Conte, Bardia Khosravi, Pouria Rouzrokh, Sanaz Vahdati, Nabile Safdar, Linda Moy, Felipe Kitamura, Amilcare Gentili, Paras Lakhani, Nina Kottler, Safwan Halabi, Joseph Yacoub, Yuankai Hou, Khaled Younis, Bradley Erickson, Elizabeth Krupinski, Shahriar Faghani

Department of Radiology Faculty Papers

The application of deep learning (DL) in medicine introduces transformative tools with the potential to enhance prognosis, diagnosis, and treatment planning. However, ensuring transparent documentation is essential for researchers to enhance reproducibility and refine techniques. Our study addresses the unique challenges presented by DL in medical imaging by developing a comprehensive checklist using the Delphi method to enhance reproducibility and reliability in this dynamic field. We compiled a preliminary checklist based on a comprehensive review of existing checklists and relevant literature. A panel of 11 experts in medical imaging and DL assessed these items using Likert scales, with two survey …


Enhancing Cybersecurity For Unmanned Systems: A Comprehensive Literature Review, Jonathan Gabriel Mardoyan Aug 2024

Enhancing Cybersecurity For Unmanned Systems: A Comprehensive Literature Review, Jonathan Gabriel Mardoyan

Electronic Theses, Projects, and Dissertations

This culminating experience project addresses the pressing cybersecurity challenges encountered by unmanned autonomous vehicles. The research provides a comprehensive literature review on how hybrid encryption techniques can improve the security of its communication systems. The chosen research questions guiding this study are: (Q1) How can we enhance cybersecurity measures to safeguard the communication and transmission of sensitive data from unmanned systems, thereby preventing unauthorized access by malicious actors? (Q2) How can we ensure the confidentiality and integrity of messages exchanged with unmanned systems to a command-and-control center operating on the tactical edge? (Q3) How can hybrid encryption tackle the consumption …


Advancing Telehealth Through Artificial Intelligence: Incorporating Emotional Intelligence And Addressing Cybersecurity Challenges, Mahima Rajendra Pulgaonkar Aug 2024

Advancing Telehealth Through Artificial Intelligence: Incorporating Emotional Intelligence And Addressing Cybersecurity Challenges, Mahima Rajendra Pulgaonkar

Electronic Theses, Projects, and Dissertations

This culminating experience project explores the integration of Emotional Artificial Intelligence (Emotional AI) into telehealth systems, addressing the dual challenges of enhancing patient care and mitigating cybersecurity risks. The research questions are: (Q1) How can Emotionally Intelligent AI improve telehealth systems' ability to recognize and respond to mental health symptoms? and (Q2) What are the specific cybersecurity challenges associated with AI in telehealth and how can they be mitigated? The findings for each question are: Q1: Emotionally Intelligent AI can significantly enhance telehealth by providing personalized, empathetic interactions that improve patient engagement, adherence to treatment plans, and early detection of …


Querymate: A Custom Llm Powered By Llamacpp, Pegah Khosravi Aug 2024

Querymate: A Custom Llm Powered By Llamacpp, Pegah Khosravi

Open Educational Resources

No abstract provided.


Establishing The Importance Of Co-Creation And Self-Efficacy In Creative Collaboration With Artificial Intelligence, Jack Mcguire, David De Cremer, Tim Van De Cruys Aug 2024

Establishing The Importance Of Co-Creation And Self-Efficacy In Creative Collaboration With Artificial Intelligence, Jack Mcguire, David De Cremer, Tim Van De Cruys

Research Collection Lee Kong Chian School Of Business

The emergence of generative AI technologies has led to an increasing number of people collaborating with AI to produce creative works. Across two experimental studies, in which we carefully designed and programmed state-of-the-art human–AI interfaces, we examine how the design of generative AI systems influences human creativity (poetry writing). First, we find that people were most creative when writing a poem on their own, compared to first receiving a poem generated by an AI system and using sophisticated tools to edit it (Study 1). Following this, we demonstrate that this creativity deficit dissipates when people co-create with—not edit—AI and establish …


We Train Ai, Why Not Humans, Too? An Exploration Of Human-Ai Team Training For Future Workplace Viability, Caitlin M. Lancaster Aug 2024

We Train Ai, Why Not Humans, Too? An Exploration Of Human-Ai Team Training For Future Workplace Viability, Caitlin M. Lancaster

All Dissertations

The integration of Artificial Intelligence (AI) in the workforce is transforming team dynamics, leading to the emergence of Human-AI Teams (HATs). These teams offer opportunities to capitalize on human strengths with AI's prowess, offering significant opportunities for innovation and efficiency. Effective HAT functioning requires aligning human expectations with AI capabilities and bridging knowledge gaps between teammates. Despite this potential, key integration challenges remain, such as developing shared mental models, addressing skill limitations, and overcoming negative AI perceptions. Existing training efforts often apply human-human teaming principles directly to HATs, overlooking AI's role as a teammate and limiting the development of HAT-specific …


Ee-Lce: An Event Extraction Framework Based On Llm-Generated Cot Explanation, Yanhua Yu, Yuanlong Wang, Yunshan Ma, Jie Li, Kangkang Lu, Zhiyong Huang, Tat-Seng Chua Aug 2024

Ee-Lce: An Event Extraction Framework Based On Llm-Generated Cot Explanation, Yanhua Yu, Yuanlong Wang, Yunshan Ma, Jie Li, Kangkang Lu, Zhiyong Huang, Tat-Seng Chua

Research Collection School Of Computing and Information Systems

Generative models have been widely used in event extraction. However, the interpretability of event extraction has not been fully investigated. In this paper, we propose an Event Extraction framework based on LLM-generated CoT Explanation EE-LCE, which can generate chain-of-thought-style (CoT-style) explanations for events. To this end, we provide each sample of event datasets with an explanation of the reasoning process using a large language model (LLM) GPT-3.5, and fine-tune the Flan-T5 lightweight language model (LM) supervised by the augmented dataset, enhancing both interpretability and performance of the event extraction. Moreover, we use a prefix tree (trie) to normalize the decoding …


Tackling Stackelberg Network Interdiction Against A Boundedly Rational Adversary, Tien Mai, Avinandan Bose, Arunesh Sinha, Thanh Nguyen, Ayushman Kumar Singh Aug 2024

Tackling Stackelberg Network Interdiction Against A Boundedly Rational Adversary, Tien Mai, Avinandan Bose, Arunesh Sinha, Thanh Nguyen, Ayushman Kumar Singh

Research Collection School Of Computing and Information Systems

This work studies Stackelberg network interdiction games --- an important class of games in which a defender first allocates (randomized) defense resources to a set of critical nodes on a graph while an adversary chooses its path to attack these nodes accordingly. We consider a boundedly rational adversary in which the adversary's response model is based on a dynamic form of classic logit-based (quantal response) discrete choice models. The resulting optimization is non-convex and additionally, involves complex terms that sum over exponentially many paths. We tackle these computational challenges by presenting new efficient algorithms with solution guarantees. First, we present …


Analyzing Temporal Complex Events With Large Language Models? A Benchmark Towards Temporal, Long Context Understanding, Zhihan Zhang, Yixin Cao, Chenchen Ye, Ma. Yunshan, Lizi Liao, Tat-Seng Chua Aug 2024

Analyzing Temporal Complex Events With Large Language Models? A Benchmark Towards Temporal, Long Context Understanding, Zhihan Zhang, Yixin Cao, Chenchen Ye, Ma. Yunshan, Lizi Liao, Tat-Seng Chua

Research Collection School Of Computing and Information Systems

The digital landscape is rapidly evolving with an ever-increasing volume of online news, emphasizing the need for swift and precise analysis of complex events.We refer to the complex events composed of many news articles over an extended period as Temporal Complex Event (TCE). This paper proposes a novel approach using Large Language Models (LLMs) to systematically extract and analyze the event chain within TCE, characterized by their key points and timestamps. We establish a benchmark, named TCELongBench, to evaluate the proficiency of LLMs in handling temporal dynamics and understanding extensive text. This benchmark encompasses three distinct tasks - reading comprehension, …


A Survey On Neural Question Generation : Methods, Applications, And Prospects, Shasha Guo, Lizi Liao, Cuiping Li, Tat-Seng Chua Aug 2024

A Survey On Neural Question Generation : Methods, Applications, And Prospects, Shasha Guo, Lizi Liao, Cuiping Li, Tat-Seng Chua

Research Collection School Of Computing and Information Systems

In this survey, we present a detailed examination of the advancements in Neural Question Generation (NQG), a field leveraging neural network techniques to generate relevant questions from diverse inputs like knowledge bases, texts, and images. The survey begins with an overview of NQG’s background, encompassing the task’s problem formulation, prevalent benchmark datasets, established evaluation metrics, and notable applications. It then methodically classifies NQG approaches into three predominant categories: structured NQG, which utilizes organized data sources, unstructured NQG, focusing on more loosely structured inputs like texts or visual content, and hybrid NQG, drawing on diverse input modalities. This classification is followed …


Empathyear : An Open-Source Avatar Multimodal Empathetic Chatbot, Hao Fei, Han Zhang, Bin Wang, Lizi Liao, Qian Liu, Erik Cambria Aug 2024

Empathyear : An Open-Source Avatar Multimodal Empathetic Chatbot, Hao Fei, Han Zhang, Bin Wang, Lizi Liao, Qian Liu, Erik Cambria

Research Collection School Of Computing and Information Systems

This paper introduces EmpathyEar, a pioneering open-source, avatar-based multimodal empathetic chatbot, to fill the gap in traditional text-only empathetic response generation (ERG) systems. Leveraging the advancements of a large language model, combined with multimodal encoders and generators, EmpathyEar supports user inputs in any combination of text, sound, and vision, and produces multimodal empathetic responses, offering users, not just textual responses but also digital avatars with talking faces and synchronized speeches. A series of emotion-aware instruction-tuning is performed for comprehensive emotional understanding and generation capabilities. In this way, EmpathyEar provides users with responses that achieve a deeper emotional resonance, closely emulating …


Planning Like Human : A Dual-Process Framework For Dialogue Planning, Tao He, Lizi Liao, Yixin Cao, Yuanxing Liu, Ming Liu, Zerui Chen, Bing Qin Aug 2024

Planning Like Human : A Dual-Process Framework For Dialogue Planning, Tao He, Lizi Liao, Yixin Cao, Yuanxing Liu, Ming Liu, Zerui Chen, Bing Qin

Research Collection School Of Computing and Information Systems

In proactive dialogue, the challenge lies not just in generating responses but in steering conversations toward predetermined goals, a task where Large Language Models (LLMs) typically struggle due to their reactive nature. Traditional approaches to enhance dialogue planning in LLMs, ranging from elaborate prompt engineering to the integration of policy networks, either face efficiency issues or deliver suboptimal performance. Inspired by the dual-process theory in psychology, which identifies two distinct modes of thinking—intuitive (fast) and analytical (slow), we propose the Dual-Process Dialogue Planning (DPDP) framework. DPDP embodies this theory through two complementary planning systems: an instinctive policy model for familiar …


Synergizing Large Language Models And Pre-Trained Smaller Models For Conversational Intent Discovery, Jinggui Liang, Lizi Liao, Hao Fei, Jing Jiang Aug 2024

Synergizing Large Language Models And Pre-Trained Smaller Models For Conversational Intent Discovery, Jinggui Liang, Lizi Liao, Hao Fei, Jing Jiang

Research Collection School Of Computing and Information Systems

In Conversational Intent Discovery (CID), Small Language Models (SLMs) struggle with overfitting to familiar intents and fail to label newly discovered ones. This issue stems from their limited grasp of semantic nuances and their intrinsically discriminative framework. Therefore, we propose Synergizing Large Language Models (LLMs) with pre-trained SLMs for CID (SynCID). It harnesses the profound semantic comprehension of LLMs alongside the operational agility of SLMs. By utilizing LLMs to refine both utterances and existing intent labels, SynCID significantly enhances the semantic depth, subsequently realigning these enriched descriptors within the SLMs’ feature space to correct cluster distortion and promote robust learning …


Unifying Global-Local Representations In Salient Object Detection With Transformers, Sucheng Ren, Nanxuan Zhao, Qiang Wen, Guoqiang Han, Shengfeng He Aug 2024

Unifying Global-Local Representations In Salient Object Detection With Transformers, Sucheng Ren, Nanxuan Zhao, Qiang Wen, Guoqiang Han, Shengfeng He

Research Collection School Of Computing and Information Systems

The fully convolutional network (FCN) has dominated salient object detection for a long period. However, the locality of CNN requires the model deep enough to have a global receptive field and such a deep model always leads to the loss of local details. In this paper, we introduce a new attention-based encoder, vision transformer, into salient object detection to ensure the globalization of the representations from shallow to deep layers. With the global view in very shallow layers, the transformer encoder preserves more local representations to recover the spatial details in final saliency maps. Besides, as each layer can capture …


Speaker Verification In Agent-Generated Conversations, Yizhe Yang, Palakorn Achananuparp, Heyan Huang, Jing Jiang, Ee-Peng Lim Aug 2024

Speaker Verification In Agent-Generated Conversations, Yizhe Yang, Palakorn Achananuparp, Heyan Huang, Jing Jiang, Ee-Peng Lim

Research Collection School Of Computing and Information Systems

The recent success of large language models (LLMs) has attracted widespread interest to develop role-playing conversational agents personalized to the characteristics and styles of different speakers to enhance their abilities to perform both general and special purpose dialogue tasks. However, the ability to personalize the generated utterances to speakers, whether conducted by human or LLM, has not been well studied. To bridge this gap, our study introduces a novel evaluation challenge: speaker verification in agent-generated conversations, which aimed to verify whether two sets of utterances originate from the same speaker. To this end, we assemble a large dataset collection encompassing …


Self-Adaptive Psro : Towards An Automatic Population-Based Game Solver, Pengdeng Li, Shuxin Li, Chang Yang, Xinrun Wang, Xiao Huang, Hau Chan, Bo An Aug 2024

Self-Adaptive Psro : Towards An Automatic Population-Based Game Solver, Pengdeng Li, Shuxin Li, Chang Yang, Xinrun Wang, Xiao Huang, Hau Chan, Bo An

Research Collection School Of Computing and Information Systems

Policy-Space Response Oracles (PSRO) as a general algorithmic framework has achieved state-of-the-art performance in learning equilibrium policies of two-player zero-sum games. However, the hand-crafted hyperparameter value selection in most of the existing works requires extensive domain knowledge, forming the main barrier to applying PSRO to different games. In this work, we make the first attempt to investigate the possibility of self-adaptively determining the optimal hyperparameter values in the PSRO framework. Our contributions are three-fold: (1) Using several hyperparameters, we propose a parametric PSRO that unifies the gradient descent ascent (GDA) and different PSRO variants. (2) We propose the self-adaptive PSRO …


A Multimodal Foundation Agent For Financial Trading : Tool-Augmented, Diversified, And Generalist, Wentao Zhang, Lingxuan Zhao, Haochong Xia, Shuo Sun, Jiaze Sun, Molei Qin, Xinyi Li, Yuqing Zhao, Yilei Zhao, Xinyu Cai, Longtao Zheng, Xinrun Wang, Bo An Aug 2024

A Multimodal Foundation Agent For Financial Trading : Tool-Augmented, Diversified, And Generalist, Wentao Zhang, Lingxuan Zhao, Haochong Xia, Shuo Sun, Jiaze Sun, Molei Qin, Xinyi Li, Yuqing Zhao, Yilei Zhao, Xinyu Cai, Longtao Zheng, Xinrun Wang, Bo An

Research Collection School Of Computing and Information Systems

Financial trading is a crucial component of the markets, informed by a multimodal information landscape encompassing news, prices, and Kline charts, and encompasses diverse tasks such as quantitative trading and high-frequency trading with various assets. While advanced AI techniques like deep learning and reinforcement learning are extensively utilized in finance, their application in financial trading tasks often faces challenges due to inadequate handling of multimodal data and limited generalizability across various tasks. To address these challenges, we present FinAgent, a multimodal foundational agent with tool augmentation for financial trading. FinAgent's market intelligence module processes a diverse range of data-numerical, textual, …


Macrohft : Memory Augmented Context-Aware Reinforcement Learning On High Frequency Trading, Chuqiao Zong, Chaojie Wang, Molei Qin, Lei Feng, Xinrun Wang, Xinrun Wang Aug 2024

Macrohft : Memory Augmented Context-Aware Reinforcement Learning On High Frequency Trading, Chuqiao Zong, Chaojie Wang, Molei Qin, Lei Feng, Xinrun Wang, Xinrun Wang

Research Collection School Of Computing and Information Systems

High-frequency trading (HFT) that executes algorithmic trading in short time scales, has recently occupied the majority of cryptocurrency market. Besides traditional quantitative trading methods, reinforcement learning (RL) has become another appealing approach for HFT due to its terrific ability of handling high-dimensional financial data and solving sophisticated sequential decision-making problems, e.g., hierarchical reinforcement learning (HRL) has shown its promising performance on second-level HFT by training a router to select only one sub-agent from the agent pool to execute the current transaction. However, existing RL methods for HFT still have some defects: 1) standard RL-based trading agents suffer from the overfitting …


Topic Modeling On Document Networks With Dirichlet Optimal Transport Barycenter, Ce Zhang, Hady Wirawan Lauw Aug 2024

Topic Modeling On Document Networks With Dirichlet Optimal Transport Barycenter, Ce Zhang, Hady Wirawan Lauw

Research Collection School Of Computing and Information Systems

Text documents are often interconnected in a network structure, e.g., academic papers via citations, Web pages via hyperlinks. On the one hand, though Graph Neural Networks (GNNs) have shown promising ability to derive effective embeddings for such networked documents, they do not assume a latent topic structure and result in uninterpretable embeddings. On the other hand, topic models can infer semantically interpretable topic distributions for documents by associating each topic with a group of understandable key words. However, most topic models mainly focus on plain text within documents and fail to leverage network structure across documents. Network connectivity reveals topic …


Reinforcement Nash Equilibrium Solver, Xinrun Wang, Chang Yang, Shuxin Li, Pengdeng Li, Xiao Huang, Hau Chan, Bo An Aug 2024

Reinforcement Nash Equilibrium Solver, Xinrun Wang, Chang Yang, Shuxin Li, Pengdeng Li, Xiao Huang, Hau Chan, Bo An

Research Collection School Of Computing and Information Systems

Nash Equilibrium (NE) is the canonical solution concept of game theory, which provides an elegant tool to understand the rationalities. Though mixed strategy NE exists in any game with finite players and actions, computing NE in two- or multi-player general-sum games is PPAD-Complete. Various alternative solutions, e.g., Correlated Equilibrium (CE), and learning methods, e.g., fictitious play (FP), are proposed to approximate NE. For convenience, we call these methods as "inexact solvers", or "solvers" for short. However, the alternative solutions differ from NE and the learning methods generally fail to converge to NE. Therefore, in this work, we propose REinforcement Nash …


Exploring A Multimodal Fusion-Based Deep Learning Network For Detecting Facial Palsy, Heng Yim Nicole Oo, Min Hun Lee, J. H. Lim Aug 2024

Exploring A Multimodal Fusion-Based Deep Learning Network For Detecting Facial Palsy, Heng Yim Nicole Oo, Min Hun Lee, J. H. Lim

Research Collection School Of Computing and Information Systems

Algorithmic detection of facial palsy offers the potential to improve current practices, which usually involve labor-intensive and subjective assessment by clinicians. In this paper, we present a multimodal fusion-based deep learning model that utilizes unstructured data (i.e. an image frame with facial line segments) and structured data (i.e. features of facial expressions) to detect facial palsy. We then contribute to a study to analyze the effect of different data modalities and the benefits of a multimodal fusion-based approach using videos of 21 facial palsy patients. Our experimental results show that among various data modalities (i.e. unstructured data - RGB images …


Towards Gradient-Based Time-Series Explanations Through A Spatiotemporal Attention Network, Min Hun Lee Aug 2024

Towards Gradient-Based Time-Series Explanations Through A Spatiotemporal Attention Network, Min Hun Lee

Research Collection School Of Computing and Information Systems

In this paper, we explore the feasibility of using a transformer-based, spatiotemporal attention network (STAN) for gradient-based time-series explanations. First, we trained the STAN model for video classifications using the global and local views of data and weakly supervised labels on time-series data (i.e. the type of an activity). We then leveraged a gradient-based XAI technique (e.g. saliency map) to identify salient frames of time-series data. According to the experiments using the datasets of four medically relevant activities, the STAN model demonstrated its potential to identify important frames of videos.


Enabling Sustainable Freight Forwarding Network Via Collaborative Games, Pang Jin Tan, Shih-Fen Cheng, Richard Chen Aug 2024

Enabling Sustainable Freight Forwarding Network Via Collaborative Games, Pang Jin Tan, Shih-Fen Cheng, Richard Chen

Research Collection School Of Computing and Information Systems

Freight forwarding plays a crucial role in facilitating global trade and logistics. However, as the freight forwarding market is extremely fragmented, freight forwarders often face the issue of not being able to fill the available shipping capacity. This recurrent issue motivates the creation of various freight forwarding networks that aim at exchanging capacities and demands so that the resource utilization of individual freight forwarders can be maximized. In this paper, we focus on how to design such a collaborative network based on collaborative game theory, with the Shapley value representing a fair scheme for profit sharing. Noting that the exact …


Is Aggregation The Only Choice? Federated Learning Via Layer-Wise Model Recombination, Ming Hu, Zhihao Yue, Xiaofei Xie, Cheng Chen Chen Aug 2024

Is Aggregation The Only Choice? Federated Learning Via Layer-Wise Model Recombination, Ming Hu, Zhihao Yue, Xiaofei Xie, Cheng Chen Chen

Research Collection School Of Computing and Information Systems

Although Federated Learning (FL) enables global model training Xiaofei Xie [email protected] Singapore Management University Singapore, Singapore Xian Wei [email protected] East China Normal University Shanghai, China Mingsong Chen∗ [email protected] East China Normal University Shanghai, China • Computing methodologies → Distributed artificial intelligence. across clients without compromising their raw data, due to the unevenly distributed data among clients, existing Federated Averaging (FedAvg)-based methods suffer from the problem of low inference performance. Specifically, different data distributions among clients lead to various optimization directions of local models. Aggregating local models usually results in a low-generalized global model, which performs worse on most of the …


Contrastive General Graph Matching With Adaptive Augmentation Sampling, Jianyuan Bo, Yuan Fang Aug 2024

Contrastive General Graph Matching With Adaptive Augmentation Sampling, Jianyuan Bo, Yuan Fang

Research Collection School Of Computing and Information Systems

Graph matching has important applications in pattern recognition and beyond. Current approaches predominantly adopt supervised learning, demanding extensive labeled data which can be limited or costly. Meanwhile, self-supervised learning methods for graph matching often require additional side information such as extra categorical information and input features, limiting their application to the general case. Moreover, designing the optimal graph augmentations for self-supervised graph matching presents another challenge to ensure robustness and effcacy. To address these issues, we introduce a novel Graph-centric Contrastive framework for Graph Matching (GCGM), capitalizing on a vast pool of graph augmentations for contrastive learning, yet without needing …


A Learned Generalized Geodesic Distance Function-Based Approach For Node Feature Augmentation On Graphs, Amitoz Azad, Yuan Fang Aug 2024

A Learned Generalized Geodesic Distance Function-Based Approach For Node Feature Augmentation On Graphs, Amitoz Azad, Yuan Fang

Research Collection School Of Computing and Information Systems

Geodesic distances on manifolds have numerous applications in image processing, computer graphics and computer vision. In this work, we introduce an approach called 'LGGD' (Learned Generalized Geodesic Distances). This method involves generating node features by learning a generalized geodesic distance function through a training pipeline that incorporates training data, graph topology and the node content features. The strength of this method lies in the proven robustness of the generalized geodesic distances to noise and outliers. Our contributions encompass improved performance in node classification tasks, competitive results with state-of-the-art methods on real-world graph datasets, the demonstration of the learnability of parameters …


Sibo : A Simple Booster For Parameter-Efficient Fine-Tuning, Zhihao Wen, Jie Zhang, Yuan Fang Aug 2024

Sibo : A Simple Booster For Parameter-Efficient Fine-Tuning, Zhihao Wen, Jie Zhang, Yuan Fang

Research Collection School Of Computing and Information Systems

Fine-tuning all parameters of large language models (LLMs) necessitates substantial computational power and extended time. Latest advancements in parameter-efficient fine-tuning (PEFT) techniques, such as Adapter tuning and LoRA, allow for adjustments to only a minor fraction of the parameters of these LLMs. Concurrently, it has been noted that the issue of over-smoothing diminishes the effectiveness of these Transformer-based LLMs, resulting in suboptimal performances in downstream tasks. In this paper, we present SIBO, which is a SImple BOoster to enhance PEFT, by injecting an initial residual. SIBO is straightforward and readily extensible to a range of state-of-the-art PEFT techniques to alleviate …


Heterogeneous Graph Transformer With Poly-Tokenization, Zhiyuan Lu, Yuan Fang, Cheng Yang, Chuan Shi Aug 2024

Heterogeneous Graph Transformer With Poly-Tokenization, Zhiyuan Lu, Yuan Fang, Cheng Yang, Chuan Shi

Research Collection School Of Computing and Information Systems

Graph neural networks have shown widespread success for learning on graphs, but they still face fundamental drawbacks, such as limited expressive power, over-smoothing, and over-squashing. Meanwhile, the transformer architecture offers a potential solution to these issues. However, existing graph transformers primarily cater to homogeneous graphs and are unable to model the intricate semantics of heterogeneous graphs. Moreover, unlike small molecular graphs where the entire graph can be considered as the receptive field in graph transformers, real-world heterogeneous graphs comprise a significantly larger number of nodes and cannot be entirely treated as such. Consequently, existing graph transformers struggle to capture the …


Cross-Problem Learning For Solving Vehicle Routing Problems, Zhuoyi Lin, Yaoxin Wu, Bangjian Zhou, Zhiguang Cao, Wen Song, Yingqian Zhang, Senthilnath Jayavelu Aug 2024

Cross-Problem Learning For Solving Vehicle Routing Problems, Zhuoyi Lin, Yaoxin Wu, Bangjian Zhou, Zhiguang Cao, Wen Song, Yingqian Zhang, Senthilnath Jayavelu

Research Collection School Of Computing and Information Systems

Existing neural heuristics often train a deep architecture from scratch for each specific vehicle routing problem (VRP), ignoring the transferable knowledge across different VRP variants. This paper proposes the cross-problem learning to assist heuristics training for different downstream VRP variants. Particularly, we modularize neural architectures for complex VRPs into 1) the backbone Transformer for tackling the travelling salesman problem (TSP), and 2) the additional lightweight modules for processing problem-specific features in complex VRPs. Accordingly, we propose to pre-train the backbone Transformer for TSP, and then apply it in the process of fine-tuning the Transformer models for each target VRP variant. …


Certified Policy Verification And Synthesis For Mdps Under Distributional Reach-Avoidance Properties, S. Akshay, Krishnendu Chatterjee, Tobias Meggendorfer, Dorde Zikelic Aug 2024

Certified Policy Verification And Synthesis For Mdps Under Distributional Reach-Avoidance Properties, S. Akshay, Krishnendu Chatterjee, Tobias Meggendorfer, Dorde Zikelic

Research Collection School Of Computing and Information Systems

Markov Decision Processes (MDPs) are a classical model for decision making in the presence of uncertainty. Often they are viewed as state transformers with planning objectives defined with respect to paths over MDP states. An increasingly popular alternative is to view them as distribution transformers, giving rise to a sequence of probability distributions over MDP states. For instance, reachability and safety properties in modeling robot swarms or chemical reaction networks are naturally defined in terms of probability distributions over states. Verifying such distributional properties is known to be hard and often beyond the reach of classical state-based verification techniques. In …