Open Access. Powered by Scholars. Published by Universities.®
Artificial Intelligence and Robotics Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Engineering (5395)
- Computer Engineering (4371)
- Operations Research, Systems Engineering and Industrial Engineering (4244)
- Numerical Analysis and Scientific Computing (4163)
- Systems Science (3895)
-
- Social and Behavioral Sciences (984)
- Databases and Information Systems (629)
- Medicine and Health Sciences (606)
- Data Science (528)
- Theory and Algorithms (487)
- Business (447)
- Graphics and Human Computer Interfaces (414)
- Electrical and Computer Engineering (403)
- Education (383)
- Software Engineering (374)
- Arts and Humanities (337)
- Public Affairs, Public Policy and Public Administration (298)
- Information Security (293)
- Other Computer Sciences (275)
- Life Sciences (270)
- Law (241)
- Statistics and Probability (204)
- Medical Specialties (180)
- Library and Information Science (165)
- Psychology (157)
- Robotics (157)
- Programming Languages and Compilers (155)
- Institution
-
- China Simulation Federation (3880)
- Singapore Management University (1897)
- Old Dominion University (642)
- San Jose State University (277)
- MBZUAI (233)
-
- City University of New York (CUNY) (184)
- Technological University Dublin (157)
- Air Force Institute of Technology (137)
- Chapman University (125)
- California Polytechnic State University, San Luis Obispo (116)
- Chinese Academy of Sciences (113)
- University of Arkansas, Fayetteville (103)
- Lindenwood University (97)
- Edith Cowan University (92)
- Embry-Riddle Aeronautical University (92)
- University of Nebraska - Lincoln (78)
- University of Kentucky (76)
- University of South Florida (71)
- Clemson University (63)
- University of Nevada, Las Vegas (63)
- Dartmouth College (62)
- University of Denver (59)
- University of Michigan Law School (57)
- Utah State University (57)
- The Texas Medical Center Library (54)
- Thomas Jefferson University (54)
- New Jersey Institute of Technology (53)
- University of Malaya (50)
- Purdue University (48)
- Missouri University of Science and Technology (47)
- Keyword
-
- Artificial intelligence (779)
- Machine learning (685)
- Deep learning (437)
- Artificial Intelligence (359)
- Machine Learning (359)
-
- AI (239)
- Deep Learning (202)
- Simulation (160)
- Computer vision (158)
- Reinforcement learning (140)
- Generative AI (135)
- Neural networks (129)
- Large language models (108)
- Natural language processing (108)
- Robotics (97)
- Natural Language Processing (91)
- ChatGPT (89)
- Path planning (89)
- Optimization (82)
- Large Language Models (78)
- Computer Vision (76)
- Classification (71)
- Neural network (67)
- Neural Networks (65)
- Virtual reality (64)
- Reinforcement Learning (63)
- Computer Science (59)
- Cybersecurity (59)
- Genetic algorithm (58)
- Algorithms (57)
- Publication Year
- Publication
-
- Journal of System Simulation (3880)
- Research Collection School Of Computing and Information Systems (1664)
- Master's Projects (248)
- Theses and Dissertations (183)
- Computer Science Faculty Publications (126)
-
- Bulletin of Chinese Academy of Sciences (Chinese Version) (113)
- Faculty Scholarship (108)
- Publications and Research (99)
- Computer Vision Faculty Publications (98)
- Master's Theses (96)
- Conference papers (92)
- Electrical & Computer Engineering Faculty Publications (90)
- Machine Learning Faculty Publications (86)
- Electronic Theses and Dissertations (85)
- Faculty Publications (77)
- Dissertations (70)
- Research outputs 2022 to 2026 (64)
- USF Tampa Graduate Theses and Dissertations (59)
- Dissertations and Theses Collection (Open Access) (57)
- Articles (54)
- Dissertations, Theses, and Capstone Projects (53)
- Theses and Dissertations--Computer Science (48)
- Natural Language Processing Faculty Publications (46)
- Teaching and Generative AI: Pedagogical Possibilities and Productive Tensions (46)
- Graduate Theses and Dissertations (45)
- Open Access Theses & Dissertations (42)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (40)
- Theses (40)
- Electrical & Computer Engineering Theses & Dissertations (39)
- Publications (39)
- Publication Type
- File Type
Articles 1141 - 1170 of 11186
Full-Text Articles in Artificial Intelligence and Robotics
Does Generative Ai Facilitate Investor Trading? Early Evidence From Chatgpt Outages, Qiang Cheng, Pengkai Lin, Yue Zhao
Does Generative Ai Facilitate Investor Trading? Early Evidence From Chatgpt Outages, Qiang Cheng, Pengkai Lin, Yue Zhao
Research Collection School Of Accountancy
In this paper, we use ChatGPT outages to provide early evidence on whether investors rely on generative artificial intelligence (GenAI) to perform professional tasks and the associated impact on stock price informativeness. We document a significant decline in stock trading volume during ChatGPT outages. The effect is stronger for firms with corporate news released immediately before or during the outages and for firms with higher ownership held by transient institutional investors. We then document declines in short-run price impact and return variance during the outage periods, consistent with reduced informed trading. Lastly, we document a positive effect of GenAI-assisted trading …
Graph Perturbations For Robust Knowledge Discovery And Retrieval, Hanhua Xiao
Graph Perturbations For Robust Knowledge Discovery And Retrieval, Hanhua Xiao
Dissertations and Theses Collection (Open Access)
Graph perturbation, rooted in classical perturbation theory, studies how small topology edits, i.e., adding or deleting edges, affects graph properties (e.g., density, centrality). This fundamental problem underpins applications like bioinformatics, privacy preservation and system defense. While much prior work targets perturbations that influence global graph statistics or model outputs, comparatively little addresses robustness for knowledge discovery and information retrieval. In these settings, graphs are attributed: nodes carry real-world semantics (e.g., locations, people) and edges encode interactions or relationships. This thesis proposes new formulations and algorithms that generate and leverage graph perturbations to make knowledge discovery and retrieval more robust. Specifically, …
Enhancing Multi-View, Multi-Modal Sensing, Perception And Actuation For Edge Intelligence, Dhanuja Tharith Wanniarachchige
Enhancing Multi-View, Multi-Modal Sensing, Perception And Actuation For Edge Intelligence, Dhanuja Tharith Wanniarachchige
Dissertations and Theses Collection (Open Access)
Artificial Intelligence of Things (AIoT) technologies have ushered in exciting new advances in intelligent sensing, perception, and actuation for many real-world cyberphysical systems (CPS) applications. These technologies have had a formidable impact in domains such as large-scale video surveillance, autonomous transportation and robotics, precision healthcare, and industrial automation. In these applications, sensors and actuators are often collocated with processing nodes, and such nodes are typically interconnected via wireless networks. Vision-based machine intelligence, exemplified by tasks such as object detection, object tracking, and activity analysis, is a very common enabler of such CPS applications. Efficient execution of Deep Neural Network (DNN) …
Scaling Up Cooperative Multi-Agent Reinforcement Learning, Minghong Geng
Scaling Up Cooperative Multi-Agent Reinforcement Learning, Minghong Geng
Dissertations and Theses Collection (Open Access)
Multi-agent systems (MAS) involve multiple autonomous agents that coordinate their actions to achieve shared or competing objectives in dynamic environments. Over the past decade, multi-agent reinforcement learning (MARL) has emerged as a powerful paradigm for enabling collaborative behaviors among autonomous agents within MAS to solve complex tasks. This dissertation discusses a critical scalability gap that exists between current MARL capabilities and real-world deployment requirements. Most existing MARL research focuses on small-scale laboratory problems, often struggling to coordinate large agent populations and facing challenges with extended decision-making horizons. In contrast, many real-world applications demand coordination among hundreds or thousands of agents …
Attorneys And Ai: How Lawyers Use Artificial Intelligence And Analyze Its Impacts, Matthew I. Hall, Christian Turner, Eddie A. Gomez Schieber, Nathaniel Kite, Ari Schlesinger
Attorneys And Ai: How Lawyers Use Artificial Intelligence And Analyze Its Impacts, Matthew I. Hall, Christian Turner, Eddie A. Gomez Schieber, Nathaniel Kite, Ari Schlesinger
Scholarly Works
AI systems are testing lawyers' professional ethics obligations of competence, confidentiality, and candor. In the legal profession, the widespread availability of AI systems presents opportunities, like improving the review of documents during the discovery stage of a lawsuit, and challenges, illustrated by the handful of high-profile incidents where lawyers submitted legal briefs in court citing and describing fictitious cases based on AI-generated output. We conducted interviews with 44 legal professionals in the U.S. to understand how attorneys are making sense of AI technology and the impacts these technologies are having on their profession, legal ethics, and legal institutions. We describe …
How Behavioral Science Can Improve The Return On Ai Investments, David De Cremer, Shane Schweitzer, Jack Mcguire, Devesh Narayanan
How Behavioral Science Can Improve The Return On Ai Investments, David De Cremer, Shane Schweitzer, Jack Mcguire, Devesh Narayanan
Research Collection Lee Kong Chian School Of Business
Many AI projects fail because leaders treat adoption as a tech purchase instead of a behavioral change problem. People resist tools that disrupt routines, overreact to visible AI errors, and prefer familiar human judgment. As a result, even good systems fail to gain purchase. Leaders can address this problem by applying “Behavioral Human-Centered AI” across the AI adoption cycle. In the design phrase, companies should co-design with diverse users, add purposeful friction where it improves scrutiny, require beta tests with subgroup results and behavioral input. During adoption, they should frame AI as an augmenter, disclose limits and safeguards, use explainability …
Building Confidence For Class Participation, Tamas Makany, Ivy Seow
Building Confidence For Class Participation, Tamas Makany, Ivy Seow
Research Collection Lee Kong Chian School Of Business
What happens to students’ critical thinking when half the class filters their thoughts through AI? During a recent debate on AI policy in education, one student mentioned they routinely run their ideas through ChatGPT before speaking up. When I asked who else did the same, more than half the class raised their hands.
Spatio-Temporal Gcn With Softmax Classifier For Skeleton-Based Human Action Recognition, Kabul Khudaybergenov, Avazjon Marakhimov, Zahriddin Muminov
Spatio-Temporal Gcn With Softmax Classifier For Skeleton-Based Human Action Recognition, Kabul Khudaybergenov, Avazjon Marakhimov, Zahriddin Muminov
Chemical Technology, Control and Management
Skeleton-based human action recognition is an important research area with many practical applications. Most existing methods rely on single representations of skeletal sequences, which cannot totally obtain all the complex features of human movements. This paper presents LFHAR (Latent Features for Human Action Recognition), a new framework that uses multiple spatio-temporal latent representations to improve the extraction of action features. Our method captures how skeletal poses change over time and combines motion information from both individual joints and connected body parts. The proposed approach applies graph-based processing to each skeleton frame in a sequence, then arranges the resulting graph features …
A Fake Friend? Ai Companions Are Exactly That, Seow Hon Tan
A Fake Friend? Ai Companions Are Exactly That, Seow Hon Tan
Research Collection Yong Pung How School Of Law
In a commentary, SMU Associate Professor of Law Tan Seow Hon discussed how AI companions, which promise emotionally intelligent companionship, have blurred the line between human and machine relationships by mimicking empathy, memory, and affection. She suggested that while such technologies may ease loneliness, they risk fostering narcissism, diminishing real human connection, and replacing authentic friendship with comforting illusions that erode the capacity for love and community.
Instructors’ Strategies In Creating And Implementing Constructivist Llm-Based Learning Activities, Emily Aurelia, Shun Yi Yeo, Michelle Lui, Effie Lai-Chong Law, Anthony Tang
Instructors’ Strategies In Creating And Implementing Constructivist Llm-Based Learning Activities, Emily Aurelia, Shun Yi Yeo, Michelle Lui, Effie Lai-Chong Law, Anthony Tang
Research Collection School Of Computing and Information Systems
Large language models (LLMs) are increasingly being integrated into educational settings, enabling more adoption of constructivist teaching and learning approaches in classrooms. This paper explores the strategies instructors are currently using to incorporate LLMs into learning activities that align with constructivist principles, which emphasize that learners actively construct their own knowledge. Through interviews with nine instructors who have designed eleven distinct LLM-based activities and using reflexive thematic analysis, this study identifies various types of learning activities with respect to four different aspects of the constructivist learning theory. The strategies employed and challenges faced to foster constructivist student-LLM interaction were also …
International Workshop On Multimodal Generative Search And Recommendation (Mmgensr@Cikm 2025), Yi Bin, Haoxuan Li, Haokai Ma, Yang Zhang, Wenjie Wang, Yunshan Ma, Yang Yang, Tat‑Seng Chua
International Workshop On Multimodal Generative Search And Recommendation (Mmgensr@Cikm 2025), Yi Bin, Haoxuan Li, Haokai Ma, Yang Zhang, Wenjie Wang, Yunshan Ma, Yang Yang, Tat‑Seng Chua
Research Collection School Of Computing and Information Systems
Recent breakthroughs in generative Artificial Intelligence (AI) have ignited a revolutionary wave across information retrieval and recommender systems. This workshop serves as a premier interdisciplinary platform to explore how generative models, particularly Large Language Models (LLMs) and Large Multimodal Models (LMMs), are transforming multimodal search and recommendation paradigms [3, 6, 9, 10, 12-14]. We aim to convene researchers and practitioners to discuss innovative architectures, methodologies, and evaluation strategies spanning generative document retrieval [5, 8] generative image retrieval [ 7, 16], grounded answer generation [17], generative recommendation [2, 4, 11], and related tasks involving multiple modalities [1,15]. The workshop will facilitate …
Designing For Novice Debuggers: A Pilot Study On An Ai-Assisted Debugging Tool, Oka Kurniawan, Erick Chandra, Christopher M. Poskitt, Yannic Noller, Kenny T.W. Choo, Cyrille Jegourel
Designing For Novice Debuggers: A Pilot Study On An Ai-Assisted Debugging Tool, Oka Kurniawan, Erick Chandra, Christopher M. Poskitt, Yannic Noller, Kenny T.W. Choo, Cyrille Jegourel
Research Collection School Of Computing and Information Systems
Debugging is a fundamental skill that novice programmers must develop. Numerous tools have been created to assist novice programmers in this process. Recently, large language models (LLMs) have been integrated with automated program repair techniques to generate fixes for students' buggy code. However, many of these tools foster an over-reliance on AI and do not actively engage students in the debugging process. In this work, we aim to design an intuitive debugging assistant, CodeHinter, that combines traditional debugging tools with LLM-based techniques to help novice debuggers fix semantic errors while promoting active engagement in the debugging process. We present findings …
Deep Reinforcement Learning For Solving The Stochastic E-Waste Collection Problem, Dang Viet Anh Nguyen, Aldy Gunawan, Mustafa Misir, Kwan Hui Lim, Pieter Vansteenwegen
Deep Reinforcement Learning For Solving The Stochastic E-Waste Collection Problem, Dang Viet Anh Nguyen, Aldy Gunawan, Mustafa Misir, Kwan Hui Lim, Pieter Vansteenwegen
Research Collection School Of Computing and Information Systems
With the growing influence of the internet and information technology, Electrical and Electronic Equipment (EEE) has become a gateway to technological innovations. However, discarded devices, also called e-waste, pose a significant threat to the environment and human health if not properly treated, disposed of, or recycled. In this study, we extend a novel model for the e-waste collection in an urban context: the Heterogeneous VRP with Multiple Time Windows and Stochastic Travel Times (HVRP-MTWSTT). We propose a solution method that employs deep reinforcement learning to guide local search heuristics (DRL-LSH). The contributions of this paper are as follows: (1) HVRP-MTWSTT …
Branch-And-Cut-And-Price For Agile Earth Observation Satellite Scheduling, Guansheng Peng, Jianjiang Wang, Guopeng Song, Aldy Gunawan, Lining Xing, Pieter Vansteenwegen
Branch-And-Cut-And-Price For Agile Earth Observation Satellite Scheduling, Guansheng Peng, Jianjiang Wang, Guopeng Song, Aldy Gunawan, Lining Xing, Pieter Vansteenwegen
Research Collection School Of Computing and Information Systems
The Agile Earth Observation Satellite scheduling selects and sequences satellite observations of possible targets on the Earth’s surface, each with a specific profit and multiple time windows. The objective is to maximize the collected profit of all observations completed under some operational constraints. The problem can be modeled as a variant of the Team Orienteering Problem with Time Windows (TOPTW). The key differences with the regular TOPTW are twofold: first, a time-dependent transition time is required for each pair of consecutive observations to adjust the camera’s look angles. Second, the time windows of each target vary during different observation cycles, …
Explainable Sentiment Analysis With Deepseek-R1: Performance, Efficiency, And Few-Shot Learning, Donghao Huang, Zhaoxia Wang
Explainable Sentiment Analysis With Deepseek-R1: Performance, Efficiency, And Few-Shot Learning, Donghao Huang, Zhaoxia Wang
Research Collection School Of Computing and Information Systems
Large language models (LLMs) have transformed sentiment analysis, yet balancing accuracy, efficiency, and explainability remains a critical challenge. This study presents the first comprehensive evaluation of DeepSeek-R1—an open-source reasoning model—against OpenAI’s GPT-4o and GPT-4o-mini. We test the full 671B model and its distilled variants, systematically documenting few-shot learning curves. Our experiments show DeepSeek-R1 achieves a 91.39% F1 score on 5-class sentiment and 99.31% accuracy on binary tasks with just 5 shots, an eightfold improvement in few-shot efficiency over GPT-4o. Architecture-specific distillation effects emerge, where a 32B Qwen2.5-based model outperforms the 70B Llama-based variant by 6.69 percentage points. While its reasoning …
Predict Social Economic Outcomes By Transferred Knowledge With Satellite Imagery, Yang Tang, Shih-Fen Cheng, Yunqiang Zhu, Yichen Yang, Zhiqiang Zou
Predict Social Economic Outcomes By Transferred Knowledge With Satellite Imagery, Yang Tang, Shih-Fen Cheng, Yunqiang Zhu, Yichen Yang, Zhiqiang Zou
Research Collection School Of Computing and Information Systems
Traditional deep learning methods and econometric models have played a crucial role in the field of data mining, particularly in the prediction of socioeconomic outcomes. However, socio-economic information is unable to be directly extracted from remote sensing data. So, in this paper, we propose a method to leverage transfer learning to predict socioeconomic indicators (outcomes) through satellite imagery. Specifically, we use road network types as a proxy for socioeconomic factors, which is more effective and stable than using nightlight. We have extracted eleven distinct road topological features to generate reasonable road network types. Given the unique characteristics of road networks, …
When Deep Learning Meets Information Retrieval-Based Bug Localization: A Survey, Feifei Niu, Chuanyi Li, Kui Liu, Xin Xia, David Lo
When Deep Learning Meets Information Retrieval-Based Bug Localization: A Survey, Feifei Niu, Chuanyi Li, Kui Liu, Xin Xia, David Lo
Research Collection School Of Computing and Information Systems
Bug localization is a crucial aspect of software maintenance, running through the entire software lifecycle. Information retrieval-based bug localization (IRBL) identifies buggy code based on bug reports, expediting the bug resolution process for developers. Recent years have witnessed significant achievements in IRBL, propelled by the widespread adoption of deep learning (DL). To provide a comprehensive overview of the current state of the art and delve into key issues, we conduct a survey encompassing 61 IRBL studies leveraging DL. We summarize best practices in each phase of the IRBL workflow, undertake a meta-analysis of prior studies, and suggest future research directions. …
Defects4c: Benchmarking Large Language Model Repair Capability With C/C++ Bugs, Jian Wang, Xiaofei Xie, Qiang Hu, Shangqing Liu, Jiongchi Yu, Jiaolong Kong, Yi Li
Defects4c: Benchmarking Large Language Model Repair Capability With C/C++ Bugs, Jian Wang, Xiaofei Xie, Qiang Hu, Shangqing Liu, Jiongchi Yu, Jiaolong Kong, Yi Li
Research Collection School Of Computing and Information Systems
Automated Program Repair (APR) plays a critical role in enhancing the quality and reliability of software systems. While substantial progress has been made in Java-based APR, largely facilitated by benchmarks like Defects4J, there remains a significant gap in research on C/C++ program repair, despite the widespread use of C/C++ and the prevalence of associated vulnerabilities. This gap is primarily due to the lack of high-quality, open-source benchmarks tailored for C/C++. To address this issue, we introduce Defects4C, a comprehensive and executable benchmark specifically designed for C/C++ program repair. Our dataset is constructed from real-world C/C++ repositories and includes a large …
Do Code Semantics Help? A Comprehensive Study On Execution Trace-Based Information For Code Large Language Models, Jian Wang, Xiaofei Xie, Qiang Hu, Shangqing Liu, Yi Li
Do Code Semantics Help? A Comprehensive Study On Execution Trace-Based Information For Code Large Language Models, Jian Wang, Xiaofei Xie, Qiang Hu, Shangqing Liu, Yi Li
Research Collection School Of Computing and Information Systems
Code Large Language Models (Code LLMs) have opened a new era in programming with their impressive capabilities. However, recent research has revealed critical limitations in their ability to reason about runtime behavior and understand the actual functionality of programs, which poses significant challenges for their post-training and practical deployment. Specifically, Code LLMs encounter two principal issues: (1) a lack of proficiency in reasoning about program execution behavior, as they struggle to interpret what programs actually do during runtime, and (2) inconsistent and fragmented representation of semantic information, such as execution traces, across existing methods, which hinders their ability to generalize …
Seeing Culture: A Benchmark For Visual Reasoning And Grounding, Burak Satar, Zhixin Ma, Patrick Amadeus Irrawan, Wilfried Ariel Mulyawan, Jing Jiang, Ee-Peng Lim, Chong-Wah Ngo
Seeing Culture: A Benchmark For Visual Reasoning And Grounding, Burak Satar, Zhixin Ma, Patrick Amadeus Irrawan, Wilfried Ariel Mulyawan, Jing Jiang, Ee-Peng Lim, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
Multimodal vision-language models (VLMs) have made substantial progress in various tasks that require a combined understanding of visual and textual content, particularly in cultural understanding tasks, with the emergence of new cultural datasets. However, these datasets frequently fall short of providing cultural reasoning while underrepresenting many cultures.In this paper, we introduce the Seeing Culture Benchmark (SCB), focusing on cultural reasoning with a novel approach that requires VLMs to reason on culturally rich images in two stages: i) selecting the correct visual option with multiple-choice visual question answering (VQA), and ii) segmenting the relevant cultural artifact as evidence of reasoning. Visual …
Efficient Integration Of External Knowledge To Llm-Based World Models Via Retrieval-Augmented Generation And Reinforcement Learning, Chang Yang, Xinrun Wang, Qinggang Zhang, Qi Jiang, Xiao Huang
Efficient Integration Of External Knowledge To Llm-Based World Models Via Retrieval-Augmented Generation And Reinforcement Learning, Chang Yang, Xinrun Wang, Qinggang Zhang, Qi Jiang, Xiao Huang
Research Collection School Of Computing and Information Systems
World models achieve remarkable success in predicting future states and planning in complex environments and Large Language Models (LLMs) serve as promising foundation to build general world models. However, their performances are usually constrained by the limited external knowledge to specific environments. Existing research attempts to enhance LLM-based world models through prompting or fine-tuning approaches, which are either requiring human knowledge or computationally extensive. Therefore, we introduce Retrieval-Augmented World Models (RAWM), a novel framework that leverages retrieval-augmented generation to efficiently integrate the external knowledge to LLM-based world models. Our main contributions are threefold: (i) We introduce a memory system and …
Seeing Is Fixing: Cross-Modal Reasoning With Multimodal Llms For Visual Software Issue Fixing, Kai Huang, Jian Zhang, Xiaofei Xie, Chunyang Chen
Seeing Is Fixing: Cross-Modal Reasoning With Multimodal Llms For Visual Software Issue Fixing, Kai Huang, Jian Zhang, Xiaofei Xie, Chunyang Chen
Research Collection School Of Computing and Information Systems
Large language model (LLM)-based automated program repair (APR) techniques have shown promising results in resolving real-world github issue tasks. Existing APR systems are primarily evaluated in unimodal settings (e.g., SWE-bench), relying solely on textual issue descriptions and source code. However, these autonomous systems struggle to resolve multimodal problem scenarios (e.g., SWE-bench M) due to limitations in interpreting and leveraging visual information. In multimodal scenarios, LLMs need to rely on visual information in the graphical user interface (GUI) to understand bugs and generate fixes. To bridge this gap, we propose GUIRepair, a cross-modal reasoning approach for resolving multimodal issue scenarios by …
Mmlu-Prox: A Multilingual Benchmark For Advanced Large Language Model Evaluation, Weihao Xuan, Et. Al.
Mmlu-Prox: A Multilingual Benchmark For Advanced Large Language Model Evaluation, Weihao Xuan, Et. Al.
Research Collection School Of Computing and Information Systems
Existing large language model (LLM) evaluation benchmarks primarily focus on English, while current multilingual tasks lack parallel questions that specifically assess cross-lingual reasoning abilities. This dual limitation makes it challenging to assess LLMs’ performance in the multilingual setting comprehensively. To fill this gap, we introduce MMLU-ProX, a comprehensive benchmark covering 29 languages, built on an English benchmark. Each language version consists of 11,829 identical questions, enabling direct cross-lingual comparisons. Additionally, to meet efficient evaluation needs, we provide a lite version containing 658 questions per language. To ensure the high quality of MMLU-ProX, we employ a rigorous development process that involves …
From Personas To Talks: Revisiting The Impact Of Personas On Llm-Synthesized Emotional Support Conversations, Shenghan Wu, Yimo Zhu, Wynne Hsu, Mong-Li Lee, Yang Deng
From Personas To Talks: Revisiting The Impact Of Personas On Llm-Synthesized Emotional Support Conversations, Shenghan Wu, Yimo Zhu, Wynne Hsu, Mong-Li Lee, Yang Deng
Research Collection School Of Computing and Information Systems
The rapid advancement of Large Language Models (LLMs) has revolutionized the generation of emotional support conversations (ESC), offering scalable solutions with reduced costs and enhanced data privacy. This paper explores the role of personas in the creation of ESC by LLMs. Our research utilizes established psychological frameworks to measure and infuse persona traits into LLMs, which then generate dialogues in the emotional support scenario. We conduct extensive evaluations to understand the stability of persona traits in dialogues, examining shifts in traits post-generation and their impact on dialogue quality and strategy distribution. Experimental results reveal several notable findings: 1) LLMs can …
Adasteer: Your Aligned Llm Is Inherently An Adaptive Jailbreak Defender, Weixiang Zhao, Jiahe Guo, Yulin Hu, Yang Deng, An Zhang, Xingyu Sui, Xinyang Han, Yanyan Zhao, Bing Qin, Tat-Seng Chua, Ting Liu
Adasteer: Your Aligned Llm Is Inherently An Adaptive Jailbreak Defender, Weixiang Zhao, Jiahe Guo, Yulin Hu, Yang Deng, An Zhang, Xingyu Sui, Xinyang Han, Yanyan Zhao, Bing Qin, Tat-Seng Chua, Ting Liu
Research Collection School Of Computing and Information Systems
Despite extensive efforts in safety alignment, large language models (LLMs) remain vulnerable to jailbreak attacks. Activation steering offers a training-free defense method but relies on fixed steering coefficients, resulting in suboptimal protection and increased false rejections of benign inputs. To address this, we propose AdaSteer, an adaptive activation steering method that dynamically adjusts model behavior based on input characteristics. We identify two key properties: Rejection Law (R-Law), which shows that stronger steering is needed for jailbreak inputs opposing the rejection direction, and Harmfulness Law (H-Law), which differentiates adversarial and benign inputs. AdaSteer steers input representations along both the Rejection Direction …
Chain Of Strategy Optimization Makes Large Language Models Better Emotional Supporter, Weixiang Zhao, Xingyu Sui, Xinyang Han, Yang Deng, Yulin Hu, Jiahe Guo, Libo Qin, Qianyun Du, Shijin Wang, Yanyan Zhao, Bing Qin, Ting Liu
Chain Of Strategy Optimization Makes Large Language Models Better Emotional Supporter, Weixiang Zhao, Xingyu Sui, Xinyang Han, Yang Deng, Yulin Hu, Jiahe Guo, Libo Qin, Qianyun Du, Shijin Wang, Yanyan Zhao, Bing Qin, Ting Liu
Research Collection School Of Computing and Information Systems
The growing emotional stress in modern society has increased the demand for Emotional Support Conversations (ESC). While Large Language Models (LLMs) show promise for ESC, they face two key challenges: (1) low strategy selection accuracy, and (2) preference bias, limiting their adaptability to users’ emotional needs. Existing supervised fine-tuning (SFT) struggles to address these issues, as it rigidly trains models on single gold-standard responses without modeling nuanced strategy trade-offs. To overcome these limitations, we propose a novel two-stage framework that optimizes strategy selection preferences at each dialogue turn. We first leverage Monte Carlo Tree Search to construct ESC-Pro, a high-quality …
Intentionframe: A Semi-Structured, Multi-Aspect Framework For Fine-Grained Conversational Intention Understanding, Zailong Tian, Zhuoheng Han, Lizi Liao, Lizi Liao
Intentionframe: A Semi-Structured, Multi-Aspect Framework For Fine-Grained Conversational Intention Understanding, Zailong Tian, Zhuoheng Han, Lizi Liao, Lizi Liao
Research Collection School Of Computing and Information Systems
Understanding user intentions in multi-turn dialogues is critical for conversational AI, yet existing approaches—relying on rigid slot-value structures or unstructured free-text—fail to fully capture conversational complexity. In this paper, we propose IntentionFrame, a semi-structured framework inspired by psychological and cognitive intention theories, which organizes conversational intents into four interrelated aspects: situation, emotion, action, and knowledge. This design not only retains interpretability but also provides LLMs with a rich context to accurately parse and respond to nuanced user inputs. To efficiently scale IntentionFrame annotations, we introduce a Weakly-supervised Reinforced Generation (WeRG) method that leverages a small set of high-quality human annotations …
One Planner To Guide Them All! Learning Adaptive Conversational Planners For Goal-Oriented Dialogues, Huy Dao, Lizi Liao
One Planner To Guide Them All! Learning Adaptive Conversational Planners For Goal-Oriented Dialogues, Huy Dao, Lizi Liao
Research Collection School Of Computing and Information Systems
Goal-oriented dialogues, such as recommendation and negotiation, often require balancing multiple, conflicting objectives. Existing methods typically involve training separate models for specific combinations of objectives, leading to computational and scalability issues. In this work, we aim to develop a new dialogue policy method that can adapt to varying objective preferences at inference time without retraining. This raises several challenges in terms of both (1) optimization strategy and (2) knowledge utilization. To address these, we propose a novel learning framework, Preference Adaptive Dialogue Policy Planner (PADPP), for multi-objective goal-oriented dialogues. Specifically, to tackle the former, we introduce a novel policy optimization …
Context-Aware Hierarchical Taxonomy Generation For Scientific Papers Via Llm-Guided Multi-Aspect Clustering, Kun Zhu, Lizi Liao, Yuxuan Gu, Lei Huang, Xiaocheng Feng, Bing Qin
Context-Aware Hierarchical Taxonomy Generation For Scientific Papers Via Llm-Guided Multi-Aspect Clustering, Kun Zhu, Lizi Liao, Yuxuan Gu, Lei Huang, Xiaocheng Feng, Bing Qin
Research Collection School Of Computing and Information Systems
The rapid growth of scientific literature demands efficient methods to organize and synthesize research findings. Existing taxonomy construction methods, leveraging unsupervised clustering or direct prompting of large language models (LLMs), often lack coherence and granularity. We propose a novel context-aware hierarchical taxonomy generation framework that integrates LLM-guided multi-aspect encoding with dynamic clustering. Our method leverages LLMs to identify key aspects of each paper (e.g., methodology, dataset, evaluation) and generates aspect-specific paper summaries, which are then encoded and clustered along each aspect to form a coherent hierarchy. In addition, we introduce a new evaluation benchmark of 156 expert-crafted taxonomies encompassing 11.6k …
Distillcaps: Enhancing Audio-Language Alignment In Captioning Via Retrieval-Augmented Knowledge Distillation, Thinh Pham, Nghiem Diep, Lizi Liao, Binh Nguyen
Distillcaps: Enhancing Audio-Language Alignment In Captioning Via Retrieval-Augmented Knowledge Distillation, Thinh Pham, Nghiem Diep, Lizi Liao, Binh Nguyen
Research Collection School Of Computing and Information Systems
Automated audio captioning (AAC) benefits from incorporatingexternal context to interpret complex sounds, but doing so withretrieval-augmented generation (RAG) at inference is sometimesinfeasible due to data availability or incurs significant latency andcomplexity. We propose DistillCaps, a novel training-time frame-work that leverages RAG to guide knowledge distillation for im-proved audio-language alignment, while lessening the relianceon retrieval during inference. In our framework, a RAG-equippedteacher model retrieves relevant textual information (e.g., simi-lar captions) for each audio clip and uses it for training to gener-ate context-enriched captions. Simultaneously, a student model istrained to imitate this teacher, learning to produce high-qualitycaptions from audio alone. We further …