Open Access. Powered by Scholars. Published by Universities.®
Artificial Intelligence and Robotics Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Databases and Information Systems (360)
- Engineering (296)
- Operations Research, Systems Engineering and Industrial Engineering (256)
- Business (177)
- Graphics and Human Computer Interfaces (176)
-
- Social and Behavioral Sciences (173)
- Software Engineering (137)
- Numerical Analysis and Scientific Computing (112)
- Theory and Algorithms (103)
- Public Affairs, Public Policy and Public Administration (75)
- Transportation (63)
- Programming Languages and Compilers (62)
- Information Security (43)
- Medicine and Health Sciences (43)
- Technology and Innovation (38)
- OS and Networks (37)
- Education (36)
- Law (36)
- Asian Studies (35)
- Computer Engineering (35)
- International and Area Studies (35)
- Health Information Technology (30)
- Science and Technology Law (22)
- Psychology (21)
- Library and Information Science (20)
- Finance and Financial Management (18)
- Higher Education (18)
- Keyword
-
- Artificial intelligence (99)
- Machine learning (55)
- Reinforcement learning (42)
- Deep learning (38)
- Artificial Intelligence (30)
-
- Large Language Models (30)
- Generative AI (29)
- Large language models (24)
- ChatGPT (23)
- Large Language Model (23)
- Singapore (22)
- Computer vision (19)
- Large language model (18)
- Optimization (18)
- Reinforcement Learning (18)
- Scheduling (18)
- Anomaly detection (17)
- Natural language processing (17)
- Deep reinforcement learning (16)
- Deep Learning (15)
- LLMs (15)
- Machine Learning (15)
- Vehicle routing problem (15)
- AI (14)
- Neural networks (13)
- Uncertainty (13)
- Software engineering (12)
- Graph neural networks (11)
- Metaverse (10)
- Multi-agent systems (10)
- Publication Year
- Publication
-
- Research Collection School Of Computing and Information Systems (1664)
- Dissertations and Theses Collection (Open Access) (57)
- Research Collection Lee Kong Chian School Of Business (33)
- Research Collection Yong Pung How School Of Law (31)
- Research Collection School of Social Sciences (22)
-
- Asian Management Insights (16)
- FORCE 2026 (14)
- Perspectives@SMU (11)
- Research Collection College of Integrative Studies (10)
- Research Collection Library (8)
- PhD Student’s Publications Collection (6)
- MITB Thought Leadership Series (4)
- 2024 AI for Research Week (3)
- CCX Research (3)
- LARC Research Publications (2)
- Research Collection School Of Accountancy (2)
- CASTLe: Collection of Articles on Scholarship for Teaching and Learning (1)
- Centre for AI & Data Governance (2019-2025) (1)
- Centre for Computational Law (2022-2025) (1)
- ROSA Journal Articles and Publications (1)
- Research Collection Office of Research (1)
- Research Collection School Of Economics (1)
- Research@SMU Infographics (1)
- Research@SMU: Connecting the Dots (1)
- SMU Press Releases and News (1)
- Sim Kee Boon Institute for Financial Economics (1)
- Student Publications (1)
- Publication Type
- File Type
Articles 241 - 270 of 1897
Full-Text Articles in Artificial Intelligence and Robotics
Uncovering The Values Of The Metaverse For Leisure Use By Individuals: A Value-Focused Thinking Approach, Ruilin Zheng, Fiona Fui-Hoon Nah
Uncovering The Values Of The Metaverse For Leisure Use By Individuals: A Value-Focused Thinking Approach, Ruilin Zheng, Fiona Fui-Hoon Nah
Research Collection School Of Computing and Information Systems
The metaverse is a computer-mediated environment where users take the form of digital avatars when participating in activities and interacting with one another. Given the popularity of the metaverse, especially among the younger population, we identified the values offered by the metaverse for leisure use by its users. Using the Value-Focused Thinking (VFT) approach, we identified these values in the form of fundamental and means objectives. The VFT approach was applied in interviewing users who conduct leisure activities in the metaverse and in analyzing the data collected. A total of 27 metaverse users were interviewed, which generated 8 fundamental objectives …
International Workshop On Multimodal Generative Search And Recommendation (Mmgensr@Cikm 2025), Yi Bin, Haoxuan Li, Haokai Ma, Yang Zhang, Wenjie Wang, Yunshan Ma, Yang Yang, Tat‑Seng Chua
International Workshop On Multimodal Generative Search And Recommendation (Mmgensr@Cikm 2025), Yi Bin, Haoxuan Li, Haokai Ma, Yang Zhang, Wenjie Wang, Yunshan Ma, Yang Yang, Tat‑Seng Chua
Research Collection School Of Computing and Information Systems
Recent breakthroughs in generative Artificial Intelligence (AI) have ignited a revolutionary wave across information retrieval and recommender systems. This workshop serves as a premier interdisciplinary platform to explore how generative models, particularly Large Language Models (LLMs) and Large Multimodal Models (LMMs), are transforming multimodal search and recommendation paradigms [3, 6, 9, 10, 12-14]. We aim to convene researchers and practitioners to discuss innovative architectures, methodologies, and evaluation strategies spanning generative document retrieval [5, 8] generative image retrieval [ 7, 16], grounded answer generation [17], generative recommendation [2, 4, 11], and related tasks involving multiple modalities [1,15]. The workshop will facilitate …
Designing For Novice Debuggers: A Pilot Study On An Ai-Assisted Debugging Tool, Oka Kurniawan, Erick Chandra, Christopher M. Poskitt, Yannic Noller, Kenny T.W. Choo, Cyrille Jegourel
Designing For Novice Debuggers: A Pilot Study On An Ai-Assisted Debugging Tool, Oka Kurniawan, Erick Chandra, Christopher M. Poskitt, Yannic Noller, Kenny T.W. Choo, Cyrille Jegourel
Research Collection School Of Computing and Information Systems
Debugging is a fundamental skill that novice programmers must develop. Numerous tools have been created to assist novice programmers in this process. Recently, large language models (LLMs) have been integrated with automated program repair techniques to generate fixes for students' buggy code. However, many of these tools foster an over-reliance on AI and do not actively engage students in the debugging process. In this work, we aim to design an intuitive debugging assistant, CodeHinter, that combines traditional debugging tools with LLM-based techniques to help novice debuggers fix semantic errors while promoting active engagement in the debugging process. We present findings …
How Behavioral Science Can Improve The Return On Ai Investments, David De Cremer, Shane Schweitzer, Jack Mcguire, Devesh Narayanan
How Behavioral Science Can Improve The Return On Ai Investments, David De Cremer, Shane Schweitzer, Jack Mcguire, Devesh Narayanan
Research Collection Lee Kong Chian School Of Business
Many AI projects fail because leaders treat adoption as a tech purchase instead of a behavioral change problem. People resist tools that disrupt routines, overreact to visible AI errors, and prefer familiar human judgment. As a result, even good systems fail to gain purchase. Leaders can address this problem by applying “Behavioral Human-Centered AI” across the AI adoption cycle. In the design phrase, companies should co-design with diverse users, add purposeful friction where it improves scrutiny, require beta tests with subgroup results and behavioral input. During adoption, they should frame AI as an augmenter, disclose limits and safeguards, use explainability …
Building Confidence For Class Participation, Tamas Makany, Ivy Seow
Building Confidence For Class Participation, Tamas Makany, Ivy Seow
Research Collection Lee Kong Chian School Of Business
What happens to students’ critical thinking when half the class filters their thoughts through AI? During a recent debate on AI policy in education, one student mentioned they routinely run their ideas through ChatGPT before speaking up. When I asked who else did the same, more than half the class raised their hands.
A Fake Friend? Ai Companions Are Exactly That, Seow Hon Tan
A Fake Friend? Ai Companions Are Exactly That, Seow Hon Tan
Research Collection Yong Pung How School Of Law
In a commentary, SMU Associate Professor of Law Tan Seow Hon discussed how AI companions, which promise emotionally intelligent companionship, have blurred the line between human and machine relationships by mimicking empathy, memory, and affection. She suggested that while such technologies may ease loneliness, they risk fostering narcissism, diminishing real human connection, and replacing authentic friendship with comforting illusions that erode the capacity for love and community.
When Deep Learning Meets Information Retrieval-Based Bug Localization: A Survey, Feifei Niu, Chuanyi Li, Kui Liu, Xin Xia, David Lo
When Deep Learning Meets Information Retrieval-Based Bug Localization: A Survey, Feifei Niu, Chuanyi Li, Kui Liu, Xin Xia, David Lo
Research Collection School Of Computing and Information Systems
Bug localization is a crucial aspect of software maintenance, running through the entire software lifecycle. Information retrieval-based bug localization (IRBL) identifies buggy code based on bug reports, expediting the bug resolution process for developers. Recent years have witnessed significant achievements in IRBL, propelled by the widespread adoption of deep learning (DL). To provide a comprehensive overview of the current state of the art and delve into key issues, we conduct a survey encompassing 61 IRBL studies leveraging DL. We summarize best practices in each phase of the IRBL workflow, undertake a meta-analysis of prior studies, and suggest future research directions. …
Defects4c: Benchmarking Large Language Model Repair Capability With C/C++ Bugs, Jian Wang, Xiaofei Xie, Qiang Hu, Shangqing Liu, Jiongchi Yu, Jiaolong Kong, Yi Li
Defects4c: Benchmarking Large Language Model Repair Capability With C/C++ Bugs, Jian Wang, Xiaofei Xie, Qiang Hu, Shangqing Liu, Jiongchi Yu, Jiaolong Kong, Yi Li
Research Collection School Of Computing and Information Systems
Automated Program Repair (APR) plays a critical role in enhancing the quality and reliability of software systems. While substantial progress has been made in Java-based APR, largely facilitated by benchmarks like Defects4J, there remains a significant gap in research on C/C++ program repair, despite the widespread use of C/C++ and the prevalence of associated vulnerabilities. This gap is primarily due to the lack of high-quality, open-source benchmarks tailored for C/C++. To address this issue, we introduce Defects4C, a comprehensive and executable benchmark specifically designed for C/C++ program repair. Our dataset is constructed from real-world C/C++ repositories and includes a large …
Do Code Semantics Help? A Comprehensive Study On Execution Trace-Based Information For Code Large Language Models, Jian Wang, Xiaofei Xie, Qiang Hu, Shangqing Liu, Yi Li
Do Code Semantics Help? A Comprehensive Study On Execution Trace-Based Information For Code Large Language Models, Jian Wang, Xiaofei Xie, Qiang Hu, Shangqing Liu, Yi Li
Research Collection School Of Computing and Information Systems
Code Large Language Models (Code LLMs) have opened a new era in programming with their impressive capabilities. However, recent research has revealed critical limitations in their ability to reason about runtime behavior and understand the actual functionality of programs, which poses significant challenges for their post-training and practical deployment. Specifically, Code LLMs encounter two principal issues: (1) a lack of proficiency in reasoning about program execution behavior, as they struggle to interpret what programs actually do during runtime, and (2) inconsistent and fragmented representation of semantic information, such as execution traces, across existing methods, which hinders their ability to generalize …
Seeing Culture: A Benchmark For Visual Reasoning And Grounding, Burak Satar, Zhixin Ma, Patrick Amadeus Irrawan, Wilfried Ariel Mulyawan, Jing Jiang, Ee-Peng Lim, Chong-Wah Ngo
Seeing Culture: A Benchmark For Visual Reasoning And Grounding, Burak Satar, Zhixin Ma, Patrick Amadeus Irrawan, Wilfried Ariel Mulyawan, Jing Jiang, Ee-Peng Lim, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
Multimodal vision-language models (VLMs) have made substantial progress in various tasks that require a combined understanding of visual and textual content, particularly in cultural understanding tasks, with the emergence of new cultural datasets. However, these datasets frequently fall short of providing cultural reasoning while underrepresenting many cultures.In this paper, we introduce the Seeing Culture Benchmark (SCB), focusing on cultural reasoning with a novel approach that requires VLMs to reason on culturally rich images in two stages: i) selecting the correct visual option with multiple-choice visual question answering (VQA), and ii) segmenting the relevant cultural artifact as evidence of reasoning. Visual …
Efficient Integration Of External Knowledge To Llm-Based World Models Via Retrieval-Augmented Generation And Reinforcement Learning, Chang Yang, Xinrun Wang, Qinggang Zhang, Qi Jiang, Xiao Huang
Efficient Integration Of External Knowledge To Llm-Based World Models Via Retrieval-Augmented Generation And Reinforcement Learning, Chang Yang, Xinrun Wang, Qinggang Zhang, Qi Jiang, Xiao Huang
Research Collection School Of Computing and Information Systems
World models achieve remarkable success in predicting future states and planning in complex environments and Large Language Models (LLMs) serve as promising foundation to build general world models. However, their performances are usually constrained by the limited external knowledge to specific environments. Existing research attempts to enhance LLM-based world models through prompting or fine-tuning approaches, which are either requiring human knowledge or computationally extensive. Therefore, we introduce Retrieval-Augmented World Models (RAWM), a novel framework that leverages retrieval-augmented generation to efficiently integrate the external knowledge to LLM-based world models. Our main contributions are threefold: (i) We introduce a memory system and …
Seeing Is Fixing: Cross-Modal Reasoning With Multimodal Llms For Visual Software Issue Fixing, Kai Huang, Jian Zhang, Xiaofei Xie, Chunyang Chen
Seeing Is Fixing: Cross-Modal Reasoning With Multimodal Llms For Visual Software Issue Fixing, Kai Huang, Jian Zhang, Xiaofei Xie, Chunyang Chen
Research Collection School Of Computing and Information Systems
Large language model (LLM)-based automated program repair (APR) techniques have shown promising results in resolving real-world github issue tasks. Existing APR systems are primarily evaluated in unimodal settings (e.g., SWE-bench), relying solely on textual issue descriptions and source code. However, these autonomous systems struggle to resolve multimodal problem scenarios (e.g., SWE-bench M) due to limitations in interpreting and leveraging visual information. In multimodal scenarios, LLMs need to rely on visual information in the graphical user interface (GUI) to understand bugs and generate fixes. To bridge this gap, we propose GUIRepair, a cross-modal reasoning approach for resolving multimodal issue scenarios by …
Mmlu-Prox: A Multilingual Benchmark For Advanced Large Language Model Evaluation, Weihao Xuan, Et. Al.
Mmlu-Prox: A Multilingual Benchmark For Advanced Large Language Model Evaluation, Weihao Xuan, Et. Al.
Research Collection School Of Computing and Information Systems
Existing large language model (LLM) evaluation benchmarks primarily focus on English, while current multilingual tasks lack parallel questions that specifically assess cross-lingual reasoning abilities. This dual limitation makes it challenging to assess LLMs’ performance in the multilingual setting comprehensively. To fill this gap, we introduce MMLU-ProX, a comprehensive benchmark covering 29 languages, built on an English benchmark. Each language version consists of 11,829 identical questions, enabling direct cross-lingual comparisons. Additionally, to meet efficient evaluation needs, we provide a lite version containing 658 questions per language. To ensure the high quality of MMLU-ProX, we employ a rigorous development process that involves …
From Personas To Talks: Revisiting The Impact Of Personas On Llm-Synthesized Emotional Support Conversations, Shenghan Wu, Yimo Zhu, Wynne Hsu, Mong-Li Lee, Yang Deng
From Personas To Talks: Revisiting The Impact Of Personas On Llm-Synthesized Emotional Support Conversations, Shenghan Wu, Yimo Zhu, Wynne Hsu, Mong-Li Lee, Yang Deng
Research Collection School Of Computing and Information Systems
The rapid advancement of Large Language Models (LLMs) has revolutionized the generation of emotional support conversations (ESC), offering scalable solutions with reduced costs and enhanced data privacy. This paper explores the role of personas in the creation of ESC by LLMs. Our research utilizes established psychological frameworks to measure and infuse persona traits into LLMs, which then generate dialogues in the emotional support scenario. We conduct extensive evaluations to understand the stability of persona traits in dialogues, examining shifts in traits post-generation and their impact on dialogue quality and strategy distribution. Experimental results reveal several notable findings: 1) LLMs can …
Adasteer: Your Aligned Llm Is Inherently An Adaptive Jailbreak Defender, Weixiang Zhao, Jiahe Guo, Yulin Hu, Yang Deng, An Zhang, Xingyu Sui, Xinyang Han, Yanyan Zhao, Bing Qin, Tat-Seng Chua, Ting Liu
Adasteer: Your Aligned Llm Is Inherently An Adaptive Jailbreak Defender, Weixiang Zhao, Jiahe Guo, Yulin Hu, Yang Deng, An Zhang, Xingyu Sui, Xinyang Han, Yanyan Zhao, Bing Qin, Tat-Seng Chua, Ting Liu
Research Collection School Of Computing and Information Systems
Despite extensive efforts in safety alignment, large language models (LLMs) remain vulnerable to jailbreak attacks. Activation steering offers a training-free defense method but relies on fixed steering coefficients, resulting in suboptimal protection and increased false rejections of benign inputs. To address this, we propose AdaSteer, an adaptive activation steering method that dynamically adjusts model behavior based on input characteristics. We identify two key properties: Rejection Law (R-Law), which shows that stronger steering is needed for jailbreak inputs opposing the rejection direction, and Harmfulness Law (H-Law), which differentiates adversarial and benign inputs. AdaSteer steers input representations along both the Rejection Direction …
Chain Of Strategy Optimization Makes Large Language Models Better Emotional Supporter, Weixiang Zhao, Xingyu Sui, Xinyang Han, Yang Deng, Yulin Hu, Jiahe Guo, Libo Qin, Qianyun Du, Shijin Wang, Yanyan Zhao, Bing Qin, Ting Liu
Chain Of Strategy Optimization Makes Large Language Models Better Emotional Supporter, Weixiang Zhao, Xingyu Sui, Xinyang Han, Yang Deng, Yulin Hu, Jiahe Guo, Libo Qin, Qianyun Du, Shijin Wang, Yanyan Zhao, Bing Qin, Ting Liu
Research Collection School Of Computing and Information Systems
The growing emotional stress in modern society has increased the demand for Emotional Support Conversations (ESC). While Large Language Models (LLMs) show promise for ESC, they face two key challenges: (1) low strategy selection accuracy, and (2) preference bias, limiting their adaptability to users’ emotional needs. Existing supervised fine-tuning (SFT) struggles to address these issues, as it rigidly trains models on single gold-standard responses without modeling nuanced strategy trade-offs. To overcome these limitations, we propose a novel two-stage framework that optimizes strategy selection preferences at each dialogue turn. We first leverage Monte Carlo Tree Search to construct ESC-Pro, a high-quality …
Intentionframe: A Semi-Structured, Multi-Aspect Framework For Fine-Grained Conversational Intention Understanding, Zailong Tian, Zhuoheng Han, Lizi Liao, Lizi Liao
Intentionframe: A Semi-Structured, Multi-Aspect Framework For Fine-Grained Conversational Intention Understanding, Zailong Tian, Zhuoheng Han, Lizi Liao, Lizi Liao
Research Collection School Of Computing and Information Systems
Understanding user intentions in multi-turn dialogues is critical for conversational AI, yet existing approaches—relying on rigid slot-value structures or unstructured free-text—fail to fully capture conversational complexity. In this paper, we propose IntentionFrame, a semi-structured framework inspired by psychological and cognitive intention theories, which organizes conversational intents into four interrelated aspects: situation, emotion, action, and knowledge. This design not only retains interpretability but also provides LLMs with a rich context to accurately parse and respond to nuanced user inputs. To efficiently scale IntentionFrame annotations, we introduce a Weakly-supervised Reinforced Generation (WeRG) method that leverages a small set of high-quality human annotations …
One Planner To Guide Them All! Learning Adaptive Conversational Planners For Goal-Oriented Dialogues, Huy Dao, Lizi Liao
One Planner To Guide Them All! Learning Adaptive Conversational Planners For Goal-Oriented Dialogues, Huy Dao, Lizi Liao
Research Collection School Of Computing and Information Systems
Goal-oriented dialogues, such as recommendation and negotiation, often require balancing multiple, conflicting objectives. Existing methods typically involve training separate models for specific combinations of objectives, leading to computational and scalability issues. In this work, we aim to develop a new dialogue policy method that can adapt to varying objective preferences at inference time without retraining. This raises several challenges in terms of both (1) optimization strategy and (2) knowledge utilization. To address these, we propose a novel learning framework, Preference Adaptive Dialogue Policy Planner (PADPP), for multi-objective goal-oriented dialogues. Specifically, to tackle the former, we introduce a novel policy optimization …
Context-Aware Hierarchical Taxonomy Generation For Scientific Papers Via Llm-Guided Multi-Aspect Clustering, Kun Zhu, Lizi Liao, Yuxuan Gu, Lei Huang, Xiaocheng Feng, Bing Qin
Context-Aware Hierarchical Taxonomy Generation For Scientific Papers Via Llm-Guided Multi-Aspect Clustering, Kun Zhu, Lizi Liao, Yuxuan Gu, Lei Huang, Xiaocheng Feng, Bing Qin
Research Collection School Of Computing and Information Systems
The rapid growth of scientific literature demands efficient methods to organize and synthesize research findings. Existing taxonomy construction methods, leveraging unsupervised clustering or direct prompting of large language models (LLMs), often lack coherence and granularity. We propose a novel context-aware hierarchical taxonomy generation framework that integrates LLM-guided multi-aspect encoding with dynamic clustering. Our method leverages LLMs to identify key aspects of each paper (e.g., methodology, dataset, evaluation) and generates aspect-specific paper summaries, which are then encoded and clustered along each aspect to form a coherent hierarchy. In addition, we introduce a new evaluation benchmark of 156 expert-crafted taxonomies encompassing 11.6k …
Distillcaps: Enhancing Audio-Language Alignment In Captioning Via Retrieval-Augmented Knowledge Distillation, Thinh Pham, Nghiem Diep, Lizi Liao, Binh Nguyen
Distillcaps: Enhancing Audio-Language Alignment In Captioning Via Retrieval-Augmented Knowledge Distillation, Thinh Pham, Nghiem Diep, Lizi Liao, Binh Nguyen
Research Collection School Of Computing and Information Systems
Automated audio captioning (AAC) benefits from incorporatingexternal context to interpret complex sounds, but doing so withretrieval-augmented generation (RAG) at inference is sometimesinfeasible due to data availability or incurs significant latency andcomplexity. We propose DistillCaps, a novel training-time frame-work that leverages RAG to guide knowledge distillation for im-proved audio-language alignment, while lessening the relianceon retrieval during inference. In our framework, a RAG-equippedteacher model retrieves relevant textual information (e.g., simi-lar captions) for each audio clip and uses it for training to gener-ate context-enriched captions. Simultaneously, a student model istrained to imitate this teacher, learning to produce high-qualitycaptions from audio alone. We further …
Instructors’ Strategies In Creating And Implementing Constructivist Llm-Based Learning Activities, Emily Aurelia, Shun Yi Yeo, Michelle Lui, Effie Lai-Chong Law, Anthony Tang
Instructors’ Strategies In Creating And Implementing Constructivist Llm-Based Learning Activities, Emily Aurelia, Shun Yi Yeo, Michelle Lui, Effie Lai-Chong Law, Anthony Tang
Research Collection School Of Computing and Information Systems
Large language models (LLMs) are increasingly being integrated into educational settings, enabling more adoption of constructivist teaching and learning approaches in classrooms. This paper explores the strategies instructors are currently using to incorporate LLMs into learning activities that align with constructivist principles, which emphasize that learners actively construct their own knowledge. Through interviews with nine instructors who have designed eleven distinct LLM-based activities and using reflexive thematic analysis, this study identifies various types of learning activities with respect to four different aspects of the constructivist learning theory. The strategies employed and challenges faced to foster constructivist student-LLM interaction were also …
Harnessing Se Community Knowledge For Developer-Centric Code Intelligence, Chengran Yang
Harnessing Se Community Knowledge For Developer-Centric Code Intelligence, Chengran Yang
Dissertations and Theses Collection (Open Access)
The integration of Large Language Models (LLMs), particularly those tailored for programming tasks—referred to as code LLMs—has created novel opportunities to enhance developer productivity. These advanced models automate routine and repetitive coding tasks, such as code generation and debugging, and enable faster prototyping and more efficient problem-solving. Despite these remarkable advantages, the current generation of code LLMs exhibits notable limitations that impact their practical effectiveness in real-world software engineering scenarios. These models frequently produce code that is inefficient or suboptimal in runtime performance, demonstrate opaque reasoning processes, and struggle to adapt effectively to diverse developer contexts and specific requirements. Moreover, …
Memory-Efficient 4-Bit Preconditioned Stochastic Optimization, Jingyang Li, Kuangyu Ding, Kim-Chuan Toh, Pan Zhou
Memory-Efficient 4-Bit Preconditioned Stochastic Optimization, Jingyang Li, Kuangyu Ding, Kim-Chuan Toh, Pan Zhou
Research Collection School Of Computing and Information Systems
Preconditioned stochastic optimization algorithms, exemplified by Shampoo, outperform first-order optimizers by offering theoretical convergence benefits and practical gains in large-scale neural network training. However, they incur substantial memory overhead due to the storage demands of non-diagonal preconditioning matrices. To address this, we introduce 4-bit quantization for Shampoo’s preconditioners. We introduce two key methods: First, we apply Cholesky decomposition followed by quantization of the Cholesky factors, reducing memory usage by leveraging their lower triangular structure while better preserving spectral properties to minimize information loss. To our knowledge, this is the first quantization approach applied to Cholesky factors of preconditioners. Second, we …
What Students Really Think: Unpacking Ai Ethics In Educational Assessments Through A Triadic Framework, Lim Ming Soon Tristan, Gottipati Swapna, Michelle L. F. Cheong
What Students Really Think: Unpacking Ai Ethics In Educational Assessments Through A Triadic Framework, Lim Ming Soon Tristan, Gottipati Swapna, Michelle L. F. Cheong
Research Collection School Of Computing and Information Systems
The rise of AI in educational assessments has significantly enhanced efficiency and accuracy. However, it also introduces critical ethical challenges, including bias in grading, data privacy risks, and accountability gaps. These issues can undermine trust in AI-driven assessments and compromise educational fairness, making a structured ethical framework essential. To address these challenges, this study empirically validates an existing triadic ethical framework for AI-assisted educational assessments, originally proposed by Lim, Gottipati and Cheong (In: Keengwe (ed) Creative AI tools and ethical implications in teaching and learning, IGI Global, 2023), grounded in student perceptions. The framework encompasses three ethical domains—physical, cognitive, and …
Cookingdiffusion: Cooking Procedural Image Generation With Stable Diffusion, Yuan Wang, Bin Zhu, Yanbin Hao, Chong-Wah Ngo, Yi Tan, Xiang Wang
Cookingdiffusion: Cooking Procedural Image Generation With Stable Diffusion, Yuan Wang, Bin Zhu, Yanbin Hao, Chong-Wah Ngo, Yi Tan, Xiang Wang
Research Collection School Of Computing and Information Systems
Recent advancements in text-to-image generation models have excelled in creating diverse and realistic images. This success extends to food imagery, where various conditional inputs like cooking styles, ingredients, and recipes are utilized. However, a yet-unexplored challenge is generating a sequence of procedural images based on cooking steps from a recipe. This could enhance the cooking experience with visual guidance and possibly lead to an intelligent cooking simulation system. To fill this gap, we introduce a novel task called cooking procedural image generation. This task is inherently demanding, as it strives to create photo-realistic images that align with cooking steps while …
Probabilistic Prototype Calibration Of Vision-Language Models For Generalized Few-Shot Semantic Segmentation, Jie Liu, Jiayi Shen, Pan Zhou, Jan-Jakob Sonke, Stratis Gavves
Probabilistic Prototype Calibration Of Vision-Language Models For Generalized Few-Shot Semantic Segmentation, Jie Liu, Jiayi Shen, Pan Zhou, Jan-Jakob Sonke, Stratis Gavves
Research Collection School Of Computing and Information Systems
Generalized Few-Shot Semantic Segmentation (GFSS) aims to extend a segmentation model to novel classes with only a few annotated examples while maintaining performance on base classes. Recently, pretrained vision-language models (VLMs) such as CLIP have been leveraged in GFSS to improve generalization on novel classes through multi-modal prototypes learning. However, existing prototype-based methods are inherently deterministic, limiting the adaptability of learned prototypes to diverse samples, particularly for novel classes with scarce annotations. To address this, we propose FewCLIP, a probabilistic prototype calibration framework over multi-modal prototypes from the pretrained CLIP, thus providing more adaptive prototype learning for GFSS. Specifically, FewCLIP …
Visual-Enhanced Multimodal Framework For Flexible Job Shop Scheduling Problem, Peng Zhao, Zhiguang Cao, Di Wang, Wen Song, Wei Pang, You Zhou, Yuan Jiang
Visual-Enhanced Multimodal Framework For Flexible Job Shop Scheduling Problem, Peng Zhao, Zhiguang Cao, Di Wang, Wen Song, Wei Pang, You Zhou, Yuan Jiang
Research Collection School Of Computing and Information Systems
Multimodal models leverage complementary information across modalities to enrich feature representations. While visual information shows potential in representing structure for some combinatorial optimization problems (COPs), its application to complex scheduling like the Flexible Job Shop Scheduling Problem (FJSP) remains underexplored. Current learning-based FJSP solvers predominantly rely on handcrafted state features. This dependence can lead to inconsistencies and may not fully capture the problem's intricate dynamics. Crucially, these methods overlook visual modalities. Visual representations offer a distinct advantage by inherently capturing the global topological structure and complex resource interactions within the FJSP state. Unlike localized handcrafted features, this holistic, structural view …
Mitigating Cross-Modal Representation Bias For Multicultural Image-To-Recipe Retrieval, Qing Wang, Chong-Wah Ngo, Yu Cao, Ee-Peng Lim
Mitigating Cross-Modal Representation Bias For Multicultural Image-To-Recipe Retrieval, Qing Wang, Chong-Wah Ngo, Yu Cao, Ee-Peng Lim
Research Collection School Of Computing and Information Systems
Existing approaches for image-to-recipe retrieval have the implicit assumption that a food image can fully capture the details textually documented in its recipe. However, a food image only reflects the visual outcome of a cooked dish and not the underlying cooking process. Consequently, learning cross-modal representations to bridge the modality gap between images and recipes tends to ignore subtle, recipe-specific details that are not visually apparent but are crucial for recipe retrieval. Specifically, the representations are biased to capture the dominant visual elements, resulting in difficulty in ranking similar recipes with subtle differences in use of ingredients and cooking methods. …
Boosting Chart-To-Code Generation In Mllm Via Dual Preference-Guided Refinement, Zhihan Zhang, Yixin Cao, Lizi Liao
Boosting Chart-To-Code Generation In Mllm Via Dual Preference-Guided Refinement, Zhihan Zhang, Yixin Cao, Lizi Liao
Research Collection School Of Computing and Information Systems
Translating chart images into executable plotting scripts-referred to as the chart-to-code generation task-requires Multimodal Large Language Models (MLLMs) to perform fine-grained visual parsing, precise code synthesis, and robust cross-modal reasoning. However, this task is inherently under-constrained: multiple valid code implementations can produce the same visual chart, and evaluation must consider both code correctness and visual fidelity across diverse dimensions. This makes it difficult to learn accurate and generalizable mappings through standard supervised fine-tuning. To address these challenges, we propose a dual preference-guided refinement framework that combines a feedback-driven, dual-modality reward mechanism with iterative preference learning. Our approach introduces a structured …
Art4math: Handwritten Mathematical Expression Recognition Via Multimodal Sketch Grounding, Yang Zhou, Jin Wang, Yuxiao Zhang, Kaixiang Huang, Guodong Lu, Jingru Yang, Shengfeng He
Art4math: Handwritten Mathematical Expression Recognition Via Multimodal Sketch Grounding, Yang Zhou, Jin Wang, Yuxiao Zhang, Kaixiang Huang, Guodong Lu, Jingru Yang, Shengfeng He
Research Collection School Of Computing and Information Systems
Handwritten Mathematical Expression Recognition (HMER) remains a challenging task due to the structural complexity of mathematical notation and the ambiguity of handwritten symbols-e.g., ''ρ'' vs. ''p'' or ''B'' vs. ''β''. While stroke-based models offer disambiguation via temporal cues, most existing methods are constrained by coarse modality fusion and a lack of fine-grained cross-modal alignment, further hindered by limited annotated data. We introduce Art for Math (Art4Math), a novel framework that leverages the structural richness of human sketches to enhance HMER through fine-grained, modality-aware learning. Art4Math follows a two-stage training paradigm: Art Grounding (A-Grd) and Math Decoding (M-Dec). In A-Grd, the …