Open Access. Powered by Scholars. Published by Universities.®

Programming Languages and Compilers Commons

Open Access. Powered by Scholars. Published by Universities.®

1,843 Full-Text Articles 3,360 Authors 748,922 Downloads 137 Institutions

All Articles in Programming Languages and Compilers

Faceted Search

1,843 full-text articles. Page 3 of 79.

Contrastrepair: Enhancing Conversation-Based Automated Program Repair Via Contrastive Test Case Pairs, Jiaolong KONG, Xiaofei XIE, Mingfei CHENG, Shangqing Liu, Xiaoning Du, Qi Guo 2025 Singapore Management University

Contrastrepair: Enhancing Conversation-Based Automated Program Repair Via Contrastive Test Case Pairs, Jiaolong Kong, Xiaofei Xie, Mingfei Cheng, Shangqing Liu, Xiaoning Du, Qi Guo

Research Collection School Of Computing and Information Systems

Automated Program Repair (APR) aims to automatically generate patches for rectifying software bugs. Recentstrides in Large Language Models (LLM), such as ChatGPT, have yielded encouraging outcomes in APR,especially within the conversation-driven APR framework. Nevertheless, the efficacy of conversation-drivenAPR is contingent on the quality of the feedback information. In this article, we propose ContrastRepair, anovel conversation-based APR approach that augments conversation-driven APR by providing LLMs withcontrastive test pairs. A test pair consists of a failing test and a passing test, which offer contrastive feedback tothe LLM. Our key insight is to minimize the difference between the generated passing test and the …


Harnessing Se Community Knowledge For Developer-Centric Code Intelligence, Chengran YANG 2025 Singapore Management University

Harnessing Se Community Knowledge For Developer-Centric Code Intelligence, Chengran Yang

Dissertations and Theses Collection (Open Access)

The integration of Large Language Models (LLMs), particularly those tailored for programming tasks—referred to as code LLMs—has created novel opportunities to enhance developer productivity. These advanced models automate routine and repetitive coding tasks, such as code generation and debugging, and enable faster prototyping and more efficient problem-solving. Despite these remarkable advantages, the current generation of code LLMs exhibits notable limitations that impact their practical effectiveness in real-world software engineering scenarios. These models frequently produce code that is inefficient or suboptimal in runtime performance, demonstrate opaque reasoning processes, and struggle to adapt effectively to diverse developer contexts and specific requirements. Moreover, …


Probabilistic Prototype Calibration Of Vision-Language Models For Generalized Few-Shot Semantic Segmentation, Jie LIU, Jiayi SHEN, Pan ZHOU, Jan-Jakob SONKE, Stratis GAVVES 2025 Singapore Management University

Probabilistic Prototype Calibration Of Vision-Language Models For Generalized Few-Shot Semantic Segmentation, Jie Liu, Jiayi Shen, Pan Zhou, Jan-Jakob Sonke, Stratis Gavves

Research Collection School Of Computing and Information Systems

Generalized Few-Shot Semantic Segmentation (GFSS) aims to extend a segmentation model to novel classes with only a few annotated examples while maintaining performance on base classes. Recently, pretrained vision-language models (VLMs) such as CLIP have been leveraged in GFSS to improve generalization on novel classes through multi-modal prototypes learning. However, existing prototype-based methods are inherently deterministic, limiting the adaptability of learned prototypes to diverse samples, particularly for novel classes with scarce annotations. To address this, we propose FewCLIP, a probabilistic prototype calibration framework over multi-modal prototypes from the pretrained CLIP, thus providing more adaptive prototype learning for GFSS. Specifically, FewCLIP …


Polymind: Parallel Visual Diagramming With Large Language Models To Support Prewriting Through Microtasks, Qian WAN, Jiannan LI, Huanchen WANG, Zhicong LU 2025 Singapore Management University

Polymind: Parallel Visual Diagramming With Large Language Models To Support Prewriting Through Microtasks, Qian Wan, Jiannan Li, Huanchen Wang, Zhicong Lu

Research Collection School Of Computing and Information Systems

Prewriting is the process of generating and organising ideas before a first draft. It consists of a combination of informal, iterative, and semi-structured strategies such as visual diagramming, which poses a challenge for collaborating with large language models (LLMs) in a turn-taking conversational manner. We present Polymind, a visual diagramming tool that leverages multiple LLM-powered agents to support prewriting. The system features a parallel collaboration workflow in place of the turn-taking conversational interactions. It defines multiple ''microtasks'' to simulate group collaboration scenarios such as collaborative writing and group brainstorming. Instead of repetitively prompting a chatbot for various purposes, Polymind enables …


Boosting Chart-To-Code Generation In Mllm Via Dual Preference-Guided Refinement, Zhihan ZHANG, Yixin CAO, Lizi LIAO 2025 Singapore Management University

Boosting Chart-To-Code Generation In Mllm Via Dual Preference-Guided Refinement, Zhihan Zhang, Yixin Cao, Lizi Liao

Research Collection School Of Computing and Information Systems

Translating chart images into executable plotting scripts-referred to as the chart-to-code generation task-requires Multimodal Large Language Models (MLLMs) to perform fine-grained visual parsing, precise code synthesis, and robust cross-modal reasoning. However, this task is inherently under-constrained: multiple valid code implementations can produce the same visual chart, and evaluation must consider both code correctness and visual fidelity across diverse dimensions. This makes it difficult to learn accurate and generalizable mappings through standard supervised fine-tuning. To address these challenges, we propose a dual preference-guided refinement framework that combines a feedback-driven, dual-modality reward mechanism with iterative preference learning. Our approach introduces a structured …


Auxiliary Prompt Tuning Of Vision‑Language Models For Few‑Shot Out‑Of‑Distribution Detection, Wenjun MIAO, Guansong PANG, Zihan WANG, Jin ZHENG, Xiao BAI 2025 Singapore Management University

Auxiliary Prompt Tuning Of Vision‑Language Models For Few‑Shot Out‑Of‑Distribution Detection, Wenjun Miao, Guansong Pang, Zihan Wang, Jin Zheng, Xiao Bai

Research Collection School Of Computing and Information Systems

Recent advancements in CLIP-based out-of-distribution (OOD) detection have shown promising results via regularization on prompt tuning, leveraging background features extracted from a few in-distribution (ID) samples as proxies for OOD features.However, these methods suffer from an inherent limitation: a lack of diversity in the extracted OOD features from the few-shot ID data.To address this issue, we propose to leverage external datasets as auxiliary outlier data (i.e., pseudo OOD samples) to extract rich, diverse OOD features, with the features from not only background regions but also foreground object regions, thereby supporting more discriminative prompt tuning for OOD detection. We further introduce …


Guiding Multiple Remote Users In Physical Tasks With Language-Driven Robotic Telepresence, Ruyi LI, Jingfei GUO, Xinyi ZHANG, Xuji ZHANG, Zeqing LI, Jiannan LI, Jiangtao GONG 2025 Singapore Management University

Guiding Multiple Remote Users In Physical Tasks With Language-Driven Robotic Telepresence, Ruyi Li, Jingfei Guo, Xinyi Zhang, Xuji Zhang, Zeqing Li, Jiannan Li, Jiangtao Gong

Research Collection School Of Computing and Information Systems

Remote assistance through robotic telepresence could involve both control and memory challenges, particularly in one expert to multiple workers situation. In this work, we proposed a novelty language-driven interface to facilitate remote collaboration through telepresence robots. Through operations and maintenance expert interviews and a scenario simulation study, we identified key pain points in executing one-expert-multiple-workers remote guidance using the telepresence robot and proposed two design goals, which together consist of five sub-design goals with corresponding features. These features were integrated into a standard telepresence robot, resulting in the development of a Collaborative LLM-based Embodied Assistant Robot, named CLEAR Robot. A …


Automatic Generation Of Introductory Programming Exercises With Large Language Models, Nguyen Binh Duong TA, Hua Gia Phuc NGUYEN, GOTTIPATI Swapna 2025 Singapore Management University

Automatic Generation Of Introductory Programming Exercises With Large Language Models, Nguyen Binh Duong Ta, Hua Gia Phuc Nguyen, Gottipati Swapna

Research Collection School Of Computing and Information Systems

Despite recent advances in code generation made possible by large language models (LLMs), programming is still an essential skill that computing students need to master now and in the foreseeable future. In learning programming, frequent practices with exercises set at an appropriate difficulty and knowledge level is of crucial importance for students. However, it’s not a trivial task for instructors to create many good quality exercises customized for each student. Programming problems found on Internet sources such as LeetCode are mostly too challenging for novice programmers with no prior coding knowledge. Recent work in AI-enabled education has been leveraging LLMs …


Exploring Parameter-Efficient Fine-Tuning Techniques For Code Generation With Large Language Models, Martin WEYSSOW, Xin ZHOU, Kisub KIM, David LO, Houari A. SAHRAOUI 2025 Singapore Management University

Exploring Parameter-Efficient Fine-Tuning Techniques For Code Generation With Large Language Models, Martin Weyssow, Xin Zhou, Kisub Kim, David Lo, Houari A. Sahraoui

Research Collection School Of Computing and Information Systems

Large language models (LLMs) demonstrate impressive capabilities to generate accurate code snippets given natural language intents in a zero-shot manner, i.e., without the need for specific fine-tuning. While prior studies have highlighted the advantages of fine-tuning LLMs, this process incurs high computational costs, making it impractical in resource-scarce environments, particularly for models with billions of parameters. To address these challenges, previous research explored in-context learning (ICL) and retrieval-augmented generation (RAG) as strategies to guide the LLM generative process with task-specific prompt examples. However, ICL and RAG introduce inconveniences, such as the need for designing contextually relevant prompts and the absence …


Detecting Defi Fraud With A Graph-Transformer Language Model, Wei MA, Junjie SHI, Jiaxi QIU, Cong WU, Jing CHEN, Lingxiao JIANG, Shangqing LIU, Yang LIU, Yang XIANG 2025 Singapore Management University

Detecting Defi Fraud With A Graph-Transformer Language Model, Wei Ma, Junjie Shi, Jiaxi Qiu, Cong Wu, Jing Chen, Lingxiao Jiang, Shangqing Liu, Yang Liu, Yang Xiang

Research Collection School Of Computing and Information Systems

With the rapid development of blockchain technology, the widespread adoption of smart contracts—particularly in decentralized finance (DeFi) applications—has introduced significant security challenges, such as reentrancy attacks, phishing, and Sybil attacks. To address these issues, we propose a novel model called TrxGNNBERT, which combines Graph Neural Network (GNN) and the Transformer architecture to effectively handle both graph-structured and textual data. This combination enhances the detection of suspicious transactions and accounts on blockchain platforms like Ethereum. TrxGNNBERT was pre-trained using a masked language model (MLM) on a dataset of 60,000 Ethereum transactions by randomly masking the attributes of nodes and edges, thereby …


How Developers Use Type-System Related Programming Language Features, Samuel W. Flint 2025 University of Nebraska-Lincoln

How Developers Use Type-System Related Programming Language Features, Samuel W. Flint

School of Computing: Dissertations, Theses, and Student Research

Optional type annotations are a popular feature of programming languages that allow developers to omit explicit type information in code while, in some cases, retaining many of the benefits of static typing, such as in-code documentation, improved detection of type errors, or enforcement of code properties. However, how developers use and understand optional type annotations is not clear. The focus of this dissertation is to understand the use and comprehension of optional type annotations.

Optional type annotations are examined through four lenses: first, by examining the evolution of usage in a statically typed programming language (Kotlin, the default language for …


How Developers Use Type-System Related Programming Language Features, Samuel W. Flint 2025 University of Nebraska-Lincoln

How Developers Use Type-System Related Programming Language Features, Samuel W. Flint

Dissertations and Doctoral Documents, University of Nebraska-Lincoln, 2023–

Optional type annotations are a popular feature of programming languages that allow developers to omit explicit type information in code while, in some cases, retaining many of the benefits of static typing, such as in-code documentation, improved detection of type errors, or enforcement of code properties. However, how developers use and understand optional type annotations is not clear. The focus of this dissertation is to understand the use and comprehension of optional type annotations.

Optional type annotations are examined through four lenses: first, by examining the evolution of usage in a statically typed programming language (Kotlin, the default language for …


Introduction To C++ (Volume I), Hussam Ghunaim Ph.D. 2025 Fort Hays State University

Introduction To C++ (Volume I), Hussam Ghunaim Ph.D.

All Open Educational Resources

This book is written as an Open Education Resource (OER) to replace expensive commercial materials currently used at the Department of Computer Science at Fort Hays State University. It has two volumes corresponding to the CSCI 121 and CSCI 221 courses. These courses are developed to introduce college freshmen students to Object-Oriented Programming utilizing C++. The author tried to bridge the gap in the current programming textbooks by avoiding lengthy and, on many occasions, unnecessary details. This book’s main feature is to present the discussed principles in the least wording possible while providing adequate examples and exercises to reinforce students’ …


Evaluating And Mitigating Linguistic Discrimination In Large Language Models: Perspectives On Safety Equity And Knowledge Equity, Guoliang DONG, Haoyu WANG, Jun SUN, Xinyu WANG 2025 Singapore Management University

Evaluating And Mitigating Linguistic Discrimination In Large Language Models: Perspectives On Safety Equity And Knowledge Equity, Guoliang Dong, Haoyu Wang, Jun Sun, Xinyu Wang

Research Collection School Of Computing and Information Systems

By training on text in various languages, large language models (LLMs) typically possess multilingual support and demonstrate remarkable capabilities in solving tasks described in different languages. However, LLMs can exhibit linguistic discrimination due to the uneven distribution of training data across languages. That is, LLMs are hard to keep the consistency of responses when faced with the same task but depicted in different languages. In this study, we first explore the consistency in the LLMs’ outputs responding to queries in various languages from two aspects: safety and quality. We conduct this analysis with two datasets (AdvBench and NQ) based on …


Mpo: Multilingual Safety Alignment Via Reward Gap Optimization, Weixiang ZHAO, Yulin HU, Yang DENG, Tongtong WU, Wenxuan ZHANG, Jiahe GUO, An ZHANG, Yanyan ZHAO, Bing QIN, Tat-Seng CHUA, Ting LIU 2025 Singapore Management University

Mpo: Multilingual Safety Alignment Via Reward Gap Optimization, Weixiang Zhao, Yulin Hu, Yang Deng, Tongtong Wu, Wenxuan Zhang, Jiahe Guo, An Zhang, Yanyan Zhao, Bing Qin, Tat-Seng Chua, Ting Liu

Research Collection School Of Computing and Information Systems

Large language models (LLMs) have become increasingly central to AI applications worldwide, necessitating robust multilingual safety alignment to ensure secure deployment across diverse linguistic contexts. Existing preference learning methods for safety alignment, such as RLHF and DPO, are primarily monolingual and struggle with noisy multilingual data. To address these limitations, we introduce Multilingual reward gaP Optimization (MPO), a novel approach that leverages the well-aligned safety capabilities of the dominant language (e.g., English) to improve safety alignment across multiple languages. MPO directly minimizes the reward gap difference between the dominant language and target languages, effectively transferring safety capabilities while preserving the …


Leveraging Reviewer Experience In Code Review Comment Generation, Hong Yi LIN, Patanamon THONGTANUNAM, Christoph TREUDE, Michael W. GODFREY, Chunhua LIU, Wachiraphan CHAROENWET 2025 Singapore Management University

Leveraging Reviewer Experience In Code Review Comment Generation, Hong Yi Lin, Patanamon Thongtanunam, Christoph Treude, Michael W. Godfrey, Chunhua Liu, Wachiraphan Charoenwet

Research Collection School Of Computing and Information Systems

Modern code review is a ubiquitous software quality assurance process aimed at identifying and resolving potential issues (e.g., functional, evolvability) within newly written code. Despite its effectiveness, the process demands large amounts of effort from the human reviewers involved. To help alleviate this workload, researchers have trained various deep learning based language models to imitate human reviewers in providing natural language code reviews for submitted code. Formally, this automation task is known as code review comment generation. Prior work has demonstrated improvements in code review comment generation by leveraging machine learning techniques and neural models, such as transfer learning and …


Computational Fact-Checking With Limited Resources, Fengzhu ZENG 2025 Singapore Management University

Computational Fact-Checking With Limited Resources, Fengzhu Zeng

Dissertations and Theses Collection (Open Access)

The rapid dissemination of information through online platforms has sparked widespread concern about the propagation of misinformation. Manual fact-checking by pro- fessional fact-checkers is time-consuming and lacks scalability to address the vast volume of daily information. Consequently, computational fact-checking, driven by automated techniques in natural language processing (NLP), has garnered interest as
a potential solution. However, computational fact-checking faces critical challenges limited resources, particularly due to the issues of data scarcity and computing resource constraints. One key challenge is data scarcity, which arises from the constant generation of new information and emerging events on social media. This scarcity manifests in …


Evowiki: Evaluating Llms On Evolving Knowledge, Wei TANG, Yixin CAO, Yang DENG, Jiahao YING, Bo WANG, Yizhe YANG, Yuyue ZHAO, Qi ZHANG, Xuanjing HUANG, Yu-Gang JIANG, Yong LIAO 2025 Singapore Management University

Evowiki: Evaluating Llms On Evolving Knowledge, Wei Tang, Yixin Cao, Yang Deng, Jiahao Ying, Bo Wang, Yizhe Yang, Yuyue Zhao, Qi Zhang, Xuanjing Huang, Yu-Gang Jiang, Yong Liao

Research Collection School Of Computing and Information Systems

Knowledge utilization is a critical aspect of LLMs, and understanding how they adapt to evolving knowledge is essential for their effective deployment. However, existing benchmarks are predominantly static, failing to capture the evolving nature of LLMs and knowledge, leading to inaccuracies and vulnerabilities such as contamination. In this paper, we introduce EvoWiki, an evolving dataset designed to reflect knowledge evolution by categorizing information into stable, evolved, and uncharted states. EvoWiki is fully auto-updated, enabling precise evaluation of continuously changing knowledge and newly released LLMs. Through experiments with Retrieval-Augmented Generation (RAG) and Continual Learning (CL), we evaluate how effectively LLMs adapt …


Focus: Evaluating Pre-Trained Vision-Language Models On Underspecification Reasoning, Kankan ZHOU, Yibin LAI, Kyriakos MOURATIDIS, Jing JIANG 2025 Singapore Management University

Focus: Evaluating Pre-Trained Vision-Language Models On Underspecification Reasoning, Kankan Zhou, Yibin Lai, Kyriakos Mouratidis, Jing Jiang

Research Collection School Of Computing and Information Systems

Humans possess a remarkable ability to interpret underspecified ambiguous statements by inferring their meanings from contexts such as visual inputs. This ability, however, may not be as developed in recent pre-trained visionlanguage models (VLMs). In this paper, we introduce a novel probing dataset called FOCUS to evaluate whether state-of-the-art VLMs have this ability. FOCUS consists of underspecified sentences paired with image contexts and carefully designed probing questions. Our experiments reveal that VLMs still fall short in handling underspecification even when visual inputs that can help resolve the ambiguities are available. To further support research in underspecification, FOCUS will be released …


Equivalence And Similarity Refutation For Probabilistic Programs, Krishnendu CHATTERJEE, Ehsan Kafshdar GOHARSHADY, Petr NOVOTNÝ, Dorde ZIKELIC 2025 Singapore Management University

Equivalence And Similarity Refutation For Probabilistic Programs, Krishnendu Chatterjee, Ehsan Kafshdar Goharshady, Petr Novotný, Dorde Zikelic

Research Collection School Of Computing and Information Systems

We consider the problems of statically refuting equivalence and similarity of output distributions defined by a pair of probabilistic programs. Equivalence and similarity are two fundamental relational properties of probabilistic programs that are essential for their correctness both in implementation and in compilation. In this work, we present a new method for static equivalence and similarity refutation. Our method refutes equivalence and similarity by computing a function over program outputs whose expected value with respect to the output distributions of two programs is different. The function is computed simultaneously with an upper expectation supermartingale and a lower expectation submartingale for …


Digital Commons powered by bepress