Open Access. Powered by Scholars. Published by Universities.®
Programming Languages and Compilers Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Artificial Intelligence and Robotics (51)
- Software Engineering (28)
- Other Computer Sciences (12)
- Databases and Information Systems (8)
- Education (7)
-
- Engineering (7)
- Theory and Algorithms (7)
- Computer Engineering (5)
- Cybersecurity (5)
- Data Science (5)
- Graphics and Human Computer Interfaces (5)
- OS and Networks (5)
- Social and Behavioral Sciences (5)
- Systems Architecture (5)
- Educational Technology (3)
- Information Security (3)
- Aeronautical Vehicles (2)
- Aerospace Engineering (2)
- Applied Mathematics (2)
- Arts and Humanities (2)
- Digital Communications and Networking (2)
- Educational Administration and Supervision (2)
- Electrical and Computer Engineering (2)
- Instructional Media Design (2)
- Navigation, Guidance, Control and Dynamics (2)
- Numerical Analysis and Scientific Computing (2)
- Other Aerospace Engineering (2)
- Institution
-
- Singapore Management University (67)
- University of South Alabama (8)
- City University of New York (CUNY) (7)
- Fort Hays State University (3)
- St. Mary's University (3)
-
- University of Nebraska - Lincoln (3)
- Chapman University (2)
- Loyola University Chicago (2)
- Bellarmine University (1)
- Bentley University (1)
- Bryant University (1)
- Clemson University (1)
- East Tennessee State University (1)
- Embry-Riddle Aeronautical University (1)
- James Madison University (1)
- Journal of Police and Legal Sciences (1)
- Liberty University (1)
- Michigan Technological University (1)
- The University of Akron (1)
- University of Minnesota Morris Digital Well (1)
- University of Nevada, Las Vegas (1)
- Virginia Commonwealth University (1)
- West Virginia University (1)
- Keyword
-
- Deep learning (6)
- Large Language Models (6)
- Refactoring (6)
- Imperative programs (5)
- Large language model (4)
-
- Large language models (4)
- Python (4)
- Graphs (3)
- Alignment (2)
- Artificial Intelligence (2)
- Code generation (2)
- Computer science (2)
- Graph execution (2)
- Hybrid programming paradigms (2)
- Large Language Model (2)
- Program verification (2)
- Programming (2)
- Software Engineering (2)
- AI (1)
- Accessibility testing (1)
- Advanced Placement (1)
- Adversarial Machine Learning (1)
- Ambiguity (1)
- Android app (1)
- Antipattern (1)
- Artifical Intelligence (1)
- Attack graph construction (1)
- Audio Captioning (1)
- Automated geospatial model (1)
- Autonomous Agents (1)
- Publication
-
- Research Collection School Of Computing and Information Systems (62)
- Publications and Research (6)
- Dissertations and Theses Collection (Open Access) (5)
- Shelby Hall Graduate Research Forum Posters (4)
- Computer Science: Faculty Publications and Other Works (2)
-
- Honors Theses (2)
- Posters - 2025 (2)
- Shelby Hall Graduate Research Forum Presentations (2)
- 2025 (1)
- All Open Educational Resources (1)
- Beyond: Undergraduate Research Journal (1)
- Department of Teaching, Learning, and Teacher Education: Faculty Publications (1)
- Dissertations and Doctoral Documents, University of Nebraska-Lincoln, 2023– (1)
- Dissertations, Master's Theses and Master's Reports (1)
- Electrical Engineering and Computer Science (MS) Theses (1)
- Electronic Theses and Dissertations (1)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (1)
- Honors Projects in Data Science (1)
- ICRE Publications (1)
- James Madison Undergraduate Research Journal (JMURJ) (1)
- Journal of Computer Science Integration (1)
- Journal of Police and Legal Sciences (1)
- Journal of South Carolina Water Resources (1)
- Library Faculty Research (1)
- Master's Theses or Doctor of Nursing Practice (1)
- Open Educational Resources (1)
- SACAD: Scholarly Activities (1)
- School of Computing: Dissertations, Theses, and Student Research (1)
- Senior Honors Theses (1)
- Systems Manuals - 2026 (1)
- Publication Type
Articles 1 - 30 of 110
Full-Text Articles in Programming Languages and Compilers
Topic Modeling And Culturomic Analysis Of 30,000 Books Over 100 Years Using Gensim, Michael A. Freeman
Topic Modeling And Culturomic Analysis Of 30,000 Books Over 100 Years Using Gensim, Michael A. Freeman
Electronic Theses and Dissertations
This thesis explores the cultural influence of historical events on English-language fiction published between 1820 and 1929. Using a corpus of 30,256 digitized books from Project Gutenberg, Latent Dirichlet Allocation (LDA) topic modeling was applied to identify recurring themes across eleven decades. The study sought to determine whether historically significant events could be detected within fictional narratives. One clear instance emerged: Napoleon Bonaparte and the Napoleonic Wars appeared explicitly in the 1820s corpus. Beyond this, several thematic patterns were observed—such as maritime language in the 1840s, national identity in the 1880s, and youth-oriented dialogue in the early 20th century—that plausibly …
Backdoorllm: A Comprehensive Benchmark For Backdoor Attacks And Defenses On Large Language Models, Yige Li, Hanxun Huang, Yunhan Zhao, Xingjun Ma, Jun Sun
Backdoorllm: A Comprehensive Benchmark For Backdoor Attacks And Defenses On Large Language Models, Yige Li, Hanxun Huang, Yunhan Zhao, Xingjun Ma, Jun Sun
Research Collection School Of Computing and Information Systems
Generative large language models (LLMs) have achieved state-of-the-art results on a wide range of tasks, yet they remain susceptible to backdoor attacks: carefully crafted triggers in the input can manipulate the model to produce adversaryspecified outputs. While prior research has predominantly focused on backdoor risks in vision and classification settings, the vulnerability of LLMs in open-ended text generation remains underexplored. To fill this gap, we introduce BackdoorLLM1 , the first comprehensive benchmark for systematically evaluating backdoor threats in text-generation LLMs. BackdoorLLM provides: (i) a unified repository of benchmarks with a standardized training and evaluation pipeline; (ii) a diverse suite of …
A Learning‑Augmented Dynamic Programming Approach For Orienteering Problem With Time Windows, Guansheng Peng, Lining Xing, Fuyan Song Ma, Aldy Gunawan, Aldy Gunawan
A Learning‑Augmented Dynamic Programming Approach For Orienteering Problem With Time Windows, Guansheng Peng, Lining Xing, Fuyan Song Ma, Aldy Gunawan, Aldy Gunawan
Research Collection School Of Computing and Information Systems
Recent years have witnessed a surge of interest in solving combinatorial optimization problems (COPs) using machine learning techniques. Motivated by this trend, we propose a learning-augmented exact approach for tackling an NP-hard COP, the Orienteering Problem with Time Windows, which aims to maximize the total score collected by visiting a subset of vertices in a graph within their time windows. Traditional exact algorithms rely heavily on domain expertise and meticulous design, making it hard to achieve further improvements. By leveraging deep learning models to learn effective relaxations of problem restrictions from data, our approach enables significant performance gains in an …
The Rise Of Parameter Specialization For Knowledge Storage In Large Language Models, Yihuai Hong, Yiran Zhao, Wei Tang, Yang Deng, Yu Rong, Wenxuan Zhang
The Rise Of Parameter Specialization For Knowledge Storage In Large Language Models, Yihuai Hong, Yiran Zhao, Wei Tang, Yang Deng, Yu Rong, Wenxuan Zhang
Research Collection School Of Computing and Information Systems
Over time, a growing wave of large language models from various series has been introduced to the community. Researchers are striving to maximize the performance of language models with constrained parameter sizes. However, from a microscopic perspective, there has been limited research on how to better store knowledge in model parameters, particularly within MLPs, to enable more effective utilization of this knowledge by the model. In this work, we analyze twenty publicly available open-source large language models to investigate the relationship between their strong performance and the way knowledge is stored in their corresponding MLP parameters. Our findings reveal that …
A Partition Cover Approach To Tokenization, Jia Peng Lim, Shawn Tan, Davin Choo, Hady Wirawan Lauw
A Partition Cover Approach To Tokenization, Jia Peng Lim, Shawn Tan, Davin Choo, Hady Wirawan Lauw
Research Collection School Of Computing and Information Systems
Tokenization is the process of encoding strings into tokens of a fixed vocabulary size, and is widely utilized in Natural Language Processing applications. The leading tokenization algorithm today is Byte Pair Encoding (BPE), which formulates the tokenization problem as a compression problem and tackles it by performing sequences of merges. In this work, we formulate tokenization as an optimization objective, show that it is NP-hard via a simple reduction from vertex cover, and propose a polynomial-time greedy algorithm GreedTok. Our formulation naturally relaxes to the well-studied weighted maximum coverage problem which has a simple -approximation algorithm GreedWMC. Through empirical evaluations …
When Less Language Is More: Language-Reasoning Disentanglement Makes Llms Better Multilingual Reasoners, Weixiang Zhao, Jiahe Guo, Yang Deng, Tongtong Wu, Wenxuan Zhang, Yulin Hu, Xingyu Sui, Yanyan Zhao, Wanxiang Che, Bing Qin, Tat-Seng Chua, Ting Liu
When Less Language Is More: Language-Reasoning Disentanglement Makes Llms Better Multilingual Reasoners, Weixiang Zhao, Jiahe Guo, Yang Deng, Tongtong Wu, Wenxuan Zhang, Yulin Hu, Xingyu Sui, Yanyan Zhao, Wanxiang Che, Bing Qin, Tat-Seng Chua, Ting Liu
Research Collection School Of Computing and Information Systems
Multilingual reasoning remains a significant challenge for large language models (LLMs), with performance disproportionately favoring high-resource languages. Drawing inspiration from cognitive neuroscience, which suggests that human reasoning functions largely independently of language processing, we hypothesize that LLMs similarly encode reasoning and language as separable components that can be disentangled to enhance multilingual reasoning. To evaluate this, we perform a causal intervention by ablating language-specific representations at inference time. Experiments on 10 open-weight LLMs spanning 11 typologically diverse languages show that this language-specific ablation consistently boosts multilingual reasoning performance. Layer-wise analyses further confirm that language and reasoning representations can be effectively …
Speculative Automated Refactoring Of Imperative Deep Learning Programs To Graph Execution, Raffi Khatchadourian Ph.D., Tatiana Castro Vélez, Mehdi Bagherzadeh, Nan Jia, Anita Raja
Speculative Automated Refactoring Of Imperative Deep Learning Programs To Graph Execution, Raffi Khatchadourian Ph.D., Tatiana Castro Vélez, Mehdi Bagherzadeh, Nan Jia, Anita Raja
Publications and Research
Efficiency is essential to support ever-growing datasets, especially for Deep Learning (DL) systems. DL frameworks have traditionally embraced deferred execution-style DL code---supporting symbolic, graph-based Deep Neural Network (DNN) computation. While scalable, such development is error-prone, non-intuitive, and difficult to debug. Consequently, more natural, imperative DL frameworks encouraging eager execution have emerged but at the expense of run-time performance. Though hybrid approaches aim for the "best of both worlds," using them effectively requires subtle considerations. Our key insight is that, while DL programs typically execute sequentially, hybridizing imperative DL code resembles parallelizing sequential code in traditional systems. Inspired by this, we …
Speculative Automated Refactoring Of Imperative Deep Learning Programs To Graph Execution, Raffi Khatchadourian Ph.D., Tatiana Castro Vélez, Mehdi Bagherzadeh, Nan Jia, Anita Raja
Speculative Automated Refactoring Of Imperative Deep Learning Programs To Graph Execution, Raffi Khatchadourian Ph.D., Tatiana Castro Vélez, Mehdi Bagherzadeh, Nan Jia, Anita Raja
Publications and Research
Efficiency is essential to support ever-growing datasets, especially for Deep Learning (DL) systems. DL frameworks have traditionally embraced deferred execution-style DL code---supporting symbolic, graph-based Deep Neural Network (DNN) computation. While scalable, such development is error-prone, non-intuitive, and difficult to debug. Consequently, more natural, imperative DL frameworks encouraging eager execution have emerged but at the expense of run-time performance. Though hybrid approaches aim for the "best of both worlds," using them effectively requires subtle considerations. Our key insight is that, while DL programs typically execute sequentially, hybridizing imperative DL code resembles parallelizing sequential code in traditional systems. Inspired by this, we …
Spatially Mapped Statewide Estimated Potential Evapotranspiration Using An Efficient Surface Interpolation Method: A Case Study Of South Carolina, Sudhanshu S. Panda, Devendra M. Amatya, Ka Kit Liu, Augustine Muwamba, Timothy J. Callahan
Spatially Mapped Statewide Estimated Potential Evapotranspiration Using An Efficient Surface Interpolation Method: A Case Study Of South Carolina, Sudhanshu S. Panda, Devendra M. Amatya, Ka Kit Liu, Augustine Muwamba, Timothy J. Callahan
Journal of South Carolina Water Resources
Potential evapotranspiration (PET) exhibits substantial spatial and temporal variability across large landscapes, necessitating site-specific estimation for accurate environmental and water resource assessments. However, obtaining PET or ET data for specific locations across an entire state remains challenging due to the limited number of weather stations and associated environmental datasets. This study aimed to develop an automated geospatial modeling framework to map PET distribution across South Carolina, USA, using PET estimated by the temperature-based Hargreaves–Samani (H–S) method with daily weather data from 59 NOAA stations. Because the accuracy of spatial interpolation depends on both the target variable and the desired spatial …
Speculative Automated Refactoring Of Imperative Deep Learning Programs To Graph Execution, Raffi T. Khatchadourian Ph.D., Tatiana Castro Vélez, Mehdi Bagherzadeh, Nan Jia, Anita Raja
Speculative Automated Refactoring Of Imperative Deep Learning Programs To Graph Execution, Raffi T. Khatchadourian Ph.D., Tatiana Castro Vélez, Mehdi Bagherzadeh, Nan Jia, Anita Raja
Publications and Research
Efficiency is essential to support ever-growing datasets, especially for Deep Learning (DL) systems. DL frameworks have traditionally embraced deferred execution-style DL code---supporting symbolic, graph-based Deep Neural Network (DNN) computation. While scalable, such development is error-prone, non-intuitive, and difficult to debug. Consequently, more natural, imperative DL frameworks encouraging eager execution have emerged but at the expense of run-time performance. Though hybrid approaches aim for the "best of both worlds," using them effectively requires subtle considerations. Our key insight is that, while DL programs typically execute sequentially, hybridizing imperative DL code resembles parallelizing sequential code in traditional systems. Inspired by this, we …
Do Code Semantics Help? A Comprehensive Study On Execution Trace-Based Information For Code Large Language Models, Jian Wang, Xiaofei Xie, Qiang Hu, Shangqing Liu, Yi Li
Do Code Semantics Help? A Comprehensive Study On Execution Trace-Based Information For Code Large Language Models, Jian Wang, Xiaofei Xie, Qiang Hu, Shangqing Liu, Yi Li
Research Collection School Of Computing and Information Systems
Code Large Language Models (Code LLMs) have opened a new era in programming with their impressive capabilities. However, recent research has revealed critical limitations in their ability to reason about runtime behavior and understand the actual functionality of programs, which poses significant challenges for their post-training and practical deployment. Specifically, Code LLMs encounter two principal issues: (1) a lack of proficiency in reasoning about program execution behavior, as they struggle to interpret what programs actually do during runtime, and (2) inconsistent and fragmented representation of semantic information, such as execution traces, across existing methods, which hinders their ability to generalize …
Mmlu-Prox: A Multilingual Benchmark For Advanced Large Language Model Evaluation, Weihao Xuan, Et. Al.
Mmlu-Prox: A Multilingual Benchmark For Advanced Large Language Model Evaluation, Weihao Xuan, Et. Al.
Research Collection School Of Computing and Information Systems
Existing large language model (LLM) evaluation benchmarks primarily focus on English, while current multilingual tasks lack parallel questions that specifically assess cross-lingual reasoning abilities. This dual limitation makes it challenging to assess LLMs’ performance in the multilingual setting comprehensively. To fill this gap, we introduce MMLU-ProX, a comprehensive benchmark covering 29 languages, built on an English benchmark. Each language version consists of 11,829 identical questions, enabling direct cross-lingual comparisons. Additionally, to meet efficient evaluation needs, we provide a lite version containing 658 questions per language. To ensure the high quality of MMLU-ProX, we employ a rigorous development process that involves …
Chain Of Strategy Optimization Makes Large Language Models Better Emotional Supporter, Weixiang Zhao, Xingyu Sui, Xinyang Han, Yang Deng, Yulin Hu, Jiahe Guo, Libo Qin, Qianyun Du, Shijin Wang, Yanyan Zhao, Bing Qin, Ting Liu
Chain Of Strategy Optimization Makes Large Language Models Better Emotional Supporter, Weixiang Zhao, Xingyu Sui, Xinyang Han, Yang Deng, Yulin Hu, Jiahe Guo, Libo Qin, Qianyun Du, Shijin Wang, Yanyan Zhao, Bing Qin, Ting Liu
Research Collection School Of Computing and Information Systems
The growing emotional stress in modern society has increased the demand for Emotional Support Conversations (ESC). While Large Language Models (LLMs) show promise for ESC, they face two key challenges: (1) low strategy selection accuracy, and (2) preference bias, limiting their adaptability to users’ emotional needs. Existing supervised fine-tuning (SFT) struggles to address these issues, as it rigidly trains models on single gold-standard responses without modeling nuanced strategy trade-offs. To overcome these limitations, we propose a novel two-stage framework that optimizes strategy selection preferences at each dialogue turn. We first leverage Monte Carlo Tree Search to construct ESC-Pro, a high-quality …
Envisioning Future Interactive Web Development: Editing Webpage With Natural Language, Truong Hai Dang, Jingyu Xiao, Yintong Huo
Envisioning Future Interactive Web Development: Editing Webpage With Natural Language, Truong Hai Dang, Jingyu Xiao, Yintong Huo
Research Collection School Of Computing and Information Systems
The evolution of web applications relies on iterative code modifications, a process that is traditionally manual and time-consuming. While Large Language Models (LLMs) can generate UI code, their ability to edit existing code from new design requirements (e.g., ”center the logo”) remains a challenge. This is largely due to the absence of large-scale, high-quality tuning data to align model performance with human expectations. In this paper, we introduce a novel, automated data generation pipeline that uses LLMs to synthesize a high-quality fine-tuning dataset for web editing, named Instruct4Edit. Our approach generates diverse instructions, applies the corresponding code modifications, and performs …
One Planner To Guide Them All! Learning Adaptive Conversational Planners For Goal-Oriented Dialogues, Huy Dao, Lizi Liao
One Planner To Guide Them All! Learning Adaptive Conversational Planners For Goal-Oriented Dialogues, Huy Dao, Lizi Liao
Research Collection School Of Computing and Information Systems
Goal-oriented dialogues, such as recommendation and negotiation, often require balancing multiple, conflicting objectives. Existing methods typically involve training separate models for specific combinations of objectives, leading to computational and scalability issues. In this work, we aim to develop a new dialogue policy method that can adapt to varying objective preferences at inference time without retraining. This raises several challenges in terms of both (1) optimization strategy and (2) knowledge utilization. To address these, we propose a novel learning framework, Preference Adaptive Dialogue Policy Planner (PADPP), for multi-objective goal-oriented dialogues. Specifically, to tackle the former, we introduce a novel policy optimization …
Context-Aware Hierarchical Taxonomy Generation For Scientific Papers Via Llm-Guided Multi-Aspect Clustering, Kun Zhu, Lizi Liao, Yuxuan Gu, Lei Huang, Xiaocheng Feng, Bing Qin
Context-Aware Hierarchical Taxonomy Generation For Scientific Papers Via Llm-Guided Multi-Aspect Clustering, Kun Zhu, Lizi Liao, Yuxuan Gu, Lei Huang, Xiaocheng Feng, Bing Qin
Research Collection School Of Computing and Information Systems
The rapid growth of scientific literature demands efficient methods to organize and synthesize research findings. Existing taxonomy construction methods, leveraging unsupervised clustering or direct prompting of large language models (LLMs), often lack coherence and granularity. We propose a novel context-aware hierarchical taxonomy generation framework that integrates LLM-guided multi-aspect encoding with dynamic clustering. Our method leverages LLMs to identify key aspects of each paper (e.g., methodology, dataset, evaluation) and generates aspect-specific paper summaries, which are then encoded and clustered along each aspect to form a coherent hierarchy. In addition, we introduce a new evaluation benchmark of 156 expert-crafted taxonomies encompassing 11.6k …
Distillcaps: Enhancing Audio-Language Alignment In Captioning Via Retrieval-Augmented Knowledge Distillation, Thinh Pham, Nghiem Diep, Lizi Liao, Binh Nguyen
Distillcaps: Enhancing Audio-Language Alignment In Captioning Via Retrieval-Augmented Knowledge Distillation, Thinh Pham, Nghiem Diep, Lizi Liao, Binh Nguyen
Research Collection School Of Computing and Information Systems
Automated audio captioning (AAC) benefits from incorporatingexternal context to interpret complex sounds, but doing so withretrieval-augmented generation (RAG) at inference is sometimesinfeasible due to data availability or incurs significant latency andcomplexity. We propose DistillCaps, a novel training-time frame-work that leverages RAG to guide knowledge distillation for im-proved audio-language alignment, while lessening the relianceon retrieval during inference. In our framework, a RAG-equippedteacher model retrieves relevant textual information (e.g., simi-lar captions) for each audio clip and uses it for training to gener-ate context-enriched captions. Simultaneously, a student model istrained to imitate this teacher, learning to produce high-qualitycaptions from audio alone. We further …
Harnessing Se Community Knowledge For Developer-Centric Code Intelligence, Chengran Yang
Harnessing Se Community Knowledge For Developer-Centric Code Intelligence, Chengran Yang
Dissertations and Theses Collection (Open Access)
The integration of Large Language Models (LLMs), particularly those tailored for programming tasks—referred to as code LLMs—has created novel opportunities to enhance developer productivity. These advanced models automate routine and repetitive coding tasks, such as code generation and debugging, and enable faster prototyping and more efficient problem-solving. Despite these remarkable advantages, the current generation of code LLMs exhibits notable limitations that impact their practical effectiveness in real-world software engineering scenarios. These models frequently produce code that is inefficient or suboptimal in runtime performance, demonstrate opaque reasoning processes, and struggle to adapt effectively to diverse developer contexts and specific requirements. Moreover, …
Probabilistic Prototype Calibration Of Vision-Language Models For Generalized Few-Shot Semantic Segmentation, Jie Liu, Jiayi Shen, Pan Zhou, Jan-Jakob Sonke, Stratis Gavves
Probabilistic Prototype Calibration Of Vision-Language Models For Generalized Few-Shot Semantic Segmentation, Jie Liu, Jiayi Shen, Pan Zhou, Jan-Jakob Sonke, Stratis Gavves
Research Collection School Of Computing and Information Systems
Generalized Few-Shot Semantic Segmentation (GFSS) aims to extend a segmentation model to novel classes with only a few annotated examples while maintaining performance on base classes. Recently, pretrained vision-language models (VLMs) such as CLIP have been leveraged in GFSS to improve generalization on novel classes through multi-modal prototypes learning. However, existing prototype-based methods are inherently deterministic, limiting the adaptability of learned prototypes to diverse samples, particularly for novel classes with scarce annotations. To address this, we propose FewCLIP, a probabilistic prototype calibration framework over multi-modal prototypes from the pretrained CLIP, thus providing more adaptive prototype learning for GFSS. Specifically, FewCLIP …
Boosting Chart-To-Code Generation In Mllm Via Dual Preference-Guided Refinement, Zhihan Zhang, Yixin Cao, Lizi Liao
Boosting Chart-To-Code Generation In Mllm Via Dual Preference-Guided Refinement, Zhihan Zhang, Yixin Cao, Lizi Liao
Research Collection School Of Computing and Information Systems
Translating chart images into executable plotting scripts-referred to as the chart-to-code generation task-requires Multimodal Large Language Models (MLLMs) to perform fine-grained visual parsing, precise code synthesis, and robust cross-modal reasoning. However, this task is inherently under-constrained: multiple valid code implementations can produce the same visual chart, and evaluation must consider both code correctness and visual fidelity across diverse dimensions. This makes it difficult to learn accurate and generalizable mappings through standard supervised fine-tuning. To address these challenges, we propose a dual preference-guided refinement framework that combines a feedback-driven, dual-modality reward mechanism with iterative preference learning. Our approach introduces a structured …
Auxiliary Prompt Tuning Of Vision‑Language Models For Few‑Shot Out‑Of‑Distribution Detection, Wenjun Miao, Guansong Pang, Zihan Wang, Jin Zheng, Xiao Bai
Auxiliary Prompt Tuning Of Vision‑Language Models For Few‑Shot Out‑Of‑Distribution Detection, Wenjun Miao, Guansong Pang, Zihan Wang, Jin Zheng, Xiao Bai
Research Collection School Of Computing and Information Systems
Recent advancements in CLIP-based out-of-distribution (OOD) detection have shown promising results via regularization on prompt tuning, leveraging background features extracted from a few in-distribution (ID) samples as proxies for OOD features.However, these methods suffer from an inherent limitation: a lack of diversity in the extracted OOD features from the few-shot ID data.To address this issue, we propose to leverage external datasets as auxiliary outlier data (i.e., pseudo OOD samples) to extract rich, diverse OOD features, with the features from not only background regions but also foreground object regions, thereby supporting more discriminative prompt tuning for OOD detection. We further introduce …
Contrastrepair: Enhancing Conversation-Based Automated Program Repair Via Contrastive Test Case Pairs, Jiaolong Kong, Xiaofei Xie, Mingfei Cheng, Shangqing Liu, Xiaoning Du, Qi Guo
Contrastrepair: Enhancing Conversation-Based Automated Program Repair Via Contrastive Test Case Pairs, Jiaolong Kong, Xiaofei Xie, Mingfei Cheng, Shangqing Liu, Xiaoning Du, Qi Guo
Research Collection School Of Computing and Information Systems
Automated Program Repair (APR) aims to automatically generate patches for rectifying software bugs. Recentstrides in Large Language Models (LLM), such as ChatGPT, have yielded encouraging outcomes in APR,especially within the conversation-driven APR framework. Nevertheless, the efficacy of conversation-drivenAPR is contingent on the quality of the feedback information. In this article, we propose ContrastRepair, anovel conversation-based APR approach that augments conversation-driven APR by providing LLMs withcontrastive test pairs. A test pair consists of a failing test and a passing test, which offer contrastive feedback tothe LLM. Our key insight is to minimize the difference between the generated passing test and the …
Polymind: Parallel Visual Diagramming With Large Language Models To Support Prewriting Through Microtasks, Qian Wan, Jiannan Li, Huanchen Wang, Zhicong Lu
Polymind: Parallel Visual Diagramming With Large Language Models To Support Prewriting Through Microtasks, Qian Wan, Jiannan Li, Huanchen Wang, Zhicong Lu
Research Collection School Of Computing and Information Systems
Prewriting is the process of generating and organising ideas before a first draft. It consists of a combination of informal, iterative, and semi-structured strategies such as visual diagramming, which poses a challenge for collaborating with large language models (LLMs) in a turn-taking conversational manner. We present Polymind, a visual diagramming tool that leverages multiple LLM-powered agents to support prewriting. The system features a parallel collaboration workflow in place of the turn-taking conversational interactions. It defines multiple ''microtasks'' to simulate group collaboration scenarios such as collaborative writing and group brainstorming. Instead of repetitively prompting a chatbot for various purposes, Polymind enables …
Exploring Parameter-Efficient Fine-Tuning Techniques For Code Generation With Large Language Models, Martin Weyssow, Xin Zhou, Kisub Kim, David Lo, Houari A. Sahraoui
Exploring Parameter-Efficient Fine-Tuning Techniques For Code Generation With Large Language Models, Martin Weyssow, Xin Zhou, Kisub Kim, David Lo, Houari A. Sahraoui
Research Collection School Of Computing and Information Systems
Large language models (LLMs) demonstrate impressive capabilities to generate accurate code snippets given natural language intents in a zero-shot manner, i.e., without the need for specific fine-tuning. While prior studies have highlighted the advantages of fine-tuning LLMs, this process incurs high computational costs, making it impractical in resource-scarce environments, particularly for models with billions of parameters. To address these challenges, previous research explored in-context learning (ICL) and retrieval-augmented generation (RAG) as strategies to guide the LLM generative process with task-specific prompt examples. However, ICL and RAG introduce inconveniences, such as the need for designing contextually relevant prompts and the absence …
Detecting Defi Fraud With A Graph-Transformer Language Model, Wei Ma, Junjie Shi, Jiaxi Qiu, Cong Wu, Jing Chen, Lingxiao Jiang, Shangqing Liu, Yang Liu, Yang Xiang
Detecting Defi Fraud With A Graph-Transformer Language Model, Wei Ma, Junjie Shi, Jiaxi Qiu, Cong Wu, Jing Chen, Lingxiao Jiang, Shangqing Liu, Yang Liu, Yang Xiang
Research Collection School Of Computing and Information Systems
With the rapid development of blockchain technology, the widespread adoption of smart contracts—particularly in decentralized finance (DeFi) applications—has introduced significant security challenges, such as reentrancy attacks, phishing, and Sybil attacks. To address these issues, we propose a novel model called TrxGNNBERT, which combines Graph Neural Network (GNN) and the Transformer architecture to effectively handle both graph-structured and textual data. This combination enhances the detection of suspicious transactions and accounts on blockchain platforms like Ethereum. TrxGNNBERT was pre-trained using a masked language model (MLM) on a dataset of 60,000 Ethereum transactions by randomly masking the attributes of nodes and edges, thereby …
Guiding Multiple Remote Users In Physical Tasks With Language-Driven Robotic Telepresence, Ruyi Li, Jingfei Guo, Xinyi Zhang, Xuji Zhang, Zeqing Li, Jiannan Li, Jiangtao Gong
Guiding Multiple Remote Users In Physical Tasks With Language-Driven Robotic Telepresence, Ruyi Li, Jingfei Guo, Xinyi Zhang, Xuji Zhang, Zeqing Li, Jiannan Li, Jiangtao Gong
Research Collection School Of Computing and Information Systems
Remote assistance through robotic telepresence could involve both control and memory challenges, particularly in one expert to multiple workers situation. In this work, we proposed a novelty language-driven interface to facilitate remote collaboration through telepresence robots. Through operations and maintenance expert interviews and a scenario simulation study, we identified key pain points in executing one-expert-multiple-workers remote guidance using the telepresence robot and proposed two design goals, which together consist of five sub-design goals with corresponding features. These features were integrated into a standard telepresence robot, resulting in the development of a Collaborative LLM-based Embodied Assistant Robot, named CLEAR Robot. A …
How Developers Use Type-System Related Programming Language Features, Samuel W. Flint
How Developers Use Type-System Related Programming Language Features, Samuel W. Flint
Dissertations and Doctoral Documents, University of Nebraska-Lincoln, 2023–
Optional type annotations are a popular feature of programming languages that allow developers to omit explicit type information in code while, in some cases, retaining many of the benefits of static typing, such as in-code documentation, improved detection of type errors, or enforcement of code properties. However, how developers use and understand optional type annotations is not clear. The focus of this dissertation is to understand the use and comprehension of optional type annotations.
Optional type annotations are examined through four lenses: first, by examining the evolution of usage in a statically typed programming language (Kotlin, the default language for …
How Developers Use Type-System Related Programming Language Features, Samuel W. Flint
How Developers Use Type-System Related Programming Language Features, Samuel W. Flint
School of Computing: Dissertations, Theses, and Student Research
Optional type annotations are a popular feature of programming languages that allow developers to omit explicit type information in code while, in some cases, retaining many of the benefits of static typing, such as in-code documentation, improved detection of type errors, or enforcement of code properties. However, how developers use and understand optional type annotations is not clear. The focus of this dissertation is to understand the use and comprehension of optional type annotations.
Optional type annotations are examined through four lenses: first, by examining the evolution of usage in a statically typed programming language (Kotlin, the default language for …
Introduction To C++ (Volume I), Hussam Ghunaim Ph.D.
Introduction To C++ (Volume I), Hussam Ghunaim Ph.D.
All Open Educational Resources
This book is written as an Open Education Resource (OER) to replace expensive commercial materials currently used at the Department of Computer Science at Fort Hays State University. It has two volumes corresponding to the CSCI 121 and CSCI 221 courses. These courses are developed to introduce college freshmen students to Object-Oriented Programming utilizing C++. The author tried to bridge the gap in the current programming textbooks by avoiding lengthy and, on many occasions, unnecessary details. This book’s main feature is to present the discussed principles in the least wording possible while providing adequate examples and exercises to reinforce students’ …
Computational Fact-Checking With Limited Resources, Fengzhu Zeng
Computational Fact-Checking With Limited Resources, Fengzhu Zeng
Dissertations and Theses Collection (Open Access)
The rapid dissemination of information through online platforms has sparked widespread concern about the propagation of misinformation. Manual fact-checking by pro- fessional fact-checkers is time-consuming and lacks scalability to address the vast volume of daily information. Consequently, computational fact-checking, driven by automated techniques in natural language processing (NLP), has garnered interest as
a potential solution. However, computational fact-checking faces critical challenges limited resources, particularly due to the issues of data scarcity and computing resource constraints. One key challenge is data scarcity, which arises from the constant generation of new information and emerging events on social media. This scarcity manifests in …