Elevating Automated Software Maintenance Tasks With Large Language Models,
2024
Singapore Management University
Elevating Automated Software Maintenance Tasks With Large Language Models, Xin Zhou
Dissertations and Theses Collection (Open Access)
Software engineering involves many tasks across different phases such as requirements, design, implementation, testing, and maintenance. Among them, software maintenance is a crucial phase, typically accounting for more than half of the software life cycle's duration.
To boost developer productivity, in recent years, numerous research endeavors in software engineering have sought to automate certain software maintenance tasks through the application of machine learning techniques.
Since 2020, the emergence of advanced Large Language Models (LLMs) of code has opened new avenues for enhancing automated solutions in software maintenance.
This dissertation presents a series of works aimed at advancing automated solutions for …
Mm‑Forecast: A Multimodal Approach To Temporal Event Forecasting With Large Language Models,
2024
Singapore Management University
Mm‑Forecast: A Multimodal Approach To Temporal Event Forecasting With Large Language Models, Haoxuan Li, Zhengmao Yang, Yunshan Ma, Yi Bin, Yang Yang, Tat-Seng Chua
Research Collection School Of Computing and Information Systems
We study an emerging and intriguing problem of multimodal temporal event forecasting with large language models. Compared to using text or graph modalities, the investigation of utilizing images for temporal event forecasting has not been fully explored, especially in the era of large language models (LLMs). To bridge this gap, we are particularly interested in two key questions of: 1) why images will help in temporal event forecasting, and 2) how to integrate images into the LLM-based forecasting framework. To answer these research questions, we propose to identify two essential functions that images play in the scenario of temporal event …
Review Of Current Trends In Information Technology Concerning Phonetic Similarity”,
2024
Faculty of Medical Sciences, Jabir ibn Hayyan University for Medical and Pharmaceutical Sciences, Najaf, Iraq
Review Of Current Trends In Information Technology Concerning Phonetic Similarity”, Zaid Rajih Mohammed, Ahmed H. Aliwy
Al-Bahir
With the increasing availability of textual information in various languages via the Internet in homes and companies through Internet and intranet services, there is an urgent need for the technologies and tools necessary to process this information, phonetic representation, and voice interaction. For example voice to voice machine translation need to phonetic mapping and similarity among the languages especially for names and foreign words. This one example of the importance of phonetic mapping and similarity. This article aims to describe, in detail, the recent surge in interest and advancements in phonetic similarity (PS), phonetic representation, and phonetic mapping researches. PS …
Wip: An Engaging Undergraduate Intro To Model Checking In Software Engineering Using Tla+,
2024
Loyola University Chicago
Wip: An Engaging Undergraduate Intro To Model Checking In Software Engineering Using Tla+, Konstantin Laufer, Gunda Mertin, George K. Thiruvathukal
Computer Science: Faculty Publications and Other Works
Background: In this Innovative Practice Work in Progress, we present our initial efforts to integrate formal methods, with a focus on model-checking specifications written in Temporal Logic of Actions (TLA+), into computer science education, targeting undergraduate juniors/seniors and graduate students. Many safety-critical systems and services crucially depend on correct and reliable behavior. Formal methods can play a key role in ensuring correct and safe system behavior, yet remain underutilized in educational and industry contexts.
Aims: We aim to (1) qualitatively assess the state of formal methods in computer science programs, (2) construct level-appropriate examples that could be included …
Interoperability In Deep Learning: A User Survey And Failure Analysis Of Onnx Model Converters,
2024
Purdue University
Interoperability In Deep Learning: A User Survey And Failure Analysis Of Onnx Model Converters, Purvish Jajal, Wenxin Jiang, Arav Tewari, Erik Kocinare, Joseph Woo, Anusha Sarraf, Yung-Hsiang Lu, George Thiruvathukal, James C. Davis
Computer Science: Faculty Publications and Other Works
Software engineers develop, fine-tune, and deploy deep learning (DL) models using a variety of development frameworks and runtime environments. DL model converters move models between frameworks and to runtime environments. Conversion errors compromise model quality and disrupt deployment. However, the failure characteristics of DL model converters are unknown, adding risk when using DL interoperability technologies. This paper analyzes failures in DL model converters. We survey software engineers about DL interoperability tools, use cases, and pain points (N=92). Then, we characterize failures in model converters associated with the main interoperability tool, ONNX (N=200 issues in PyTorch and TensorFlow). Finally, we formulate …
Recasting The Mould – Librarianship Of The Future: Leveraging Automation, Apis, And Ai,
2024
Singapore Management University
Recasting The Mould – Librarianship Of The Future: Leveraging Automation, Apis, And Ai, Samantha Seah
Research Collection Library
With leaps in artificial intelligence made in recent years redefining the information landscape and introducing new means of information production, librarianship also must evolve to include new literacies. One way librarians can equip and empower ourselves is by understanding the building blocks of how machines and automation work. Perhaps more important than learning specific programming languages, learning computational thinking provides us with more ways to spot and evaluate problems and devise solutions without extensive coding knowledge. My presentation will take the improvement of membership processing as an example using Power Automate, a low-code Microsoft tool mimicking block programming. The tool …
Sound And Complete Witnesses For Template-Based Verification Of Ltl Properties On Polynomial Programs,
2024
Singapore Management University
Sound And Complete Witnesses For Template-Based Verification Of Ltl Properties On Polynomial Programs, Krishnendu Chatterjee, Amir Goharshady, Ehsan Goharshady, Mehrdad Karrabi, Dorde Zikelic
Research Collection School Of Computing and Information Systems
We study the classical problem of verifying programs with respect to formal specifications given in the linear temporal logic (LTL). We first present novel sound and complete witnesses for LTL verification over imperative programs. Our witnesses are applicable to both verification (proving) and refutation (finding bugs) settings. We then consider LTL formulas in which atomic propositions can be polynomial constraints and turn our focus to polynomial arithmetic programs, i.e. programs in which every assignment and guard consists only of polynomial expressions. For this setting, we provide an efficient algorithm to automatically synthesize such LTL witnesses. Our synthesis procedure is both …
Style: Improving Domain Transferability Of Asking Clarification Questions In Large Language Model Powered Conversational Agents,
2024
Singapore Management University
Style: Improving Domain Transferability Of Asking Clarification Questions In Large Language Model Powered Conversational Agents, Yue Chen, Chen Huang, Yang Deng, Wenqiang Lei, Dingnan Jin, Jia Liu, Tat-Seng Chua
Research Collection School Of Computing and Information Systems
Equipping a conversational search engine with strategies regarding when to ask clarification questions is becoming increasingly important across various domains. Attributing to the context understanding capability of LLMs and their access to domain-specific sources of knowledge, LLM-based clarification strategies feature rapid transfer to various domains in a posthoc manner. However, they still struggle to deliver promising performance on unseen domains, struggling to achieve effective domain transferability. We take the first step to investigate this issue and existing methods tend to produce one-size-fits-all strategies across diverse domains, limiting their search effectiveness. In response, we introduce a novel method, called STYLE, to …
Chain-Of-Exemplar: Enhancing Distractor Generation For Multimodal Educational Question Generation,
2024
Singapore Management University
Chain-Of-Exemplar: Enhancing Distractor Generation For Multimodal Educational Question Generation, Haohao Luo, Yang Deng, Ying Shen, See-Kiong Ng, Tat-Seng Chua
Research Collection School Of Computing and Information Systems
Multiple-choice questions (MCQs) are important in enhancing concept learning and student engagement for educational purposes. Despite the multimodal nature of educational content, current methods focus mainly on text-based inputs and often neglect the integration of visual information. In this work, we study the problem of multimodal educational question generation, which aims at generating subject-specific educational questions with plausible yet incorrect distractors based on multimodal educational content. To tackle this problem, we introduce a novel framework, named Chain-of-Exemplar (CoE), which utilizes multimodal large language models (MLLMs) with Chain-of-Thought reasoning to improve the generation of challenging distractors. Furthermore, CoE leverages three-stage contextualized …
On The Multi-Turn Instruction Following For Conversational Web Agents,
2024
Singapore Management University
On The Multi-Turn Instruction Following For Conversational Web Agents, Yang Deng, Xuan Zhang, Wenxuan Zhang, Yifei Yuan, See-Kiong Ng, Tat-Seng Chua
Research Collection School Of Computing and Information Systems
Web agents powered by Large Language Models (LLMs) have demonstrated remarkable abilities in planning and executing multi-step interactions within complex web-based environments, fulfilling a wide range of web navigation tasks. Despite these advancements, the potential for LLM-powered agents to effectively engage with sequential user instructions in real-world scenarios has not been fully explored. In this work, we introduce a new task of Conversational Web Navigation, which necessitates sophisticated interactions that span multiple turns with both the users and the environment, supported by a specially developed dataset named Multi-Turn Mind2Web (MT-Mind2Web). To tackle the limited context length of LLMs and the …
Watme: Towards Lossless Watermarking Through Lexical Redundancy,
2024
Singapore Management University
Watme: Towards Lossless Watermarking Through Lexical Redundancy, Liang Chen, Yatao Bian, Yang Deng, Deng Cai, Shuaiyi Li, Peilin Zhao, Kam-Fai Wong
Research Collection School Of Computing and Information Systems
Text watermarking has emerged as a pivotal technique for identifying machine-generated text. However, existing methods often rely on arbitrary vocabulary partitioning during decoding to embed watermarks, which compromises the availability of suitable tokens and significantly degrades the quality of responses. This study assesses the impact of watermarking on different capabilities of large language models (LLMs) from a cognitive science lens. Our finding highlights a significant disparity; knowledge recall and logical reasoning are more adversely affected than language generation. These results suggest a more profound effect of watermarking on LLMs than previously understood. To address these challenges, we introduce Watermarking with …
Clamber: A Benchmark Of Identifying And Clarifying Ambiguous Information Needs In Large Language Models,
2024
Singapore Management University
Clamber: A Benchmark Of Identifying And Clarifying Ambiguous Information Needs In Large Language Models, Tong Zhang, Peixin Qin, Yang Deng, Chen Huang, Wenqiang Lei, Junhong Liu, Dingnan Jin, Hongru Liang, Tat-Seng Chua
Research Collection School Of Computing and Information Systems
Large language models (LLMs) are increasingly used to meet user information needs, but their effectiveness in dealing with user queries that contain various types of ambiguity remains unknown, ultimately risking user trust and satisfaction. To this end, we introduce CLAMBER, a benchmark for evaluating LLMs using a well-organized taxonomy. Building upon the taxonomy, we construct 12K high-quality data to assess the strengths, weaknesses, and potential risks of various off-the-shelf LLMs.Our findings indicate the limited practical utility of current LLMs in identifying and clarifying ambiguous user queries, even enhanced by chain-of-thought (CoT) and few-shot prompting. These techniques may result in overconfidence …
Self-Chats From Large Language Models Make Small Emotional Support Chatbot Better,
2024
Singapore Management University
Self-Chats From Large Language Models Make Small Emotional Support Chatbot Better, Zhonghua Zheng, Lizi Liao, Yang Deng, Libo Qin, Liqiang Nie
Research Collection School Of Computing and Information Systems
Large Language Models (LLMs) have shown strong generalization abilities to excel in various tasks, including emotion support conversations. However, deploying such LLMs like GPT-3 (175B parameters) is resource-intensive and challenging at scale. In this study, we utilize LLMs as “Counseling Teacher” to enhance smaller models’ emotion support response abilities, significantly reducing the necessity of scaling up model size. To this end, we first introduce an iterative expansion framework, aiming to prompt the large teacher model to curate an expansive emotion support dialogue dataset. This curated dataset, termed ExTES, encompasses a broad spectrum of scenarios and is crafted with meticulous strategies …
Reinforcement Tuning For Detecting Stances And Debunking Rumors Jointly With Large Language Models,
2024
Singapore Management University
Reinforcement Tuning For Detecting Stances And Debunking Rumors Jointly With Large Language Models, Ruichao Yang, Wei Gao, Jing Ma, Hongzhan Ling, Bo Wang
Research Collection School Of Computing and Information Systems
Learning multi-task models for jointly detecting stance and verifying rumors poses challenges due to the need for training data of stance at post level and rumor veracity at claim level, which are difficult to obtain. To address this issue, we leverage large language models (LLMs) as the foundation annotators for the joint stance detection (SD) and rumor verification (RV) tasks, dubbed as JSDRV. We introduce a novel reinforcement tuning framework to enhance the joint predictive capabilities of LLM-based SD and RV components. Specifically, we devise a policy for selecting LLM-annotated data at the two levels, employing a hybrid reward mechanism …
Larp: Language Audio Relational Pre‑Training For Cold‑Start Playlist Continuation,
2024
Singapore Management University
Larp: Language Audio Relational Pre‑Training For Cold‑Start Playlist Continuation, Rebecca Salganik, Xiaohao Liu, Yunshan Ma, Jian Kang, Tat‑Seng Chua
Research Collection School Of Computing and Information Systems
As online music consumption increasingly shifts towards playlist-based listening, the task of playlist continuation, in which an algorithm suggests songs to extend a playlist in a personalized and musically cohesive manner, has become vital to the success of music streaming services. Currently, many existing playlist continuation approaches rely on collaborative filtering methods to perform their recommendations. However, such methods will struggle to recommend songs that lack interaction data, an issue known as the cold-start problem. Current approaches to this challenge design complex mechanisms for extracting relational signals from sparse collaborative signals and integrating them into content representations. However, these approaches …
Introduction To Programming And Applied Analytics Using Python,
2024
Arkansas Tech University
Introduction To Programming And Applied Analytics Using Python, Matt Brown
ATU Faculty OER Books and Materials
This open electronic textbook is a collection of lecture notes, assignments, and additional background material for a junior level analytics course targeted for business students, it is free to use and copy. The text assumes readers have not had prior programming or computing courses, but have had at least one analytics course. The textbook differs from other textbooks because it serves a dual purpose, to first introduce to students the Python programming language and secondly to introduce analytics programming in Python. It is not meant to be a comprehensive book on the Python language or data analytics, rather a semester’s …
Development Of An Algorithm To Identify And Calculate The Amount Of File Slack On An Image Of A Given Drive,
2024
University of South Alabama
Development Of An Algorithm To Identify And Calculate The Amount Of File Slack On An Image Of A Given Drive, Nicholas Flynn
Honors Theses
As society increasingly relies on technology, the rates of cyber crime have been increasing at exponential rates. Cyber criminals are also discovering new ways to hide evidence of their crimes. This study develops a forensic analysis algorithm to evaluate the amount of file slack on an image of a drive. Slack space, leftover drive space on a disk sector after a file has been written, can be exploited to hide data. The algorithm aims to detect and calculate this slack space to help direct forensic investigations. The algorithm was evaluated on a population dataset of 100,000 files with random data …
Large Language Model Powered Agents For Information Retrieval,
2024
Singapore Management University
Large Language Model Powered Agents For Information Retrieval, An Zhang, Yang Deng, Yankai Lin, Xu Chen, Ji-Rong Wen, Tat-Seng Chua
Research Collection School Of Computing and Information Systems
The vital goal of information retrieval today extends beyond merely connecting users with relevant information they search for. It also aims to enrich the diversity, personalization, and interactivity of that connection, ensuring the information retrieval process is as seamless, beneficial, and supportive as possible in the global digital era. Current information retrieval systems often encounter challenges like a constrained understanding of queries, static and inflexible responses, limited personalization, and restricted interactivity. With the advent of large language models (LLMs), there's a transformative paradigm shift as we integrate LLM-powered agents into these systems. These agents bring forth crucial human capabilities like …
Who Wrote The Scientific News? Improving The Discernibility Of Llms To Human-Written Scientific News,
2024
Old Dominion University
Who Wrote The Scientific News? Improving The Discernibility Of Llms To Human-Written Scientific News, Dominik Soós
Computer Science Theses & Dissertations
Large Language Models (LLMs) have rapidly advanced the field of Natural Language Processing and become powerful tools for generating and evaluating scientific text. Although LLMs have demonstrated promising as evaluators for certain text generation tasks, there is still a gap until they are used as reliable text evaluators for general purposes. In this thesis project, I attempted to fill this gap by examining the discernibility of LLMs from human-written and LLM-generated scientific news. This research demonstrated that although it was relatively straightforward for humans to discern scientific news written by humans from scientific news generated by GPT-3.5 using basic prompts, …
How Do Preservice Teachers Learn To Teach Integrated Computational Thinking?: Evidence From Planning, Enactment, And Reflection,
2024
University of California, Santa Cruz
How Do Preservice Teachers Learn To Teach Integrated Computational Thinking?: Evidence From Planning, Enactment, And Reflection, Rachael Dektor, Samuel Severance, Kip Téllez
Journal of Computer Science Integration
This study examines pre-service teachers’ (PSTs) beliefs and understandings about computational thinking (CT) integration and lesson implementation over time. Utilizing a design-based research approach, 3 PSTs led the co-design of integrated CT lessons with support from researchers and enacted these CT integrated lessons with K-5 students. All PSTs participated in a whole-group CT workshop and engaged in one-on-one lesson design sessions with a researcher. We utilized a grounded theory approach to qualitatively analyze pre-surveys, semi-structured interviews, and video data of three PSTs enacting their lessons. We found that PSTs’ initial beliefs about CT instruction – including the importance of it …
