The Amethyst Compiler,
2024
The University of Akron
The Amethyst Compiler, David Britton
Williams Honors College, Honors Research Projects
Amethyst is a custom programming language, whose compiler translates Amethyst source code into textual LLVM IR, which can, in turn, be compiled by LLVM back-ends to produce executable binaries. This paper explores compiler concepts and implementation details for the Amethyst compiler, describes the LLVM architecture, and provides an overview of some Amethyst language features.
Ung: A Diagnostic Standard C Library,
2024
West Virginia University
Ung: A Diagnostic Standard C Library, Jakob Knute Kaivo
Graduate Theses, Dissertations, and Problem Reports (ETD)
Undefined behavior in C programs is a major source of unreliable software. Many of the most common exploitable software vulnerabilities can be traced directly to undefined behavior. In the increasingly connected world, a successful attack can cost the victim millions of dollars to recover from. While static program analysis aids in identifying undefined behavior, testing indicates that even the best static analysis tools correctly identifies about 35% of these defects. This dissertation introduces UNG’s Not GNU (UNG), a standard C library designed to mitigate undefined behavior. Where others have seen opportunities for optimization, UNG makes every effort to identify undefined …
A Critical Computing Curriculum Design Case: Exploring Tribal Sovereignty For Middle School Students,
2023
Utah State University
A Critical Computing Curriculum Design Case: Exploring Tribal Sovereignty For Middle School Students, Kristin A. Searle, Aubrey Rogowski, Colby Tofel-Grehl, Mengying Jiang
Journal of Computer Science Integration
We report on our efforts to design an integrated computing curriculum for middle school students in Montana that is in line with the Kapor Center’s focus on culturally sustaining-revitalizing pedagogies. Montana provides a unique context for doing this work because a state constitutional mandate requires all K-12 students to learn about tribal histories and cultures through Indian Education For All (IEFA). IEFA centers around seven essential understandings about Indigenous peoples in Montana that are integrated across content areas. In addition, implementation of Montana’s CS standards began in the 2021–2022 school year. In the curricular design, we sought to bring together …
Employing An Abolitionist, Critical Race Pedagogy In Cs: Centering The Voices, Experiences And Technological Innovations Of Black Youth,
2023
University of California, Los Angeles
Employing An Abolitionist, Critical Race Pedagogy In Cs: Centering The Voices, Experiences And Technological Innovations Of Black Youth, Tiera Tanksley
Journal of Computer Science Integration
This paper proposes a pedagogical extension of culturally responsive praxis called abolitionist, critical race pedagogy in CS. To showcase the power and potentiality of this pedagogy, this paper examines the experiences of 2 cohorts of Black high school students (n = 30) who participated in a critical race technology course that was taught during the dual pandemic of COVID-19 and anti-Black racism. The goal of this summer course was to employ an abolitionist, critical race pedagogy in CS to foster Black students’ ability to critically examine the ubiquity of anti-Black racism within the socio-technical architectures (e.g. code, data, algorithms and …
Culturally Responsive-Sustaining Computational Thinking: Enactment In Elementary Classrooms,
2023
Massey University
Culturally Responsive-Sustaining Computational Thinking: Enactment In Elementary Classrooms, Victoria Macann, Aman Yadav
Journal of Computer Science Integration
Technology has increasingly permeated many aspects of everyday life and this evolution raises the need for individuals to understand how the digital world works and what opportunities and risks it brings (Nouri, Zhang, Mannila & Norén, 2019). For this to be an experience for everyone, we need to rethink how we integrate computational thinking (CT) and provide teachers with tools to center their students’ identities, experiences, and cultures in the classroom. In this paper, we present two case studies of primary (elementary) teachers from a full primary (student ages 5–13) semi-rural school in the North Island of New Zealand that …
Choosing A Sophisticated, Robust, And Secure Programming Language,
2023
Cleveland State University
Choosing A Sophisticated, Robust, And Secure Programming Language, J. Simon Richard
The Downtown Review: An Interdisciplinary Journal Written and Peer-Reviewed by Mandel Honors College Students at Cleveland State University
This paper explores which programming languages maximize the quality and efficiency of software development projects requiring high levels of sophistication, security, and stability. Of the four languages discussed in this paper—C, C++, Java, and Rust—we conclude that Rust is the best for this application.
Μakka: Mutation Testing For Actor Concurrency In Akka Using Real-World Bugs,
2023
Oakland University
Μakka: Mutation Testing For Actor Concurrency In Akka Using Real-World Bugs, Mohsen Moradi Moghadam, Mehdi Bagherzadeh, Raffi Khatchadourian, Hamid Bagheri
Publications and Research
Actor concurrency is becoming increasingly important in the real-world and mission-critical software. This requires these applications to be free from actor bugs, that occur in the real world, and have tests that are effective in finding these bugs. Mutation testing is a well-established technique that transforms an application to induce its likely bugs and evaluate the effectiveness of its tests in finding these bugs. Mutation testing is available for a broad spectrum of applications and their bugs, ranging from web to mobile to machine learning, and is used at scale in companies like Google and Facebook. However, there still is …
Ensuring Non-Repudiation In Long-Distance Constrained Devices,
2023
University of South Alabama
Ensuring Non-Repudiation In Long-Distance Constrained Devices, Ethan Blum
Honors Theses
Satellite communication is essential for the exploration and study of space. Satellites allow communications with many devices and systems residing in space and on the surface of celestial bodies from ground stations on Earth. However, with the rise of Ground Station as a Service (GsaaS), the ability to efficiently send action commands to distant satellites must ensure non-repudiation such that an attacker is unable to send malicious commands to distant satellites. Distant satellites are also constrained devices and rely on limited power, meaning security on these devices is minimal. Therefore, this study attempted to propose a novel algorithm to allow …
Llm-Adapters: An Adapter Family For Parameter-Efficient Fine-Tuning Of Large Language Models,
2023
Singapore Management University
Llm-Adapters: An Adapter Family For Parameter-Efficient Fine-Tuning Of Large Language Models, Zhiqiang Hu, Lei Wang, Yihuai Lan, Wanyu Xu, Ee-Peng Lim, Lidong Bing, Xing Xu, Soujanya Poria, Roy Ka-Wei Lee
Research Collection School Of Computing and Information Systems
The success of large language models (LLMs), like GPT-4 and ChatGPT, has led to the development of numerous cost-effective and accessible alternatives that are created by finetuning open-access LLMs with task-specific data (e.g., ChatDoctor) or instruction data (e.g., Alpaca). Among the various fine-tuning methods, adapter-based parameter-efficient fine-tuning (PEFT) is undoubtedly one of the most attractive topics, as it only requires fine-tuning a few external parameters instead of the entire LLMs while achieving comparable or even better performance. To enable further research on PEFT methods of LLMs, this paper presents LLMAdapters, an easy-to-use framework that integrates various adapters into LLMs and …
Large Language Model Is Not A Good Few-Shot Information Extractor, But A Good Reranker For Hard Samples!,
2023
Singapore Management University
Large Language Model Is Not A Good Few-Shot Information Extractor, But A Good Reranker For Hard Samples!, Yubo Ma, Yixin Cao, Yongchin Hong, Aixin Sun
Research Collection School Of Computing and Information Systems
Large Language Models (LLMs) have made remarkable strides in various tasks. However, whether they are competitive few-shot solvers for information extraction (IE) tasks and surpass fine-tuned small Pre-trained Language Models (SLMs) remains an open problem. This paper aims to provide a thorough answer to this problem, and moreover, to explore an approach towards effective and economical IE systems that combine the strengths of LLMs and SLMs. Through extensive experiments on nine datasets across four IE tasks, we show that LLMs are not effective few-shot information extractors in general, given their unsatisfactory performance in most settings and the high latency and …
Examining The Inter-Consistency Of Large Language Models: An In-Depth Analysis Via Debate,
2023
Singapore Management University
Examining The Inter-Consistency Of Large Language Models: An In-Depth Analysis Via Debate, Kai Xiong, Xiao Ding, Yixin Cao, Ting Liu, Bing Qin
Research Collection School Of Computing and Information Systems
Large Language Models (LLMs) have shown impressive capabilities in various applications, but they still face various inconsistency issues. Existing works primarily focus on the inconsistency issues within a single LLM, while we complementarily explore the inter-consistency among multiple LLMs for collaboration. To examine whether LLMs can collaborate effectively to achieve a consensus for a shared goal, we focus on commonsense reasoning, and introduce a formal debate framework (FORD) to conduct a three-stage debate among LLMs with real-world scenarios alignment: fair debate, mismatched debate, and roundtable debate. Through extensive experiments on various datasets, LLMs can effectively collaborate to reach a consensus …
Benchmarking Foundation Models With Language-Model-As-An-Examiner,
2023
Singapore Management University
Benchmarking Foundation Models With Language-Model-As-An-Examiner, Yushi Bai, Jiahao Ying, Yixin Cao, Xin Lv, Yuze He, Xiaozhi Wang, Jifan Yu, Kaisheng Zeng, Yijia Xiao, Haozhe Lyu, Jiayin Zhang, Juanzi Li, Lei Hou
Research Collection School Of Computing and Information Systems
Numerous benchmarks have been established to assess the performance of foundation models on open-ended question answering, which serves as a comprehensive test of a model’s ability to understand and generate language in a manner similar to humans. Most of these works focus on proposing new datasets, however, we see two main issues within previous benchmarking pipelines, namely testing leakage and evaluation automation. In this paper, we propose a novel benchmarking framework, Language-Model-as-an-Examiner, where the LM serves as a knowledgeable examiner that formulates questions based on its knowledge and evaluates responses in a reference-free manner. Our framework allows for effortless extensibility …
Molca: Molecular Graph-Language Modeling With Cross-Modal Projector And Uni-Modal Adapter,
2023
Singapore Management University
Molca: Molecular Graph-Language Modeling With Cross-Modal Projector And Uni-Modal Adapter, Zhiyuan Liu, Sihang Li, Yanchen Luo, Hao Fei, Yixin Cao, Kenji Kawaguchi, Xiang Wang, Tat-Seng Chua
Research Collection School Of Computing and Information Systems
Language Models (LMs) have demonstrated impressive molecule understanding ability on various 1D text-related tasks. However, they inherently lack 2D graph perception — a critical ability of human professionals in comprehending molecules’ topological structures. To bridge this gap, we propose MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal Adapter. MolCA enables an LM (i.e., Galactica) to understand both text- and graph-based molecular contents via the cross-modal projector. Specifically, the cross-modal projector is implemented as a QFormer to connect a graph encoder’s representation space and an LM’s text space. Further, MolCA employs a uni-modal adapter (i.e., LoRA) for the LM’s efficient …
A Comprehensive Evaluation Of Large Language Models On Legal Judgment Prediction,
2023
Singapore Management University
A Comprehensive Evaluation Of Large Language Models On Legal Judgment Prediction, Ruihao Shui, Yixin Cao, Xiang Wang, Tat-Seng Chua
Research Collection School Of Computing and Information Systems
Large language models (LLMs) have demonstrated great potential for domain-specific applications, such as the law domain. However, recent disputes over GPT-4’s law evaluation raise questions concerning their performance in real-world legal tasks. To systematically investigate their competency in the law, we design practical baseline solutions based on LLMs and test on the task of legal judgment prediction. In our solutions, LLMs can work alone to answer open questions or coordinate with an information retrieval (IR) system to learn from similar cases or solve simplified multi-choice questions. We show that similar cases and multi-choice options, namely label candidates, included in prompts …
Wsdms: Debunk Fake News Via Weakly Supervised Detection Of Misinforming Sentences With Contextualized Social Wisdom,
2023
Hong Kong Baptist University
Wsdms: Debunk Fake News Via Weakly Supervised Detection Of Misinforming Sentences With Contextualized Social Wisdom, Ruichao Yang, Wei Gao, Jing Ma, Hongzhan Lin, Zhiwei Yang
Research Collection School Of Computing and Information Systems
In recent years, we witness the explosion of false and unconfirmed information (i.e., rumors) that went viral on social media and shocked the public. Rumors can trigger versatile, mostly controversial stance expressions among social media users. Rumor verification and stance detection are different yet relevant tasks. Fake news debunking primarily focuses on determining the truthfulness of news articles, which oversimplifies the issue as fake news often combines elements of both truth and falsehood. Thus, it becomes crucial to identify specific instances of misinformation within the articles. In this research, we investigate a novel task in the field of fake news …
Disentangling Transformer Language Models As Superposed Topic Models,
2023
Singapore Management University
Disentangling Transformer Language Models As Superposed Topic Models, Jia Peng Lim, Hady Wirawan Lauw
Research Collection School Of Computing and Information Systems
Topic Modelling is an established research area where the quality of a given topic is measured using coherence metrics. Often, we infer topics from Neural Topic Models (NTM) by interpreting their decoder weights, consisting of top-activated words projected from individual neurons. Transformer-based Language Models (TLM) similarly consist of decoder weights. However, due to its hypothesised superposition properties, the final logits originating from the residual path are considered uninterpretable. Therefore, we posit that we can interpret TLM as superposed NTM by proposing a novel weight-based, model-agnostic and corpus-agnostic approach to search and disentangle decoder-only TLM, potentially mapping individual neurons to multiple …
Kape: Knn-Based Performance Testing For Deep Code Search,
2023
Singapore Management University
Kape: Knn-Based Performance Testing For Deep Code Search, Yuejun Guo, Qiang Hu, Xiaofei Xie, Cordy Maxime, Mike Papadakis, Yves Le Traon
Research Collection School Of Computing and Information Systems
Code search is a common yet important activity of software developers. An efficient code search model can largely facilitate the development process and improve the programming quality. Given the superb performance of learning the contextual representations, deep learning models, especially pre-trained language models, have been widely explored for the code search task. However, studies mainly focus on proposing new architectures for ever-better performance on designed test sets but ignore the performance on unseen test data where only natural language queries are available. The same problem in other domains, e.g., CV and NLP, is usually solved by test input selection that …
Attack Prompt Generation For Red Teaming And Defending Large Language Models,
2023
Singapore Management University
Attack Prompt Generation For Red Teaming And Defending Large Language Models, Boyi Deng, Wenjie Wang, Fuli Feng, Yang Deng, Qifan Wang, Xiangnan He
Research Collection School Of Computing and Information Systems
Large language models (LLMs) are susceptible to red teaming attacks, which can induce LLMs to generate harmful content. Previous research constructs attack prompts via manual or automatic methods, which have their own limitations on construction cost and quality. To address these issues, we propose an integrated approach that combines manual and automatic methods to economically generate high-quality attack prompts. Specifically, considering the impressive capabilities of newly emerged LLMs, we propose an attack framework to instruct LLMs to mimic human-generated prompts through in-context learning. Furthermore, we propose a defense framework that fine-tunes victim LLMs through iterative interactions with the attack framework …
Large Language Models As Source Planner For Personalized Knowledge-Grounded Dialogues,
2023
Singapore Management University
Large Language Models As Source Planner For Personalized Knowledge-Grounded Dialogues, Hongru Wang, Minda Hu, Yang Deng, Rui Wang, Fei Mi, Weichao Wang, Yasheng Wang, Wai-Chung Kwan, Irwin King, Kam-Fai Wong
Research Collection School Of Computing and Information Systems
Open-domain dialogue system usually requires different sources of knowledge to generate more informative and evidential responses. However, existing knowledge-grounded dialogue systems either focus on a single knowledge source or overlook the dependency between multiple sources of knowledge, which may result in generating inconsistent or even paradoxical responses. To incorporate multiple knowledge sources and dependencies between them, we propose SAFARI, a novel framework that leverages the exceptional capabilities of large language models (LLMs) in planning, understanding, and incorporating under both supervised and unsupervised settings. Specifically, SAFARI decouples the knowledge grounding into multiple sources and response generation, which allows easy extension to …
Supporting Software Engineers With Large Language Model-Based Automation,
2023
Singapore Management University
Supporting Software Engineers With Large Language Model-Based Automation, Ting Zhang
Dissertations and Theses Collection (Open Access)
In recent years, software engineering (SE) has witnessed significant growth, leading to the creation and sharing of an abundance of software artifacts such as source code, bug reports, and pull requests. Analyzing these artifacts is crucial for comprehending the sentiments of software developers and automating various SE tasks, ultimately leading to more human-centered automated SE and enhancing software development efficiency. However, the diverse and unstructured nature of software text poses a significant challenge to this analysis. In response, researchers have investigated a variety of approaches, including the utilization of natural language processing techniques. The advent of large language models (LLMs), …
