Open Access. Powered by Scholars. Published by Universities.®

2024

Discipline
Institution
Keyword
Publication
Publication Type

Articles 1 - 30 of 67

Full-Text Articles in Programming Languages and Compilers

Calculation And Statistical Analysis Of Wins Above Replacement, Joshua Taylor Dec 2024

Calculation And Statistical Analysis Of Wins Above Replacement, Joshua Taylor

Departmental Honors & Graduate Capstone Projects

The Wins Above Replacement (WAR) statistic in Major League Baseball is a prominent metric used to estimate player value by quantifying all aspects of play in terms of wins added to a baseball team. We will use R to calculate WAR for all players from 1871 to 2012 and use data from those years to construct multivariate predictive models to attempt to estimate WAR for players from 2013 to 2024. We find strong correlations between predicted and actual WAR values for most models, with the exception of the polynomial predictive model for non-qualified pitchers.


Towards Robust, Secure, And Privacy-Aware Large Language Models Of Code, Zhou Yang Dec 2024

Towards Robust, Secure, And Privacy-Aware Large Language Models Of Code, Zhou Yang

Dissertations and Theses Collection (Open Access)

The field of software engineering has witnessed a surge in large language models specifically tailored to understand and process code, which we call large language models for code (LLM4Code). The increasing popularity of LLM4Code is inseparable from three key factors: the availability of extensive datasets compiled from diverse data sources, the advancements in deep learning algorithms and computational power that facilitate the training of these powerful models, and the active engagement and collaboration within the research community fostering innovation and the rapid exchange of ideas and methodologies. As evidenced by a series of studies, LLM4Code has been experiencing rapid development …


Reevo: Large Language Models As Hyper-Heuristics With Reflective Evolution, Haoran Ye, Jiarui Wang, Zhiguang Cao, Federico Berto, Chuanbo Hua, Haeyeon Kim, Jinkyoo Park, Guojie Song Dec 2024

Reevo: Large Language Models As Hyper-Heuristics With Reflective Evolution, Haoran Ye, Jiarui Wang, Zhiguang Cao, Federico Berto, Chuanbo Hua, Haeyeon Kim, Jinkyoo Park, Guojie Song

Research Collection School Of Computing and Information Systems

The omnipresence of NP-hard combinatorial optimization problems (COPs) compels domain experts to engage in trial-and-error heuristic design process. The long-standing endeavor of design automation has gained new momentum with the rise of large language models (LLMs). This paper introduces Language Hyper-Heuristics (LHHs), an emerging variant of Hyper-Heuristics that leverages LLMs for heuristic generation, featuring minimal manual intervention and open-ended heuristic spaces. To empower LHHs, we present Reflective Evolution (ReEvo), a generic searching framework that emulates the reflective design approach of human experts while far surpassing human capabilities with its scalable LLM inference, Internet-scale domain knowledge, and powerful evolutionary search. Evaluations …


Revisiting Masked Auto-Encoders For Ecg-Language Representation Learning, Hung Manh Pham, Aaqib Saeed, Dong Ma Dec 2024

Revisiting Masked Auto-Encoders For Ecg-Language Representation Learning, Hung Manh Pham, Aaqib Saeed, Dong Ma

Research Collection School Of Computing and Information Systems

We propose C-MELT, a novel framework for multimodal self-supervised learning of Electrocardiogram (ECG) and text encoders. C-MELT pre-trains a contrastive-enhanced masked auto-encoder architecture using ECG-text paired data. It exploits the generative strengths with improved discriminative capabilities to enable robust cross-modal alignment. This is accomplished through a carefully designed model, loss functions, and a novel negative sampling strategy. Our preliminary experiments demonstrate significant performance improvements with up to 12% in downstream cardiac arrhythmia classification and patient identification tasks. Our findings demonstrate C-MELT's capacity to extract rich, clinically relevant features from ECG-text pairs, paving the way for more accurate and efficient cardiac …


Divlog: Log Parsing With Prompt Enhanced In-Context Learning, Junjielong Xu, Ruichun Yang, Yintong Huo, Chengyu Zhang, Pinjia He Dec 2024

Divlog: Log Parsing With Prompt Enhanced In-Context Learning, Junjielong Xu, Ruichun Yang, Yintong Huo, Chengyu Zhang, Pinjia He

Research Collection School Of Computing and Information Systems

Log parsing, which involves log template extraction from semistructured logs to produce structured logs, is the first and the most critical step in automated log analysis. However, current log parsers suffer from limited effectiveness for two reasons. First, traditional data-driven log parsers solely rely on heuristics or handcrafted features designed by domain experts, which may not consistently perform well on logs from diverse systems. Second, existing supervised log parsers require model tuning, which is often limited to fixed training samples and causes sub-optimal performance across the entire log source. To address this limitation, we propose DivLog, an effective log parsing …


Elevating Automated Software Maintenance Tasks With Large Language Models, Xin Zhou Nov 2024

Elevating Automated Software Maintenance Tasks With Large Language Models, Xin Zhou

Dissertations and Theses Collection (Open Access)

Software engineering involves many tasks across different phases such as requirements, design, implementation, testing, and maintenance. Among them, software maintenance is a crucial phase, typically accounting for more than half of the software life cycle's duration.
To boost developer productivity, in recent years, numerous research endeavors in software engineering have sought to automate certain software maintenance tasks through the application of machine learning techniques.
Since 2020, the emergence of advanced Large Language Models (LLMs) of code has opened new avenues for enhancing automated solutions in software maintenance.
This dissertation presents a series of works aimed at advancing automated solutions for …


Mm‑Forecast: A Multimodal Approach To Temporal Event Forecasting With Large Language Models, Haoxuan Li, Zhengmao Yang, Yunshan Ma, Yi Bin, Yang Yang, Tat-Seng Chua Nov 2024

Mm‑Forecast: A Multimodal Approach To Temporal Event Forecasting With Large Language Models, Haoxuan Li, Zhengmao Yang, Yunshan Ma, Yi Bin, Yang Yang, Tat-Seng Chua

Research Collection School Of Computing and Information Systems

We study an emerging and intriguing problem of multimodal temporal event forecasting with large language models. Compared to using text or graph modalities, the investigation of utilizing images for temporal event forecasting has not been fully explored, especially in the era of large language models (LLMs). To bridge this gap, we are particularly interested in two key questions of: 1) why images will help in temporal event forecasting, and 2) how to integrate images into the LLM-based forecasting framework. To answer these research questions, we propose to identify two essential functions that images play in the scenario of temporal event …


Review Of Current Trends In Information Technology Concerning Phonetic Similarity”, Zaid Rajih Mohammed, Ahmed H. Aliwy Oct 2024

Review Of Current Trends In Information Technology Concerning Phonetic Similarity”, Zaid Rajih Mohammed, Ahmed H. Aliwy

Al-Bahir

With the increasing availability of textual information in various languages via the Internet in homes and companies through Internet and intranet services, there is an urgent need for the technologies and tools necessary to process this information, phonetic representation, and voice interaction. For example voice to voice machine translation need to phonetic mapping and similarity among the languages especially for names and foreign words. This one example of the importance of phonetic mapping and similarity. This article aims to describe, in detail, the recent surge in interest and advancements in phonetic similarity (PS), phonetic representation, and phonetic mapping researches. PS …


Wip: An Engaging Undergraduate Intro To Model Checking In Software Engineering Using Tla+, Konstantin Laufer, Gunda Mertin, George K. Thiruvathukal Oct 2024

Wip: An Engaging Undergraduate Intro To Model Checking In Software Engineering Using Tla+, Konstantin Laufer, Gunda Mertin, George K. Thiruvathukal

Computer Science: Faculty Publications and Other Works

Background: In this Innovative Practice Work in Progress, we present our initial efforts to integrate formal methods, with a focus on model-checking specifications written in Temporal Logic of Actions (TLA+), into computer science education, targeting undergraduate juniors/seniors and graduate students. Many safety-critical systems and services crucially depend on correct and reliable behavior. Formal methods can play a key role in ensuring correct and safe system behavior, yet remain underutilized in educational and industry contexts.

Aims: We aim to (1) qualitatively assess the state of formal methods in computer science programs, (2) construct level-appropriate examples that could be included …


Interoperability In Deep Learning: A User Survey And Failure Analysis Of Onnx Model Converters, Purvish Jajal, Wenxin Jiang, Arav Tewari, Erik Kocinare, Joseph Woo, Anusha Sarraf, Yung-Hsiang Lu, George Thiruvathukal, James C. Davis Sep 2024

Interoperability In Deep Learning: A User Survey And Failure Analysis Of Onnx Model Converters, Purvish Jajal, Wenxin Jiang, Arav Tewari, Erik Kocinare, Joseph Woo, Anusha Sarraf, Yung-Hsiang Lu, George Thiruvathukal, James C. Davis

Computer Science: Faculty Publications and Other Works

Software engineers develop, fine-tune, and deploy deep learning (DL) models using a variety of development frameworks and runtime environments. DL model converters move models between frameworks and to runtime environments. Conversion errors compromise model quality and disrupt deployment. However, the failure characteristics of DL model converters are unknown, adding risk when using DL interoperability technologies. This paper analyzes failures in DL model converters. We survey software engineers about DL interoperability tools, use cases, and pain points (N=92). Then, we characterize failures in model converters associated with the main interoperability tool, ONNX (N=200 issues in PyTorch and TensorFlow). Finally, we formulate …


Recasting The Mould – Librarianship Of The Future: Leveraging Automation, Apis, And Ai, Samantha Seah Sep 2024

Recasting The Mould – Librarianship Of The Future: Leveraging Automation, Apis, And Ai, Samantha Seah

Research Collection Library

With leaps in artificial intelligence made in recent years redefining the information landscape and introducing new means of information production, librarianship also must evolve to include new literacies. One way librarians can equip and empower ourselves is by understanding the building blocks of how machines and automation work. Perhaps more important than learning specific programming languages, learning computational thinking provides us with more ways to spot and evaluate problems and devise solutions without extensive coding knowledge. My presentation will take the improvement of membership processing as an example using Power Automate, a low-code Microsoft tool mimicking block programming. The tool …


Sound And Complete Witnesses For Template-Based Verification Of Ltl Properties On Polynomial Programs, Krishnendu Chatterjee, Amir Goharshady, Ehsan Goharshady, Mehrdad Karrabi, Dorde Zikelic Sep 2024

Sound And Complete Witnesses For Template-Based Verification Of Ltl Properties On Polynomial Programs, Krishnendu Chatterjee, Amir Goharshady, Ehsan Goharshady, Mehrdad Karrabi, Dorde Zikelic

Research Collection School Of Computing and Information Systems

We study the classical problem of verifying programs with respect to formal specifications given in the linear temporal logic (LTL). We first present novel sound and complete witnesses for LTL verification over imperative programs. Our witnesses are applicable to both verification (proving) and refutation (finding bugs) settings. We then consider LTL formulas in which atomic propositions can be polynomial constraints and turn our focus to polynomial arithmetic programs, i.e. programs in which every assignment and guard consists only of polynomial expressions. For this setting, we provide an efficient algorithm to automatically synthesize such LTL witnesses. Our synthesis procedure is both …


Style: Improving Domain Transferability Of Asking Clarification Questions In Large Language Model Powered Conversational Agents, Yue Chen, Chen Huang, Yang Deng, Wenqiang Lei, Dingnan Jin, Jia Liu, Tat-Seng Chua Aug 2024

Style: Improving Domain Transferability Of Asking Clarification Questions In Large Language Model Powered Conversational Agents, Yue Chen, Chen Huang, Yang Deng, Wenqiang Lei, Dingnan Jin, Jia Liu, Tat-Seng Chua

Research Collection School Of Computing and Information Systems

Equipping a conversational search engine with strategies regarding when to ask clarification questions is becoming increasingly important across various domains. Attributing to the context understanding capability of LLMs and their access to domain-specific sources of knowledge, LLM-based clarification strategies feature rapid transfer to various domains in a posthoc manner. However, they still struggle to deliver promising performance on unseen domains, struggling to achieve effective domain transferability. We take the first step to investigate this issue and existing methods tend to produce one-size-fits-all strategies across diverse domains, limiting their search effectiveness. In response, we introduce a novel method, called STYLE, to …


Chain-Of-Exemplar: Enhancing Distractor Generation For Multimodal Educational Question Generation, Haohao Luo, Yang Deng, Ying Shen, See-Kiong Ng, Tat-Seng Chua Aug 2024

Chain-Of-Exemplar: Enhancing Distractor Generation For Multimodal Educational Question Generation, Haohao Luo, Yang Deng, Ying Shen, See-Kiong Ng, Tat-Seng Chua

Research Collection School Of Computing and Information Systems

Multiple-choice questions (MCQs) are important in enhancing concept learning and student engagement for educational purposes. Despite the multimodal nature of educational content, current methods focus mainly on text-based inputs and often neglect the integration of visual information. In this work, we study the problem of multimodal educational question generation, which aims at generating subject-specific educational questions with plausible yet incorrect distractors based on multimodal educational content. To tackle this problem, we introduce a novel framework, named Chain-of-Exemplar (CoE), which utilizes multimodal large language models (MLLMs) with Chain-of-Thought reasoning to improve the generation of challenging distractors. Furthermore, CoE leverages three-stage contextualized …


On The Multi-Turn Instruction Following For Conversational Web Agents, Yang Deng, Xuan Zhang, Wenxuan Zhang, Yifei Yuan, See-Kiong Ng, Tat-Seng Chua Aug 2024

On The Multi-Turn Instruction Following For Conversational Web Agents, Yang Deng, Xuan Zhang, Wenxuan Zhang, Yifei Yuan, See-Kiong Ng, Tat-Seng Chua

Research Collection School Of Computing and Information Systems

Web agents powered by Large Language Models (LLMs) have demonstrated remarkable abilities in planning and executing multi-step interactions within complex web-based environments, fulfilling a wide range of web navigation tasks. Despite these advancements, the potential for LLM-powered agents to effectively engage with sequential user instructions in real-world scenarios has not been fully explored. In this work, we introduce a new task of Conversational Web Navigation, which necessitates sophisticated interactions that span multiple turns with both the users and the environment, supported by a specially developed dataset named Multi-Turn Mind2Web (MT-Mind2Web). To tackle the limited context length of LLMs and the …


Watme: Towards Lossless Watermarking Through Lexical Redundancy, Liang Chen, Yatao Bian, Yang Deng, Deng Cai, Shuaiyi Li, Peilin Zhao, Kam-Fai Wong Aug 2024

Watme: Towards Lossless Watermarking Through Lexical Redundancy, Liang Chen, Yatao Bian, Yang Deng, Deng Cai, Shuaiyi Li, Peilin Zhao, Kam-Fai Wong

Research Collection School Of Computing and Information Systems

Text watermarking has emerged as a pivotal technique for identifying machine-generated text. However, existing methods often rely on arbitrary vocabulary partitioning during decoding to embed watermarks, which compromises the availability of suitable tokens and significantly degrades the quality of responses. This study assesses the impact of watermarking on different capabilities of large language models (LLMs) from a cognitive science lens. Our finding highlights a significant disparity; knowledge recall and logical reasoning are more adversely affected than language generation. These results suggest a more profound effect of watermarking on LLMs than previously understood. To address these challenges, we introduce Watermarking with …


Clamber: A Benchmark Of Identifying And Clarifying Ambiguous Information Needs In Large Language Models, Tong Zhang, Peixin Qin, Yang Deng, Chen Huang, Wenqiang Lei, Junhong Liu, Dingnan Jin, Hongru Liang, Tat-Seng Chua Aug 2024

Clamber: A Benchmark Of Identifying And Clarifying Ambiguous Information Needs In Large Language Models, Tong Zhang, Peixin Qin, Yang Deng, Chen Huang, Wenqiang Lei, Junhong Liu, Dingnan Jin, Hongru Liang, Tat-Seng Chua

Research Collection School Of Computing and Information Systems

Large language models (LLMs) are increasingly used to meet user information needs, but their effectiveness in dealing with user queries that contain various types of ambiguity remains unknown, ultimately risking user trust and satisfaction. To this end, we introduce CLAMBER, a benchmark for evaluating LLMs using a well-organized taxonomy. Building upon the taxonomy, we construct 12K high-quality data to assess the strengths, weaknesses, and potential risks of various off-the-shelf LLMs.Our findings indicate the limited practical utility of current LLMs in identifying and clarifying ambiguous user queries, even enhanced by chain-of-thought (CoT) and few-shot prompting. These techniques may result in overconfidence …


Self-Chats From Large Language Models Make Small Emotional Support Chatbot Better, Zhonghua Zheng, Lizi Liao, Yang Deng, Libo Qin, Liqiang Nie Aug 2024

Self-Chats From Large Language Models Make Small Emotional Support Chatbot Better, Zhonghua Zheng, Lizi Liao, Yang Deng, Libo Qin, Liqiang Nie

Research Collection School Of Computing and Information Systems

Large Language Models (LLMs) have shown strong generalization abilities to excel in various tasks, including emotion support conversations. However, deploying such LLMs like GPT-3 (175B parameters) is resource-intensive and challenging at scale. In this study, we utilize LLMs as “Counseling Teacher” to enhance smaller models’ emotion support response abilities, significantly reducing the necessity of scaling up model size. To this end, we first introduce an iterative expansion framework, aiming to prompt the large teacher model to curate an expansive emotion support dialogue dataset. This curated dataset, termed ExTES, encompasses a broad spectrum of scenarios and is crafted with meticulous strategies …


Reinforcement Tuning For Detecting Stances And Debunking Rumors Jointly With Large Language Models, Ruichao Yang, Wei Gao, Jing Ma, Hongzhan Ling, Bo Wang Aug 2024

Reinforcement Tuning For Detecting Stances And Debunking Rumors Jointly With Large Language Models, Ruichao Yang, Wei Gao, Jing Ma, Hongzhan Ling, Bo Wang

Research Collection School Of Computing and Information Systems

Learning multi-task models for jointly detecting stance and verifying rumors poses challenges due to the need for training data of stance at post level and rumor veracity at claim level, which are difficult to obtain. To address this issue, we leverage large language models (LLMs) as the foundation annotators for the joint stance detection (SD) and rumor verification (RV) tasks, dubbed as JSDRV. We introduce a novel reinforcement tuning framework to enhance the joint predictive capabilities of LLM-based SD and RV components. Specifically, we devise a policy for selecting LLM-annotated data at the two levels, employing a hybrid reward mechanism …


Larp: Language Audio Relational Pre‑Training For Cold‑Start Playlist Continuation, Rebecca Salganik, Xiaohao Liu, Yunshan Ma, Jian Kang, Tat‑Seng Chua Aug 2024

Larp: Language Audio Relational Pre‑Training For Cold‑Start Playlist Continuation, Rebecca Salganik, Xiaohao Liu, Yunshan Ma, Jian Kang, Tat‑Seng Chua

Research Collection School Of Computing and Information Systems

As online music consumption increasingly shifts towards playlist-based listening, the task of playlist continuation, in which an algorithm suggests songs to extend a playlist in a personalized and musically cohesive manner, has become vital to the success of music streaming services. Currently, many existing playlist continuation approaches rely on collaborative filtering methods to perform their recommendations. However, such methods will struggle to recommend songs that lack interaction data, an issue known as the cold-start problem. Current approaches to this challenge design complex mechanisms for extracting relational signals from sparse collaborative signals and integrating them into content representations. However, these approaches …


Introduction To Programming And Applied Analytics Using Python, Matt Brown Jul 2024

Introduction To Programming And Applied Analytics Using Python, Matt Brown

ATU Faculty OER Books and Materials

This open electronic textbook is a collection of lecture notes, assignments, and additional background material for a junior level analytics course targeted for business students, it is free to use and copy. The text assumes readers have not had prior programming or computing courses, but have had at least one analytics course. The textbook differs from other textbooks because it serves a dual purpose, to first introduce to students the Python programming language and secondly to introduce analytics programming in Python. It is not meant to be a comprehensive book on the Python language or data analytics, rather a semester’s …


Development Of An Algorithm To Identify And Calculate The Amount Of File Slack On An Image Of A Given Drive, Nicholas Flynn Jul 2024

Development Of An Algorithm To Identify And Calculate The Amount Of File Slack On An Image Of A Given Drive, Nicholas Flynn

Honors Theses

As society increasingly relies on technology, the rates of cyber crime have been increasing at exponential rates. Cyber criminals are also discovering new ways to hide evidence of their crimes. This study develops a forensic analysis algorithm to evaluate the amount of file slack on an image of a drive. Slack space, leftover drive space on a disk sector after a file has been written, can be exploited to hide data. The algorithm aims to detect and calculate this slack space to help direct forensic investigations. The algorithm was evaluated on a population dataset of 100,000 files with random data …


Large Language Model Powered Agents For Information Retrieval, An Zhang, Yang Deng, Yankai Lin, Xu Chen, Ji-Rong Wen, Tat-Seng Chua Jul 2024

Large Language Model Powered Agents For Information Retrieval, An Zhang, Yang Deng, Yankai Lin, Xu Chen, Ji-Rong Wen, Tat-Seng Chua

Research Collection School Of Computing and Information Systems

The vital goal of information retrieval today extends beyond merely connecting users with relevant information they search for. It also aims to enrich the diversity, personalization, and interactivity of that connection, ensuring the information retrieval process is as seamless, beneficial, and supportive as possible in the global digital era. Current information retrieval systems often encounter challenges like a constrained understanding of queries, static and inflexible responses, limited personalization, and restricted interactivity. With the advent of large language models (LLMs), there's a transformative paradigm shift as we integrate LLM-powered agents into these systems. These agents bring forth crucial human capabilities like …


Who Wrote The Scientific News? Improving The Discernibility Of Llms To Human-Written Scientific News, Dominik Soós Jul 2024

Who Wrote The Scientific News? Improving The Discernibility Of Llms To Human-Written Scientific News, Dominik Soós

Computer Science Theses & Dissertations

Large Language Models (LLMs) have rapidly advanced the field of Natural Language Processing and become powerful tools for generating and evaluating scientific text. Although LLMs have demonstrated promising as evaluators for certain text generation tasks, there is still a gap until they are used as reliable text evaluators for general purposes. In this thesis project, I attempted to fill this gap by examining the discernibility of LLMs from human-written and LLM-generated scientific news. This research demonstrated that although it was relatively straightforward for humans to discern scientific news written by humans from scientific news generated by GPT-3.5 using basic prompts, …


How Do Preservice Teachers Learn To Teach Integrated Computational Thinking?: Evidence From Planning, Enactment, And Reflection, Rachael Dektor, Samuel Severance, Kip Téllez Jun 2024

How Do Preservice Teachers Learn To Teach Integrated Computational Thinking?: Evidence From Planning, Enactment, And Reflection, Rachael Dektor, Samuel Severance, Kip Téllez

Journal of Computer Science Integration

This study examines pre-service teachers’ (PSTs) beliefs and understandings about computational thinking (CT) integration and lesson implementation over time. Utilizing a design-based research approach, 3 PSTs led the co-design of integrated CT lessons with support from researchers and enacted these CT integrated lessons with K-5 students. All PSTs participated in a whole-group CT workshop and engaged in one-on-one lesson design sessions with a researcher. We utilized a grounded theory approach to qualitatively analyze pre-surveys, semi-structured interviews, and video data of three PSTs enacting their lessons. We found that PSTs’ initial beliefs about CT instruction – including the importance of it …


Towards Faster Inference Of Transformers: Strategies For Accelerating Decoding Processes, Cunxiao Du Jun 2024

Towards Faster Inference Of Transformers: Strategies For Accelerating Decoding Processes, Cunxiao Du

Dissertations and Theses Collection (Open Access)

This thesis delves into the acceleration and optimization of Transformer inference, a subject of increasing importance with the emergence of Large Language Models (LLMs). The study primarily addresses the challenges posed by two inherent properties of Transformers during inference: the quadratic complexity of the attention mechanism and the sequential nature of autoregressive inference. The research is structured into three main parts. The first part enhances the learning capabilities of non-autoregressive Transformers, achieving a remarkable 15.0x acceleration on machine translation tasks. The following section focuses on lossless acceleration through speculative decoding, where the proposed algorithm, Glide with CAPE, is shown to …


Let’S Think Outside The Box: Exploring Leap-Of-Thought In Large Language Models With Multimodal Humor Generation, Shanshan Zhong, Zhongzhan Huang, Shanghua Gao, Wushao Wen, Liang Lin, Marinka Zitnik, Pan Zhou Jun 2024

Let’S Think Outside The Box: Exploring Leap-Of-Thought In Large Language Models With Multimodal Humor Generation, Shanshan Zhong, Zhongzhan Huang, Shanghua Gao, Wushao Wen, Liang Lin, Marinka Zitnik, Pan Zhou

Research Collection School Of Computing and Information Systems

Chain-of-Thought (CoT) [2, 3] guides large language models (LLMs) to reason step-by-step, and can motivate their logical reasoning ability. While effective for logical tasks, CoT is not conducive to creative problem-solving which often requires out-of-box thoughts and is crucial for innovation advancements. In this paper, we explore the Leap-of-Thought (LoT) abilities within LLMs — a nonsequential, creative paradigm involving strong associations and knowledge leaps. To this end, we study LLMs on the popular Oogiri game which needs participants to have good creativity and strong associative thinking for responding unexpectedly and humorously to the given image, text, or both, and thus …


A Survey Of Practical Haskell: Parsing, Interpreting, And Testing, Parker Landon May 2024

A Survey Of Practical Haskell: Parsing, Interpreting, And Testing, Parker Landon

Honors Projects

Strongly typed pure functional programming languages like Haskell have historically been confined to academia as vehicles for programming language research. While features of functional programming have greatly influenced mainstream programming languages, the imperative programming style remains pervasive in practical software development. This paper illustrates the practical utility of Haskell and pure functional programming by exploring “hson,” a scripting language for processing JSON developed in Haskell. After introducing the relevant features of Haskell to the unfamiliar reader, this paper reveals how hson leverages functional programming to implement parsing, interpreting, and testing. By showcasing how Haskell’s language features enable the creation of …


Program Analysis Of C For Conversion To Memory-Safe Rust, Dylan Cassidy May 2024

Program Analysis Of C For Conversion To Memory-Safe Rust, Dylan Cassidy

Honors Scholar Theses

C is a memory-unsafe language, which can cause software security issues. Rust is a more recent high-performance language that has memory-safe features, which motivates developers to move software to Rust. However, given the large existing C codebase, this is a tedious task, and current approaches result in memory-unsafe blocks of code remaining unsafe after conversion. We seek to use program analysis techniques to create software that identifies blocks of C code that could be safely converted to memory-safe Rust, despite using seemingly memory- unsafe access patterns. We performed manual translation of functions within the libGeoIP C library to Rust, ensuring …


Machine Learning: Face Recognition, Mohammed E. Amin May 2024

Machine Learning: Face Recognition, Mohammed E. Amin

Publications and Research

This project explores the cutting-edge intersection of machine learning (ML) and face recognition (FR) technology, utilizing the OpenCV library to pioneer innovative applications in real-time security and user interface enhancement. By processing live video feeds, our system encodes visual inputs and employs advanced face recognition algorithms to accurately identify individuals from a database of photos. This integration of machine learning with OpenCV not only showcases the potential for bolstering security systems but also enriches user experiences across various technological platforms. Through a meticulous examination of unique facial features and the application of sophisticated ML algorithms and neural networks, our project …