Don’T Complete It! Preventing Unhelpful Code Completion For Productive And Sustainable Neural Code Completion Systems,
2025
Singapore Management University
Don’T Complete It! Preventing Unhelpful Code Completion For Productive And Sustainable Neural Code Completion Systems, Zhensu Sun, Xiaoning Du, Fu Song, Shangwen Wang, Mingze Ni, Li Li, David Lo
Research Collection School Of Computing and Information Systems
Currently, large pre-trained language models are widely applied in neural code completion systems. Though large code models significantly outperform their smaller counterparts, around 70% of displayed code completions from Github Copilot are not accepted by developers. Being reviewed but not accepted, their help to developer productivity is considerably limited and may conversely aggravate the workload of developers, as the code completions are automatically and actively generated in state-of-the-art code completion systems as developers type out once the service is enabled. Even worse, considering the high cost of the large code models, it is a huge waste of computing resources and …
Large Scale Machine Learning Over Knowledge Graphs,
2025
University at Albany, State University of New York
Large Scale Machine Learning Over Knowledge Graphs, Bedirhan Gergin
Electronic Theses & Dissertations (2024 - present)
Knowledge graphs (KGs) have become popular across various fields, providing convenient access to web-based knowledge while storing and formalizing domain-specific information. By analyzing KGs, patterns, connections, and dependencies can be identified across different data sources, enabling the inference of new knowledge from given facts. As the use of KGs expands, the size of modern KGs has grown significantly, making them impossible to process within the main memory of a single computer. Distributed computing offers a viable solution to this challenge by leveraging the combined capabilities of multiple servers within a cluster. This thesis explores how distributed computing can be effectively …
Openmuse: Integrating Open-Source Models Into Music Creation Workflows,
2025
Dartmouth College
Openmuse: Integrating Open-Source Models Into Music Creation Workflows, Tyler K. Vergho
Dartmouth College Master’s Theses
This master's thesis introduces OpenMUSE (Open Multimodal Unified Sound Engine), a platform that demonstrates the potential of open-source AI music generation by integrating state-of-the-art deep learning models into a unified system. By unifying ten different open-source models, including MusicGen, AudioLDM2, and custom-trained text-to-symbolic music generation models, OpenMUSE aims to create a user-friendly interface that empowers artists to produce complex, adaptive musical compositions. The system enhances accessibility by providing a simple web interface and natural language controls, while improving controllability through features like melody conditioning and semantic audio editing. Specifically, OpenMUSE offers a digital audio workstation (DAW)-inspired interface that lowers the …
Signal-Based Error Handling: Case Study Using The Bathymetric Attributed Grid Library,
2025
University of New Hampshire - Main Campus
Signal-Based Error Handling: Case Study Using The Bathymetric Attributed Grid Library, Anthony R. Papetti
Honors Theses and Capstones
No abstract provided.
Neural And Computational Approach To Understanding Environmental Modulation Of Behavioral Identity In Zebrafish,
2025
West Virginia University
Neural And Computational Approach To Understanding Environmental Modulation Of Behavioral Identity In Zebrafish, John W. Hageter
Graduate Theses, Dissertations, and Problem Reports (ETD)
Organisms rely on behavior for survival. Animals engage in behaviors that allow for feeding, mating, exploring and navigating their environment among others. Necessary for these behaviors to develop are the environmental factors and underlying circuitry which make behavior possible. Specifically, how the environment guides underlying neural circuitry to develop unique facets or phenotypes of a larger behavior are key to understanding why unique behaviors exist. In this thesis, I build foundational evidence for determining these mechanisms through the use of the zebrafish local search behavior. This is a behavior that zebrafish employ following the loss of environmental illumination where they …
Automated Program Refinement: Guide And Verify Code Large Language Model With Refinement Calculus,
2025
Singapore Management University
Automated Program Refinement: Guide And Verify Code Large Language Model With Refinement Calculus, Yufan Cai, Zhe Hou, David Sanan, Xiaokun Luan, Yun Lin, Jun Sun, Jin Song Dong
Research Collection School Of Computing and Information Systems
Recently, the rise of code-centric large language models (LLMs) appears to have reshaped the software engineering world with low-barrier tools like Copilot that can generate code easily. However, there is no correctness guarantee for the code generated by LLMs, which suffer from the hallucination problem, and their output is fraught with risks. Besides, the end-to-end process from specification to code through LLMs is a non-transparent and uncontrolled black box. This opacity makes it difficult for users to understand and trust the generated code. Addressing these challenges is both necessary and critical. In contrast, program refinement transforms high-level specification statements into …
Performance Evaluation Of Newsql Databases In A Distributed Architecture,
2025
Singapore Management University
Performance Evaluation Of Newsql Databases In A Distributed Architecture, Zhiyao Zhang, Alan @ Ali Madjelisi Megargel, Lingxiao Jiang
Research Collection School Of Computing and Information Systems
In the last decade, application architectures have evolved drastically, moving from monolithic architectures to distributed architectures where deployment has shifted from dedicated on-premises servers to the cloud. Distributed architectures and cloud computing has enabled businesses to scale their application components across different geographical locations. While it is easy to scale the application layer, scaling its database layer that relies on traditional SQL databases is challenging and often is a common source of bottlenecks when it comes to application performance. This paper evaluates the performance characteristics between two NewSQL databases solutions, MySQL NDB Cluster vs. TIBCO ActiveSpaces IMDG. Serving as an …
Neuron Semantic-Guided Test Generation For Deep Neural Networks Fuzzing,
2025
Singapore Management University
Neuron Semantic-Guided Test Generation For Deep Neural Networks Fuzzing, Li Huang, Weifeng Sun, Meng Yan, Zhongxin Liu, Yan Lei, David Lo
Research Collection School Of Computing and Information Systems
In recent years, significant progress has been made in testing methods for deep neural networks (DNNs) to ensure their correctness and robustness. Coverage-guided criteria, such as neuron-wise, layer-wise, and path-/trace-wise, have been proposed for DNN fuzzing. However, existing coverage-based criteria encounter performance bottlenecks for several reasons: Testing Adequacy: Partial neural coverage criteria have been observed to achieve full coverage using only a small number of test inputs. In this case, increasing the number of test inputs does not consistently improve the quality of models. Interpretability: The current coverage criteria lack interpretability. Consequently, testers are unable to identify and understand which …
The Gender Wage Gap In An Online Labor Market: The Cost Of Interruptions,
2025
University of Oxford
The Gender Wage Gap In An Online Labor Market: The Cost Of Interruptions, Abi Adams, Kotaro Hara, Kristy Milland, Chris Callison-Burch
Research Collection School Of Computing and Information Systems
This paper analyses gender differences in working patterns and wages on Amazon Mechanical Turk, a popular online labour platform. Using information on 2 million tasks, we find no gender differences in task selection nor experience. Nonetheless, women earn 20% less per hour on average. Gender differences in working patterns are a significant driver of this wage gap. Women are more likely to interrupt their working time on the platform with consequences for their task completion speed. A follow-up survey shows that the gender differences in working patterns and hourly wages are concentrated amongst workers with children.
Measuring And Improving Api Usability And Quality: A Comprehensive Framework And Empirical Study,
2024
Southern Methodist University
Measuring And Improving Api Usability And Quality: A Comprehensive Framework And Empirical Study, Sultan Alanazy
Computer Science and Engineering Theses and Dissertations
Cloud computing provides on-demand access to flexible computing resources, enabling rapid application deployment without substantial infrastructure investment. Application Programming Interfaces (APIs) play an important role in ensuring the success of cloud applications. The primary users of APIs are the extensive community of application programmers who search, read, and understand APIs before integrating them into their applications or systems. In addition, developers often turn to online API support when seeking help. Problems in such support can result in incorrect API usage and integration problems. There is an urgent need to measure API usability and support issues to identify, characterize, and assess …
Exploration Of The Gap Between The Secure Web Application Development Competencies Needed By Industry And Those Competencies Provided By Graduates Of U.S. Undergraduate Software Engineering Programs,
2024
University of Arkansas Little Rock
Exploration Of The Gap Between The Secure Web Application Development Competencies Needed By Industry And Those Competencies Provided By Graduates Of U.S. Undergraduate Software Engineering Programs, Gary Allen Harris
Theses and Dissertations
Literature demonstrates that threats and attacks on computer systems and networks have been around since the beginning of computing, and the number, severity, sophistication, and costs of attacks and data breaches are continuing to grow. Several studies suggest that one of the most common causes of data breaches is insecure web applications that contain vulnerable application code. These studies suggest that poor secure web application development practices are a prime cause of the susceptible web applications. Additionally, studies suggest that higher education is not meeting industry’s secure software/web application development needs. Employers have reported that they are not getting the …
Dancetag: Using Sensors To Improve Feedback Given To Dance Students,
2024
Chapman University
Dancetag: Using Sensors To Improve Feedback Given To Dance Students, Yanelly Mego, Franceli L. Cibrian
Student Scholar Symposium Abstracts and Posters
The structure of dance classrooms has remained largely unchanged for years, with minimal integration of technology to enhance teaching. This has motivated our research project, which aims to capture dance movements using wearable sensors and translate the information into meaningful visualizations to help dancers improve their skills. As the first step in addressing the research question—can data from commercial wearables differentiate between the movements of dancers and non-dancers?—we developed DANCETAG (Data Analytics and Notation with Captured Event Tagging), a platform designed for data collection and movement annotation. We utilized Sony’s Mocopi sensors, a motion capture system with six sensors attached …
Software Implementations And Analyses Of The Emotional Impact Of Various Binaural Beat Classifications Layered Into Music.,
2024
Chapman University
Software Implementations And Analyses Of The Emotional Impact Of Various Binaural Beat Classifications Layered Into Music., Neil Azimi
Student Scholar Symposium Abstracts and Posters
This study explores the psychoacoustic effects of binaural beats, which are produced when sinusoidal waves of slightly differing frequencies are played into each ear, leading to brainwave entrainment. Binaural beats are categorized by frequency bands (e.g., Beta: 14–30 Hz for energy, Theta: 4–8 Hz for relaxation), each associated with different psychological effects. This research contributes to the field by empirically analyzing whether binaural beats alter the emotional impact of music. Past studies have investigated the potential benefits of binaural beats in relaxation and energy stimulation. However, their effects, when combined with music, especially regarding a song's perceived emotional quality or …
Visualization Of Paleocurrents On A Web Application Using Gplates,
2024
Southern Adventist University
Visualization Of Paleocurrents On A Web Application Using Gplates, Anjan Sapkota
MS in Computer Science Theses
Paleocurrents are flow directions derived from features of sedimentary rocks that reveal the direction of the current of wind or water that deposited the sediment. In 2015, Brand et al. created a global database of paleocurrents, which contains over 1,000,000 measurements worldwide: North America, South America, Australia, Great Britain, parts of Western Europe, China, Africa are fairly well represented; Antarctica, Eastern Europe, and Asia are modestly represented and Russia is poorly represented. The contribution of this thesis is a web application that uses the GPlates’ Application Programming Interface (API) to visualize global paleocurrents through time in an interactive way based …
Graph Neural Networks Powered Scientific Paper Recommendation,
2024
Southern Methodist University
Graph Neural Networks Powered Scientific Paper Recommendation, Junhao Shen
Computer Science and Engineering Theses and Dissertations
Scientific paper recommendation systems aim to help researchers discover relevant papers amidst the vast and ever-growing body of literature. With the exponential yearly increase in scientific publications, the demand for effective paper recommendation solutions has become both critical and increasingly challenging. In recent years, deep learning techniques have revolutionized recommender systems, and scientific paper recommendations have naturally integrated these advancements. In this dissertation, we address these challenges through three progressive contributions.
First, we enhance traditional content-based methods using Graph Neural Networks (GNNs) by introducing a Graph Convolutional Network-strengthened Topic Modeling (GCN-TM) approach. This method improves upon conventional topic modeling techniques …
Triadic Temporal-Semantic Alignment For Weakly-Supervised Video Moment Retrieval,
2024
Singapore Management University
Triadic Temporal-Semantic Alignment For Weakly-Supervised Video Moment Retrieval, Jin Liu, Jialong Xie, Fengyu Zhou, Shengfeng He
Research Collection School Of Computing and Information Systems
Video Moment Retrieval (VMR) aims to identify specific event moments within untrimmed videos based on natural language queries. Existing VMR methods have been criticized for relying heavily on moment annotation bias rather than true multi-modal alignment reasoning. Weakly supervised VMR approaches inherently overcome this issue by training without precise temporal location information. However, they struggle with fine-grained semantic alignment and often yield multiple speculative predictions with prolonged video spans. In this paper, we take a step forward in the context of weakly supervised VMR by proposing a triadic temporalsemantic alignment model. Our proposed approach augments weak supervision by comprehensively addressing …
Ali-Agent: Assessing Llms’ Alignment With Human Values Via Agent-Based Evaluation,
2024
Singapore Management University
Ali-Agent: Assessing Llms’ Alignment With Human Values Via Agent-Based Evaluation, Jingnan Zheng, Han Wang, Tai D. Nguyen, An Zhang, Jun Sun, Tat-Seng Chua
Research Collection School Of Computing and Information Systems
Large Language Models (LLMs) can elicit unintended and even harmful content when misaligned with human values, posing severe risks to users and society. To mitigate these risks, current evaluation benchmarks predominantly employ expertdesigned contextual scenarios to assess how well LLMs align with human values. However, the labor-intensive nature of these benchmarks limits their test scope, hindering their ability to generalize to the extensive variety of open-world use cases and identify rare but crucial long-tail risks. Additionally, these static tests fail to adapt to the rapid evolution of LLMs, making it hard to evaluate timely alignment issues. To address these challenges, …
A Comprehensive Study On Static Application Security Testing (Sast) Tools For Android,
2024
Singapore Management University
A Comprehensive Study On Static Application Security Testing (Sast) Tools For Android, Jingyun Zhu, Kaixuan Li, Sen Chen, Lingling Fan, Junjie Wang, Xiaofei Xie
Research Collection School Of Computing and Information Systems
To identify security vulnerabilities in Android applications, numerous static application security testing (SAST) tools have been proposed. However, it poses significant challenges to assess their overall performance on diverse vulnerability types. The task is non-trivial and poses considerable challenges. Firstly, the absence of a unified evaluation platform for defining and describing tools’ supported vulnerability types, coupled with the lack of normalization for the intricate and varied reports generated by different tools, significantly adds to the complexity. Secondly, there is a scarcity of adequate benchmarks, particularly those derived from real-world scenarios. To address these problems, we are the first to propose …
Mining Work Items To Streamline Software Maintenance Tasks,
2024
University of Nebraska-Lincoln
Mining Work Items To Streamline Software Maintenance Tasks, Salomé Perez-Rosero
Dissertations and Doctoral Documents, University of Nebraska-Lincoln, 2023–
Software engineering maintenance tasks often require associating code changes into groupings of related units of work to have as much information as possible about the developments toward addressing a specific code task. A comprehensive understanding of how a code task has evolved helps developers make better decisions about changes in the overall codebase, where a commit represents the set of code changes made to the codebase at a specific time. While the concept of work items as logically related code changes has been primarily theoretical, its impact on software maintenance tasks, such as tracing the origins of bugs or fixes …
Towards General Conceptual Model Editing Via Adversarial Representation Engineering,
2024
Singapore Management University
Towards General Conceptual Model Editing Via Adversarial Representation Engineering, Yihao Zhang, Zeming Wei, Jun Sun, Meng Sun
Research Collection School Of Computing and Information Systems
Since the rapid development of Large Language Models (LLMs) has achieved remarkable success, understanding and rectifying their internal complex mechanisms has become an urgent issue. Recent research has attempted to interpret their behaviors through the lens of inner representation. However, developing practical and efficient methods for applying these representations for general and flexible model editing remains challenging. In this work, we explore how to leverage insights from representation engineering to guide the editing of LLMs by deploying a representation sensor as an editing oracle. We first identify the importance of a robust and reliable sensor during editing, then propose an …
