Open Access. Powered by Scholars. Published by Universities.®
Programming Languages and Compilers Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
- Keyword
-
- Deep learning (2)
- Software engineering (2)
- Adversarial Robustness (1)
- Anytime neural network (1)
- Capsule network (1)
-
- Code instrumentation (1)
- Code learning (1)
- Cross-lingual (1)
- Data-driven Optimization (1)
- Deep neural network (1)
- Domain Adaptation (1)
- Duplicate bug reports (1)
- Dynamic Pickup and Delivery Problem (1)
- Dynamic Vehicle Routing Problem (1)
- Empirical study (1)
- Fact Checking (1)
- General AI (1)
- Hardware virtualization (1)
- Hateful meme detection (1)
- Indexing (1)
- Interpretability (1)
- LLM (1)
- Language Processing (1)
- Large Language Model (1)
- Large language models (1)
- Machine learning (1)
- Mobile deep learning (1)
- Mobile sensing (1)
- Model approximation (1)
- Model pruning (1)
Articles 1 - 20 of 20
Full-Text Articles in Programming Languages and Compilers
Harnessing Se Community Knowledge For Developer-Centric Code Intelligence, Chengran Yang
Harnessing Se Community Knowledge For Developer-Centric Code Intelligence, Chengran Yang
Dissertations and Theses Collection (Open Access)
The integration of Large Language Models (LLMs), particularly those tailored for programming tasks—referred to as code LLMs—has created novel opportunities to enhance developer productivity. These advanced models automate routine and repetitive coding tasks, such as code generation and debugging, and enable faster prototyping and more efficient problem-solving. Despite these remarkable advantages, the current generation of code LLMs exhibits notable limitations that impact their practical effectiveness in real-world software engineering scenarios. These models frequently produce code that is inefficient or suboptimal in runtime performance, demonstrate opaque reasoning processes, and struggle to adapt effectively to diverse developer contexts and specific requirements. Moreover, …
Computational Fact-Checking With Limited Resources, Fengzhu Zeng
Computational Fact-Checking With Limited Resources, Fengzhu Zeng
Dissertations and Theses Collection (Open Access)
The rapid dissemination of information through online platforms has sparked widespread concern about the propagation of misinformation. Manual fact-checking by pro- fessional fact-checkers is time-consuming and lacks scalability to address the vast volume of daily information. Consequently, computational fact-checking, driven by automated techniques in natural language processing (NLP), has garnered interest as
a potential solution. However, computational fact-checking faces critical challenges limited resources, particularly due to the issues of data scarcity and computing resource constraints. One key challenge is data scarcity, which arises from the constant generation of new information and emerging events on social media. This scarcity manifests in …
Learning And Optimization Under Human-Centric Considerations, Qian Shao
Learning And Optimization Under Human-Centric Considerations, Qian Shao
Dissertations and Theses Collection (Open Access)
This dissertation investigates learning and optimization problems shaped by humancentric considerations, such as preferences, demonstrations, behavioral patterns, and resource constraints. As real-world decision-making increasingly involves interaction with human agents, data, and limitations, modeling these factors becomes critical for building practical, adaptive, and robust systems.
The research spans four domains. First, we study preference-aware delivery routing by learning implicit practitioner preferences and incorporating them into a hierarchical route optimization framework. Second, we develop imitation learning methods for cost-constrained settings, enabling agents to mimic expert behavior while respecting safety and resource limitations. Third,we explore early rumor detection in data-limited environments, integrating large …
Evaluation Of Pre-Trained Vision Language Models In Challenging Contexts, Kankan Zhou
Evaluation Of Pre-Trained Vision Language Models In Challenging Contexts, Kankan Zhou
Dissertations and Theses Collection (Open Access)
The rapid advancement and proliferation of pre-trained vision-language models (VLMs) have heralded a new era in the realm of artificial intelligence (AI), opening up unprecedented opportunities and challenges alike. This dissertation sets forth on an ambitious and comprehensive journey to critically evaluate the performance and limitations of pre-trained VLMs, particularly in complex and challenging contexts that test the bounds of their capabilities. Our focus is twofold: to rigorously assess the extent of bias embedded in these models, and to meticulously scrutinize their reasoning abilities, highlighting parallels and disparities between machine and human cognition.
We initiate our exploration with a targeted …
Tailoring Transformer-Based Deep Learning For Code Generation And Translation, Imam Nur Bani Yusuf
Tailoring Transformer-Based Deep Learning For Code Generation And Translation, Imam Nur Bani Yusuf
Dissertations and Theses Collection (Open Access)
Software is increasingly pervasive in modern society, making the effective translation of human intent into code essential. Novice programmers often struggle with domain-specific code due to limited background knowledge, while experienced developers face challenges in maintaining evolving largescale codebases. Traditional pattern-based approaches address these issues, but such approaches are task-specific and require significant adaptation for different tasks. Transformer-based models offer a more flexible alternative, as the same architecture can be tailored for diverse programming tasks.
This dissertation investigates how Transformer-based models can be customized for various code generation and translation tasks. First, it introduces Transformer-based approaches that assist end-users with …
Towards Robust, Secure, And Privacy-Aware Large Language Models Of Code, Zhou Yang
Towards Robust, Secure, And Privacy-Aware Large Language Models Of Code, Zhou Yang
Dissertations and Theses Collection (Open Access)
The field of software engineering has witnessed a surge in large language models specifically tailored to understand and process code, which we call large language models for code (LLM4Code). The increasing popularity of LLM4Code is inseparable from three key factors: the availability of extensive datasets compiled from diverse data sources, the advancements in deep learning algorithms and computational power that facilitate the training of these powerful models, and the active engagement and collaboration within the research community fostering innovation and the rapid exchange of ideas and methodologies. As evidenced by a series of studies, LLM4Code has been experiencing rapid development …
Elevating Automated Software Maintenance Tasks With Large Language Models, Xin Zhou
Elevating Automated Software Maintenance Tasks With Large Language Models, Xin Zhou
Dissertations and Theses Collection (Open Access)
Software engineering involves many tasks across different phases such as requirements, design, implementation, testing, and maintenance. Among them, software maintenance is a crucial phase, typically accounting for more than half of the software life cycle's duration.
To boost developer productivity, in recent years, numerous research endeavors in software engineering have sought to automate certain software maintenance tasks through the application of machine learning techniques.
Since 2020, the emergence of advanced Large Language Models (LLMs) of code has opened new avenues for enhancing automated solutions in software maintenance.
This dissertation presents a series of works aimed at advancing automated solutions for …
Towards Faster Inference Of Transformers: Strategies For Accelerating Decoding Processes, Cunxiao Du
Towards Faster Inference Of Transformers: Strategies For Accelerating Decoding Processes, Cunxiao Du
Dissertations and Theses Collection (Open Access)
This thesis delves into the acceleration and optimization of Transformer inference, a subject of increasing importance with the emergence of Large Language Models (LLMs). The study primarily addresses the challenges posed by two inherent properties of Transformers during inference: the quadratic complexity of the attention mechanism and the sequential nature of autoregressive inference. The research is structured into three main parts. The first part enhances the learning capabilities of non-autoregressive Transformers, achieving a remarkable 15.0x acceleration on machine translation tasks. The following section focuses on lossless acceleration through speculative decoding, where the proposed algorithm, Glide with CAPE, is shown to …
Using Pre-Trained Models For Vision-Language Understanding Tasks, Rui Cao
Using Pre-Trained Models For Vision-Language Understanding Tasks, Rui Cao
Dissertations and Theses Collection (Open Access)
In recent years, remarkable progress has been made in Artificial Intelligence (AI), with an increasing focus on integrating AI systems into people’s daily lives. In the context of our diverse world, research attention has shifted towards applying AI to multimodal understanding tasks. This thesis specifically addresses two key modalities, namely, vision and language, and explores Vision-Language Understanding (VLU).
In the past, addressing VLU tasks involved training distinct models from scratch using task-specific data. However, limited by the amount of training data, models may easily overfit the training data and fail to generalize. A recent breakthrough is the development of Pre-trained …
Supporting Software Engineers With Large Language Model-Based Automation, Ting Zhang
Supporting Software Engineers With Large Language Model-Based Automation, Ting Zhang
Dissertations and Theses Collection (Open Access)
In recent years, software engineering (SE) has witnessed significant growth, leading to the creation and sharing of an abundance of software artifacts such as source code, bug reports, and pull requests. Analyzing these artifacts is crucial for comprehending the sentiments of software developers and automating various SE tasks, ultimately leading to more human-centered automated SE and enhancing software development efficiency. However, the diverse and unstructured nature of software text poses a significant challenge to this analysis. In response, researchers have investigated a variety of approaches, including the utilization of natural language processing techniques. The advent of large language models (LLMs), …
Effective And Efficient Semantic Representations And Their Applications, Chong Cher Chia
Effective And Efficient Semantic Representations And Their Applications, Chong Cher Chia
Dissertations and Theses Collection (Open Access)
The proliferation of affordable and compact digital storage has also led to the creation of enormous databases of information, and much attention has been focused on the problem of processing unorganized and unstructured information into some form from which additional value can be extracted. Contemporary approaches to this problem virtually necessitate the use of complex models running on computational systems due to the sheer volume of information to be processed. While it is possible for the model to be fed the actual data as input, typically a representation of the data is used instead. These representations are therefore of interest, …
Reinforcement Learning Approach To Coordinate Real-World Multi-Agent Dynamic Routing And Scheduling, Joe Waldy
Reinforcement Learning Approach To Coordinate Real-World Multi-Agent Dynamic Routing And Scheduling, Joe Waldy
Dissertations and Theses Collection (Open Access)
In this thesis, we study new variants of routing and scheduling problems motivated by real-world problems from the urban logistics and law enforcement domains. In particular, we focus on two key aspects: dynamic and multi-agent. While routing problems such as the Vehicle Routing Problem (VRP) is well-studied in the Operations Research (OR) community, we know that in real-world route planning today, initially-planned route plans and schedules may be disrupted by dynamically-occurring events. In addition, routing and scheduling plans cannot be done in silos due to the presence of other agents which may be independent and self-interested. These requirements create …
Robustness And Cross-Lingual Transfer: An Exploration Of Out-Of-Distribution Scenario In Natural Language Processing, Yu, Sicheng
Robustness And Cross-Lingual Transfer: An Exploration Of Out-Of-Distribution Scenario In Natural Language Processing, Yu, Sicheng
Dissertations and Theses Collection (Open Access)
Most traditional machine learning or deep learning methods are based on the premise that training data and test data are independent and identical distributed, i.e., IID. However, it is just an ideal situation. In real-world applications, test set and training data often follow different distributions, which we refer to as the out of distribution, i.e., OOD, setting. As a result, models trained with traditional methods always suffer from an undesirable performance drop on the OOD test set. It's necessary to develop techniques to solve this problem for real applications. In this dissertation, we present four pieces of work in the …
Novel Deep Learning Methods Combined With Static Analysis For Source Code Processing, Duy Quoc Nghi Bui
Novel Deep Learning Methods Combined With Static Analysis For Source Code Processing, Duy Quoc Nghi Bui
Dissertations and Theses Collection (Open Access)
It is desirable to combine machine learning and program analysis so that one can leverage the best of both to increase the performance of software analytics. On one side, machine learning can analyze the source code of thousands of well-written software projects that can uncover patterns that partially characterize software that is reliable, easy to read, and easy to maintain. On the other side, the program analysis can be used to define rigorous and unique rules that are only available in programming languages, which enrich the representation of source code and help the machine learning to capture the patterns better. …
A Virtualization Based System Infrastructure For Dynamic Program Analysis, Jiaqi Hong
A Virtualization Based System Infrastructure For Dynamic Program Analysis, Jiaqi Hong
Dissertations and Theses Collection (Open Access)
Dynamic malware analysis schemes either run the target program as is in an isolated environment assisted by additional hardware facilities or modify it with instrumentation code statically or dynamically. The hardware-assisted schemes usually trap the target during its execution to a more privileged environment based on the available hardware events. The more privileged environment is not accessible by the untrusted kernel, thus this approach is often applied for transparent and secure kernel analysis. Nevertheless, the isolated environment induces a virtual address gap between the analyzer and the target, which hinders effective and efficient memory introspection and undermines the correctness of …
Multimodal Mobile Sensing Systems For Physiological And Psychological Assessment, Nguyen Phan Sinh Huynh
Multimodal Mobile Sensing Systems For Physiological And Psychological Assessment, Nguyen Phan Sinh Huynh
Dissertations and Theses Collection (Open Access)
Sensing systems for monitoring physiological and psychological states have been studied extensively in both academic and industry research for different applications across various domains. However, most of the studies have been done in the lab environment with controlled and complicated sensor setup, which is only suitable for serious healthcare applications in which the obtrusiveness and immobility can be compromised in a trade-off for accurate clinical screening or diagnosing. The recent substantial development of mobile devices with embedded miniaturized sensors are now allowing new opportunities to adapt and develop such sensing systems in the mobile context. The ability to sense physiological …
Preference Learning And Similarity Learning Perspectives On Personalized Recommendation, Duy Dung Le
Preference Learning And Similarity Learning Perspectives On Personalized Recommendation, Duy Dung Le
Dissertations and Theses Collection (Open Access)
Personalized recommendation, whose objective is to generate a limited list of items (e.g., products on Amazon, movies on Netflix, or pins on Pinterest, etc.) for each user, has gained extensive attention from both researchers and practitioners in the last decade. The necessity of personalized recommendation is driven by the explosion of available options online, which makes it difficult, if not downright impossible, for each user to investigate every option. Product and service providers rely on recommendation algorithms to identify manageable number of the most likely or preferred options to be presented to each user. Also, due to the limited screen …
Exploiting Approximation, Caching And Specialization To Accelerate Vision Sensing Applications, Nguyen Loc Huynh
Exploiting Approximation, Caching And Specialization To Accelerate Vision Sensing Applications, Nguyen Loc Huynh
Dissertations and Theses Collection (Open Access)
Over the past few years, deep learning has emerged as state-of-the-art solutions for many challenging computer vision tasks such as face recognition, object detection, etc. Despite of its outstanding performance, deep neural networks (DNNs) are computational intensive, which prevent them to be widely adopted on billions of mobile and embedded devices with scarce resources. To address that limitation, we
focus on building systems and optimization algorithms to accelerate those models, making them more computational-efficient.
First, this thesis explores the computational capabilities of different existing processors (or co-processors) on modern mobile devices. It recognizes that by leveraging the mobile Graphics Processing …
Feature-Based Transfer Learning In Natural Language Processing, Jianfei Yu
Feature-Based Transfer Learning In Natural Language Processing, Jianfei Yu
Dissertations and Theses Collection (Open Access)
In the past few decades, supervised machine learning approach is one of the most important methodologies in the Natural Language Processing (NLP) community. Although various kinds of supervised learning methods have been proposed to obtain the state-of-the-art performance across most NLP tasks, the bottleneck of them lies in the heavy reliance on the large amount of manually annotated data, which is not always available in our desired target domain/task. To alleviate the data sparsity issue in the target domain/task, an attractive solution is to find sufficient labeled data from a related source domain/task. However, for most NLP applications, due to …
Overfitting In Automated Program Repair: Challenges And Solutions, Dinh Xuan Bach Le
Overfitting In Automated Program Repair: Challenges And Solutions, Dinh Xuan Bach Le
Dissertations and Theses Collection (Open Access)
This chapter discusses the main problem and motivation of this dissertation. It also discusses a quantification of various research issues directly related to the dissertation. A summary of works done will also be presented along with the structure of the dissertation.