Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Databases and Information Systems (598)
- Engineering (554)
- Artificial Intelligence and Robotics (453)
- Programming Languages and Compilers (435)
- Computer Engineering (369)
-
- Graphics and Human Computer Interfaces (308)
- Other Computer Sciences (238)
- Information Security (227)
- Theory and Algorithms (204)
- Systems Architecture (197)
- Social and Behavioral Sciences (194)
- OS and Networks (174)
- Business (157)
- Numerical Analysis and Scientific Computing (146)
- Education (140)
- Electrical and Computer Engineering (104)
- Medicine and Health Sciences (97)
- Computer and Systems Architecture (94)
- Data Science (86)
- Digital Communications and Networking (71)
- Operations Research, Systems Engineering and Industrial Engineering (69)
- Communication (54)
- Life Sciences (54)
- Environmental Sciences (53)
- Arts and Humanities (48)
- Technology and Innovation (42)
- Systems Engineering (41)
- Institution
-
- Singapore Management University (2211)
- California Polytechnic State University, San Luis Obispo (206)
- Western University (130)
- Air Force Institute of Technology (124)
- University of Malaya (114)
-
- City University of New York (CUNY) (100)
- California State University, San Bernardino (88)
- MMU Press (74)
- Old Dominion University (72)
- Portland State University (50)
- Edith Cowan University (48)
- United Arab Emirates University (48)
- University of Nevada, Las Vegas (48)
- University of Arkansas, Fayetteville (42)
- Loyola University Chicago (40)
- Chapman University (36)
- San Jose State University (36)
- University of Nebraska - Lincoln (35)
- Kennesaw State University (34)
- Embry-Riddle Aeronautical University (32)
- St. Mary's University (31)
- Rochester Institute of Technology (29)
- The University of Akron (23)
- Purdue University (22)
- University of Dayton (22)
- Technological University Dublin (21)
- Dakota State University (18)
- Universitas Negeri Yogyakarta (17)
- University of Nebraska at Omaha (17)
- Institute of Business Administration (16)
- Keyword
-
- Software engineering (152)
- Software (83)
- Deep learning (80)
- Machine learning (77)
- Software Engineering (62)
-
- Android (60)
- Machine Learning (59)
- Computer Science (52)
- Deep Learning (49)
- Empirical study (47)
- Software development (44)
- Refactoring (42)
- Computer science (38)
- Security (37)
- Programming (36)
- Java (35)
- Software maintenance (34)
- Software testing (34)
- Collaboration (32)
- Model Check (29)
- Testing (29)
- GitHub (27)
- Python (26)
- Stack Overflow (25)
- Data mining (24)
- Visualization (24)
- Artificial Intelligence (23)
- Computer software -- Development (23)
- Large language models (23)
- Empirical software engineering (22)
- Publication Year
- Publication
-
- Research Collection School Of Computing and Information Systems (2149)
- Theses and Dissertations (144)
- Electrical and Computer Engineering Publications (130)
- Collaborative Agent Design (CAD) Research Center (103)
- Student Works (2000-2009) (103)
-
- Journal of Informatics and Web Engineering (74)
- Theses Digitization Project (73)
- Publications and Research (67)
- Master's Theses (47)
- Dissertations and Theses Collection (Open Access) (40)
- Computer Science: Faculty Publications and Other Works (39)
- Theses (35)
- Computer Science Faculty Publications (31)
- Theses : Honours (28)
- Articles (27)
- Computer Science and Software Engineering (27)
- Computer Engineering (24)
- Open Educational Resources (24)
- Separations Campaign (TRP) (24)
- Williams Honors College, Honors Research Projects (23)
- Computer Science Faculty Publications and Presentations (21)
- Electronic Theses and Dissertations (21)
- Honors Theses (21)
- Computer Science and Computer Engineering Undergraduate Honors Theses (20)
- Faculty Publications (19)
- Dissertations (18)
- Master's Projects (18)
- University Honors Theses (18)
- Elinvo (Electronics, Informatics, and Vocational Education) (17)
- School of Computing: Dissertations, Theses, and Student Research (17)
- Publication Type
- File Type
Articles 481 - 510 of 4404
Full-Text Articles in Software Engineering
How Effective Are They? Exploring Large Language Model Based Fuzz Driver Generation, Cen Zhang, Yaowen Zheng, Mingqiang Bai, Yeting Li, Wei Ma, Xiaofei Xie
How Effective Are They? Exploring Large Language Model Based Fuzz Driver Generation, Cen Zhang, Yaowen Zheng, Mingqiang Bai, Yeting Li, Wei Ma, Xiaofei Xie
Research Collection School Of Computing and Information Systems
Fuzz drivers are essential for library API fuzzing. However, automatically generating fuzz drivers is a complex task, as it demands the creation of high-quality, correct, and robust API usage code. An LLM-based (Large Language Model) approach for generating fuzz drivers is a promising area of research. Unlike traditional program analysis-based generators, this text-based approach is more generalized and capable of harnessing a variety of API usage information, resulting in code that is friendly for human readers. However, there is still a lack of understanding regarding the fundamental issues on this direction, such as its e ectiveness and potential challenges. To …
Integrating Blockchain Technology Into The Software Development Life Cycle To Satisfy The Software Bill Of Materials Requirement For Government Software Systems, Walter T. Scott Ii
Integrating Blockchain Technology Into The Software Development Life Cycle To Satisfy The Software Bill Of Materials Requirement For Government Software Systems, Walter T. Scott Ii
Theses and Dissertations
This thesis explores the integration of Blockchain Technology (BT) into the Software Development Life Cycle (SDLC) to satisfy the Software Bill of Materials (SBOM) requirement for government software systems. This study begins by synthesizing a standard SDLC definition from various government and industry references, which may provide the foundation for future efforts to standardize software development practices across the government software development community. This study proceeds to define working definitions for the software supply chain (SSC) and software supply chain management (SCM) before introducing and detailing the SBOM requirement as well as providing an overview of prior research regarding SBOMs …
Granular3d: Delving Into Multi-Granularity 3d Scene Graph Prediction, Kaixiang Huang, Jingru Yang, Jin Wang, Shengfeng He, Zhan Wang, Haiyan He, Qifeng Zhang, Guodong Lu
Granular3d: Delving Into Multi-Granularity 3d Scene Graph Prediction, Kaixiang Huang, Jingru Yang, Jin Wang, Shengfeng He, Zhan Wang, Haiyan He, Qifeng Zhang, Guodong Lu
Research Collection School Of Computing and Information Systems
This paper addresses the significant challenges in 3D Semantic Scene Graph (3DSSG) prediction, essential for understanding complex 3D environments. Traditional approaches, primarily using PointNet and Graph Convolutional Networks, struggle with effectively extracting multi-grained features from intricate 3D scenes, largely due to a focus on global scene processing and single-scale feature extraction. To overcome these limitations, we introduce Granular3D, a novel approach that shifts the focus towards multi-granularity analysis by predicting relation triplets from specific sub-scenes. One key is the Adaptive Instance Enveloping Method (AIEM), which establishes an approximate envelope structure around irregular instances, providing shape-adaptive local point cloud sampling, thereby …
An Empirical Study Of Static Analysis Tools For Secure Code Review, Wachiraphan Charoenwet, Patanamon Thongtanunam, Van-Thuan Pham, Christoph Treude
An Empirical Study Of Static Analysis Tools For Secure Code Review, Wachiraphan Charoenwet, Patanamon Thongtanunam, Van-Thuan Pham, Christoph Treude
Research Collection School Of Computing and Information Systems
Early identification of security issues in software development is vital to minimize their unanticipated impacts. Code review is a widely used manual analysis method that aims to uncover security issues along with other coding issues in software projects. While some studies suggest that automated static application security testing tools (SASTs) could enhance security issue identification, there is limited understanding of SAST’s practical effectiveness in supporting secure code review. Moreover, most SAST studies rely on synthetic or fully vulnerable versions of the subject program, which may not accurately represent real-world code changes in the code review process. To address this gap, …
Evaluating Szz Implementations : An Empirical Study On The Linux Kernel, Yunbo Lyu, Hong Jin Kang, Ratnadira Widyasari, Julia Lawall, David Lo
Evaluating Szz Implementations : An Empirical Study On The Linux Kernel, Yunbo Lyu, Hong Jin Kang, Ratnadira Widyasari, Julia Lawall, David Lo
Research Collection School Of Computing and Information Systems
The SZZ algorithm is used to connect bug-fixing commits to the earlier commits that introduced bugs. This algorithm has many applications and many variants have been devised. However, there are some types of commits that cannot be traced by the SZZ algorithm, referred to as “ghost commits”. The evaluation of how these ghost commits impact the SZZ implementations remains limited. Moreover, these implementations have been evaluated on datasets created by software engineering researchers from information in bug trackers and version controlled histories. Since Oct 2013, the Linux kernel developers have started labelling bug-fixing patches with the commit identifiers of the …
Ai Coders Are Among Us : Rethinking Programming Language Grammar Towards Efficient Code Generation, Sun Zhensu, Du Xiaoning, Yang Zhou, Li Li, David Lo
Ai Coders Are Among Us : Rethinking Programming Language Grammar Towards Efficient Code Generation, Sun Zhensu, Du Xiaoning, Yang Zhou, Li Li, David Lo
Research Collection School Of Computing and Information Systems
Artificial Intelligence (AI) models have emerged as another important audience for programming languages alongside humans and machines, as we enter the era of large language models (LLMs). LLMs can now perform well in coding competitions and even write programs like developers to solve various tasks, including mathematical problems. However, the grammar and layout of current programs are designed to cater the needs of human developers -- with many grammar tokens and formatting tokens being used to make the code easier for humans to read. While this is helpful, such a design adds unnecessary computational work for LLMs, as each token …
Meta-Learning For Multi-Family Android Malware Classification, Yao Li, Dawei Yuan, Tao Zhang, Haipeng Cai, David Lo, Cuiyun Gao, Xiapu Luo, He Jiang
Meta-Learning For Multi-Family Android Malware Classification, Yao Li, Dawei Yuan, Tao Zhang, Haipeng Cai, David Lo, Cuiyun Gao, Xiapu Luo, He Jiang
Research Collection School Of Computing and Information Systems
With the emergence of smartphones, Android has become a widely used mobile operating system. However, it is vulnerable when encountering various types of attacks. Every day, new malware threatens the security of users' devices and private data. Many methods have been proposed to classify malicious applications, utilizing static or dynamic analysis for classification. However, previous methods still suffer from unsatisfactory performance due to two challenges. First, they are unable to address the imbalanced data distribution problem, leading to poor performance for malware families with few members. Second, they are unable to address the zero-day malware (zero-day malware refers to malicious …
Real Time Pii Scanning, John David
Real Time Pii Scanning, John David
Electronic Theses and Dissertations
The increased amount of web applications and internet software solutions utilizing cloud frameworks has contributed to large data sets of system log messages being generated constantly. These messages may contain sensitive data, creating an additional security risk for the systems and contributing to the need for analysis of such large volumes of data in real time. Large commercial data monitoring systems can solve for these analysis requirements, but they can be costly. We present a solution to analyzing web application log data which ingests it, processes it and visualizes sensitive data found within in real time. Our solution utilizes an …
Engineering Ecological Analysis Of Rise Of Huawei’S Harmonyos And Its Implications, Dazhou Wang, Yishi Lyu, Zhihuan Fu
Engineering Ecological Analysis Of Rise Of Huawei’S Harmonyos And Its Implications, Dazhou Wang, Yishi Lyu, Zhihuan Fu
Bulletin of Chinese Academy of Sciences (Chinese Version)
This study aims to explore the rise of the Harmony operating system from the perspective of engineering ecology. In the face of increasingly fierce global technological competition, Huawei, as a leading Chinese information technology enterprise, launched its self-developed HarmonyOS and has progressively built the Harmony ecosystem, providing foundational support for the development of all sectors. The development path of HarmonyOS can be divided into three main phases based on its strategic goals, technological and product characteristics, and ecological construction: the initial phase, the acceleration phase, and the transformation phase. It is revealed that, throughout this development, various construction strategies for …
Challenges And Practices Of Deep Learning Model Reengineering: A Case Study On Computer Vision, Wenxin Jiang, Vishnu Banna, Naveen Vivek, Abhinav Goel, Nicholas Synovic, George K. Thiruvathukal, James C. Davis
Challenges And Practices Of Deep Learning Model Reengineering: A Case Study On Computer Vision, Wenxin Jiang, Vishnu Banna, Naveen Vivek, Abhinav Goel, Nicholas Synovic, George K. Thiruvathukal, James C. Davis
Computer Science: Faculty Publications and Other Works
Many engineering organizations are reimplementing and extending deep neural networks from the research community. We describe this process as deep learning model reengineering. Deep learning model reengineering — reusing, replicating, adapting, and enhancing state-of-the-art deep learning approaches — is challenging for reasons including under-documented reference models, changing requirements, and the cost of implementation and testing.
Reless: A Framework For Assessing Safety In Deep Learning Systems, Nan Jia, Anita Raja, Raffi T. Khatchadourian
Reless: A Framework For Assessing Safety In Deep Learning Systems, Nan Jia, Anita Raja, Raffi T. Khatchadourian
Publications and Research
Traditionally, software refactoring helps to improve a system's internal structure and enhance its non-functional features, such as reliability and run-time performance, while preserving external behavior including original program semantics. However, in the context of learning-enabled software systems (LESS), e.g., Machine Learning (ML) systems, it is unclear which portions of a software's semantics require preservation at the development phase. This is mainly because (a) the behavior of the LESS is not defined until run-time; and (b) the inherently iterative and non-deterministic nature of ML algorithms. Consequently, there is a knowledge gap in what refactoring truly means in the context of LESS …
Feature Importance In The Context Of Traditional And Just-In-Time Software Defect Prediction Models, Susmita Haldar, Luiz Fernando Capretz
Feature Importance In The Context Of Traditional And Just-In-Time Software Defect Prediction Models, Susmita Haldar, Luiz Fernando Capretz
Electrical and Computer Engineering Publications
Software defect prediction models can assist software testing initiatives by prioritizing testing error-prone modules. In recent years, in addition to the traditional defect prediction model approach of predicting defects from class, modules, etc., Just-In- Time defect prediction research, which focuses on the change history of software products is getting prominent. For building these defect prediction models, it is important to understand which features are primary contributors to these classifiers. This study considered developing defect prediction models incorporating the traditional and the Just-In-Time approaches from the publicly available dataset of the Apache Camel project. A multi-layer deep learning algorithm was applied …
Materials Data Science Ontology (Mds-Onto): Unifying Domain Knowledge In Materials And Applied Data Science, Van D. Tran, Jonathan E. Gordon, Alexander Harding Bradley, Balashanmuga Priyan Rajamohan, Quynh D. Tran, Gabriel Ponón, Yinghui Wu, Laura S. Bruckman, Erika I. Barcelos, Roger H. French
Materials Data Science Ontology (Mds-Onto): Unifying Domain Knowledge In Materials And Applied Data Science, Van D. Tran, Jonathan E. Gordon, Alexander Harding Bradley, Balashanmuga Priyan Rajamohan, Quynh D. Tran, Gabriel Ponón, Yinghui Wu, Laura S. Bruckman, Erika I. Barcelos, Roger H. French
Student Scholarship
Ontologies have gained popularity in the scientific community as a means of standardizing concepts and terminology used in metadata across different institutions to facilitate data comprehension, sharing, and reuse. Despite the existence of frameworks and guidelines for building ontologies, the processes and standards used to develop ontologies still differ significantly, particularly in Materials Science. Our goal with the MDS-Onto Framework is to provide a unified and automated system for ontology development in the Materials and Data Sciences. This framework offers recommendations on where to publish ontologies online, how to best integrate them within the semantic web, and which formats to …
Neural Network Semantic Backdoor Detection And Mitigation: A Causality-Based Approach, Bing Sun, Jun Sun, Wayne Koh, Jie Shi
Neural Network Semantic Backdoor Detection And Mitigation: A Causality-Based Approach, Bing Sun, Jun Sun, Wayne Koh, Jie Shi
Research Collection School Of Computing and Information Systems
Different from ordinary backdoors in neural networks which are introduced with artificial triggers (e.g., certain specific patch) and/or by tampering the samples, semantic backdoors are introduced by simply manipulating the semantic, e.g., by labeling green cars as frogs in the training set. By focusing on samples with rare semantic features (such as green cars), the accuracy of the model is often minimally affected. Since the attacker is not required to modify the input sample during training nor inference time, semantic backdoors are challenging to detect and remove. Existing backdoor detection and mitigation techniques are shown to be ineffective with respect …
A New Hope: Contextual Privacy Policies For Mobile Applications And An Approach Toward Automated Generation, Shidong Pan, Zhen Tao, Thong Hoang, Dawen Zhang, Tianshi Li, Zhenchang Xing, Xiwei Xu, Mark Staples, Thierry Rakotoarivelo, David Lo
A New Hope: Contextual Privacy Policies For Mobile Applications And An Approach Toward Automated Generation, Shidong Pan, Zhen Tao, Thong Hoang, Dawen Zhang, Tianshi Li, Zhenchang Xing, Xiwei Xu, Mark Staples, Thierry Rakotoarivelo, David Lo
Research Collection School Of Computing and Information Systems
Privacy policies have emerged as the predominant approach to conveying privacy notices to mobile application users. In an effort to enhance both readability and user engagement, the concept of contextual privacy policies (CPPs) has been proposed by researchers. The aim of CPPs is to fragment privacy policies into concise snippets, displaying them only within the corresponding contexts within the application’s graphical user interfaces (GUIs). In this paper, we first formulate CPP in mobile application scenario, and then present a novel multimodal framework, named SEEPRIVACY, specifically designed to automatically generate CPPs for mobile applications. This method uniquely integrates vision-based GUI understanding …
Exponential Qubit Reduction In Optimization For Financial Transaction Settlement, Elias X. Huber, Benjamin Y. L. Tan, Paul Robert Griffin, Dimitris G. Angelakis
Exponential Qubit Reduction In Optimization For Financial Transaction Settlement, Elias X. Huber, Benjamin Y. L. Tan, Paul Robert Griffin, Dimitris G. Angelakis
Research Collection School Of Computing and Information Systems
We extend the qubit-efficient encoding presented in (Tan et al. in Quantum 5:454, 2021) and apply it to instances of the financial transaction settlement problem constructed from data provided by a regulated financial exchange. Our methods are directly applicable to any QUBO problem with linear inequality constraints. Our extension of previously proposed methods consists of a simplification in varying the number of qubits used to encode correlations as well as a new class of variational circuits which incorporate symmetries thereby reducing sampling overhead, improving numerical stability and recovering the expression of the cost objective as a Hermitian observable. We also …
Remote Onboarding Of Software Developers: Leveraging Virtual Reality And Ai Tools, James Dominic
Remote Onboarding Of Software Developers: Leveraging Virtual Reality And Ai Tools, James Dominic
All Dissertations
Software development teams add newcomers to accommodate the increasing demand, complexity of software solutions, and turnover. Onboarding newcomers is expensive and error-prone. It can take up to three years for a newcomer to become an expert on a project. Onboarding techniques described in the current literature focus on collocated teams. As more teams are adopting remote and distributed team structures, I address this research gap in understanding remote onboarding for software developers. I present my research on the use of Virtual Reality (VR) for remote software developer onboarding. I discuss a VR remote pair programming environment. With positive outcomes, pair …
Vysion Software, Isaias Hernandez-Dominguez Jr, Chander Luderman Miller
Vysion Software, Isaias Hernandez-Dominguez Jr, Chander Luderman Miller
2024 Symposium
Vision loss presents significant challenges in daily life. Existing solutions for blind and visually impaired individuals are often limited in functionality, expensive, or complex to use. Vysion Software addresses this gap by developing a user-friendly, all-in-one AI companion app that provides features including text summarization, real-time audio descriptions, and AI-enhanced navigation. This project details the development plan, initial functionalities, and future vision for Vysion Software.
Microservices Architecture: Evolution, Realizing Benefits, And Addressing Challenges In The Modern Software Era -A Systematic Literature Review, Linah M. Elnaghi, Ramadan Moawad
Microservices Architecture: Evolution, Realizing Benefits, And Addressing Challenges In The Modern Software Era -A Systematic Literature Review, Linah M. Elnaghi, Ramadan Moawad
Future Computing and Informatics Journal
This paper explores the world of modern software development and the rising popularity of microservices architecture. Microservices, a modern approach, brings benefits like scalability, Reusability, and fault tolerance. challenging traditional monolithic approaches.This survey involves a detailed comparison, unraveling the motivations behind the wide usage of microservices. This paper extracts insights from a diverse range of studies, presenting a clear and accessible synthesis of the key benefits and challenges associated with microservices architecture. Through a methodical analysis of these factors, the study aims to discern the most pivotal advantages and challenges within the domain of microservices. Steering away from complicated terminology, …
Peatmoss: A Dataset And Initial Analysis Of Pre-Trained Models In Open-Source Software, Wenxin Jiang, Jerin Yasmin, Jason Jones, Nicholas Synovic, Jiashen Kuo, Nathaniel Bielanski, Yuan Tian, George K. Thiruvathukal, James C. Davis
Peatmoss: A Dataset And Initial Analysis Of Pre-Trained Models In Open-Source Software, Wenxin Jiang, Jerin Yasmin, Jason Jones, Nicholas Synovic, Jiashen Kuo, Nathaniel Bielanski, Yuan Tian, George K. Thiruvathukal, James C. Davis
Computer Science: Faculty Publications and Other Works
The development and training of deep learning models have become increasingly costly and complex. Consequently, software engineers are adopting pre-trained models (PTMs) for their downstream applications. The dynamics of the PTM supply chain remain largely unexplored, signaling a clear need for structured datasets that document not only the metadata but also the subsequent applications of these models. Without such data, the MSR community cannot comprehensively understand the impact of PTM adoption and reuse. This paper presents the PeaTMOSS dataset, which comprises metadata for 281,638 PTMs and detailed snapshots for all PTMs with over 50 monthly downloads (14,296 PTMs), along with …
The Impacts Of Dimensionality, Diffusion, And Directedness On Intrinsic Cross-Model Simulation In Tile-Based Self-Assembly, Daniel Hader, Matthew J. Patitz
The Impacts Of Dimensionality, Diffusion, And Directedness On Intrinsic Cross-Model Simulation In Tile-Based Self-Assembly, Daniel Hader, Matthew J. Patitz
Computer Science and Computer Engineering Faculty Publications and Presentations
Motivated by applications in DNA-nanotechnology, theoretical investigations in algorithmic tile-assembly have blossomed into a mature theory. In addition to computational universality, the abstract Tile Assembly Model (aTAM) was shown to be intrinsically universal (FOCS 2012), a strong notion of completeness where a single tile set is capable of simulating the full dynamics of all systems within the model; however, this construction fundamentally required non-deterministic tile attachments. This was confirmed necessary when it was shown that the class of directed aTAM systems, those where all possible sequences of tile attachments result in the same terminal assembly, is not intrinsically universal (FOCS …
Reproducibility Debt: Challenges And Future Pathways, Zara Hassan, Christoph Treude, Michael Norrish, Graham Williams, Alex Potanin
Reproducibility Debt: Challenges And Future Pathways, Zara Hassan, Christoph Treude, Michael Norrish, Graham Williams, Alex Potanin
Research Collection School Of Computing and Information Systems
Reproducibility of scientic computation is a critical factor in validating its underlying process, but it is often elusive. Complexity and continuous evolution in software systems have introduced new challenges for reproducibility across a myriad of computational sciences, resulting in growing debt. This requires a comprehensive domain-agnostic study to dene and asses Reproducibility Debt (RpD) in scientic software, thus uncovering and classifying all underlying factors attributed towards its emergence and identication i.e., causes and eects. Moreover, an organised map of prevention strategies is imperative to guide researchers for its proactive management. This vision paper highlights the challenges that hinder eective management …
On The Sustainability Of Deep Learning Projects: Maintainers' Perspective, Junxiao Han, Jiakun Liu, David Lo, Chen Zhi, Yishan Chen, Shuiguang Deng
On The Sustainability Of Deep Learning Projects: Maintainers' Perspective, Junxiao Han, Jiakun Liu, David Lo, Chen Zhi, Yishan Chen, Shuiguang Deng
Research Collection School Of Computing and Information Systems
Deep learning (DL) techniques have grown in leaps and bounds in both academia and industry over the past few years. Despite the growth of DL projects, there has been little study on how DL projects evolve, whether maintainers in this domain encounter a dramatic increase in workload and whether or not existing maintainers can guarantee the sustained development of projects. To address this gap, we perform an empirical study to investigate the sustainability of DL projects, understand maintainers' workloads and workloads growth in DL projects, and compare them with traditional open-source software (OSS) projects. In this regard, we first investigate …
Toward Effective Secure Code Reviews: An Empirical Study Of Security-Related Coding Weaknesses, Wachiraphan Charoenwet, Patanamon Thongtanunam, Thuan Pham, Christoph Treude
Toward Effective Secure Code Reviews: An Empirical Study Of Security-Related Coding Weaknesses, Wachiraphan Charoenwet, Patanamon Thongtanunam, Thuan Pham, Christoph Treude
Research Collection School Of Computing and Information Systems
Identifying security issues early is encouraged to reduce the latent negative impacts on software systems. Code review is a widely-used method that allows developers to manually inspect modified code, catching security issues during a software development cycle. However, existing code review studies often focus on known vulnerabilities, neglecting coding weaknesses, which can introduce real-world security issues that are more visible through code review. The practices of code reviews in identifying such coding weaknesses are not yet fully investigated. To better understand this, we conducted an empirical case study in two large open-source projects, OpenSSL and PHP. Based on 135,560 code …
Generative Ai For Pull Request Descriptions: Adoption, Impact, And Developer Interventions, Tao Xiao, Hideaki Hata, Christoph Treude, Kenichi Matsumoto
Generative Ai For Pull Request Descriptions: Adoption, Impact, And Developer Interventions, Tao Xiao, Hideaki Hata, Christoph Treude, Kenichi Matsumoto
Research Collection School Of Computing and Information Systems
GitHub's Copilot for Pull Requests (PRs) is a promising service aiming to automate various developer tasks related to PRs, such as generating summaries of changes or providing complete walkthroughs with links to the relevant code. As this innovative technology gains traction in the Open Source Software (OSS) community, it is crucial to examine its early adoption and its impact on the development process. Additionally, it offers a unique opportunity to observe how developers respond when they disagree with the generated content. In our study, we employ a mixed-methods approach, blending quantitative analysis with qualitative insights, to examine 18,256 PRs in …
Certified Robust Accuracy Of Neural Networks Are Bounded Due To Bayes Errors, Ruihan Zhang, Jun Sun
Certified Robust Accuracy Of Neural Networks Are Bounded Due To Bayes Errors, Ruihan Zhang, Jun Sun
Research Collection School Of Computing and Information Systems
Adversarial examples pose a security threat to many critical systems built on neural networks. While certified training improves robustness, it also decreases accuracy noticeably. Despite various proposals for addressing this issue, the significant accuracy drop remains. More importantly, it is not clear whether there is a certain fundamental limit on achieving robustness whilst maintaining accuracy. In this work, we offer a novel perspective based on Bayes errors. By adopting Bayes error to robustness analysis, we investigate the limit of certified robust accuracy, taking into account data distribution uncertainties. We first show that the accuracy inevitably decreases in the pursuit of …
Partial Solution Based Constraint Solving Cache In Symbolic Execution, Ziqi Shuai, Zhenbang Chen, Kelin Ma, Kunlin Liu, Yufeng Zhang, Jun Sun, Ji Wang
Partial Solution Based Constraint Solving Cache In Symbolic Execution, Ziqi Shuai, Zhenbang Chen, Kelin Ma, Kunlin Liu, Yufeng Zhang, Jun Sun, Ji Wang
Research Collection School Of Computing and Information Systems
Constraint solving is one of the main challenges for symbolic execution. Caching is an effective mechanism to reduce the number of the solver invocations in symbolic execution and is adopted by many mainstream symbolic execution engines. However, caching can not perform well on all programs. How to improve caching’s effectiveness is challenging in general. In this work, we propose a partial solution-based caching method for improving caching’s effectiveness. Our key idea is to utilize the partial solutions inside the constraint solving to generate more cache entries. A partial solution may satisfy other constraints of symbolic execution. Hence, our partial solution-based …
Prioritising Github Priority Labels, James Caddy, Christoph Treude
Prioritising Github Priority Labels, James Caddy, Christoph Treude
Research Collection School Of Computing and Information Systems
Communities on GitHub often use issue labels as a way of triaging issues by assigning them priority ratings based on how urgently they should be addressed. The labels used are determined by the repository contributors and notstandardisedbyGitHub.Thismakes it difficult for priority-related reasoning across repositories for both researchers and contributors. Previous work shows interest in how issues are labelled and what the consequences for those labels are. For instance, some previous work has used clustering models and natural language processing to categorise labels without a particular emphasis on priority. With this publication, we introduce a unique data set of 812 manually …
Transformer Models With Explainability For It Telemetry And Business Events, Shiau Hong Lim, Laura Wynter
Transformer Models With Explainability For It Telemetry And Business Events, Shiau Hong Lim, Laura Wynter
Research Collection School Of Computing and Information Systems
Temporal event data are commonly encountered in software applications across a wide range of domains, from IT telemetry and system logs to business process automation. Temporal event data in software applications carry information in various forms: both structured and unstructured, and with both regular and irregular occurrence frequency, making the representation of temporal event data challenging. Further chal-lenges come from the diversity in terms of the scale and volume of events that need to be summarized by the representation. We propose a general and unified approach to handle temporal event data for the purpose of learning a predictive transformer model. …
Hierarchical Damage Correlations For Old Photo Restoration, Weiwei Cai, Xuemiao Xu, Jiajia Xu, Huaidong Zhang, Haoxin Yang, Kun Zhang, Shengfeng He
Hierarchical Damage Correlations For Old Photo Restoration, Weiwei Cai, Xuemiao Xu, Jiajia Xu, Huaidong Zhang, Haoxin Yang, Kun Zhang, Shengfeng He
Research Collection School Of Computing and Information Systems
Restoring old photographs can preserve cherished memories. Previous methods handled diverse damages within the same network structure, which proved impractical. In addition, these methods cannot exploit correlations among artifacts, especially in scratches versus patch-misses issues. Hence, a tailored network is particularly crucial. In light of this, we propose a unified framework consisting of two key components: ScratchNet and PatchNet. In detail, ScratchNet employs the parallel Multi-scale Partial Convolution Module to effectively repair scratches, learning from multi-scale local receptive fields. In contrast, the patch-misses necessitate the network to emphasize global information. To this end, we incorporate a transformer-based encoder and decoder …