Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Artificial Intelligence and Robotics (1009)
- Engineering (915)
- Computer Engineering (516)
- Databases and Information Systems (386)
- Numerical Analysis and Scientific Computing (341)
-
- Information Security (303)
- Operations Research, Systems Engineering and Industrial Engineering (268)
- Software Engineering (260)
- Social and Behavioral Sciences (259)
- Systems Science (232)
- Electrical and Computer Engineering (195)
- Graphics and Human Computer Interfaces (191)
- Data Science (170)
- Business (158)
- Theory and Algorithms (153)
- Medicine and Health Sciences (144)
- Mathematics (140)
- Other Computer Sciences (119)
- Life Sciences (114)
- Education (113)
- Programming Languages and Compilers (95)
- OS and Networks (79)
- Physics (78)
- Arts and Humanities (66)
- Public Affairs, Public Policy and Public Administration (62)
- Applied Mathematics (60)
- Technology and Innovation (59)
- Statistics and Probability (57)
- Institution
-
- Singapore Management University (576)
- China Simulation Federation (223)
- Old Dominion University (160)
- Neutrosophic Systems with Applications (142)
- MBZUAI (115)
-
- Kennesaw State University (112)
- Zayed University (89)
- Chulalongkorn University (88)
- Missouri University of Science and Technology (86)
- University of Texas at El Paso (86)
- San Jose State University (82)
- TÜBİTAK (78)
- Technological University Dublin (57)
- Edith Cowan University (56)
- Air Force Institute of Technology (53)
- University of Nebraska - Lincoln (49)
- City University of New York (CUNY) (45)
- Michigan Technological University (44)
- Utah State University (40)
- University of Texas Rio Grande Valley (39)
- Dartmouth College (37)
- Walden University (35)
- University of Central Florida (32)
- United Arab Emirates University (31)
- Wright State University (31)
- Boise State University (29)
- University of South Florida (28)
- Chapman University (27)
- Portland State University (27)
- University of Texas at Arlington (27)
- Keyword
-
- Machine learning (200)
- Deep learning (150)
- Artificial intelligence (128)
- Machine Learning (108)
- Technical Reports (70)
-
- UTEP Computer Science Department (70)
- Artificial Intelligence (61)
- Deep Learning (61)
- Cybersecurity (58)
- Security (46)
- Computer Science (42)
- AI (38)
- Computer vision (38)
- Natural language processing (37)
- Neural networks (35)
- Blockchain (34)
- Classification (34)
- Reinforcement learning (34)
- ChatGPT (30)
- Computer science (30)
- Optimization (29)
- Privacy (29)
- MCDM (24)
- Neutrosophic Set (24)
- Visualization (24)
- Algorithms (22)
- COVID-19 (22)
- Task analysis (22)
- Training (22)
- Department of Computer Science (21)
- Publication
-
- Research Collection School Of Computing and Information Systems (531)
- Journal of System Simulation (223)
- Neutrosophic Systems with Applications (142)
- Theses and Dissertations (121)
- All Works (89)
-
- Chulalongkorn University Theses and Dissertations (Chula ETD) (88)
- Turkish Journal of Electrical Engineering and Computer Sciences (78)
- Master's Projects (71)
- Departmental Technical Reports (CS) (70)
- Computer Science Faculty Research & Creative Works (66)
- Machine Learning Faculty Publications (62)
- C-Day Computing Showcase (60)
- Computer Science Faculty Publications (54)
- Research outputs 2022 to 2026 (53)
- Walden Dissertations and Doctoral Studies (35)
- Journal of Cybersecurity Education, Research and Practice (34)
- School of Computing: Faculty Publications (31)
- Electronic Theses and Dissertations (30)
- Academic Posters Collection (29)
- Computer Vision Faculty Publications (28)
- Dissertations (28)
- Faculty Scholarship (26)
- Computer Science Faculty Publications and Presentations (25)
- Cybersecurity Undergraduate Research Showcase (25)
- Natural Language Processing Faculty Publications (25)
- USF Tampa Graduate Theses and Dissertations (25)
- Karbala International Journal of Modern Science (23)
- Computer Science: Faculty Publications and Other Works (22)
- Electrical & Computer Engineering Faculty Publications (21)
- Boise State University Theses and Dissertations (19)
- Publication Type
- File Type
Articles 1411 - 1440 of 3503
Full-Text Articles in Computer Sciences
Using Deep Learning For Encrypted Traffic Analysis Of Amazon Echo, Surendra Pathak
Using Deep Learning For Encrypted Traffic Analysis Of Amazon Echo, Surendra Pathak
Theses and Dissertations
The adoption of the Amazon Echo family of devices in modern homes has become very widespread at the current time, with hundreds of millions of devices sold. Moreover, the global smart speaker market size is growing vigorously and is projected to continue to bigger. Smart speakers allow users hands-free interaction by allowing voice control, promoting human-computer interaction to greater avenues. Though smart speaker can be useful assistant, it has some serious security concerns that need to be studied. In this study, an analysis of the security and privacy concerns of smart speakers is presented along with a passive attack, namely …
Socialz: Multi-Feature Social Fuzz Testing, Francisco Zanartu, Christoph Treude, Markus Wagner
Socialz: Multi-Feature Social Fuzz Testing, Francisco Zanartu, Christoph Treude, Markus Wagner
Research Collection School Of Computing and Information Systems
Online social networks have become an integral aspect of our daily lives and play a crucial role in shaping our relationships with others. However, bugs and glitches, even minor ones, can cause anything from frustrating problems to serious data leaks that can have farreaching impacts on millions of users. To mitigate these risks, fuzz testing, a method of testing with randomised inputs, can provide increased confidence in the correct functioning of a social network. However, implementing traditional fuzz testing methods can be prohibitively difficult or impractical for programmers outside of the network’s development team. To tackle this challenge, we present …
Barriers And Self-Efficacy: A Large-Scale Study On The Impact Of Oss Courses On Student Perceptions, Larissa Salerno, Simone De França Tonhão, Igor Steinmacher, Christoph Treude
Barriers And Self-Efficacy: A Large-Scale Study On The Impact Of Oss Courses On Student Perceptions, Larissa Salerno, Simone De França Tonhão, Igor Steinmacher, Christoph Treude
Research Collection School Of Computing and Information Systems
Open source software (OSS) development offers a unique opportunity for students in Software Engineering to experience and participate in large-scale software development, however, the impact of such courses on students’ self-efficacy and the challenges faced by students are not well understood. This paper aims to address this gap by analyzing data from multiple instances of OSS development courses at universities in different countries and reporting on how students’ self-efficacy changed as a result of taking the course, as well as the barriers and challenges faced by students
A Comparative Effectiveness Study On Opioid Use Disorder Prediction Using Artificial Intelligence And Existing Risk Models, Sajjad Fouladvand, Jeffery Talbert, Linda Phyliss Dwoskin, Heather M. Bush, Amy L. Meadows, Lars E. Peterson, Yash R. Mishra, Steven K. Roggenkamp, Fei Wang, Ramakanth Kavuluru, Jin Chen
A Comparative Effectiveness Study On Opioid Use Disorder Prediction Using Artificial Intelligence And Existing Risk Models, Sajjad Fouladvand, Jeffery Talbert, Linda Phyliss Dwoskin, Heather M. Bush, Amy L. Meadows, Lars E. Peterson, Yash R. Mishra, Steven K. Roggenkamp, Fei Wang, Ramakanth Kavuluru, Jin Chen
Markey Cancer Center Faculty Publications
Opioid use disorder (OUD) is a leading cause of death in the United States placing a tremendous burden on patients, their families, and health care systems. Artificial intelligence (AI) can be harnessed with available healthcare data to produce automated OUD prediction tools. In this retrospective study, we developed AI based models for OUD prediction and showed that AI can predict OUD more effectively than existing clinical tools including the unweighted opioid risk tool (ORT). Data include 474,208 patients’ data over 10 years; 269,748 were females with an average age of 56.78 years. Cases are prescription opioid users with at least …
The Metabolomics Workbench File Status Website: A Metadata Repository Promoting Fair Principles Of Metabolomics Data, Christian D. Powell, Hunter N. B. Moseley
The Metabolomics Workbench File Status Website: A Metadata Repository Promoting Fair Principles Of Metabolomics Data, Christian D. Powell, Hunter N. B. Moseley
Markey Cancer Center Faculty Publications
Background: An updated version of the mwtab Python package for programmatic access to the Metabolomics Workbench (MetabolomicsWB) data repository was released at the beginning of 2021. Along with updating the package to match the changes to MetabolomicsWB’s ‘mwTab’ file format specification and enhancing the package’s functionality, the included validation facilities were used to detect and catalog file inconsistencies and errors across all publicly available datasets in MetabolomicsWB.
Results: The MetabolomicsWB File Status website was developed to provide continuous validation of MetabolomicsWB data files and a useful interface to all found inconsistencies and errors. This list of detectable issues/errors include format …
Lecture Notes On Cloud Computing (Ver. Summer 2023), Jun Li
Lecture Notes On Cloud Computing (Ver. Summer 2023), Jun Li
Open Educational Resources
No abstract provided.
Synthesizing Speech Test Cases With Text-To-Speech? An Empirical Study On The False Alarms In Automated Speech Recognition Testing, Julia Kaiwen Lau, Kelvin Kai Wen Kong, Julian Hao Yong, Per Hoong Tan, Zhou Yang, Zi Qian Yong, Joshua Chern Wey Low, Chun Yong Chong, Mei Kuan Lim, David Lo
Synthesizing Speech Test Cases With Text-To-Speech? An Empirical Study On The False Alarms In Automated Speech Recognition Testing, Julia Kaiwen Lau, Kelvin Kai Wen Kong, Julian Hao Yong, Per Hoong Tan, Zhou Yang, Zi Qian Yong, Joshua Chern Wey Low, Chun Yong Chong, Mei Kuan Lim, David Lo
Research Collection School Of Computing and Information Systems
Recent studies have proposed the use of Text-To-Speech (TTS) systems to automatically synthesise speech test cases on a scale and uncover a large number of failures in ASR systems. However, the failures uncovered by synthetic test cases may not reflect the actual performance of an ASR system when it transcribes human audio, which we refer to as false alarms. Given a failed test case synthesised from TTS systems, which consists of TTS-generated audio and the corresponding ground truth text, we feed the human audio stating the same text to an ASR system. If human audio can be correctly transcribed, an …
18 Million Links In Commit Messages: Purpose, Evolution, And Decay, Tao Xiao, Sebastian Baltes, Hideaki Hata, Christoph Treude, Raula Kula, Takashi Ishio, Kenichi Matsumoto
18 Million Links In Commit Messages: Purpose, Evolution, And Decay, Tao Xiao, Sebastian Baltes, Hideaki Hata, Christoph Treude, Raula Kula, Takashi Ishio, Kenichi Matsumoto
Research Collection School Of Computing and Information Systems
Commit messages contain diverse and valuable types of knowledge in all aspects of software maintenance and evolution. Links are an example of such knowledge. Previous work on “9.6 million links in source code comments” showed that links are prone to decay, become outdated, and lack bidirectional traceability. We conducted a large-scale study of 18,201,165 links from commits in 23,110 GitHub repositories to investigate whether they suffer the same fate. Results show that referencing external resources is prevalent and that the most frequent domains other than github.com are the external domains of Stack Overflow and Google Code. Similarly, links serve as …
Chatgpt, Can You Generate Solutions For My Coding Exercises? An Evaluation On Its Effectiveness In An Undergraduate Java Programming Course, Eng Lieh Ouh, Benjamin Gan, Kyong Jin Shim, Swavek Wlodkowski
Chatgpt, Can You Generate Solutions For My Coding Exercises? An Evaluation On Its Effectiveness In An Undergraduate Java Programming Course, Eng Lieh Ouh, Benjamin Gan, Kyong Jin Shim, Swavek Wlodkowski
Research Collection School Of Computing and Information Systems
In this study, we assess the efficacy of employing the ChatGPT language model to generate solutions for coding exercises within an undergraduate Java programming course. ChatGPT, a large-scale, deep learning-driven natural language processing model, is capable of producing programming code based on textual input. Our evaluation involves analyzing ChatGPT-generated solutions for 80 diverse programming exercises and comparing them to the correct solutions. Our findings indicate that ChatGPT accurately generates Java programming solutions, which are characterized by high readability and well-structured organization. Additionally, the model can produce alternative, memory-efficient solutions. However, as a natural language processing model, ChatGPT struggles with coding …
An Efficient Hybrid Genetic Algorithm For The Quadratic Traveling Salesman Problem, Quang Anh Pham, Hoong Chuin Lau, Minh Hoang Ha, Lam Vu
An Efficient Hybrid Genetic Algorithm For The Quadratic Traveling Salesman Problem, Quang Anh Pham, Hoong Chuin Lau, Minh Hoang Ha, Lam Vu
Research Collection School Of Computing and Information Systems
The traveling salesman problem (TSP) is the most well-known problem in combinatorial optimization which hasbeen studied for many decades. This paper focuses on dealing with one of the most difficult TSP variants named thequadratic traveling salesman problem (QTSP) that has numerous planning applications in robotics and bioinformatics.The goal of QTSP is similar to TSP which finds a cycle visiting all nodes exactly once with minimum total costs. However, the costs in QTSP are associated with three vertices traversed in succession (instead of two like in TSP). This leadsto a quadratic objective function that is much harder to solve.To efficiently solve …
Few-Shot Event Detection: An Empirical Study And A Unified View, Yubo Ma, Zehao Wang, Yixin Cao, Aixin Sun
Few-Shot Event Detection: An Empirical Study And A Unified View, Yubo Ma, Zehao Wang, Yixin Cao, Aixin Sun
Research Collection School Of Computing and Information Systems
Few-shot event detection (ED) has been widely studied, while this brings noticeable discrepancies, e.g., various motivations, tasks, and experimental settings, that hinder the understanding of models for future progress. This paper presents a thorough empirical study, a unified view of ED models, and a better unified baseline. For fair evaluation, we compare 12 representative methods on three datasets, which are roughly grouped into prompt-based and prototype-based models for detailed analysis. Experiments consistently demonstrate that prompt-based methods, including ChatGPT, still significantly trail prototype-based methods in terms of overall performance. To investigate their superior performance, we break down their design elements along …
Discriminative Reasoning With Sparse Event Representation For Document-Level Event-Event Relation Extraction, Changsen Yuan, Heyan Huang, Yixin Cao, Yonggang Wen
Discriminative Reasoning With Sparse Event Representation For Document-Level Event-Event Relation Extraction, Changsen Yuan, Heyan Huang, Yixin Cao, Yonggang Wen
Research Collection School Of Computing and Information Systems
Document-level Event-Event Relation Extraction (DERE) aims to extract relations between events in a document. It challenges conventional sentence-level task (SERE) with difficult long-text understanding. In this paper, we propose a novel DERE model (SENDIR) for better document-level reasoning. Different from existing works that build an event graph via linguistic tools, SENDIR does not require any prior knowledge. The basic idea is to discriminate event pairs in the same sentence or span multiple sentences by assuming their different information density: 1) low density in the document suggests sparse attention to skip irrelevant information. Our module 1 designs various types of attention …
Context-Aware Neural Fault Localization, Zhuo Zhang, Xiaoguang Mao, Meng Yan, Xin Xia, David Lo, David Lo
Context-Aware Neural Fault Localization, Zhuo Zhang, Xiaoguang Mao, Meng Yan, Xin Xia, David Lo, David Lo
Research Collection School Of Computing and Information Systems
Numerous fault localization techniques identify suspicious statements potentially responsible for program failures by discovering the statistical correlation between test results (i.e., failing or passing) and the executions of the different statements of a program (i.e., covered or not covered). They rarely incorporate a failure context into their suspiciousness evaluation despite the fact that a failure context showing how a failure is produced is useful for analyzing and locating faults. Since a failure context usually contains the transitive relationships among the statements of causing a failure, its relationship complexity becomes one major obstacle for the context incorporation in suspiciousness evaluation of …
Prompt To Be Consistent Is Better Than Self-Consistent? Few-Shot And Zero-Shot Fact Verification With Pre-Trained Language Models, Fengzhu Zeng, Wei Gao
Prompt To Be Consistent Is Better Than Self-Consistent? Few-Shot And Zero-Shot Fact Verification With Pre-Trained Language Models, Fengzhu Zeng, Wei Gao
Research Collection School Of Computing and Information Systems
Few-shot or zero-shot fact verification only relies on a few or no labeled training examples. In this paper, we propose a novel method called ProToCo, to Prompt pre-trained language models (PLMs) To be Consistent, for improving the factuality assessment capability of PLMs in the few-shot and zero-shot settings. Given a claim-evidence pair, ProToCo generates multiple variants of the claim with different relations and frames a simple consistency mechanism as constraints for making compatible predictions across these variants. We update PLMs by using parameter-efficient fine-tuning (PEFT), leading to more accurate predictions in few-shot and zero-shot fact verification tasks. Our experiments on …
Analyzing Taxi Drivers’ Decision-Making And Recommending Strategies For Enhanced Performance: A Data-Driven Approach, Mengyu Ji
Dissertations and Theses Collection (Open Access)
This thesis focuses on analyzing the decision-making process of taxi drivers and providing data-driven strategies to enhance their performance. By examin- ing comprehensive historical data encompassing passenger demand patterns, drivers’ spatial dynamics, and fare structures, valuable insights are gained into drivers’ choices regarding optimal routes, timing, and areas with high demand. Integrating real-time information sources, such as GPS data and passenger updates, allows drivers to adapt their strategies dynamically to changing traffic conditions and emerging demand patterns. Predictive analytics models, includ- ing ARIMA, XGBoost, and Linear Regression, are utilized to forecast demand flow at key locations, enabling proactive decision-making and …
Balanced Blended Space: Foundational Human–Ai Dialogues In A Symmetry-Based Mediation Framework, David Smith
Balanced Blended Space: Foundational Human–Ai Dialogues In A Symmetry-Based Mediation Framework, David Smith
Publications and Research
This working paper documents the early development of the Balanced Blended Space (BBS) framework through a series of iterative interactions between a cognitive agent (human researcher) and a computational agent (AI system) conducted in 2023. The work is motivated by the need for a universal theoretical model capable of describing the integration of physical, virtual, and conceptual spaces, particularly in response to increasing fragmentation across contemporary communication systems.
BBS is proposed as a symmetry-based mediation framework in which relationships between domains—such as physical and virtual space, cognition and computation, and multiple sensory modalities—are treated as structurally equivalent and mappable. Central …
Surveillance And The Future Of Work: Exploring Employees’ Attitudes Toward Monitoring In A Post-Covid Workplace, Jessica Vitak, Michael Zimmer
Surveillance And The Future Of Work: Exploring Employees’ Attitudes Toward Monitoring In A Post-Covid Workplace, Jessica Vitak, Michael Zimmer
Computer Science Faculty Research and Publications
The future of work increasingly focuses on the collection and analysis of worker data to monitor communication, ensure productivity, reduce security threats, and assist in decision-making. The COVID-19 pandemic increased employer reliance on these technologies; however, the blurring of home and work boundaries meant these monitoring tools might also surveil private spaces. To explore workers’ attitudes toward increased monitoring practices, we present findings from a factorial vignette survey of 645 U.S. adults who worked from home during the early months of the pandemic. Using the theory of privacy as contextual integrity to guide the survey design and analysis, we unpack …
Security Of Text To Image Conversions, Zobaida Alssadi
Security Of Text To Image Conversions, Zobaida Alssadi
Theses and Dissertations
The use of images and icons to represent news or narratives has grown in popularity. Still, one critical problem is that they are not equivalent to language, making them vulnerable to adversary attacks. This study examines the impact of image-poisoning attacks based on polysemantic words and of image attacks based on cultural differences when converting text to images. Such attacks can lead to the loss of important information and create confusion and incorrect interpretations of the intended meaning, misinforming the general public. The study specifically focuses on possible effects in a news and story context. This study highlights the significance …
Reinforcement Learning For Sequential Decision Making With Constraints, Jiajing Ling
Reinforcement Learning For Sequential Decision Making With Constraints, Jiajing Ling
Dissertations and Theses Collection (Open Access)
Reinforcement learning is a widely used approach to tackle problems in sequential decision making where an agent learns from rewards or penalties. However, in decision-making problems that involve safety or limited resources, the agent's exploration is often limited by constraints. To model such problems, constrained Markov decision processes and constrained decentralized partially observable Markov decision processes have been proposed for single-agent and multi-agent settings, respectively. A significant challenge in solving constrained Dec-POMDP is determining the contribution of each agent to the primary objective and constraint violations. To address this issue, we propose a fictitious play-based method that uses Lagrangian Relaxation …
A Hierarchical Optimization Approach For Dynamic Pickup And Delivery Problem With Lifo Constraints, Jianhui Du, Zhiqin Zhang, Xu Wang, Hoong Chuin Lau
A Hierarchical Optimization Approach For Dynamic Pickup And Delivery Problem With Lifo Constraints, Jianhui Du, Zhiqin Zhang, Xu Wang, Hoong Chuin Lau
Research Collection School Of Computing and Information Systems
We consider a dynamic pickup and delivery problem (DPDP) where loading and unloading operations must follow a last in first out (LIFO) sequence. A fleet of vehicles will pick up orders in pickup points and deliver them to destinations. The objective is to minimize the total over-time (that is the amount of time that exceeds the committed delivery time) and total travel distance. Given the dynamics of orders and vehicles, this paper proposes a hierarchical optimization approach based on multiple intuitive yet often-neglected strategies, namely what we term as the urgent strategy, hitchhike strategy and packing-bags strategy. These multiple strategies …
Safe Mdp Planning By Learning Temporal Patterns Of Undesirable Trajectories And Averting Negative Side Effects, Siow Meng Low, Akshat Kumar, Scott Sanner
Safe Mdp Planning By Learning Temporal Patterns Of Undesirable Trajectories And Averting Negative Side Effects, Siow Meng Low, Akshat Kumar, Scott Sanner
Research Collection School Of Computing and Information Systems
In safe MDP planning, a cost function based on the current state and action is often used to specify safety aspects. In real world, often the state representation used may lack sufficient fidelity to specify such safety constraints. Operating based on an incomplete model can often produce unintended negative side effects (NSEs). To address these challenges, first, we associate safety signals with state-action trajectories (rather than just immediate state-action). This makes our safety model highly general. We also assume categorical safety labels are given for different trajectories, rather than a numerical cost function, which is harder to specify by the …
Learning Deep Time-Index Models For Time Series Forecasting, Jiale Gerald Woo, Chenghao Liu, Doyen Sahoo, Akshat Kumar, Steven Hoi
Learning Deep Time-Index Models For Time Series Forecasting, Jiale Gerald Woo, Chenghao Liu, Doyen Sahoo, Akshat Kumar, Steven Hoi
Research Collection School Of Computing and Information Systems
Deep learning has been actively applied to time series forecasting, leading to a deluge of new methods, belonging to the class of historicalvalue models. Yet, despite the attractive properties of time-index models, such as being able to model the continuous nature of underlying time series dynamics, little attention has been given to them. Indeed, while naive deep timeindex models are far more expressive than the manually predefined function representations of classical time-index models, they are inadequate for forecasting, being unable to generalize to unseen time steps due to the lack of inductive bias. In this paper, we propose DeepTime, a …
Understanding The Role Of External Pull Requests In The Npm Ecosystem, Vittunyuta Maeprasart, Supatsara Wattanakriengkrai, Raula Gaikovina Kula, Christoph Treude, Kenichi Matsumoto
Understanding The Role Of External Pull Requests In The Npm Ecosystem, Vittunyuta Maeprasart, Supatsara Wattanakriengkrai, Raula Gaikovina Kula, Christoph Treude, Kenichi Matsumoto
Research Collection School Of Computing and Information Systems
The risk to using third-party libraries in a software application is that much needed maintenance is solely carried out by library maintainers. These libraries may rely on a core team of maintainers (who might be a single maintainer that is unpaid and overworked) to serve a massive client user-base. On the other hand, being open source has the benefit of receiving contributions (in the form of External PRs) to help fix bugs and add new features. In this paper, we investigate the role by which External PRs (contributions from outside the core team of maintainers) contribute to a library. Through …
A Comprehensive Study On Quality Assurance Tools For Java, Han Liu, Sen Chen, Ruitao Feng, Chengwei Liu, Kaixuan Li, Zhengzi Xu, Liming Nie, Yang Liu, Yixiang Chen
A Comprehensive Study On Quality Assurance Tools For Java, Han Liu, Sen Chen, Ruitao Feng, Chengwei Liu, Kaixuan Li, Zhengzi Xu, Liming Nie, Yang Liu, Yixiang Chen
Research Collection School Of Computing and Information Systems
Quality assurance (QA) tools are receiving more and more attention and are widely used by developers. Given the wide range of solutions for QA technology, it is still a question of evaluating QA tools. Most existing research is limited in the following ways: (i) They compare tools without considering scanning rules analysis. (ii) They disagree on the effectiveness of tools due to the study methodology and benchmark dataset. (iii) They do not separately analyze the role of the warnings. (iv) There is no large-scale study on the analysis of time performance. To address these problems, in the paper, we systematically …
Impact Of Difficult Negatives On Twitter Crisis Detection, Yuhao Zhang, Siaw Ling Lo, Phyo Yi Win Myint
Impact Of Difficult Negatives On Twitter Crisis Detection, Yuhao Zhang, Siaw Ling Lo, Phyo Yi Win Myint
Research Collection School Of Computing and Information Systems
Twitter has become an alternative information source during a crisis. However, the short, noisy nature of tweets hinders information extraction. While models trained with standard Twitter crisis datasets accomplished decent performance, it remained a challenge to generalize to unseen crisis events. Thus, we proposed adding “difficult” negative examples during training to improve model generalization for Twitter crisis detection. Although adding random noise is a common practice, the impact of difficult negatives, i.e., negative data semantically similar to true examples, was never examined in NLP. Most of existing research focuses on the classification task, without considering the primary information need of …
Imitation Improvement Learning For Large-Scale Capacitated Vehicle Routing Problems, The Viet Bui, Tien Mai
Imitation Improvement Learning For Large-Scale Capacitated Vehicle Routing Problems, The Viet Bui, Tien Mai
Research Collection School Of Computing and Information Systems
Recent works using deep reinforcement learning (RL) to solve routing problems such as the capacitated vehicle routing problem (CVRP) have focused on improvement learning-based methods, which involve improving a given solution until it becomes near-optimal. Although adequate solutions can be achieved for small problem instances, their efficiency degrades for large-scale ones. In this work, we propose a newimprovement learning-based framework based on imitation learning where classical heuristics serve as experts to encourage the policy model to mimic and produce similar or better solutions. Moreover, to improve scalability, we propose Clockwise Clustering, a novel augmented framework for decomposing large-scale CVRP into …
Machine-Learning Approach To Automated Doubt Identification On Stack Overflow Comments To Guide Programming Learners, Tianhao Chen, Eng Lieh Ouh, Kar Way Tan, Siaw Ling Lo
Machine-Learning Approach To Automated Doubt Identification On Stack Overflow Comments To Guide Programming Learners, Tianhao Chen, Eng Lieh Ouh, Kar Way Tan, Siaw Ling Lo
Research Collection School Of Computing and Information Systems
Stack Overflow is a popular Q&A platform for developers to find solutions to programming problems. However, due to the varying quality of user-generated answers, there is a need for ways to help users find high-quality answers. While Stack Overflow's community-based approach can be effective, important technical aspects of the answer need to be captured, and users’ comments might contain doubts regarding these aspects. In this paper, we showed the feasibility of using a machine learning model to identify doubts and conducted data analysis. We found that highly reputed users tend to raise more doubts; most answers have doubt in the …
Augmenting Low-Resource Text Classification With Graph-Grounded Pre-Training And Prompting, Zhihao Wen, Yuan Fang
Augmenting Low-Resource Text Classification With Graph-Grounded Pre-Training And Prompting, Zhihao Wen, Yuan Fang
Research Collection School Of Computing and Information Systems
ext classification is a fundamental problem in information retrieval with many real-world applications, such as predicting the topics of online articles and the categories of e-commerce product descriptions. However, low-resource text classification, with few or no labeled samples, poses a serious concern for supervised learning. Meanwhile, many text data are inherently grounded on a network structure, such as a hyperlink/citation network for online articles, and a user-item purchase network for e-commerce products. These graph structures capture rich semantic relationships, which can potentially augment low-resource text classification. In this paper, we propose a novel model called Graph-Grounded Pre-training and Prompting (G2P2) …
Do-Good: Towards Distribution Shift Evaluation For Pre-Trained Visual Document Understanding Models, Jiabang He, Yi Hu, Lei Wang, Xing Xu, Ning Liu, Hui Liu
Do-Good: Towards Distribution Shift Evaluation For Pre-Trained Visual Document Understanding Models, Jiabang He, Yi Hu, Lei Wang, Xing Xu, Ning Liu, Hui Liu
Research Collection School Of Computing and Information Systems
Numerous pre-training techniques for visual document understanding (VDU) have recently shown substantial improvements in performance across a wide range of document tasks. However, these pre-trained VDU models cannot guarantee continued success when the distribution of test data differs from the distribution of training data. In this paper, to investigate how robust existing pre-trained VDU models are to various distribution shifts, we first develop an out-of-distribution (OOD) benchmark termed Do-GOOD for the fine-Grained analysis on Document image-related tasks specifically. The Do-GOOD benchmark defines the underlying mechanisms that result in different distribution shifts and contains 9 OOD datasets covering 3 VDU related …
Silent Compiler Bug De-Duplication Via Three-Dimensional Analysis, Chen Yang, Junjie Chen, Xingyu Fan, Jiajun Jiang, Jun Sun
Silent Compiler Bug De-Duplication Via Three-Dimensional Analysis, Chen Yang, Junjie Chen, Xingyu Fan, Jiajun Jiang, Jun Sun
Research Collection School Of Computing and Information Systems
Compiler testing is an important task for assuring the quality of compilers, but investigating test failures is very time-consuming. This is because many test failures are caused by the same compiler bug (known as bug duplication problem). In particular, this problem becomes much more challenging on silent compiler bugs (also called wrong code bugs), since these bugs can provide little information (unlike crash bugs that can produce error messages) for bug de-duplication. In this work, we propose a novel technique (called D3) to solve the duplication problem on silent compiler bugs. Its key insight is to characterize the silent bugs …