Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences Commons

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 16891 - 16920 of 63016

Full-Text Articles in Computer Sciences

Internet Of Things: Cybersecurity In Small Businesses, Zobair Wali Nov 2021

Internet Of Things: Cybersecurity In Small Businesses, Zobair Wali

Cybersecurity Undergraduate Research Showcase

Small businesses are the most vital part of a nation’s economy. In today’s world, as we are moving towards digitizing almost everything around us, cybersecurity is essential and vital for our digitalized world to function. Small businesses are no exception. All businesses collect, use, and store information. They store employees’ information, tax information, customers’ information, business transaction information, and all other operational information that is needed for a business to function. Without an appropriate cybersecurity program, these businesses are vulnerable and can be easily impacted by cyber incidents and malicious attacks. Businesses are putting resources to protect their systems against …


Transfer-Learned Pruned Deep Convolutional Neural Networks For Efficient Plant Classification In Resource-Constrained Environments, Martinson Ofori Nov 2021

Transfer-Learned Pruned Deep Convolutional Neural Networks For Efficient Plant Classification In Resource-Constrained Environments, Martinson Ofori

Masters Theses & Doctoral Dissertations

Traditional means of on-farm weed control mostly rely on manual labor. This process is time-consuming, costly, and contributes to major yield losses. Further, the conventional application of chemical weed control can be economically and environmentally inefficient. Site-specific weed management (SSWM) counteracts this by reducing the amount of chemical application with localized spraying of weed species. To solve this using computer vision, precision agriculture researchers have used remote sensing weed maps, but this has been largely ineffective for early season weed control due to problems such as solar reflectance and cloud cover in satellite imagery. With the current advances in artificial …


Intercept Graph: An Interactive Radial Visualization For Comparison Of State Changes, Shaolun Ruan, Yong Wang, Qiang Guan Nov 2021

Intercept Graph: An Interactive Radial Visualization For Comparison Of State Changes, Shaolun Ruan, Yong Wang, Qiang Guan

Research Collection School Of Computing and Information Systems

State change comparison of multiple data items is often necessary in multiple application domains, such as medical science, financial engineering, sociology, biological science, etc. Slope graphs and grouped bar charts have been widely used to show a “before-and-after” story of different data states and indicate their changes. However, they visualize state changes as either slope or difference of bars, which has been proved less effective for quantitative comparison. Also, both visual designs suffer from visual clutter issues with an increasing number of data items. In this paper, we propose Intercept Graph, a novel visual design to facilitate effective interactive comparison …


Stock Market Trend Forecasting Based On Multiple Textual Features: A Deep Learning Method, Zhenda Hu, Zhaoxia Wang, Seng-Beng Ho, Ah-Hwee Tan Nov 2021

Stock Market Trend Forecasting Based On Multiple Textual Features: A Deep Learning Method, Zhenda Hu, Zhaoxia Wang, Seng-Beng Ho, Ah-Hwee Tan

Research Collection School Of Computing and Information Systems

Stock market trend forecasting is a valuable and challenging research task for both industry and academia. In order to explore the influence of stock news information on the stock market trend, a textual embedding construction method is proposed to encode multiple textual features, including topic features, sentiment features, and semantic features extracted from stock news textual content. In addition, a deep learning method is designed by using financial data and multiple textual features obtained from multiple news textual embeddings for short-term stock market trend prediction. For evaluation, extensive experiments on real stock market data are conducted. The experimental results illustrate …


Incbl: Incremental Bug Localization, Zhou Yang, Jieke Shi, Wang Shaowei, David Lo Nov 2021

Incbl: Incremental Bug Localization, Zhou Yang, Jieke Shi, Wang Shaowei, David Lo

Research Collection School Of Computing and Information Systems

Numerous efforts have been invested in improving the effectiveness of bug localization techniques, whereas little attention is paid to making these tools run more efficiently in continuously evolving software repositories. This paper first analyzes the information retrieval model behind a classic bug localization tool, BugLocator, and builds a mathematical foundation illustrating that the model can be updated incrementally when codebase or bug reports evolve. Then, we present IncBL, a tool for Incremental Bug Localization in evolving software repositories. IncBL is evaluated on the Bugzbook dataset, and the results show that IncBL can significantly reduce the running time by 77.79% on …


A Bert-Based Two-Stage Model For Chinese Chengyu Recommendation, Minghuan Tan, Jing Jiang, Bingtian Dai Nov 2021

A Bert-Based Two-Stage Model For Chinese Chengyu Recommendation, Minghuan Tan, Jing Jiang, Bingtian Dai

Research Collection School Of Computing and Information Systems

In Chinese, Chengyu are fixed phrases consisting of four characters. As a type of idioms, their meanings usually cannot be derived from their component characters. In this paper, we study the task of recommending a Chengyu given a textual context. Observing some of the limitations with existing work, we propose a two-stage model, where during the first stage we re-train a Chinese BERT model by masking out Chengyu from a large Chinese corpus with a wide coverage of Chengyu. During the second stage, we fine-tune the retrained, Chengyu-oriented BERT on a specific Chengyu recommendation dataset. We evaluate this method on …


Information Extraction And Classification On Journal Papers, Lei Yu Nov 2021

Information Extraction And Classification On Journal Papers, Lei Yu

School of Computing: Dissertations, Theses, and Student Research

The importance of journals for diffusing the results of scientific research has increased considerably. In the digital era, Portable Document Format (PDF) became the established format of electronic journal articles. This structured form, combined with a regular and wide dissemination, spread scientific advancements easily and quickly. However, the rapidly increasing numbers of published scientific articles requires more time and effort on systematic literature reviews, searches and screens. The comprehension and extraction of useful information from the digital documents is also a challenging task, due to the complex structure of PDF.

To help a soil science team from the United States …


Mapping E-Commerce Locally And Beyond: Citt K12 Special Investigation Project, Thomas O’Brien, Deanna Matsumoto Nov 2021

Mapping E-Commerce Locally And Beyond: Citt K12 Special Investigation Project, Thomas O’Brien, Deanna Matsumoto

Mineta Transportation Institute

As all aspects of the American workplace become automated or digitally enhanced to some degree, K12 educators have an increasing responsibility to help their students acquire the technical skills necessary to organize and interpret information. Increasingly, this is done through Geographic Information Systems (GIS), especially in careers related to transportation and logistics. The Center for International Trade & Transportation (CITT) at CSU Long Beach has developed this K12 Special Investigation Project to introduce ArcGIS StoryMaps, an engaging, accessible and sophisticated web-based GIS application. The lessons center on e-commerce and its accompanying environmental and economic impact. Still, the activities can be …


Cuts: Scaling Subgraph Isomorphism On Distributed Multi-Gpu Systems Using Trie Based Data Structure, Lizhi Xiang, Arif Khan, Edoardo Serra, Mahantesh Halappanavar, Aravind Sukumaran-Rajam Nov 2021

Cuts: Scaling Subgraph Isomorphism On Distributed Multi-Gpu Systems Using Trie Based Data Structure, Lizhi Xiang, Arif Khan, Edoardo Serra, Mahantesh Halappanavar, Aravind Sukumaran-Rajam

Computer Science Faculty Publications and Presentations

Subgraph isomorphism is a pattern-matching algorithm widely used in many domains such as chem-informatics, bioinformatics, databases, and social network analysis. It is computationally expensive and is a proven NP-hard problem. The massive parallelism in GPUs is well suited for solving subgraph isomorphism. However, current GPU implementations are far from the achievable performance. Moreover, the enormous memory requirement of current approaches limits the problem size that can be handled. This work analyzes the fundamental challenges associated with processing subgraph isomorphism on GPUs and develops an efficient GPU implementation. We also develop a GPU-friendly trie-based data structure to drastically reduce the intermediate …


On Lexicographic Proof Rules For Probabilistic Termination, Krishnendu Chatterjee, Ehsan Kafshdar Goharshady, Petr Novotný, Jiří Zárevucký, Dorde Zikelic Nov 2021

On Lexicographic Proof Rules For Probabilistic Termination, Krishnendu Chatterjee, Ehsan Kafshdar Goharshady, Petr Novotný, Jiří Zárevucký, Dorde Zikelic

Research Collection School Of Computing and Information Systems

We consider the almost-sure (a.s.) termination problem for probabilistic programs, which are a stochastic extension of classical imperative programs. Lexicographic ranking functions provide a sound and practical approach for termination of non-probabilistic programs, and their extension to probabilistic programs is achieved via lexicographic ranking supermartingales (LexRSMs). However, LexRSMs introduced in the previous work have a limitation that impedes their automation: all of their components have to be non-negative in all reachable states. This might result in LexRSM not existing even for simple terminating programs. Our contributions are twofold: First, we introduce a generalization of LexRSMs which allows for some components …


Topic Modeling For Multi-Aspect Listwise Comparison, Delvin Ce Zhang, Hady W. Lauw Nov 2021

Topic Modeling For Multi-Aspect Listwise Comparison, Delvin Ce Zhang, Hady W. Lauw

Research Collection School Of Computing and Information Systems

As a well-established probabilistic method, topic models seek to uncover latent semantics from plain text. In addition to having textual content, we observe that documents are usually compared in listwise rankings based on their content. For instance, world-wide countries are compared in an international ranking in terms of electricity production based on their national reports. Such document comparisons constitute additional information that reveal documents' relative similarities. Incorporating them into topic modeling could yield comparative topics that help to differentiate and rank documents. Furthermore, based on different comparison criteria, the observed document comparisons usually cover multiple aspects, each expressing a distinct …


Self-Supervised Multi-Class Pre-Training For Unsupervised Anomaly Detection And Segmentation In Medical Images, Yu Tian, Fengbei Liu, Guansong Pang, Yuanhong Chen, Yuyuan Liu, Johan W. Verjans, Rajvinder Singh Nov 2021

Self-Supervised Multi-Class Pre-Training For Unsupervised Anomaly Detection And Segmentation In Medical Images, Yu Tian, Fengbei Liu, Guansong Pang, Yuanhong Chen, Yuyuan Liu, Johan W. Verjans, Rajvinder Singh

Research Collection School Of Computing and Information Systems

Unsupervised anomaly detection (UAD) that requires only normal (healthy) training images is an important tool for enabling the development of medical image analysis (MIA) applications, such as disease screening, since it is often difficult to collect and annotate abnormal (or disease) images in MIA. However, heavily relying on the normal images may cause the model training to overfit the normal class. Self-supervised pre-training is an effective solution to this problem. Unfortunately, current self-supervision methods adapted from computer vision are sub-optimal for MIA applications because they do not explore MIA domain knowledge for designing the pretext tasks or the training process. …


Wav-Bert: Cooperative Acoustic And Linguistic Representation Learning For Low-Resource Speech Recognition, Guolin Zheng, Yubei Xiao, Ke Gong, Pan Zhou, Xiaodan Liang, Liang Lin Nov 2021

Wav-Bert: Cooperative Acoustic And Linguistic Representation Learning For Low-Resource Speech Recognition, Guolin Zheng, Yubei Xiao, Ke Gong, Pan Zhou, Xiaodan Liang, Liang Lin

Research Collection School Of Computing and Information Systems

Unifying acoustic and linguistic representation learning has become increasingly crucial to transfer the knowledge learned on the abundance of high-resource language data for low-resource speech recognition. Existing approaches simply cascade pre-trained acoustic and language models to learn the transfer from speech to text. However, how to solve the representation discrepancy of speech and text is unexplored, which hinders the utilization of acoustic and linguistic information. Moreover, previous works simply replace the embedding layer of the pre-trained language model with the acoustic features, which may cause the catastrophic forgetting problem. In this work, we introduce Wav-BERT, a cooperative acoustic and linguistic …


The Forestecology R Package For Fitting And Assessing Neighborhood Models Of The Effect Of Interspecific Competition On The Growth Of Trees, Albert Y. Kim, David N. Allen, Simon P. Couch Nov 2021

The Forestecology R Package For Fitting And Assessing Neighborhood Models Of The Effect Of Interspecific Competition On The Growth Of Trees, Albert Y. Kim, David N. Allen, Simon P. Couch

Statistical and Data Sciences: Faculty Publications

Neighborhood competition models are powerful tools to measure the effect of interspecific competition. Statistical methods to ease the application of these models are currently lacking. We present the forestecology package providing methods to (a) specify neighborhood competition models, (b) evaluate the effect of competitor species identity using permutation tests, and (cs) measure model performance using spatial cross-validation. Following Allen and Kim (PLoS One, 15, 2020, e0229930), we implement a Bayesian linear regression neighborhood competition model. We demonstrate the package's functionality using data from the Smithsonian Conservation Biology Institute's large forest dynamics plot, part of the ForestGEO global network of research …


Frames For Justice Consciousness, Colin M. Gray, Rua M. Williams, Paul Parsons, Austin L. Toombs, Abbee Westbrook Nov 2021

Frames For Justice Consciousness, Colin M. Gray, Rua M. Williams, Paul Parsons, Austin L. Toombs, Abbee Westbrook

Computer Graphics Technology Open Educational Resources

We describe how UX design students become aware of citizen-engaged design work, and indicate the extent to which a progression toward social justice-focused design work might be possible in a single project cycle. Our study site is a sophomore-level UX design studio at a large Midwestern US university—part of a five-semester sequence in which students engage in a range of projects that address competence in user research, prototyping, and evaluation. The project cycle we focus on directly challenges the apolitical framing in most foundational UX methods literature, explicitly asking students to engage with issues of power disparities. We analyzed …


Transforming Businesses With E-Commerce Intelligence, Yuanto Kusnadi, Gary Pan Nov 2021

Transforming Businesses With E-Commerce Intelligence, Yuanto Kusnadi, Gary Pan

Research Collection School Of Accountancy

2020 had been an extraordinary year as the Covid-19 pandemic struck almost all countries in the world and created an extraordinary impact on businesses worldwide. Singapore and many other Southeast Asian countries were not spared and had to implement lockdowns swiftly. To cope with physical store closures and the increased volume of online transactions, most businesses tried to revamp their business models and set up online stores to capitalise on the rise of the e-commerce wave. With the growing trend of online transactions, it has become imperative for companies operating in the Fast Moving Consumer Goods (FMCG) industry to track …


Factual Consistency Evaluation For Text Summarization Via Counterfactual Estimation, Yuexiang Xie, Fei Sun, Yang Deng, Yaliang Li, Bolin Ding Nov 2021

Factual Consistency Evaluation For Text Summarization Via Counterfactual Estimation, Yuexiang Xie, Fei Sun, Yang Deng, Yaliang Li, Bolin Ding

Research Collection School Of Computing and Information Systems

Despite significant progress has been achieved in text summarization, factual inconsistency in generated summaries still severely limits its practical applications. Among the key factors to ensure factual consistency, a reliable automatic evaluation metric is the first and the most crucial one. However, existing metrics either neglect the intrinsic cause of the factual inconsistency or rely on auxiliary tasks, leading to an unsatisfied correlation with human judgments or increasing the inconvenience of usage in practice. In light of these challenges, we propose a novel metric to evaluate the factual consistency in text summarization via counterfactual estimation, which formulates the causal relationship …


Fingerlings Mass Estimation: A Comparison Between Deep And Shallow Learning Algorithms, Adair Da Silva Oliveira Junior, Diego André Sant’Ana, Marcio Carneiro Brito Pache, Vanir Garcia, Vanessa Aparecida De Moares Weber, Gilberto Astolfi, Fabricio De Lima Weber, Geazy Vilharva Menezes, Gabriel Kirsten Menezes, Pedro Lucas França Albuquerque, Celso Soares Costa, Eduardo Quirino Arguelho De Queiroz, João Victor Araújo Rozales, Milena Wolff Ferreira, Marco Hiroshi Naka, Hemerson Pistori Nov 2021

Fingerlings Mass Estimation: A Comparison Between Deep And Shallow Learning Algorithms, Adair Da Silva Oliveira Junior, Diego André Sant’Ana, Marcio Carneiro Brito Pache, Vanir Garcia, Vanessa Aparecida De Moares Weber, Gilberto Astolfi, Fabricio De Lima Weber, Geazy Vilharva Menezes, Gabriel Kirsten Menezes, Pedro Lucas França Albuquerque, Celso Soares Costa, Eduardo Quirino Arguelho De Queiroz, João Victor Araújo Rozales, Milena Wolff Ferreira, Marco Hiroshi Naka, Hemerson Pistori

School of Computing: Faculty Publications

The paper presents some results regarding the automatic mass estimation of Pintado Real fingerlings, using machine learning techniques to support the fish production process. For this purpose, an image dataset called FISHCV1206FSEG, was created which is composed of 1206 images of fingerlings with their respective annotated masses. Through the fish contours, the area and perimeter were extracted, and submitted to the J48, SVM, and KNN classification algorithms and a linear regression algorithm. The images were also submitted to ResNet50, In- ceptionV3, Exception, VGG16, and VGG19 convolutional neural networks. As a result, the classification algorithm J48 reached an accuracy of 58.2% …


K-Sums Clustering: A Stochastic Optimization Approach, Zhao Wan-Lei, Shi Ying Lan, Run-Qing Chen, Chong-Wah Ngo Nov 2021

K-Sums Clustering: A Stochastic Optimization Approach, Zhao Wan-Lei, Shi Ying Lan, Run-Qing Chen, Chong-Wah Ngo

Research Collection School Of Computing and Information Systems

In this paper, we revisit the decades-old clustering method k-means. The egg-chicken loop in traditional k-means has been replaced by a pure stochastic optimization procedure. The optimization is undertaken from the perspective of each individual sample. Different from existing incremental k-means, an individual sample is tentatively joined into a new cluster to evaluate its distance to the corresponding new centroid, in which the contribution from this sample is accounted. The sample is moved to this new cluster concretely only after we find the reallocation makes the sample closer to the new centroid than it is to the current one. Compared …


On A Multistage Discrete Stochastic Optimization Problem With Stochastic Constraints And Nested Sampling, Thuy Anh Ta, Tien Mai, Fabian Bastin, Pierre L'Ecuyer Nov 2021

On A Multistage Discrete Stochastic Optimization Problem With Stochastic Constraints And Nested Sampling, Thuy Anh Ta, Tien Mai, Fabian Bastin, Pierre L'Ecuyer

Research Collection School Of Computing and Information Systems

We consider a multistage stochastic discrete program in which constraints on any stage might involve expectations that cannot be computed easily and are approximated by simulation. We study a sample average approximation (SAA) approach that uses nested sampling, in which at each stage, a number of scenarios are examined and a number of simulation replications are performed for each scenario to estimate the next-stage constraints. This approach provides an approximate solution to the multistage problem. To establish the consistency of the SAA approach, we first consider a two-stage problem and show that in the second-stage problem, given a scenario, the …


Where2change: Change Request Localization For App Reviews, Tao Zhang, Jiachi Chen, Xian Zhan, Xiapu Luo, David Lo, He Jiang Nov 2021

Where2change: Change Request Localization For App Reviews, Tao Zhang, Jiachi Chen, Xian Zhan, Xiapu Luo, David Lo, He Jiang

Research Collection School Of Computing and Information Systems

Million of mobile apps have been released to the market. Developers need to maintain these apps so that they can continue to benefit end users. Developers usually extract useful information from user reviews to maintain and evolve mobile apps. One of the important activities that developers need to do while reading user reviews is to locate the source code related to requested changes. Unfortunately, this manual work is costly and time consuming since: (1) an app can receive thousands of reviews, and (2) a mobile app can consist of hundreds of source code files. To address this challenge, Palomba et …


Information Technology And Organizational Learning: Managing Behavioral Change In The Digital Age By Arthur M. Langer, Siu Loon Hoe Nov 2021

Information Technology And Organizational Learning: Managing Behavioral Change In The Digital Age By Arthur M. Langer, Siu Loon Hoe

Research Collection School Of Computing and Information Systems

As the world battles yet another crisis because of the spread of COVID-19, the idea of digitalization brings about a whole new meaning. Many professionals and information technology (IT) managers have remarked that the spread of the coronavirus has accelerated the pace of digital transformation much more so than any effort put forth by C-suite executives. While it is true that most organizations do not accept new technology readily because of embedded legacy systems, changing the corporate cultures does play an important role in affecting the rate of IT adoption. Very often, leaders and senior executives focus on the technological …


Efficient Server-Aided Secure Two-Party Computation In Heterogeneous Mobile Cloud Computing, Yulin Wu, Xuan Wang, Willy Susilo, Guomin Yang, Zoe L. Jiang, Qian Chen, Peng Xu Nov 2021

Efficient Server-Aided Secure Two-Party Computation In Heterogeneous Mobile Cloud Computing, Yulin Wu, Xuan Wang, Willy Susilo, Guomin Yang, Zoe L. Jiang, Qian Chen, Peng Xu

Research Collection School Of Computing and Information Systems

With the ubiquity of mobile devices and rapid development of cloud computing, mobile cloud computing (MCC) has been considered as an essential computation setting to support complicated, scalable and flexible mobile applications by overcoming the physical limitations of mobile devices with the aid of cloud. In the MCC setting, since many mobile applications (e.g., map apps) interacting with cloud server and application server need to perform computation with the private data of users, it is important to realize secure computation for MCC. In this article, we propose an efficient server-aided secure two-party computation (2PC) protocol for MCC. This is the …


Cs-Light: Camera Sensing Based Occupancy-Aware Robust Smart Building Lighting Control, Anuradha Ravi, Kasun Pramuditha Gamlath, Siyan Hu, Archan Misra Nov 2021

Cs-Light: Camera Sensing Based Occupancy-Aware Robust Smart Building Lighting Control, Anuradha Ravi, Kasun Pramuditha Gamlath, Siyan Hu, Archan Misra

Research Collection School Of Computing and Information Systems

We describe the practical development of a smart lighting control system, CS-Light, that uses a preexisting surveillance camera infrastructure as the sole sensing substrate. At a high level, the camera feeds are used to both (a) estimate the illuminance of individual, fine-grained (roughly 12m2) sub-regions, and (b) identify sub-regions that have non-transient human occupancy. Subsequently, these estimates are used to perform fine-grained (non-binary) power optimization of a set of LED luminaires, collectively minimizing energy consumption while assuring comfort to human occupants. The key to our approach is the ability to tackle the challenging problem of translating the luminance (pixel intensity) …


Flip & Slack – Active Flipped Classroom Learning With Collaborative Slack Interactions, Kyong Jin Shim, Gottipati Swapna, Yi Meng Lau Nov 2021

Flip & Slack – Active Flipped Classroom Learning With Collaborative Slack Interactions, Kyong Jin Shim, Gottipati Swapna, Yi Meng Lau

Research Collection School Of Computing and Information Systems

Active flipped classroom learning is stipulated with faculty structuring the activities involving constructive interactions, either formal or informal. Sharing ideas and responding to ideas improve the cognitive skills of the students. Encouraging peers to contribute to class activities and respecting peers contribute to the development of affective skills. We present an integrated platform for cognitive and affective skills development. A flipped classroom arrangement allows the faculty to focus more on in-class activities such as programming and lab exercises to support active learning in computing courses. We share the design of an innovative flipped classroom model integrated with Slack and present …


Generating Music With Sentiments, Chunhui Bao Nov 2021

Generating Music With Sentiments, Chunhui Bao

Dissertations and Theses Collection (Open Access)

In this thesis, I focus on the music generation conditional on human sentiments such as positive and negative. As there are no existing large-scale music datasets annotated with sentiment labels, generating high-quality music conditioned on sentiments is hard. I thus build a new dataset consisting of the triplets of lyric, melody and sentiment, without requiring any manual annotations. I utilize an automated sentiment recognition model (based on the BERT trained on Edmonds Dance dataset) to "label'' the music according to the sentiments recognized from its lyrics. I then train the model of generating sentimental music and call the method Sentimental …


Development Of Sensor, Sensory System And Signal Processing Algorithm For Intelligent Sensing Applications, Xingzhe Zhang Nov 2021

Development Of Sensor, Sensory System And Signal Processing Algorithm For Intelligent Sensing Applications, Xingzhe Zhang

Dissertations

Sensors have been receiving significant attention in the last decade and the demand for sensory systems has increased in recent years due to the rapid growth in the field of artificial intelligence (AI). Sensors can improve people’s awareness by providing them with real-time information on the environment and their immediate health conditions. This dissertation presents the fulfilment of three main projects and focuses on the development of a sensor, a sensory system, and a sensor signal recognition system for AI applications by employing printed electronics, analog circuit design, and digital signal processing techniques.

In the first project, a multi-channel stethograph …


Can We Make It Better? Assessing And Improving Quality Of Github Repositories, Gede Artha Azriadi Prana Nov 2021

Can We Make It Better? Assessing And Improving Quality Of Github Repositories, Gede Artha Azriadi Prana

Dissertations and Theses Collection (Open Access)

The code hosting platform GitHub has gained immense popularity worldwide in recent years, with over 200 million repositories hosted as of June 2021. Due to its popularity, it has great potential to facilitate widespread improvements across many software projects. Naturally, GitHub has attracted much research attention, and the source code in the various repositories it hosts also provide opportunity to apply techniques and tools developed by software engineering researchers over the years. However, much of existing body of research applicable to GitHub focuses on code quality of the software projects and ways to improve them. Fewer work focus on potential …


Finding A Needle In A Haystack: Automatic Mining Of Silent Vulnerability Fixes, Jiayuan Zhou, Michael Pacheco, Zhiyuan Wan, Xin Xia, David Lo, Yuan Wang, Ahmed E. Hassan Nov 2021

Finding A Needle In A Haystack: Automatic Mining Of Silent Vulnerability Fixes, Jiayuan Zhou, Michael Pacheco, Zhiyuan Wan, Xin Xia, David Lo, Yuan Wang, Ahmed E. Hassan

Research Collection School Of Computing and Information Systems

Following the coordinated vulnerability disclosure model, a vulnerability in open source software (OSS) is suggested to be fixed “silently”, without disclosing the fix until the vulnerability is disclosed. Yet, it is crucial for OSS users to be aware of vulnerability fixes as early as possible, as once a vulnerability fix is pushed to the source code repository, a malicious party could probe for the corresponding vulnerability to exploit it. In practice, OSS users often rely on the vulnerability disclosure information from security advisories (e.g., National Vulnerability Database) to sense vulnerability fixes. However, the time between the availability of a vulnerability …


Learning To Teach And Learn For Semi-Supervised Few-Shot Image Classification, Xinzhe Li, Jianqiang Huang, Yaoyao Liu, Qin Zhou, Shibao Zheng, Bernt Schiele, Qianru Sun Nov 2021

Learning To Teach And Learn For Semi-Supervised Few-Shot Image Classification, Xinzhe Li, Jianqiang Huang, Yaoyao Liu, Qin Zhou, Shibao Zheng, Bernt Schiele, Qianru Sun

Research Collection School Of Computing and Information Systems

This paper presents a novel semi-supervised few-shot image classification method named Learning to Teach and Learn (LTTL) to effectively leverage unlabeled samples in small-data regimes. Our method is based on self-training, which assigns pseudo labels to unlabeled data. However, the conventional pseudo-labeling operation heavily relies on the initial model trained by using a handful of labeled data and may produce many noisy labeled samples. We propose to solve the problem with three steps: firstly, cherry-picking searches valuable samples from pseudo-labeled data by using a soft weighting network; and then, cross-teaching allows the classifiers to teach mutually for rejecting more noisy …