Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences Commons™

Open Access. Powered by Scholars. Published by Universities.®

Singapore Management University

Discipline
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 1981 - 2010 of 9003

Full-Text Articles in Computer Sciences

Techsumbot: A Stack Overflow Answer Summarization Tool For Technical Query, Chengran Yang, Bowen Xu, Jiakun Liu, David Lo May 2023

Techsumbot: A Stack Overflow Answer Summarization Tool For Technical Query, Chengran Yang, Bowen Xu, Jiakun Liu, David Lo

Research Collection School Of Computing and Information Systems

Stack Overflow is a popular platform for developers to seek solutions to programming-related problems. However, prior studies identified that developers may suffer from the redundant, useless, and incomplete information retrieved by the Stack Overflow search engine. To help developers better utilize the Stack Overflow knowledge, researchers proposed tools to summarize answers to a Stack Overflow question. However, existing tools use hand-craft features to assess the usefulness of each answer sentence and fail to remove semantically redundant information in the result. Besides, existing tools only focus on a certain programming language and cannot retrieve up-to-date new posted knowledge from Stack Overflow. …


Chronos: Time-Aware Zero-Shot Identification Of Libraries From Vulnerability Reports, Yunbo Lyu, Thanh Le Cong, Hong Jin Kang, Ratnadira Widyasari, Zhipeng Zhao, Xuan-Bach Dinh Le, Ming Li, David Lo May 2023

Chronos: Time-Aware Zero-Shot Identification Of Libraries From Vulnerability Reports, Yunbo Lyu, Thanh Le Cong, Hong Jin Kang, Ratnadira Widyasari, Zhipeng Zhao, Xuan-Bach Dinh Le, Ming Li, David Lo

Research Collection School Of Computing and Information Systems

Tools that alert developers about library vulnerabilities depend on accurate, up-to-date vulnerability databases which are maintained by security researchers. These databases record the libraries related to each vulnerability. However, the vulnerability reports may not explicitly list every library and human analysis is required to determine all the relevant libraries. Human analysis may be slow and expensive, which motivates the need for automated approaches. Researchers and practitioners have proposed to automatically identify libraries from vulnerability reports using extreme multi-label learning (XML). While state-of-the-art XML techniques showed promising performance, their experimental settings do not practically fit what happens in reality. Previous studies …


Understanding The Role Of Images On Stack Overflow, Dong Wang, Tao Xiao, Christoph Treude, Raula Kula, Hideaki Hata, Yasutaka Kamei May 2023

Understanding The Role Of Images On Stack Overflow, Dong Wang, Tao Xiao, Christoph Treude, Raula Kula, Hideaki Hata, Yasutaka Kamei

Research Collection School Of Computing and Information Systems

Images are increasingly being shared by software developers in diverse channels including question-and-answer forums like Stack Overflow. Although prior work has pointed out that these images are meaningful and provide complementary information compared to their associated text, how images are used to support questions is empirically unknown. To address this knowledge gap, in this paper we specifically conduct an empirical study to investigate (I) the characteristics of images, (II) the extent to which images are used in different question types, and (III) the role of images on receiving answers. Our results first show that user interface is the most common …


Re-Evaluating Natural Intelligence In The Face Of Chatgpt, Elvin T. Lim, Tze K Koh May 2023

Re-Evaluating Natural Intelligence In The Face Of Chatgpt, Elvin T. Lim, Tze K Koh

Research Collection College of Integrative Studies

How will new technologies impact the nature of higher education? Before ChatGPT, the world witnessed major shifts led by innovations in information storage and transmission. Papyrus in ancient Egypt, the Gutenberg press in 15th-century Europe, and the internet in the 20th century were all milestones in the mass dissemination of knowledge.


The Persuasive Design Of Ai-Synthesized Voices, Hannah H. Chang, Anirban Mukherjee May 2023

The Persuasive Design Of Ai-Synthesized Voices, Hannah H. Chang, Anirban Mukherjee

Research Collection Lee Kong Chian School Of Business

We investigate the impact of AI-based, machine-synthesized narrating voices on consumer cognitions and behavior in media-rich environment. Across four studies (plus pretests), we show that the design of AI voices systematically and predictably affects consumer cognition and behavior. Specifically, the designs of AI voices have differential effects in early versus later stages of consumer purchase journey. In situations where the consumers’ attention is already directed to the message, we find that marcomm with more AI voices generates a smaller proportion of favorable thoughts, which leads to a lower purchase likelihood. These results support our conceptualization that hearing more AI voices …


Algorithms, Leadership, And Morality: Why A Mere Human Effect Drives The Preference For Human Over Algorithmic Leadership, Jack Mcguire, David De Cremer May 2023

Algorithms, Leadership, And Morality: Why A Mere Human Effect Drives The Preference For Human Over Algorithmic Leadership, Jack Mcguire, David De Cremer

Research Collection Lee Kong Chian School Of Business

Algorithms are increasingly making decisions in organizations that carry moral consequences and such decisions are considered to be ordinarily made by leaders. An important consideration to be made by organizations is therefore whether adopting algorithms in this domain will be accepted by employees and whether this practice will harm their reputation. Considering this emergent phenomenon, we set out to examine employees’ perceptions about (a) algorithmic decision-making systems employed to occupy leadership roles and make moral decisions in organizations, and (b) the reputation of organizations that employ such systems. Furthermore, we examine the extent to which the decision agent needs to …


Are You Cloud-Certified? Preparing Computing Undergraduates For Cloud Certification With Experiential Learning, Eng Lieh Ouh, Benjamin Gan May 2023

Are You Cloud-Certified? Preparing Computing Undergraduates For Cloud Certification With Experiential Learning, Eng Lieh Ouh, Benjamin Gan

Research Collection School Of Computing and Information Systems

Cloud Computing skills have been increasing in demand. Many software engineers are learning these skills and taking cloud certification examinations to be job competitive. Preparing undergraduates to be cloud-certified remains challenging as cloud computing is a relatively new topic in the computing curriculum, and many of these certifications require working experience. In this paper, we report our experiences designing a course with experiential learning to prepare our computing undergraduates to take the cloud certification. We adopt a university project-based experiential learning framework to engage industry partners who provide project requirements for students to develop cloud solutions and an experiential risk …


Reinforced Adaptation Network For Partial Domain Adaptation, Keyu Wu, Min Wu, Zhenghua Chen, Ruibing Jin, Wei Cui, Zhiguang Cao, Xiaoli Li May 2023

Reinforced Adaptation Network For Partial Domain Adaptation, Keyu Wu, Min Wu, Zhenghua Chen, Ruibing Jin, Wei Cui, Zhiguang Cao, Xiaoli Li

Research Collection School Of Computing and Information Systems

Domain adaptation enables generalized learning in new environments by transferring knowledge from label-rich source domains to label-scarce target domains. As a more realistic extension, partial domain adaptation (PDA) relaxes the assumption of fully shared label space, and instead deals with the scenario where the target label space is a subset of the source label space. In this paper, we propose a Reinforced Adaptation Network (RAN) to address the challenging PDA problem. Specifically, a deep reinforcement learning model is proposed to learn source data selection policies. Meanwhile, a domain adaptation model is presented to simultaneously determine rewards and learn domain-invariant feature …


Link Prediction On Latent Heterogeneous Graphs, Trung Kien Nguyen, Zemin Liu, Yuan Fang May 2023

Link Prediction On Latent Heterogeneous Graphs, Trung Kien Nguyen, Zemin Liu, Yuan Fang

Research Collection School Of Computing and Information Systems

On graph data, the multitude of node or edge types gives rise to heterogeneous information networks (HINs). To preserve the heterogeneous semantics on HINs, the rich node/edge types become a cornerstone of HIN representation learning. However, in real-world scenarios, type information is often noisy, missing or inaccessible. Assuming no type information is given, we define a so-called latent heterogeneous graph (LHG), which carries latent heterogeneous semantics as the node/edge types cannot be observed. In this paper, we study the challenging and unexplored problem of link prediction on an LHG. As existing approaches depend heavily on type-based information, they are suboptimal …


Contrabert: Enhancing Code Pre-Trained Models Via Contrastive Learning, Shangqing Liu, Bozhi Wu, Xiaofei Xie, Guozhu Meng, Yang. Liu May 2023

Contrabert: Enhancing Code Pre-Trained Models Via Contrastive Learning, Shangqing Liu, Bozhi Wu, Xiaofei Xie, Guozhu Meng, Yang. Liu

Research Collection School Of Computing and Information Systems

Large-scale pre-trained models such as CodeBERT, GraphCodeBERT have earned widespread attention from both academia and industry. Attributed to the superior ability in code representation, they have been further applied in multiple downstream tasks such as clone detection, code search and code translation. However, it is also observed that these state-of-the-art pre-trained models are susceptible to adversarial attacks. The performance of these pre-trained models drops significantly with simple perturbations such as renaming variable names. This weakness may be inherited by their downstream models and thereby amplified at an unprecedented scale. To this end, we propose an approach namely ContraBERT that aims …


Widget Detection-Based Testing For Industrial Mobile Games, Xiongfei Wu, Jiaming Ye, Ke Chen, Xiaofei Xie, Ruochen Huang, Lei Ma, Jianjun Zhao May 2023

Widget Detection-Based Testing For Industrial Mobile Games, Xiongfei Wu, Jiaming Ye, Ke Chen, Xiaofei Xie, Ruochen Huang, Lei Ma, Jianjun Zhao

Research Collection School Of Computing and Information Systems

The fast advances in mobile hardware and widespread smartphone usage have fueled the growth of global mobile gaming in the past decade. As a result, the need for quality assurance of mobile gaming has become increasingly pressing. While general-purpose testing methods have been developed for mobile applications, they become struggling when being applied to mobile games due to the unique characteristics of mobile games, such as dynamic loading and stunning visual effects. There comes a growing industrial demand for automated testing techniques with high compatibility (compatible with various resolutions, and platforms) and non-intrusive characteristics (without packaging external modules into the …


Multi-Lingual Multi-Partite Product Title Matching, Huan Lin Tay, Wei Jie Tay, Hady Wirawan Lauw May 2023

Multi-Lingual Multi-Partite Product Title Matching, Huan Lin Tay, Wei Jie Tay, Hady Wirawan Lauw

Research Collection School Of Computing and Information Systems

In a globalized marketplace, one could access products or services from almost anywhere. However, resolving which product in one language corresponds to another product in a different language remains an under-explored problem. We explore this from two perspectives. First, given two products of different languages, how to assess their similarity that could signal a potential match. Second, given products from various languages, how to arrive at a multi-partite clustering that respects cardinality constraints efficiently. We describe algorithms for each perspective and integrate them into a promising solution validated on real-world datasets.


Resale Hdb Price Prediction Considering Covid-19 Through Sentiment Analysis, Srinaath Anbu Durai, Zhaoxia Wang May 2023

Resale Hdb Price Prediction Considering Covid-19 Through Sentiment Analysis, Srinaath Anbu Durai, Zhaoxia Wang

Research Collection School Of Computing and Information Systems

Twitter sentiment has been used as a predictor to predict price values or trends in both the stock market and housing market. The pioneering works in this stream of research drew upon works in behavioural economics to show that sentiment or emotions impact economic decisions. Latest works in this stream focus on the algorithm used as opposed to the data used. A literature review of works in this stream through the lens of data used shows that there is a paucity of work that considers the impact of sentiments caused due to an external factor on either the stock or …


Wearing Masks Implies Refuting Trump?: Towards Target-Specific User Stance Prediction Across Events In Covid-19 And Us Election 2020, Hong Zhang, Haewoon Kwak, Wei Gao, Jisun An May 2023

Wearing Masks Implies Refuting Trump?: Towards Target-Specific User Stance Prediction Across Events In Covid-19 And Us Election 2020, Hong Zhang, Haewoon Kwak, Wei Gao, Jisun An

Research Collection School Of Computing and Information Systems

People who share similar opinions towards controversial topics could form an echo chamber and may share similar political views toward other topics as well. The existence of such connections, which we call connected behavior, gives researchers a unique opportunity to predict how one would behave for a future event given their past behaviors. In this work, we propose a framework to conduct connected behavior analysis. Neural stance detection models are trained on Twitter data collected on three seemingly independent topics, i.e., wearing a mask, racial equality, and Trump, to detect people’s stance, which we consider as their online behavior in …


Fine-Grained Commit-Level Vulnerability Type Prediction By Cwe Tree Structure, Shengyi Pan, Lingfeng Bao, Xin Xia, David Lo, Shanping Li May 2023

Fine-Grained Commit-Level Vulnerability Type Prediction By Cwe Tree Structure, Shengyi Pan, Lingfeng Bao, Xin Xia, David Lo, Shanping Li

Research Collection School Of Computing and Information Systems

Identifying security patches via code commits to allow early warnings and timely fixes for Open Source Software (OSS) has received increasing attention. However, the existing detection methods can only identify the presence of a patch (i.e., a binary classification) but fail to pinpoint the vulnerability type. In this work, we take the first step to categorize the security patches into fine-grained vulnerability types. Specifically, we use the Common Weakness Enumeration (CWE) as the label and perform fine-grained classification using categories at the third level of the CWE tree. We first formulate the task as a Hierarchical Multi-label Classification (HMC) problem, …


Colefunda: Explainable Silent Vulnerability Fix Identification, Jiayuan Zhou, Michael Pacheco, Jinfu Chen, Xing Hu, Xin Xia, David Lo, Ahmed E. Hassan May 2023

Colefunda: Explainable Silent Vulnerability Fix Identification, Jiayuan Zhou, Michael Pacheco, Jinfu Chen, Xing Hu, Xin Xia, David Lo, Ahmed E. Hassan

Research Collection School Of Computing and Information Systems

It is common practice for OSS users to leverage and monitor security advisories to discover newly disclosed OSS vulnerabilities and their corresponding patches for vulnerability remediation. It is common for vulnerability fixes to be publicly available one week earlier than their disclosure. This gap in time provides an opportunity for attackers to exploit the vulnerability. Hence, OSS users need to sense the fix as early as possible so that the vulnerability can be remediated before it is exploited. However, it is common for OSS to adopt a vulnerability disclosure policy which causes the majority of vulnerabilities to be fixed silently, …


A Study Of Variable-Role-Based Feature Enrichment In Neural Models Of Code, Aftab. Hussain, Md. Rafiqul Islam. Rabin, Bowen. Xu, David Lo, Mohammad Amin. Alipour May 2023

A Study Of Variable-Role-Based Feature Enrichment In Neural Models Of Code, Aftab. Hussain, Md. Rafiqul Islam. Rabin, Bowen. Xu, David Lo, Mohammad Amin. Alipour

Research Collection School Of Computing and Information Systems

Although deep neural models substantially reduce the overhead of feature engineering, the features readily available in the inputs might significantly impact training cost and the performance of the models. In this paper, we explore the impact of an unsuperivsed feature enrichment approach based on variable roles on the performance of neural models of code. The notion of variable roles (as introduced in the works of Sajaniemi et al. [1], [2]) has been found to help students' abilities in programming. In this paper, we investigate if this notion would improve the performance of neural models of code. To the best of …


Niche: A Curated Dataset Of Engineered Machine Learning Projects In Python, Ratnadira Widyasari, Zhou Yang, Ferdian Thung, Sheng Qin Sim, Fiona Wee, Camellia Lok, Jack Phan, Haodi Qi, Constance Tan, David Lo, David Lo May 2023

Niche: A Curated Dataset Of Engineered Machine Learning Projects In Python, Ratnadira Widyasari, Zhou Yang, Ferdian Thung, Sheng Qin Sim, Fiona Wee, Camellia Lok, Jack Phan, Haodi Qi, Constance Tan, David Lo, David Lo

Research Collection School Of Computing and Information Systems

Machine learning (ML) has gained much attention and has been incorporated into our daily lives. While there are numerous publicly available ML projects on open source platforms such as GitHub, there have been limited attempts in filtering those projects to curate ML projects of high quality. The limited availability of such a high-quality dataset poses an obstacle to understanding ML projects. To help clear this obstacle, we present NICHE, a manually labelled dataset consisting of 572 ML projects. Based on the evidence of good software engineering practices, we label 441 of these projects as engineered and 131 as non-engineered. This …


Msrl-Net: A Multi-Level Semantic Relation-Enhanced Learning Network For Aspect-Based Sentiment Analysis, Zhenda Hu, Zhaoxia Wang, Yinglin Wang, Ah-Hwee Tan May 2023

Msrl-Net: A Multi-Level Semantic Relation-Enhanced Learning Network For Aspect-Based Sentiment Analysis, Zhenda Hu, Zhaoxia Wang, Yinglin Wang, Ah-Hwee Tan

Research Collection School Of Computing and Information Systems

Aspect-based sentiment analysis (ABSA) aims to analyze the sentiment polarity of a given text towards several specific aspects. For implementing the ABSA, one way is to convert the original problem into a sentence semantic matching task, using pre-trained language models, such as BERT. However, for such a task, the intra- and inter-semantic relations among input sentence pairs are often not considered. Specifically, the semantic information and guidance of relations revealed in the labels, such as positive, negative and neutral, have not been completely exploited. To address this issue, we introduce a self-supervised sentence pair relation classification task and propose a …


Learning-Based Stock Trending Prediction By Incorporating Technical Indicators And Social Media Sentiment, Zhaoxia Wang, Zhenda Hu, Fang Li, Seng-Beng Ho, Erik Cambria May 2023

Learning-Based Stock Trending Prediction By Incorporating Technical Indicators And Social Media Sentiment, Zhaoxia Wang, Zhenda Hu, Fang Li, Seng-Beng Ho, Erik Cambria

Research Collection School Of Computing and Information Systems

Stock trending prediction is a challenging task due to its dynamic and nonlinear characteristics. With the development of social platform and artificial intelligence (AI), incorporating timely news and social media information into stock trending models becomes possible. However, most of the existing works focus on classification or regression problems when predicting stock market trending without fully considering the effects of different influence factors in different phases. To address this gap, this research solves stock trending prediction problem utilizing both technical indicators and sentiments of the social media text as influence factors in different situations. A 3-phase hybrid model is proposed …


Boosting Just-In-Time Defect Prediction With Specific Features Of C/C++ Programming Languages In Code Changes, Chao Ni, Xiaodan Xu, Kaiwen Yang, David Lo May 2023

Boosting Just-In-Time Defect Prediction With Specific Features Of C/C++ Programming Languages In Code Changes, Chao Ni, Xiaodan Xu, Kaiwen Yang, David Lo

Research Collection School Of Computing and Information Systems

Just-in-time (JIT) defect prediction can identify changes as defect-inducing ones or clean ones and many approaches are proposed based on several programming language-independent change-level features. However, different programming languages have different characteristics and consequently may affect the quality of software projects. Meanwhile, the C programming language, one of the most popular ones, is widely used to develop foundation applications (i.e., operating system, database, compiler, etc.) in IT companies and its change-level characteristics on project quality have not been fully investigated. Additionally, whether open-source C projects have similar important features to commercial projects has not been studied much.To address the aforementioned …


Trustworthy And Synergistic Artificial Intelligence For Software Engineering: Vision And Roadmaps, David Lo May 2023

Trustworthy And Synergistic Artificial Intelligence For Software Engineering: Vision And Roadmaps, David Lo

Research Collection School Of Computing and Information Systems

For decades, much software engineering research has been dedicated to devising automated solutions aimed at enhancing developer productivity and elevating software quality. The past two decades have witnessed an unparalleled surge in the development of intelligent solutions tailored for software engineering tasks. This momentum established the Artificial Intelligence for Software Engineering (AI4SE) area, which has swiftly become one of the most active and popular areas within the software engiueering field. This Future of Software Engineering (FoSE) paper navigates through several focal points. It commences with a succinct introduction and history of AI4SE. Thereafter, it underscores the core challenges inherent to …


What's Behind Tight Deadlines? Business Causes Of Technical Debt, Rodrigo Rebouças De Almeida, Christoph Treude, Uirá Kulesza May 2023

What's Behind Tight Deadlines? Business Causes Of Technical Debt, Rodrigo Rebouças De Almeida, Christoph Treude, Uirá Kulesza

Research Collection School Of Computing and Information Systems

What are the business causes behind tight deadlines? What drives the prioritization of features that pushes quality matters to the back burner? We conducted a survey with 71 experienced practitioners and did a thematic analysis of the openended answers to the question: “Could you give examples of how business may contribute to technical debt?” Business-related causes were organized into two categories: pure-business and business/IT gap, and they were related to ‘tight deadlines’ and ‘features over quality’, the most frequently cited management reasons for technical debt. We contribute a cause-effect model which relates the various business causes of tight deadlines and …


She Elicits Requirements And He Tests: Software Engineering Gender Bias In Large Language Models, Christoph Treude, Hideaki Hata May 2023

She Elicits Requirements And He Tests: Software Engineering Gender Bias In Large Language Models, Christoph Treude, Hideaki Hata

Research Collection School Of Computing and Information Systems

Implicit gender bias in software development is a well-documented issue, such as the association of technical roles with men. To address this bias, it is important to understand it in more detail. This study uses data mining techniques to investigate the extent to which 56 tasks related to software development, such as assigning GitHub issues and testing, are affected by implicit gender bias embedded in large language models. We systematically translated each task from English into a genderless language and back, and investigated the pronouns associated with each task. Based on translating each task 100 times in different permutations, we …


Towards Understanding The Open Source Interest In Gender-Related Github Projects, Rita Garcia, Christoph Treude, Wendy La May 2023

Towards Understanding The Open Source Interest In Gender-Related Github Projects, Rita Garcia, Christoph Treude, Wendy La

Research Collection School Of Computing and Information Systems

The open-source community uses the GitHub platform to exchange and share software applications and services of interest. This paper aims to identify the open-source community’s interest in gender-related projects on GitHub. Our findings create research opportunities and identify resources by the open-source community that promote diversity, equity, and inclusion. We use data mining to identify GitHub projects that focus on gender-related topics. We apply quantitative and qualitative methodologies to examine the projects’ attributes and to classify them within a gender social structure and a gender bias taxonomy. We aim to understand the open-source community’s efforts and interests in gender topics …


Applying Information Theory To Software Evolution, Adriano Torres, Sebastian Baltes, Christoph Treude, Markus Wagner May 2023

Applying Information Theory To Software Evolution, Adriano Torres, Sebastian Baltes, Christoph Treude, Markus Wagner

Research Collection School Of Computing and Information Systems

Although information theory has found success in disciplines, the literature on its applications to software evolution is limit. We are still missing artifacts that leverage the data and tooling available to measure how the information content of a project can be a proxy for its complexity. In this work, we explore two definitions of entropy, one structural and one textual, and apply it to the historical progression of the commit history of 25 open source projects. We produce evidence that they generally are highly correlated. We also observed that they display weak and unstable correlations with other complexity metrics. Our …


Overcoming Challenges In Devops Education Through Teaching Methods, Samuel Ferino, Marcelo Fernandes, Elder Cirilo, Lucas Agnez, Bruno Batista, Uirá Kulesza, Eduardo Aranha, Christoph Treude May 2023

Overcoming Challenges In Devops Education Through Teaching Methods, Samuel Ferino, Marcelo Fernandes, Elder Cirilo, Lucas Agnez, Bruno Batista, Uirá Kulesza, Eduardo Aranha, Christoph Treude

Research Collection School Of Computing and Information Systems

DevOps is a set of practices that deals with coordination between development and operation teams and ensures rapid and reliable new software releases that are essential in industry. DevOps education assumes the vital task of preparing new professionals in these practices using appropriate teaching methods. However, there are insufficient studies investigating teaching methods in DevOps. We performed an analysis based on interviews to identify teaching methods and their relationship with DevOps educational challenges. Our findings show that project-based learning and collaborative learning are emerging as the most relevant teaching methods.


Stop Words For Processing Software Engineering Documents: Do They Matter, Yaohou Fan, Chetan Arora, Christoph Treude May 2023

Stop Words For Processing Software Engineering Documents: Do They Matter, Yaohou Fan, Chetan Arora, Christoph Treude

Research Collection School Of Computing and Information Systems

Stop words, which are considered non-predictive, are often eliminated in natural language processing tasks. However, the definition of uninformative vocabulary is vague, so most algorithms use general knowledge-based stop lists to remove stop words. There is an ongoing debate among academics about the usefulness of stop word elimination, especially in domainspecific settings. In this work, we investigate the usefulness of stop word removal in a software engineering context. To do this, we replicate and experiment with three software engineering research tools from related work. Additionally, we construct a corpus of software engineering domain-related text from 10,000 Stack Overflow questions and …


Lpt: Long-Tailed Prompt Tuning For Image Classification, Bowen Dong, Pan Zhou, Shuicheng Yan, Wangmeng Zuo May 2023

Lpt: Long-Tailed Prompt Tuning For Image Classification, Bowen Dong, Pan Zhou, Shuicheng Yan, Wangmeng Zuo

Research Collection School Of Computing and Information Systems

For long-tailed classification tasks, most works often pretrain a big model on a large-scale (unlabeled) dataset, and then fine-tune the whole pretrained model for adapting to long-tailed data. Though promising, fine-tuning the whole pretrained model tends to suffer from high cost in computation and deployment of different models for different tasks, as well as weakened generalization capability for overfitting to certain features of long-tailed data. To alleviate these issues, we propose an effective Long-tailed Prompt Tuning (LPT) method for long-tailed classification tasks. LPT introduces several trainable prompts into a frozen pretrained model to adapt it to long-tailed data. For better …


Towards Understanding Why Mask Reconstruction Pretraining Helps In Downstream Tasks, Jiachun Pan, Pan Zhou, Shuicheng Yan May 2023

Towards Understanding Why Mask Reconstruction Pretraining Helps In Downstream Tasks, Jiachun Pan, Pan Zhou, Shuicheng Yan

Research Collection School Of Computing and Information Systems

For unsupervised pretraining, mask-reconstruction pretraining (MRP) approaches, e.g. MAE (He et al., 2021) and data2vec (Baevski et al., 2022), randomly mask input patches and then reconstruct the pixels or semantic features of these masked patches via an auto-encoder. Then for a downstream task, supervised fine-tuning the pretrained encoder remarkably surpasses the conventional “supervised learning" (SL) trained from scratch. However, it is still unclear 1) how MRP performs semantic feature learning in the pretraining phase and 2) why it helps in downstream tasks. To solve these problems, we first theoretically show that on an auto-encoder of a two/one-layered convolution encoder/decoder, MRP …