Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences Commons™

Open Access. Powered by Scholars. Published by Universities.®

Databases and Information Systems

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 1651 - 1680 of 7256

Full-Text Articles in Computer Sciences

Effective Digital Learning Practices For Is Design Courses During Covid-19, Eng Lieh Ouh, Benjamin Gan Aug 2021

Effective Digital Learning Practices For Is Design Courses During Covid-19, Eng Lieh Ouh, Benjamin Gan

Research Collection School Of Computing and Information Systems

The COVID-19 pandemic has pushed educational institutions to adopt digital learning for an extended period. This research studies the effectiveness of digital learning practices based on student feedback data collected for two Information Systems design courses: human interaction design and solution architecture design. This paper leverages the data to analyze the effectiveness of a set of digital learning practices: ZOOM lectures, polling or Kahoot questions, self-reflection, virtual exercises and virtual mentorship. Our research questions are on the effectiveness of these learning practices to keep the student’s interest and learn the course materials. The research compares each learning practice and the …


Towards Generative Aspect-Based Sentiment Analysis, Wenxuan Zhang, Xin Li, Yang Deng, Lidong Bing, Wai Lam Aug 2021

Towards Generative Aspect-Based Sentiment Analysis, Wenxuan Zhang, Xin Li, Yang Deng, Lidong Bing, Wai Lam

Research Collection School Of Computing and Information Systems

Aspect-based sentiment analysis (ABSA) has received increasing attention recently. Most existing work tackles ABSA in a discriminative manner, designing various task-specific classification networks for the prediction. Despite their effectiveness, these methods ignore the rich label semantics in ABSA problems and require extensive task-specific designs. In this paper, we propose to tackle various ABSA tasks in a unified generative framework. Two types of paradigms, namely annotation-style and extraction-style modeling, are designed to enable the training process by formulating each ABSA task as a text generation problem. We conduct experiments on four ABSA tasks across multiple benchmark datasets where our proposed generative …


An Empirical Study Of The Discreteness Prior In Low-Rank Matrix Completion, Rodrigo Alves, Antoine Ledent, Renato Assunção, Marius And Kloft Aug 2021

An Empirical Study Of The Discreteness Prior In Low-Rank Matrix Completion, Rodrigo Alves, Antoine Ledent, Renato Assunção, Marius And Kloft

Research Collection School Of Computing and Information Systems

A reasonable assumption in recommender systems is that the rows (users) and columns (items) of the rating matrix can be split into groups (communities) with the following property: each entry of the matrix is the sum of components corresponding to community behavior and a purely low-rank component corresponding to individual behavior. We investigate (1) whether such a structure is present in real-world datasets, (2) whether the knowledge of the existence of such structure alone can improve performance, without explicit information about the community memberships. To these ends, we formulate a joint optimization problem over all (completed matrix, set of communities) …


Are Missing Links Predictable? An Inferential Benchmark For Knowledge Graph Completion, Yixin Cao, Xiang Ji, Xin Lv, Juanzi Li, Yonggang Wen, Hanwang Zhang Aug 2021

Are Missing Links Predictable? An Inferential Benchmark For Knowledge Graph Completion, Yixin Cao, Xiang Ji, Xin Lv, Juanzi Li, Yonggang Wen, Hanwang Zhang

Research Collection School Of Computing and Information Systems

We present InferWiki, a Knowledge Graph Completion (KGC) dataset that improves upon existing benchmarks in inferential ability, assumptions, and patterns. First, each testing sample is predictable with supportive data in the training set. To ensure it, we propose to utilize rule-guided train/test generation, instead of conventional random split. Second, InferWiki initiates the evaluation following the open-world assumption and improves the inferential difficulty of the closed-world assumption, by providing manually annotated negative and unknown triples. Third, we include various inference patterns (e.g., reasoning path length and types) for comprehensive evaluation. In experiments, we curate two settings of InferWiki varying in sizes …


Learning From Miscellaneous Other-Class Words For Few-Shot Named Entity Recognition, Meihan Tong, Shuai Wang, Bin Xu, Yixin Cao, Minghui Liu, Lei Hou, Juanzi Li Aug 2021

Learning From Miscellaneous Other-Class Words For Few-Shot Named Entity Recognition, Meihan Tong, Shuai Wang, Bin Xu, Yixin Cao, Minghui Liu, Lei Hou, Juanzi Li

Research Collection School Of Computing and Information Systems

Few-shot Named Entity Recognition (NER) exploits only a handful of annotations to identify and classify named entity mentions. Prototypical network shows superior performance on few-shot NER. However, existing prototypical methods fail to differentiate rich semantics in other-class words, which will aggravate overfitting under few shot scenario. To address the issue, we propose a novel model, Mining Undefined Classes from Other-class (MUCO), that can automatically induce different undefined classes from the other class to improve few-shot NER. With these extra-labeled undefined classes, our method will improve the discriminative ability of NER classifier and enhance the understanding of predefined classes with stand-by …


How Knowledge Graph And Attention Help? A Qualitative Analysis Into Bag-Level Relation Extraction, Zikun Hu, Yixin Cao, Lifu Huang, Tat-Seng Chua Aug 2021

How Knowledge Graph And Attention Help? A Qualitative Analysis Into Bag-Level Relation Extraction, Zikun Hu, Yixin Cao, Lifu Huang, Tat-Seng Chua

Research Collection School Of Computing and Information Systems

Knowledge Graph (KG) and attention mechanism have been demonstrated effective in introducing and selecting useful information for weakly supervised methods. However, only qualitative analysis and ablation study are provided as evidence. In this paper, we contribute a dataset and propose a paradigm to quantitatively evaluate the effect of attention and KG on bag-level relation extraction (RE). We find that (1) higher attention accuracy may lead to worse performance as it may harm the model’s ability to extract entity mention features; (2) the performance of attention is largely influenced by various noise distribution patterns, which is closely related to real-world datasets; …


The 4th Workshop On Heterogeneous Information Network Analysis And Applications (Hena 2021), Chuan Shi, Yuan Fang, Yanfang Ye, Jiawei Zhang Aug 2021

The 4th Workshop On Heterogeneous Information Network Analysis And Applications (Hena 2021), Chuan Shi, Yuan Fang, Yanfang Ye, Jiawei Zhang

Research Collection School Of Computing and Information Systems

The 4th Workshop on Heterogeneous Information Network Analysis and Applications (HENA 2021) is co-located with the 27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. The goal of this workshop is to bring together researchers and practitioners in the field and provide a forum for sharing new techniques and applications in heterogeneous information network analysis. This workshop has an exciting program that spans a number of subtopics, such as heterogeneous network embedding and graph neural networks, data mining techniques on heterogeneous information networks, and applications of heterogeneous information network analysis. The workshop program includes several invited speakers, lively discussion …


A Survey On Ml4vis: Applying Machine Learning Advances To Data Visualization, Qianwen Wang, Zhutian Chen, Yong Wang, Huamin Qu Aug 2021

A Survey On Ml4vis: Applying Machine Learning Advances To Data Visualization, Qianwen Wang, Zhutian Chen, Yong Wang, Huamin Qu

Research Collection School Of Computing and Information Systems

Inspired by the great success of machine learning (ML), researchers have applied ML techniques to visualizations to achieve a better design, development, and evaluation of visualizations. This branch of studies, known as ML4VIS, is gaining increasing research attention in recent years. To successfully adapt ML techniques for visualizations, a structured understanding of the integration of ML4VIS is needed. In this article, we systematically survey 88 ML4VIS studies, aiming to answer two motivating questions: “what visualization processes can be assisted by ML?” and “how ML techniques can be used to solve visualization problems? ” This survey reveals seven main processes where …


Multilateration Index., Chip Lynch Aug 2021

Multilateration Index., Chip Lynch

Electronic Theses and Dissertations

We present an alternative method for pre-processing and storing point data, particularly for Geospatial points, by storing multilateration distances to fixed points rather than coordinates such as Latitude and Longitude. We explore the use of this data to improve query performance for some distance related queries such as nearest neighbor and query-within-radius (i.e. “find all points in a set P within distance d of query point q”). Further, we discuss the problem of “Network Adequacy” common to medical and communications businesses, to analyze questions such as “are at least 90% of patients living within 50 miles of a covered emergency …


Spatial Analyses Of Gray Fossil Site Vertebrate Remains: Implications For Depositional Setting And Site Formation Processes, David Carney Aug 2021

Spatial Analyses Of Gray Fossil Site Vertebrate Remains: Implications For Depositional Setting And Site Formation Processes, David Carney

Electronic Theses and Dissertations

This project uses exploratory 3D geospatial analyses to assess the taphonomy of the Gray Fossil Site (GFS). During the Pliocene, the GFS was a forested, inundated sinkhole that accumulated biological materials between 4.9-4.5 mya. This deposit contains fossils exhibiting different preservation modes: from low energy lacustrine settings to high energy colluvial deposits. All macro-paleontological materials have been mapped in situ using survey-grade instrumentation. Vertebrate skeletal material from the site is well-preserved, but the degree of skeletal articulation varies spatially within the deposit. This analysis uses geographic information systems (GIS) to analyze the distribution of mapped specimens at different spatial scales. …


Leveraging Two Types Of Global Graph For Sequential Fashion Recommendation, Yujuan Ding, Yunshan Ma, Wai Keung Wong, Tat‑Seng Chua Aug 2021

Leveraging Two Types Of Global Graph For Sequential Fashion Recommendation, Yujuan Ding, Yunshan Ma, Wai Keung Wong, Tat‑Seng Chua

Research Collection School Of Computing and Information Systems

Sequential fashion recommendation is of great significance in online fashion shopping, which accounts for an increasing portion of either fashion retailing or online e-commerce. The key to building an effective sequential fashion recommendation model lies in capturing two types of patterns: the personal fashion preference of users and the transitional relationships between adjacent items. The two types of patterns are usually related to user-item interaction and item-item transition modeling respectively. However, due to the large sets of users and items as well as the sparse historical interactions, it is difficult to train an effective and efficient sequential fashion recommendation model. …


Metaxmorph: Hierarchical Transformation Of Data With Metadata, Shubham Airan Aug 2021

Metaxmorph: Hierarchical Transformation Of Data With Metadata, Shubham Airan

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

This research is about transforming data. Data comes in different shapes; it can be structured as a graph, a tree, a collection of tables, or some other shape. In this thesis, we focus on data structured as a tree, which is known as hierarchical data. The same data could be structured in many different tree shapes. Previously it was shown how to transform data from one tree shape, one hierarchy to another without losing any information. But sometimes the pieces of the hierarchy are annotated or associated with metadata, that is, with data about the data itself. The metadata can …


Deep Learning Data And Indexes In A Database, Vishal Sharma Aug 2021

Deep Learning Data And Indexes In A Database, Vishal Sharma

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

A database is used to store and retrieve data, which is a critical component for any software application. Databases requires configuration for efficiency, however, there are tens of configuration parameters. It is a challenging task to manually configure a database. Furthermore, a database must be reconfigured on a regular basis to keep up with newer data and workload. The goal of this thesis is to use the query workload history to autonomously configure the database and improve its performance. We achieve proposed work in four stages: (i) we develop an index recommender using deep reinforcement learning for a standalone database. …


Socio-Technical Perspective For Electronic Tax Information System In Tanzania, Lucas Ngowi, Ellen Kalinga Jul 2021

Socio-Technical Perspective For Electronic Tax Information System In Tanzania, Lucas Ngowi, Ellen Kalinga

Tanzania Journal of Engineering and Technology (TJET)

Socio-technical systems theory has rarely been used by system architects in setting up computing systems. However, the role of socio-technical concepts in computing, which is becoming social in nature, has made the concepts more relevant and commercial. Tax information systems are examples of such systems because they are influenced by external variables such as the political environment, technological trends, and social environment, introducing complexity in their deployment and determining the type of e-services and their delivery to a diverse group of people. It was observed that in Tanzania there is resistance, reluctance and minimal use of electronic tax system because …


The Is Social Continuance Model: Using Conversational Agents To Support Co-Creation, Naif Alawi Jul 2021

The Is Social Continuance Model: Using Conversational Agents To Support Co-Creation, Naif Alawi

USF Tampa Graduate Theses and Dissertations

With the rise of Agentic IS Artifact and the increasing integration of this technology within organizations, our understanding of the impact of this technology on individuals remains limited. Although IS use literature provides important guidance for organization to increase employees’ willingness to work with new technology implementations, the utilitarian view of prior IS use limits its application in light of the new evolving social interaction between humans and Agentic IS Artifacts. To that end, we contribute to the IS use literature by implementing a social view to understand the impact of Agentic IS Artifacts on an individual’s perception and behavior. …


Dan Farkas, Dan Farkas Jul 2021

Dan Farkas, Dan Farkas

Oral History

Dan Farkas has taught on the Pleasantville campus of Pace University since 1977.


Variational Learning From Implicit Bandit Feedback, Quoc Tuan Truong, Hady W. Lauw Jul 2021

Variational Learning From Implicit Bandit Feedback, Quoc Tuan Truong, Hady W. Lauw

Research Collection School Of Computing and Information Systems

Recommendations are prevalent in Web applications (e.g., search ranking, item recommendation, advertisement placement). Learning from bandit feedback is challenging due to the sparsity of feedback limited to system-provided actions. In this work, we focus on batch learning from logs of recommender systems involving both bandit and organic feedbacks. We develop a probabilistic framework with a likelihood function for estimating not only explicit positive observations but also implicit negative observations inferred from the data. Moreover, we introduce a latent variable model for organic-bandit feedbacks to robustly capture user preference distributions. Next, we analyze the behavior of the new likelihood under two …


A Differentially Private Task Planning Framework For Spatial Crowdsourcing, Qian Tao, Yongxin Tong, Shuyuan Li, Yuxiang Zeng, Zimu Zhou, Ke Xu Jul 2021

A Differentially Private Task Planning Framework For Spatial Crowdsourcing, Qian Tao, Yongxin Tong, Shuyuan Li, Yuxiang Zeng, Zimu Zhou, Ke Xu

Research Collection School Of Computing and Information Systems

Spatial crowdsourcing has stimulated various new applications such as taxi calling and food delivery. A key enabler for these spatial crowdsourcing based applications is to plan routes for crowd workers to execute tasks given diverse requirements of workers and the spatial crowdsourcing platform. Despite extensive studies on task planning in spatial crowdsourcing, few have accounted for the location privacy of tasks, which may be misused by an untrustworthy platform. In this paper, we explore efficient task planning for workers while protecting the locations of tasks. Specifically, we define the Privacy-Preserving Task Planning (PPTP) problem, which aims at both total revenue …


A Coprocessor-Based Introspection Framework Via Intel Management Engine, Lei Zhou, Fengwei Zhang, Jidong Xiao, Kevin Leach, Westley Weimer, Xuhua Ding, Guojun Wang Jul 2021

A Coprocessor-Based Introspection Framework Via Intel Management Engine, Lei Zhou, Fengwei Zhang, Jidong Xiao, Kevin Leach, Westley Weimer, Xuhua Ding, Guojun Wang

Research Collection School Of Computing and Information Systems

During the past decade, virtualization-based (e.g., virtual machine introspection) and hardware-assisted approaches (e.g., x86 SMM and ARM TrustZone) have been used to defend against low-level malware such as rootkits. However, these approaches either require a large Trusted Computing Base (TCB) or they must share CPU time with the operating system, disrupting normal execution. In this article, we propose an introspection framework called NIGHTHAWK that transparently checks system integrity and monitor the runtime state of target system. NIGHTHAWK leverages the Intel Management Engine (IME), a co-processor that runs in isolation from the main CPU. By using the IME, our approach has …


Dehumor: Visual Analytics For Decomposing Humor, Xingbo Wang, Yao Ming, Tongshuang Wu, Haipeng Zeng, Yong Wang, Huamin Qu Jul 2021

Dehumor: Visual Analytics For Decomposing Humor, Xingbo Wang, Yao Ming, Tongshuang Wu, Haipeng Zeng, Yong Wang, Huamin Qu

Research Collection School Of Computing and Information Systems

Despite being a critical communication skill, grasping humor is challenginga successful use of humor requires a mixture of both engaging content build-up and an appropriate vocal delivery (e.g., pause). Prior studies on computational humor emphasize the textual and audio features immediately next to the punchline, yet overlooking longer-term context setup. Moreover, the theories are usually too abstract for understanding each concrete humor snippet. To fill in the gap, we develop DeHumor, a visual analytical system for analyzing humorous behaviors in public speaking. To intuitively reveal the building blocks of each concrete example, DeHumor decomposes each humorous video into multimodal features …


A Mean-Field Markov Decision Process Model For Spatial-Temporal Subsidies In Ride-Sourcing Markets, Zheng Zhu, Jintao Ke, Hai Wang Jul 2021

A Mean-Field Markov Decision Process Model For Spatial-Temporal Subsidies In Ride-Sourcing Markets, Zheng Zhu, Jintao Ke, Hai Wang

Research Collection School Of Computing and Information Systems

Ride-sourcing services are increasingly popular because of their ability to accommodate on-demand travel needs. A critical issue faced by ride-sourcing platforms is the supply-demand imbalance, as a result of which drivers may spend substantial time on idle cruising and picking up remote passengers. Some platforms attempt to mitigate the imbalance by providing relocation guidance for idle drivers who may have their own self-relocation strategies and decline to follow the suggestions. Platforms then seek to induce drivers to system-desirable locations by offering them subsidies. This paper proposes a mean-field Markov decision process (MF-MDP) model to depict the dynamics in ride-sourcing markets …


Integrated Framework For Developing Instructional Videos For Foundational Computing Courses, Kyong Jin Shim, Gottipati Swapna, Yi Meng Lau Jul 2021

Integrated Framework For Developing Instructional Videos For Foundational Computing Courses, Kyong Jin Shim, Gottipati Swapna, Yi Meng Lau

Research Collection School Of Computing and Information Systems

Instructional videos are widely used in higher education due to their effectiveness and flexibility of personalized learning features. Computing courses usually focuses on programming, user interface design, server connectivity, data storage, and architecture, among others. The design of instructional videos varies in not only the course content but also the style of content creation. We propose an integrated framework, Computing Videos Design Framework (CVDF), for designing and developing instructional videos for computing courses. CVDF combines the cognitive skills from Bloom’s taxonomy, video design principles, and course learning outcomes for designing different types of instructional videos. We apply the framework to …


Meta-Inductive Node Classification Across Graphs, Zhihao Wen, Yuan Fang, Zemin Liu Jul 2021

Meta-Inductive Node Classification Across Graphs, Zhihao Wen, Yuan Fang, Zemin Liu

Research Collection School Of Computing and Information Systems

Semi-supervised node classification on graphs is an important research problem, with many real-world applications in information retrieval such as content classification on a social network and query intent classification on an e-commerce query graph. While traditional approaches are largely transductive, recent graph neural networks (GNNs) integrate node features with network structures, thus enabling inductive node classification models that can be applied to new nodes or even new graphs in the same feature space. However, inter-graph differences still exist across graphs within the same domain. Thus, training just one global model (e.g., a state-of-the-art GNN) to handle all new graphs, whilst …


Paying Attention To Video Object Pattern Understanding, Wenguan Wang, Jianbing Shen, Xiankai Lu, Steven C. H. Hoi, Haibin Ling Jul 2021

Paying Attention To Video Object Pattern Understanding, Wenguan Wang, Jianbing Shen, Xiankai Lu, Steven C. H. Hoi, Haibin Ling

Research Collection School Of Computing and Information Systems

This paper conducts a systematic study on the role of visual attention in video object pattern understanding. By elaborately annotating three popular video segmentation datasets (DAVIS) with dynamic eye-tracking data in the unsupervised video object segmentation (UVOS) setting. For the first time, we quantitatively verified the high consistency of visual attention behavior among human observers, and found strong correlation between human attention and explicit primary object judgments during dynamic, task-driven viewing. Such novel observations provide an in-depth insight of the underlying rationale behind video object pattens. Inspired by these findings, we decouple UVOS into two sub-tasks: UVOS-driven Dynamic Visual Attention …


Oesense: Employing Occlusion Effect For In-Ear Human Sensing, Dong Ma, Andrea Ferlini, Cecilia Mascolo Jul 2021

Oesense: Employing Occlusion Effect For In-Ear Human Sensing, Dong Ma, Andrea Ferlini, Cecilia Mascolo

Research Collection School Of Computing and Information Systems

Smart earbuds are recognized as a new wearable platform for personal-scale human motion sensing. However, due to the interference from head movement or background noise, commonly-used modalities (e.g. accelerometer and microphone) fail to reliably detect both intense and light motions. To obviate this, we propose OESense, an acoustic-based in-ear system for general human motion sensing. The core idea behind OESense is the joint use of the occlusion effect (i.e., the enhancement of low-frequency components of bone-conducted sounds in an occluded ear canal) and inward-facing microphone, which naturally boosts the sensing signal and suppresses external interference. We prototype OESense as an …


Exploring Cross-Modality Utilization In Recommender Systems, Quoc Tuan Truong, Aghiles Salah, Thanh-Binh Tran, Jingyao Guo, Hady W. Lauw Jul 2021

Exploring Cross-Modality Utilization In Recommender Systems, Quoc Tuan Truong, Aghiles Salah, Thanh-Binh Tran, Jingyao Guo, Hady W. Lauw

Research Collection School Of Computing and Information Systems

Multimodal recommender systems alleviate the sparsity of historical user-item interactions. They are commonly catalogued based on the type of auxiliary data (modality) they leverage, such as preference data plus user-network (social), user/item texts (textual), or item images (visual) respectively. One consequence of this categorization is the tendency for virtual walls to arise between modalities. For instance, a study involving images would compare to only baselines ostensibly designed for images. However, a closer look at existing models' statistical assumptions about any one modality would reveal that many could work just as well with other modalities. Therefore, we pursue a systematic investigation …


Make It Easy: An Effective End-To-End Entity Alignment Framework, Congcong Ge, Xiaoze Liu, Lu Chen Chen, Baihua Zheng, Yunjun Gao Jul 2021

Make It Easy: An Effective End-To-End Entity Alignment Framework, Congcong Ge, Xiaoze Liu, Lu Chen Chen, Baihua Zheng, Yunjun Gao

Research Collection School Of Computing and Information Systems

Entity alignment (EA) is a prerequisite for enlarging the coverage of a unified knowledge graph. Previous EA approaches either restrain the performance due to inadequate information utilization or need labor-intensive pre-processing to get external or reliable information to perform the EA task. This paper proposes EASY, an effective end-to-end EA framework, which is able to (i) remove the labor-intensive pre-processing by fully discovering the name information provided by the entities themselves; and (ii) jointly fuse the features captured by the names of entities and the structural information of the graph to improve the EA results. Specifically, EASY first introduces NEAP, …


Frameaxis: Characterizing Microframe Bias And Intensity With Word Embedding, Haewoon Kwak, Jisun An, Elise Jing Jing, Yong-Yeol Ahn Jul 2021

Frameaxis: Characterizing Microframe Bias And Intensity With Word Embedding, Haewoon Kwak, Jisun An, Elise Jing Jing, Yong-Yeol Ahn

Research Collection School Of Computing and Information Systems

Framing is a process of emphasizing a certain aspect of an issue over the others, nudging readers or listeners towards different positions on the issue even without making a biased argument. Here, we propose FrameAxis, a method for characterizing documents by identifying the most relevant semantic axes (“microframes”) that are overrepresented in the text using word embedding. Our unsupervised approach can be readily applied to large datasets because it does not require manual annotations. It can also provide nuanced insights by considering a rich set of semantic axes. FrameAxis is designed to quantitatively tease out two important dimensions of how …


Unified Conversational Recommendation Policy Learning Via Graph-Based Reinforcement Learning, Yang Deng, Yaliang Li, Fei Sun, Bolin Ding, Wai Lam Jul 2021

Unified Conversational Recommendation Policy Learning Via Graph-Based Reinforcement Learning, Yang Deng, Yaliang Li, Fei Sun, Bolin Ding, Wai Lam

Research Collection School Of Computing and Information Systems

Conversational recommender systems (CRS) enable the traditional recommender systems to explicitly acquire user preferences towards items and attributes through interactive conversations. Reinforcement learning (RL) is widely adopted to learn conversational recommendation policies to decide what attributes to ask, which items to recommend, and when to ask or recommend, at each conversation turn. However, existing methods mainly target at solving one or two of these three decision-making problems in CRS with separated conversation and recommendation components, which restrict the scalability and generality of CRS and fall short of preserving a stable training procedure. In the light of these challenges, we propose …


Page: A Simple And Optimal Probabilistic Gradient Estimator For Nonconvex Optimization, Zhize Li, Hongyan Bao, Xiangliang Zhang, Peter Richtarik Jul 2021

Page: A Simple And Optimal Probabilistic Gradient Estimator For Nonconvex Optimization, Zhize Li, Hongyan Bao, Xiangliang Zhang, Peter Richtarik

Research Collection School Of Computing and Information Systems

In this paper, we propose a novel stochastic gradient estimator---ProbAbilistic Gradient Estimator (PAGE)---for nonconvex optimization. PAGE is easy to implement as it is designed via a small adjustment to vanilla SGD: in each iteration, PAGE uses the vanilla minibatch SGD update with probability $p_t$ or reuses the previous gradient with a small adjustment, at a much lower computational cost, with probability $1-p_t$. We give a simple formula for the optimal choice of $p_t$. Moreover, we prove the first tight lower bound $\Omega(n+\frac{\sqrt{n}}{\epsilon^2})$ for nonconvex finite-sum problems, which also leads to a tight lower bound $\Omega(b+\frac{\sqrt{b}}{\epsilon^2})$ for nonconvex online problems, where …