Open Access. Powered by Scholars. Published by Universities.®
Databases and Information Systems Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Social and Behavioral Sciences (1505)
- Communication (945)
- Business (886)
- OS and Networks (858)
- Numerical Analysis and Scientific Computing (808)
-
- Life Sciences (727)
- Communication Technology and New Media (691)
- Bioinformatics (656)
- Science and Technology Studies (650)
- Artificial Intelligence and Robotics (625)
- Engineering (547)
- Software Engineering (524)
- Graphics and Human Computer Interfaces (473)
- Information Security (394)
- Theory and Algorithms (355)
- Computer Engineering (344)
- Management Information Systems (294)
- Medicine and Health Sciences (261)
- Social Media (220)
- Education (216)
- Other Computer Sciences (215)
- Data Science (212)
- Systems Architecture (187)
- Library and Information Science (181)
- Programming Languages and Compilers (172)
- Public Affairs, Public Policy and Public Administration (167)
- Business Administration, Management, and Operations (164)
- Institution
-
- Singapore Management University (3555)
- Wright State University (631)
- Walden University (447)
- New Jersey Institute of Technology (143)
- University of Malaya (130)
-
- University of Nebraska at Omaha (119)
- Old Dominion University (108)
- California State University, San Bernardino (100)
- San Jose State University (89)
- University of Dayton (82)
- City University of New York (CUNY) (70)
- University of Dar es Salaam (63)
- Air Force Institute of Technology (61)
- University of Nebraska - Lincoln (60)
- University of South Florida (56)
- Kennesaw State University (54)
- Nova Southeastern University (52)
- Technological University Dublin (51)
- University of Arkansas, Fayetteville (46)
- Dakota State University (43)
- Claremont Colleges (42)
- California Polytechnic State University, San Luis Obispo (41)
- Institute of Business Administration (38)
- Western Kentucky University (36)
- Purdue University (35)
- Ateneo de Manila University (34)
- Governors State University (34)
- Portland State University (34)
- University of Arkansas Little Rock (33)
- University of Nevada, Las Vegas (32)
- Keyword
-
- Machine learning (122)
- Information technology (91)
- Data mining (90)
- Social media (83)
- Machine Learning (64)
-
- Cybersecurity (63)
- Deep learning (60)
- Twitter (60)
- Artificial intelligence (58)
- Semantic Web (53)
- Online learning (51)
- Databases (46)
- Cloud computing (45)
- Information Technology (45)
- Information retrieval (45)
- Classification (43)
- Database (42)
- Blockchain (41)
- Natural language processing (41)
- Ontology (41)
- Big data (40)
- Security (39)
- Technology (39)
- Computer science (38)
- Privacy (38)
- Algorithms (37)
- Clustering (37)
- Deep Learning (37)
- Information systems (37)
- Management (37)
- Publication Year
- Publication
-
- Research Collection School Of Computing and Information Systems (3436)
- Kno.e.sis Publications (540)
- Walden Dissertations and Doctoral Studies (447)
- Theses and Dissertations (129)
- Student Works (2000-2009) (120)
-
- Dissertations (113)
- Computer Science Faculty Publications (95)
- Computer Science and Engineering Faculty Publications (91)
- Theses Digitization Project (86)
- Master's Projects (68)
- Information Systems and Quantitative Analysis Faculty Proceedings & Presentations (64)
- Tanzania Journal of Engineering and Technology (TJET) (60)
- Dissertations and Theses Collection (Open Access) (58)
- USF Tampa Graduate Theses and Dissertations (51)
- Theses (48)
- CCAC Theses and Dissertations (43)
- Information Systems and Quantitative Analysis Faculty Publications (41)
- CGU Faculty Publications and Research (37)
- International Conference on Information and Communication Technologies (36)
- Open Educational Resources (35)
- Graduate Theses and Dissertations (34)
- Department of Information Systems & Computer Science Faculty Publications (33)
- All Capstone Projects (32)
- Masters Theses & Doctoral Dissertations (32)
- Conference papers (28)
- All Maxine Goodman Levin School of Urban Affairs Publications (27)
- UBT International Conference (23)
- Electronic Theses and Dissertations (22)
- Faculty Articles (22)
- Master's Theses (22)
- Publication Type
- File Type
Articles 1741 - 1770 of 7250
Full-Text Articles in Databases and Information Systems
On The Root Of Trust Identification Problem, Ivan De Oliveira Nunes, Xuhua Ding, Gene Tsudik
On The Root Of Trust Identification Problem, Ivan De Oliveira Nunes, Xuhua Ding, Gene Tsudik
Research Collection School Of Computing and Information Systems
Trusted Execution Environments (TEEs) are becoming ubiquitous and are currently used in many security applications: from personal IoT gadgets to banking and databases. Prominent examples of such architectures are Intel SGX, ARM TrustZone, and Trusted Platform Modules (TPMs). A typical TEE relies on a dynamic Root of Trust (RoT) to provide security services such as code/data confidentiality and integrity, isolated secure software execution, remote attestation, and sensor auditing. Despite their usefulness, there is currently no secure means to determine whether a given security service or task is being performed by the particular RoT within a specific physical device. We refer …
An Empirical Study Of The Landscape Of Open Source Projects In Baidu, Alibaba, And Tencent, Junxiao Han, Shuiguang Deng, David Lo, Chen Zhi, Jianwei Yin, Xin Xia
An Empirical Study Of The Landscape Of Open Source Projects In Baidu, Alibaba, And Tencent, Junxiao Han, Shuiguang Deng, David Lo, Chen Zhi, Jianwei Yin, Xin Xia
Research Collection School Of Computing and Information Systems
Open source software has drawn more and more attention from researchers, developers and companies nowadays. Meanwhile, many Chinese technology companies are embracing open source and choosing to open source their projects. Nevertheless, most previous studies are concentrated on international companies such as Microsoft or Google, while the practical values of open source projects of Chinese technology companies remain unclear. To address this issue, we conduct a mixed-method study to investigate the landscape of projects open sourced by three large Chinese technology companies, namely Baidu, Alibaba, and Tencent (BAT). We study the categories and characteristics of open source projects, the developer's …
A Differential Testing Approach For Evaluating Abstract Syntax Tree Mapping Algorithms, Yuanrui Fan, Xin Xia, David Lo, Ahmed E. Hassan, Yuan Wang, Shanping Li
A Differential Testing Approach For Evaluating Abstract Syntax Tree Mapping Algorithms, Yuanrui Fan, Xin Xia, David Lo, Ahmed E. Hassan, Yuan Wang, Shanping Li
Research Collection School Of Computing and Information Systems
Abstract syntax tree (AST) mapping algorithms are widely used to analyze changes in source code. Despite the foundational role of AST mapping algorithms, little effort has been made to evaluate the accuracy of AST mapping algorithms, i.e., the extent to which an algorithm captures the evolution of code. We observe that a program element often has only one best-mapped program element. Based on this observation, we propose a hierarchical approach to automatically compare the similarity of mapped statements and tokens by different algorithms. By performing the comparison, we determine if eachof the compared algorithms generates inaccurate mappings for a statement …
Unveiling The Mystery Of Api Evolution In Deep Learning Frameworks: A Case Study Of Tensorflow 2, Zejun Zhang, Yanming Yang, Xin Xia, David Lo, Xiaoxue Ren, John C. Grundy
Unveiling The Mystery Of Api Evolution In Deep Learning Frameworks: A Case Study Of Tensorflow 2, Zejun Zhang, Yanming Yang, Xin Xia, David Lo, Xiaoxue Ren, John C. Grundy
Research Collection School Of Computing and Information Systems
API developers have been working hard to evolve APIs to provide more simple, powerful, and robust API libraries. Although API evolution has been studied for multiple domains, such as Web and Android development, API evolution for deep learning frameworks has not yet been studied. It is not very clear how and why APIs evolve in deep learning frameworks, and yet these are being more and more heavily used in industry. To fill this gap, we conduct a large-scale and in-depth study on the API evolution of Tensorflow 2, which is currently the most popular deep learning framework. We first extract …
Action Selection For Composable Modular Deep Reinforcement Learning, Vaibhav Gupta, Daksh Anand, Praveen Paruchuri, Akshat Kumar
Action Selection For Composable Modular Deep Reinforcement Learning, Vaibhav Gupta, Daksh Anand, Praveen Paruchuri, Akshat Kumar
Research Collection School Of Computing and Information Systems
In modular reinforcement learning (MRL), a complex decision making problem is decomposed into multiple simpler subproblems each solved by a separate module. Often, these subproblems have conflicting goals, and incomparable reward scales. A composable decision making architecture requires that even the modules authored separately with possibly misaligned reward scales can be combined coherently. An arbitrator should consider different module's action preferences to learn effective global action selection. We present a novel framework called GRACIAS that assigns fine-grained importance to the different modules based on their relevance in a given state, and enables composable decision making based on modern deep RL …
Guest Editorial: Non-Iid Outlier Detection In Complex Contexts, Guansong Pang, Fabrizio Angiulli, Mihai Cucuringu, Huan Liu
Guest Editorial: Non-Iid Outlier Detection In Complex Contexts, Guansong Pang, Fabrizio Angiulli, Mihai Cucuringu, Huan Liu
Research Collection School Of Computing and Information Systems
Outlier detection, also known as anomaly detection, aims at identifying data instances that are rare or significantly different from the majority of instances. Due to its significance in many critical domains like cybersecurity, fintech, healthcare, public security, and AI safety, outlier detection has been one of the most active research areas in various communities, such as machine learning, data mining, computer vision, and statistics. Traditional outlier-detection techniques generally assume that data are independent and identically distributed (IID), which are significantly challenged in complex contexts where data are actually non-IID. These contexts are ubiquitous in not only graph data, sequence data, …
How Do Software Developers Use Github Actions To Automate Their Workflows?, Timothy Kinsman, Mairieli Wessel, Marco Gerosa, Christoph Treude
How Do Software Developers Use Github Actions To Automate Their Workflows?, Timothy Kinsman, Mairieli Wessel, Marco Gerosa, Christoph Treude
Research Collection School Of Computing and Information Systems
Automated tools are frequently used in social coding repositories to perform repetitive activities that are part of the distributed software development process. Recently, GitHub introduced GitHub Actions, a feature providing automated work-flows for repository maintainers. Although several Actions have been built and used by practitioners, relatively little has been done to evaluate them. Understanding and anticipating the effects of adopting such kind of technology is important for planning and management. Our research is the first to investigate how developers use Actions and how several activity indicators change after their adoption. Our results indicate that, although only a small subset of …
Data-Driven Recommendation Of Academic Options Based On Personality Traits, Aashish Ghimire
Data-Driven Recommendation Of Academic Options Based On Personality Traits, Aashish Ghimire
All Graduate Theses and Dissertations, Spring 1920 to Summer 2023
The choice of academic major and, subsequently, an academic institution has a massive effect on a person’s career. It not only determines their career path but their earning potential, professional happiness, etc. [1] About 40% of people who are admitted to a college do not graduate within six years. Yet, very limited resources are available for students to help make those decisions, and each guidance counselor is responsible for roughly 400 to 900 students across the United States. A tool to help these decisions would benefit students, parents, and guidance counselors.
Various research studies have shown that personality traits affect …
Artificial Neural Network Based Approach For Malware Detection, Matthew Fletcher
Artificial Neural Network Based Approach For Malware Detection, Matthew Fletcher
Honors Capstone Projects and Theses
No abstract provided.
A Deep Analysis And Algorithmic Approach To Solving Complex Fitness Issues In Collegiate Student Athletes, Holly N. Puckett
A Deep Analysis And Algorithmic Approach To Solving Complex Fitness Issues In Collegiate Student Athletes, Holly N. Puckett
Honors College Theses
Sports are not simply an entertainment source. For many, it creates a sense of community, support, and trust among both fans and athletes alike. In order to continue the sense of community sports provides, athletes must be properly cared for in order to perform at the highest level possible. Thus, their fitness and health must be monitored continuously. In a professional sense, one can expect individualized attention to athletes daily due to an abundance of funding and resources. However, when looking at college communities and student athletes within them, the number of athletes per athletic trainer increases due to both …
Mapping Renewal: How An Unexpected Interdisciplinary Collaboration Transformed A Digital Humanities Project, Elise Tanner, Geoffrey Joseph
Mapping Renewal: How An Unexpected Interdisciplinary Collaboration Transformed A Digital Humanities Project, Elise Tanner, Geoffrey Joseph
Digital Initiatives Symposium
Funded by a National Endowment for Humanities (NEH) Humanities Collections and Reference Resources Foundations Grant, the UA Little Rock Center for Arkansas History and Culture’s “Mapping Renewal” pilot project focused on creating access to and providing spatial context to archival materials related to racial segregation and urban renewal in the city of Little Rock, Arkansas, from 1954-1989. An unplanned interdisciplinary collaboration with the UA Little Rock Arkansas Economic Development Institute (AEDI) has proven to be an invaluable partnership. One team member from each department will demonstrate the Mapping Renewal website and discuss how the collaborative process has changed and shaped …
Building A Data Washing Machine For Unsupervised Entity Resolution Of Unstandardized References Sources, Awaad K. Al Sarkhi
Building A Data Washing Machine For Unsupervised Entity Resolution Of Unstandardized References Sources, Awaad K. Al Sarkhi
Theses and Dissertations
This dissertation describes a first attempt to build a data washing machine, a system able to take dirty data and through an unsupervised process, output clean data. The washing machine design described here focuses on two main aspects of the data curation process, token correction and data redundancy. It aims to simplify and automate the preparation of data used to create information products. In this approach, all these steps would be automated, thus saving the time and effort of the data analysts who ordinarily perform these actions. In other words, this is the opposite of the current approach to first …
Exploring Ai And Multiplayer In Java, Ronni Kurtzhals
Exploring Ai And Multiplayer In Java, Ronni Kurtzhals
Student Academic Conference
I conducted research into three topics: artificial intelligence, package deployment, and multiplayer servers in Java. This research came together to form my project presentation on the implementation of these topics, which I felt accurately demonstrated the various things I have learned from my courses at Moorhead State University. Several resources were consulted throughout the project, including the work of W3Schools and StackOverflow as well as relevant assignments and textbooks from previous classes. I found this project relevant to computer science and information systems for several reasons, such as the AI component and use of SQL data tables; but it was …
Non-Hazardous Industrial Solid Waste Tracking System, Justin Tank
Non-Hazardous Industrial Solid Waste Tracking System, Justin Tank
Masters Theses & Doctoral Dissertations
The Olmsted Non-Hazardous Industrial Solid Waste Tracking System allows waste generators of certain materials to electronically have their waste assessments evaluated, approved, and tracked through a simple online process. The current process of manually requesting evaluations, prepopulating tracking forms, and filling them out on triplicate carbonless forms is out of sync with other processes in the department. Complying with audit requirements requires pulling physical copies and providing them physically to fulfill information requests.
Waste generators in Minnesota are required to track their waste disposals for certain types of industrial waste streams. This ensures waste is accounted for at the point …
Sql Injection & Web Application Security: A Python-Based Network Traffic Detection Model, Nyki Anderson
Sql Injection & Web Application Security: A Python-Based Network Traffic Detection Model, Nyki Anderson
Cybersecurity Undergraduate Research Showcase
The Internet of Things (IoT) presents a great many challenges in cybersecurity as the world grows more and more digitally dependent. Personally identifiable information (PII) (i,e., names, addresses, emails, credit card numbers) is stored in databases across websites the world over. The greatest threat to privacy, according to the Open Worldwide Application Security Project (OWASP) is SQL injection attacks (SQLIA) [1]. In these sorts of attacks, hackers use malicious statements entered into forms, search bars, and other browser input mediums to trick the web application server into divulging database assets. A proposed technique against such exploitation is convolution neural network …
Implementing A Registry Federation For Materials Science Data Discovery, Raymond L. Plante, Chandler A. Becker, Andrea Medina-Smith, Kevin Brady, Alden Dima, Benjamin Long, Laura M. Bartolo, James A. Warren, Robert J. Hanisch
Implementing A Registry Federation For Materials Science Data Discovery, Raymond L. Plante, Chandler A. Becker, Andrea Medina-Smith, Kevin Brady, Alden Dima, Benjamin Long, Laura M. Bartolo, James A. Warren, Robert J. Hanisch
Copyright, Fair Use, Scholarly Communication, etc.
As a result of a number of national initiatives, we are seeing rapid growth in the data important to materials science that are available over the web. Consequently, it is becoming increasingly difficult for researchers to learn what data are available and how to access them. To address this problem, the Research Data Alliance (RDA) Working Group for International Materials Science Registries (IMRR) was established to bring together materials science and information technology experts to develop an international federation of registries that can be used for global discovery of data resources for materials science. A resource registry collects high-level metadata …
Predicting The Outcome Of Nba Games, Matthew Houde
Predicting The Outcome Of Nba Games, Matthew Houde
Honors Projects in Data Science
The aim of the project is to create a machine learning model to predict NBA games. The purpose is to build upon and improve existing models. Research into other predictive sports models and machine learning techniques was conducted to understand what is currently being done to predict NBA games and how effective it is in doing so. After a thorough literary review, the model was created using Python and a variety of machine learning techniques. The dataset used had an array of team statistics for both the home and away team for each corresponding matchup and two supporting features were …
Buffer Overflow And Sql Injection In C++, Noah Warren Kapley
Buffer Overflow And Sql Injection In C++, Noah Warren Kapley
Masters Theses & Specialist Projects
Buffer overflows and SQL Injection have plagued programmers for many years. A successful buffer overflow, innocuous or not, damages a computer’s permanent memory. Safer buffer overflow programs are presented in this thesis for the C programs characterizing string concatenation, string copy, and format get string, a C program which takes input and output from a keyboard, in most cases. Safer string concatenation and string copy programs presented in this thesis require the programmer to specify the amount of storage space necessary for the program’s execution. This safety mechanism is designed to help programmers avoid over specifying the amount of storage …
The Role Of Privacy Within The Realm Of Healthcare Wearables' Acceptance And Use, Thomas Jernejcic
The Role Of Privacy Within The Realm Of Healthcare Wearables' Acceptance And Use, Thomas Jernejcic
Masters Theses & Doctoral Dissertations
The flexibility and vitality of the Internet along with technological innovation have fueled an industry focused on the design of portable devices capable of supporting personal activities and wellbeing. These compute devices, known as wearables, are unique from other computers in that they are portable, specific in function, and worn or carried by the user. While there are definite benefits attributable to wearables, there are also notable risks, especially in the realm of security where personal information and/or activities are often accessible to third parties. In addition, protecting one’s private information is regularly an afterthought and thus lacking in maturity. …
Newslink: Empowering Intuitive News Search With Knowledge Graphs, Yueji Yang, Yuchen Li, Anthony Tung
Newslink: Empowering Intuitive News Search With Knowledge Graphs, Yueji Yang, Yuchen Li, Anthony Tung
Research Collection School Of Computing and Information Systems
News search tools help end users to identify relevant news stories. However, existing search approaches often carry out in a "black-box" process. There is little intuition that helps users understand how the results are related to the query. In this paper, we propose a novel news search framework, called NEWSLINK, to empower intuitive news search by using relationship paths discovered from open Knowledge Graphs (KGs). Specifically, NEWSLINK embeds both a query and news documents to subgraphs, called subgraph embeddings, in the KG. Their embeddings' overlap induces relationship paths between the involving entities. Two major advantages are obtained by incorporating subgraph …
Tour: Dynamic Topic And Sentiment Analysis Of User Reviews For Assisting App Release, Tianyi Yang, Cuiyun Gao, Jingya Zang, David Lo, Michael R. Lyu
Tour: Dynamic Topic And Sentiment Analysis Of User Reviews For Assisting App Release, Tianyi Yang, Cuiyun Gao, Jingya Zang, David Lo, Michael R. Lyu
Research Collection School Of Computing and Information Systems
App reviews deliver user opinions and emerging issues (e.g., new bugs) about the app releases. Due to the dynamic nature of app reviews, topics and sentiment of the reviews would change along with app release versions. Although several studies have focused on summarizing user opinions by analyzing user sentiment towards app features, no practical tool is released. The large quantity of reviews and noise words also necessitates an automated tool for monitoring user reviews. In this paper, we introduce TOUR for dynamic TOpic and sentiment analysis of User Reviews. TOUR is able to (i) detect and summarize emerging app issues …
Homophily Outlier Detection In Non-Iid Categorical Data, Guansong Pang, Longbing Cao, Ling Chen
Homophily Outlier Detection In Non-Iid Categorical Data, Guansong Pang, Longbing Cao, Ling Chen
Research Collection School Of Computing and Information Systems
Most of existing outlier detection methods assume that the outlier factors (i.e., outlierness scoring measures) of data entities (e.g., feature values and data objects) are Independent and Identically Distributed (IID). This assumption does not hold in real-world applications where the outlierness of different entities is dependent on each other and/or taken from different probability distributions (non-IID). This may lead to the failure of detecting important outliers that are too subtle to be identified without considering the non-IID nature. The issue is even intensified in more challenging contexts, e.g., high-dimensional data with many noisy features. This work introduces a novel outlier …
Spectral Tensor Train Parameterization Of Deep Learning Layers, A. Obukhov, M. Rakhuba, A. Liniger, Zhiwu Huang, S. Georgoulis, D. Dai, Van Gool L.
Spectral Tensor Train Parameterization Of Deep Learning Layers, A. Obukhov, M. Rakhuba, A. Liniger, Zhiwu Huang, S. Georgoulis, D. Dai, Van Gool L.
Research Collection School Of Computing and Information Systems
We study low-rank parameterizations of weight matrices with embedded spectral properties in the Deep Learning context. The low-rank property leads to parameter efficiency and permits taking computational shortcuts when computing mappings. Spectral properties are often subject to constraints in optimization problems, leading to better models and stability of optimization. We start by looking at the compact SVD parameterization of weight matrices and identifying redundancy sources in the parameterization. We further apply the Tensor Train (TT) decomposition to the compact SVD components, and propose a non-redundant differentiable parameterization of fixed TT-rank tensor manifolds, termed the Spectral Tensor Train Parameterization (STTP). We …
Mixed Dish Recognition With Contextual Relation And Domain Alignment, Lixi Deng, Jingjing Chen, Chong-Wah Ngo, Qianru Sun, Sheng Tang, Yongdong Zhang, Tat-Seng Chua
Mixed Dish Recognition With Contextual Relation And Domain Alignment, Lixi Deng, Jingjing Chen, Chong-Wah Ngo, Qianru Sun, Sheng Tang, Yongdong Zhang, Tat-Seng Chua
Research Collection School Of Computing and Information Systems
Mixed dish is a food category that contains different dishes mixed in one plate, and is popular in Eastern and Southeast Asia. Recognizing the individual dishes in a mixed dish image is important for health related applications, e.g. to calculate the nutrition values of the dish. However, most existing methods that focus on single dish classification are not applicable to the recognition of mixed dish images. The main challenge of mixed dish recognition comes from three aspects: a wide range of dish types, the complex dish combination with severe overlap between different dishes and the large visual variances of same …
Boundary Precedence Image Inpainting Method Based On Self-Organizing Maps, Haibo Pen, Quan Wang, Zhaoxia Wang
Boundary Precedence Image Inpainting Method Based On Self-Organizing Maps, Haibo Pen, Quan Wang, Zhaoxia Wang
Research Collection School Of Computing and Information Systems
In addition to text data analysis, image analysis is an area that has increasingly gained importance in recent years because more and more image data have spread throughout the internet and real life. As an important segment of image analysis techniques, image restoration has been attracting a lot of researchers’ attention. As one of AI methodologies, Self-organizing Maps (SOMs) have been applied to a great number of useful applications. However, it has rarely been applied to the domain of image restoration. In this paper, we propose a novel image restoration method by leveraging the capability of SOMs, and we name …
Do Users Care About Ad's Performance Costs? Exploring The Effects Of The Performance Costs Of In-App Ads On User Experience, Cuiyun Gao, Jichuan Zeng, Federica Sarro, David Lo, Irwin King, Michael R. Lyu
Do Users Care About Ad's Performance Costs? Exploring The Effects Of The Performance Costs Of In-App Ads On User Experience, Cuiyun Gao, Jichuan Zeng, Federica Sarro, David Lo, Irwin King, Michael R. Lyu
Research Collection School Of Computing and Information Systems
Context: In-app advertising is the primary source of revenue for many mobile apps. The cost of advertising (ad cost) is non-negligible for app developers to ensure a good user experience and continuous profits. Previous studies mainly focus on addressing the hidden performance costs generated by ads, including consumption of memory, CPU, data traffic, and battery. However, there is no research on analyzing users’ perceptions of ads’ performance costs to our knowledge.Objective: To fill this gap and better understand the effects of performance costs of in-app ads on user experience, we conduct a study on analyzing user concerns about ads’ performance …
Efficient Retrieval Of Matrix Factorization-Based Top-K Recommendations: A Survey Of Recent Approaches, Dung D. Le, Hady W. Lauw
Efficient Retrieval Of Matrix Factorization-Based Top-K Recommendations: A Survey Of Recent Approaches, Dung D. Le, Hady W. Lauw
Research Collection School Of Computing and Information Systems
Top-k recommendation seeks to deliver a personalized list of k items to each individual user. An established methodology in the literature based on matrix factorization (MF), which usually represents users and items as vectors in low-dimensional space, is an effective approach to recommender systems, thanks to its superior performance in terms of recommendation quality and scalability. A typical matrix factorization recommender system has two main phases: preference elicitation and recommendation retrieval. The former analyzes user-generated data to learn user preferences and item characteristics in the form of latent feature vectors, whereas the latter ranks the candidate items based on the …
Dbl: Efficient Reachability Queries On Dynamic Graphs, Qiuyi Lyu, Yuchen Li, Bingsheng He, Bin Gong
Dbl: Efficient Reachability Queries On Dynamic Graphs, Qiuyi Lyu, Yuchen Li, Bingsheng He, Bin Gong
Research Collection School Of Computing and Information Systems
Reachability query is a fundamental problem on graphs, which has been extensively studied in academia and industry. Since graphs are subject to frequent updates in many applications, it is essential to support efficient graph updates while offering good performance in reachability queries. Existing solutions compress the original graph with the Directed Acyclic Graph (DAG) and propose efficient query processing and index update techniques. However, they focus on optimizing the scenarios where the Strong Connected Components (SCCs) remain unchanged and have overlooked the prohibitively high cost of the DAG maintenance when SCCs are updated. In this paper, we propose DBL, an …
Towards Efficient Motif-Based Graph Partitioning: An Adaptive Sampling Approach, Shixun Huang, Yuchen Li, Zhifeng Bao, Zhao Li
Towards Efficient Motif-Based Graph Partitioning: An Adaptive Sampling Approach, Shixun Huang, Yuchen Li, Zhifeng Bao, Zhao Li
Research Collection School Of Computing and Information Systems
In this paper, we study the problem of efficient motif-based graph partitioning (MGP). We observe that existing methods require to enumerate all motif instances to compute the exact edge weights for partitioning. However, the enumeration is prohibitively expensive against large graphs. We thus propose a sampling-based MGP (SMGP) framework that employs an unbiased sampling mechanism to efficiently estimate the edge weights while trying to preserve the partitioning quality. To further improve the effectiveness, we propose a novel adaptive sampling framework called SMGP+. SMGP+ iteratively partitions the input graph based on up-to-date estimated edge weights, and adaptively adjusts the sampling distribution …
Dismastd: An Efficient Distributed Multi-Aspect Streaming Tensor Decomposition, Keyu Yang, Yunjun Gao, Yifeng Shen, Baihua Zheng, Lu Chen
Dismastd: An Efficient Distributed Multi-Aspect Streaming Tensor Decomposition, Keyu Yang, Yunjun Gao, Yifeng Shen, Baihua Zheng, Lu Chen
Research Collection School Of Computing and Information Systems
Tensor decomposition is a fundamental multidimensional data analysis tool for many data-driven applications, such as social computing, computer vision, and bioinformatics, to name but a few. However, the rapidly increasing streaming data nowadays introduces new challenges to traditional static tensor decomposition. It requires an efficient distributed dynamic tensor decomposition without re-computing the whole tensor from scratch. In this paper, we propose DisMASTD, an efficient distributed multi-aspect streaming tensor decomposition. First, we prove the optimal tensor partitioning problem is NP-hard. Second, we present two heuristic tensor partitioning approaches to ensure the load balancing. Third, we develop a distributed multi-aspect streaming tensor …