Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- Singapore Management University (3560)
- Wright State University (631)
- Walden University (447)
- New Jersey Institute of Technology (143)
- University of Malaya (131)
-
- University of Nebraska at Omaha (119)
- Old Dominion University (109)
- California State University, San Bernardino (100)
- San Jose State University (89)
- University of Dayton (82)
- MMU Press (74)
- City University of New York (CUNY) (70)
- University of Dar es Salaam (65)
- Air Force Institute of Technology (61)
- University of Nebraska - Lincoln (60)
- University of South Florida (56)
- Kennesaw State University (54)
- Nova Southeastern University (52)
- Technological University Dublin (51)
- University of Arkansas, Fayetteville (46)
- Dakota State University (43)
- Claremont Colleges (42)
- California Polytechnic State University, San Luis Obispo (41)
- Institute of Business Administration (38)
- Western Kentucky University (36)
- Purdue University (35)
- Ateneo de Manila University (34)
- Governors State University (34)
- Portland State University (34)
- University of Arkansas Little Rock (33)
- Keyword
-
- Machine learning (123)
- Information technology (91)
- Data mining (90)
- Social media (84)
- Machine Learning (71)
-
- Cybersecurity (63)
- Deep learning (61)
- Twitter (61)
- Artificial intelligence (58)
- Semantic Web (53)
- Online learning (52)
- Databases (46)
- Cloud computing (45)
- Deep Learning (45)
- Information Technology (45)
- Information retrieval (45)
- Classification (44)
- Blockchain (42)
- Database (42)
- Natural language processing (41)
- Ontology (41)
- Big data (40)
- Technology (40)
- Security (39)
- Computer science (38)
- Privacy (38)
- Algorithms (37)
- Clustering (37)
- Information systems (37)
- Management (37)
- Publication Year
- Publication
-
- Research Collection School Of Computing and Information Systems (3441)
- Kno.e.sis Publications (540)
- Walden Dissertations and Doctoral Studies (447)
- Theses and Dissertations (129)
- Student Works (2000-2009) (120)
-
- Dissertations (114)
- Computer Science Faculty Publications (95)
- Computer Science and Engineering Faculty Publications (91)
- Theses Digitization Project (86)
- Journal of Informatics and Web Engineering (74)
- Master's Projects (68)
- Information Systems and Quantitative Analysis Faculty Proceedings & Presentations (64)
- Tanzania Journal of Engineering and Technology (TJET) (62)
- Dissertations and Theses Collection (Open Access) (58)
- USF Tampa Graduate Theses and Dissertations (51)
- Theses (48)
- CCAC Theses and Dissertations (43)
- Information Systems and Quantitative Analysis Faculty Publications (41)
- CGU Faculty Publications and Research (37)
- International Conference on Information and Communication Technologies (36)
- Open Educational Resources (35)
- Graduate Theses and Dissertations (34)
- Department of Information Systems & Computer Science Faculty Publications (33)
- All Capstone Projects (32)
- Masters Theses & Doctoral Dissertations (32)
- Conference papers (28)
- All Maxine Goodman Levin School of Urban Affairs Publications (27)
- UBT International Conference (23)
- Electronic Theses and Dissertations (22)
- Faculty Articles (22)
- Publication Type
- File Type
Articles 3211 - 3240 of 7334
Full-Text Articles in Computer Sciences
A System For Stratified Sampling Of Entity Resolution Results To Assess And Improve Accuracy With Minimal Clerical Review Effort, Daniel L. Pullen
A System For Stratified Sampling Of Entity Resolution Results To Assess And Improve Accuracy With Minimal Clerical Review Effort, Daniel L. Pullen
Theses and Dissertations
Organizations across many industries from banking to medicine depend on master data that describe their customers, products and services. Data integration and unique representations of master data are supported by Entity Resolution processes to link records for the same master entity. Master data is key to the core operations of these organizations. Unfortunately, many of these organizations do not use a systematic method, or in some cases, use no method for measuring and improving the quality of record linkage provided by Entity Resolution. This research presents an implementation and proposed methodology for routinely and systematically measuring the quality of Entity …
Micro-Review Synthesis For Multi-Entity Summarization, Thanh-Son Nguyen, Hady W. Lauw, Panayiotis Tsaparas
Micro-Review Synthesis For Multi-Entity Summarization, Thanh-Son Nguyen, Hady W. Lauw, Panayiotis Tsaparas
Research Collection School Of Computing and Information Systems
Location-based social networks (LBSNs), exemplified by Foursquare, are fast gaining popularity. One important feature of LBSNs is micro-review. Upon check-in at a particular venue, a user may leave a short review (up to 200 characters long), also known as a tip. These tips are an important source of information for others to know more about various aspects of an entity (e.g., restaurant), such as food, waiting time, or service. However, a user is often interested not in one particular entity, but rather in several entities collectively, for instance within a neighborhood or a category. In this paper, we address the …
Vcksm: Verifiable Conjunctive Keyword Search Over Mobile E-Health Cloud In Shared Multi-Owner Settings, Yinbin Miao, Jianfeng Ma, Ximeng Liu, Qi Jiang, Junwei Zhang, Limin Shen, Zhiquan Liu
Vcksm: Verifiable Conjunctive Keyword Search Over Mobile E-Health Cloud In Shared Multi-Owner Settings, Yinbin Miao, Jianfeng Ma, Ximeng Liu, Qi Jiang, Junwei Zhang, Limin Shen, Zhiquan Liu
Research Collection School Of Computing and Information Systems
Searchable encryption (SE) is a promising technique which enables cloud users to conduct search over encrypted cloud data in a privacy-preserving way, especially for the electronic health record (EHR) system that contains plenty of medical history, diagnosis, radiology images, etc. In this paper, we focus on a more practical scenario, also named as the shared multi-owner settings, where each e-health record is co-owned by a fixed number of parties. Although the existing SE schemes under the unshared multi-owner settings can be adapted to this shared scenario, these schemes have to build multiple indexes,which definitely incur higher computational overhead. To save …
Personalized Microtopic Recommendation On Microblogs, Yang Li, Jing Jiang, Ting Liu, Minghui Qiu, Xiaofei Sun
Personalized Microtopic Recommendation On Microblogs, Yang Li, Jing Jiang, Ting Liu, Minghui Qiu, Xiaofei Sun
Research Collection School Of Computing and Information Systems
Microblogging services such as Sina Weibo and Twitter allow users to create tags explicitly indicated by the # symbol. In Sina Weibo, these tags are called microtopics, and in Twitter, they are called hashtags. In Sina Weibo, each microtopic has a designate page and can be directly visited or commented on. Recommending these microtopics to users based on their interests can help users efficiently acquire information. However, it is non-trivial to recommend microtopics to users to satisfy their information needs. In this article, we investigate the task of personalized microtopic recommendation, which exhibits two challenges. First, users usually do not …
Clsters: A General System For Reducing Errors Of Trajectories Under Challenging Localization Situations, Hao Wu, Weiwei Sun, Baihua Zheng, Li Yang, Wei Zhou
Clsters: A General System For Reducing Errors Of Trajectories Under Challenging Localization Situations, Hao Wu, Weiwei Sun, Baihua Zheng, Li Yang, Wei Zhou
Research Collection School Of Computing and Information Systems
Trajectory data generated by outdoor activities have great potential for location based services. However, depending on the localization technique used, certain trajectory data could contain large errors. For example, the error of trajectories generated by cellular-based localization techniques is around 100m which is ten times larger than that of GPS-based trajectories. Hence, enhancing the utility of those large-error trajectories becomes a challenge. In this paper we show how to improve the quality of trajectory data having large errors. Some existing works reduce the error through hardware which requires information such as the time of arrival (TOA), received signal strength indication …
Vurle: Automatic Vulnerability Detection And Repair By Learning From Examples, Ma Siqi, Ferdian Thung, David Lo, Cong Sun, Robert H. Deng
Vurle: Automatic Vulnerability Detection And Repair By Learning From Examples, Ma Siqi, Ferdian Thung, David Lo, Cong Sun, Robert H. Deng
Research Collection School Of Computing and Information Systems
Vulnerability becomes a major threat to the security of many systems. Attackers can steal private information and perform harmful actions by exploiting unpatched vulnerabilities. Vulnerabilities often remain undetected for a long time as they may not affect typical systems’ functionalities. Furthermore, it is often difficult for a developer to fix a vulnerability correctly if he/she is not a security expert. To assist developers to deal with multiple types of vulnerabilities, we propose a new tool, called VuRLE, for automatic detection and repair of vulnerabilities. VuRLE (1) learns transformative edits and their contexts (i.e., code characterizing edit locations) from examples of …
On-Demand Developer Documentation, Martin P. Robillard, Andrian Marcus, Christoph Treude, Gabriele Bavota, Oscar Chaparro, Neil Ernst, Marco Aurélio Gerosa, Michael Godfrey, Michele Lanza, Mario Linares-Vasquez, Gail C. Murphy, Laura Moreno, David Shepherd, Edmund Wong
On-Demand Developer Documentation, Martin P. Robillard, Andrian Marcus, Christoph Treude, Gabriele Bavota, Oscar Chaparro, Neil Ernst, Marco Aurélio Gerosa, Michael Godfrey, Michele Lanza, Mario Linares-Vasquez, Gail C. Murphy, Laura Moreno, David Shepherd, Edmund Wong
Research Collection School Of Computing and Information Systems
We advocate for a paradigm shift in supporting the information needs of developers, centered around the concept of automated on-demand developer documentation. Currently, developer information needs are fulfilled by asking experts or consulting documentation. Unfortunately, traditional documentation practices are inefficient because of, among others, the manual nature of its creation and the gap between the creators and consumers. We discuss the major challenges we face in realizing such a paradigm shift, highlight existing research that can be leveraged to this end, and promote opportunities for increased convergence in research on software documentation.
Combining Machine-Based And Econometrics Methods For Policy Analytics Insights, Robert J. Kauffman, Kwansoo Kim, Sang-Yong Tom Lee, Ai Phuong Hoang, Jing Ren
Combining Machine-Based And Econometrics Methods For Policy Analytics Insights, Robert J. Kauffman, Kwansoo Kim, Sang-Yong Tom Lee, Ai Phuong Hoang, Jing Ren
Research Collection School Of Computing and Information Systems
Computational Social Science (CSS) has become a mainstream approach in the empirical study of policy analytics issues in various domains of e-commerce research. This article is intended to represent recent advances that have been made for the discovery of new policy-related insights in business, consumer, and social settings. The approach discussed is fusion analytics, which combines machine-based methods from Computer Science (CS) and explanatory empiricism involving advanced Econometrics and Statistics. It explores several efforts to conduct research inquiry in different functional areas of Electronic Commerce and Information Systems (IS), with applications that represent different functional areas of business, as well …
Understanding Stack Overflow Code Fragments, Christoph Treude, Martin P. Robillard
Understanding Stack Overflow Code Fragments, Christoph Treude, Martin P. Robillard
Research Collection School Of Computing and Information Systems
Code fragments posted in answers on Q&A forums can form an important source of developer knowledge. However, effective reuse of code fragments found online often requires information other than the code fragment alone. We report on the results of a survey-based study to investigate to what extent developers perceive Stack Overflow code fragments to be self-explanatory. As part of the study, we also investigated the types of information missing from fragments that were not self-explanatory. We find that less than half of the Stack Overflow code fragments in our sample are considered to be self-explanatory by the 321 participants who …
Machine Learning Based Protein Sequence To (Un)Structure Mapping And Interaction Prediction, Sumaiya Iqbal
Machine Learning Based Protein Sequence To (Un)Structure Mapping And Interaction Prediction, Sumaiya Iqbal
LSU New Orleans Theses and Dissertations
Proteins are the fundamental macromolecules within a cell that carry out most of the biological functions. The computational study of protein structure and its functions, using machine learning and data analytics, is elemental in advancing the life-science research due to the fast-growing biological data and the extensive complexities involved in their analyses towards discovering meaningful insights. Mapping of protein’s primary sequence is not only limited to its structure, we extend that to its disordered component known as Intrinsically Disordered Proteins or Regions in proteins (IDPs/IDRs), and hence the involved dynamics, which help us explain complex interaction within a cell that …
Ancr—An Adaptive Network Coding Routing Scheme For Wsns With Different-Success-Rate Links †, Xiang Ji, Anwen Wang, Chunyu Li, Chun Ma, Yao Peng, Dajin Wang, Qingyi Hua, Feng Chen, Dingyi Fang
Ancr—An Adaptive Network Coding Routing Scheme For Wsns With Different-Success-Rate Links †, Xiang Ji, Anwen Wang, Chunyu Li, Chun Ma, Yao Peng, Dajin Wang, Qingyi Hua, Feng Chen, Dingyi Fang
Department of Computer Science Faculty Scholarship and Creative Works
As the underlying infrastructure of the Internet of Things (IoT), wireless sensor networks (WSNs) have been widely used in many applications. Network coding is a technique in WSNs to combine multiple channels of data in one transmission, wherever possible, to save node’s energy as well as increase the network throughput. So far most works on network coding are based on two assumptions to determine coding opportunities: (1) All the links in the network have the same transmission success rate; (2) Each link is bidirectional, and has the same transmission success rate on both ways. However, these assumptions may not be …
Detect Rumors In Microblog Posts Using Propagation Structure Via Kernel Learning, Jing Ma, Wei Gao, Kam-Fai Wong
Detect Rumors In Microblog Posts Using Propagation Structure Via Kernel Learning, Jing Ma, Wei Gao, Kam-Fai Wong
Research Collection School Of Computing and Information Systems
How fake news goes viral via social media? How does its propagation pattern differ from real stories? In this paper, we attempt to address the problem of identifying rumors, i.e., fake information, out of microblog posts based on their propagation structure. We firstly model microblog posts diffusion with propagation trees, which provide valuable clues on how an original message is transmitted and developed over time. We then propose a kernel-based method called Propagation Tree Kernel, which captures high-order patterns differentiating different types of rumors by evaluating the similarities between their propagation tree structures. Experimental results on two real-world datasets demonstrate …
Predicting Locations Of Pollution Sources Using Convolutional Neural Networks, Yiheng Chi, Nickolas D. Winovich, Guang Lin
Predicting Locations Of Pollution Sources Using Convolutional Neural Networks, Yiheng Chi, Nickolas D. Winovich, Guang Lin
The Summer Undergraduate Research Fellowship (SURF) Symposium
Pollution is a severe problem today, and the main challenge in water and air pollution controls and eliminations is detecting and locating pollution sources. This research project aims to predict the locations of pollution sources given diffusion information of pollution in the form of array or image data. These predictions are done using machine learning. The relations between time, location, and pollution concentration are first formulated as pollution diffusion equations, which are partial differential equations (PDEs), and then deep convolutional neural networks are built and trained to solve these PDEs. The convolutional neural networks consist of convolutional layers, reLU layers …
From Retweet To Believability: Utilizing Trust To Identify Rumor Spreaders On Twitter, Bhavtosh Rath, Wei Gao, Jing Ma, Jaideep Srivastava
From Retweet To Believability: Utilizing Trust To Identify Rumor Spreaders On Twitter, Bhavtosh Rath, Wei Gao, Jing Ma, Jaideep Srivastava
Research Collection School Of Computing and Information Systems
Ubiquitous use of social media such as microblogging platforms brings about ample opportunities for the false information to diffuse online. It is very important not just to determine the veracity of information but also the authenticity of the users who spread the information, especially in time-critical situations like real-world emergencies, where urgent measures have to be taken for stopping the spread of fake information. In this work, we propose a novel machine learning based approach for automatic identification of the users spreading rumorous information by leveraging the concept of believability, i.e., the extent to which the propagated information is likely …
A Case Study For Ecampus Spatial: Business Data Exploration, James Carswell, Thanh Thao Pham Ti, Andrea Ballatore, Junjun Yin, Linh Truong-Hong
A Case Study For Ecampus Spatial: Business Data Exploration, James Carswell, Thanh Thao Pham Ti, Andrea Ballatore, Junjun Yin, Linh Truong-Hong
Books/Book chapters
Location based querying is the core interaction paradigm between mobile citizens and the Internet of Things, so providing users with intelligent web-services that interact efficiently with web and wireless devices to recommend personalised services is a key goal. With today's popular Web Map Services, users can ask for general information at a specific location, but not detailed information such as related functionality or environments. This shortcoming comes from a lack of connection between non-spatial “business” data and spatial “map” data. This chapter presents a novel approach for location-based querying in web and wireless environments, in which non-spatial business data is …
Database Management System For Byuh Jonathan Napela Center, Olivia K. F. Moleni
Database Management System For Byuh Jonathan Napela Center, Olivia K. F. Moleni
Masters Theses & Doctoral Dissertations
The purpose of this project is to build a database management system (DBMS) for the Jonathan Napela Center department. The Napela Center is a department for students who are majoring or minoring in Hawaiian Studies and/or Pacific Island Studies. Currently the Napela Center uses Microsoft Excel as their DBMS to store and track both current and past student information. Unfortunately, this system hasn’t been working well for them due to unreliable information, limited user access and sometime can get too complex with too much data. So the director of the department decided to seek for another system.
This paper will …
Zero Textbook Cost Syllabus For Cis 3367 (Spreadsheet Applications In Business), Soniya Dsouza
Zero Textbook Cost Syllabus For Cis 3367 (Spreadsheet Applications In Business), Soniya Dsouza
Open Educational Resources
The primary focus of this course is to learn how to construct and use powerful spreadsheets for effective managerial decision-making. This course is mostly project- oriented with a dual focus on spreadsheet engineering and quantitative modeling of financial applications. Students will learn to develop powerful spreadsheet models and perform data analysis using Pivot Tables, VLookUp, Data Validation techniques and Sub Total functions. Students will also learn how to enhance spreadsheets by creating dashboards on financial data. The Visual Basic (macro) concepts will also be introduced to students. With the knowledge and hands-on experience of these concepts, students will be prepared …
Estimating Accuracy Of Personal Identifiable Information In Integrated Data Systems, Amani "Mohammad Jum'h" Amin Shatnawi
Estimating Accuracy Of Personal Identifiable Information In Integrated Data Systems, Amani "Mohammad Jum'h" Amin Shatnawi
All Graduate Theses and Dissertations, Spring 1920 to Summer 2023
Both government agencies and private companies rely on the collection of personal data on an ever-increasing scale. Out of necessity, person data include Personal Identifiable Information (PII), which is information that could potentially identify a specific individual. Many of these data would be integrated, so data analyst, policy makers or corporate officers can use it to make decisions or get a conclusion. Integrating data in a heterogeneous database environment create a need to estimate the accuracy of that data; without a valid assessment of accuracy there is a risk of coming with incorrect conclusions or making bad decision based on …
Ged: Moving Into The Electronic Age, Kateri Montileaux
Ged: Moving Into The Electronic Age, Kateri Montileaux
Masters Theses & Doctoral Dissertations
The purpose of this study is to find a direction as the Community Continuing Education/General Education Diploma (CCE/GED) department goes into the electronic age. Not only has the General Education Diploma test become computer based, the process of studying, preparing and communicating has also required one to use desktop computers, laptops, tablets, smart phones, email, and webinars daily. The goal is to promote the department and its services to the younger generation (18-25 years old) who are completely comfortable using electronic devices, and to the older generation (40+years) who may know a little bit of electronic communicating but who are …
Dynamic Adversarial Mining - Effectively Applying Machine Learning In Adversarial Non-Stationary Environments., Tegjyot Singh Sethi
Dynamic Adversarial Mining - Effectively Applying Machine Learning In Adversarial Non-Stationary Environments., Tegjyot Singh Sethi
Electronic Theses and Dissertations
While understanding of machine learning and data mining is still in its budding stages, the engineering applications of the same has found immense acceptance and success. Cybersecurity applications such as intrusion detection systems, spam filtering, and CAPTCHA authentication, have all begun adopting machine learning as a viable technique to deal with large scale adversarial activity. However, the naive usage of machine learning in an adversarial setting is prone to reverse engineering and evasion attacks, as most of these techniques were designed primarily for a static setting. The security domain is a dynamic landscape, with an ongoing never ending arms race …
Encoding And Recall Of Spatio-Temporal Episodic Memory In Real Time, Poo-Hee Chang, Ah-Hwee Tan
Encoding And Recall Of Spatio-Temporal Episodic Memory In Real Time, Poo-Hee Chang, Ah-Hwee Tan
Research Collection School Of Computing and Information Systems
Episodic memory enables a cognitive system to improve its performance by reflecting upon past events. In this paper, we propose a computational model called STEM for encoding and recall of episodic events together with the associated contextual information in real time. Based on a class of self-organizing neural networks, STEM is designed to learn memory chunks or cognitive nodes, each encoding a set of co-occurring multi-modal activity patterns across multiple pattern channels. We present algorithms for recall of events based on partial and inexact input patterns. Our empirical results based on a public domain data set show that STEM displays …
Don’T Bury Your Head In Warnings: A Game-Theoretic Approach For Intelligent Allocation Of Cyber-Security Alerts, Aaron Schlenker, Haifeng Xu, Mina Guirguis, Christopher Kiekintveld, Arunesh Sinha, Milind Tambe, Solomon Sonya, Darryl Balderas, Noah Dunstatter
Don’T Bury Your Head In Warnings: A Game-Theoretic Approach For Intelligent Allocation Of Cyber-Security Alerts, Aaron Schlenker, Haifeng Xu, Mina Guirguis, Christopher Kiekintveld, Arunesh Sinha, Milind Tambe, Solomon Sonya, Darryl Balderas, Noah Dunstatter
Research Collection School Of Computing and Information Systems
In recent years, there have been a number of successful cyber attacks on enterprise networks by malicious actors which have caused severe damage. These networks have Intrusion Detection and Prevention Systems in place to protect them, but they are notorious for producing a high volume of alerts. These alerts must be investigated by cyber analysts to determine whether they are an attack or benign. Unfortunately, there are magnitude more alerts generated than there are cyber analysts to investigate them. This trend is expected to continue into the future creating a need for tools which find optimal assignments of the incoming …
Bridge Text And Knowledge By Learning Multi-Prototype Entity Mention Embedding, Yixin Cao, Lifu Huang, Heng Ji, Xu Chen, Juanzi Li
Bridge Text And Knowledge By Learning Multi-Prototype Entity Mention Embedding, Yixin Cao, Lifu Huang, Heng Ji, Xu Chen, Juanzi Li
Research Collection School Of Computing and Information Systems
Integrating text and knowledge into a unified semantic space has attracted significant research interests recently. However, the ambiguity in the common space remains a challenge, namely that the same mention phrase usually refers to various entities. In this paper, to deal with the ambiguity of entity mentions, we propose a novel Multi-Prototype Mention Embedding model, which learns multiple sense embeddings for each mention by jointly modeling words from textual contexts and entities derived from a knowledge base. In addition, we further design an efficient language model based approach to disambiguate each mention to a specific sense. In experiments, both qualitative …
Representativeness-Aware Aspect Analysis For Brand Monitoring In Social Media, Lizi Liao, Xiangnan He, Zhaochun Ren, Liqiang Nie, Huan Xu, Ta-Seng Chua
Representativeness-Aware Aspect Analysis For Brand Monitoring In Social Media, Lizi Liao, Xiangnan He, Zhaochun Ren, Liqiang Nie, Huan Xu, Ta-Seng Chua
Research Collection School Of Computing and Information Systems
Owing to the fast-responding nature and extreme success of social media, many companies resort to social media sites for monitoring their brands’ reputation and the opinions of general public. To help companies monitor their brands, in this work, we delve into the task of extracting representative aspects and posts from users’ free-text posts in social media. Previous efforts have treated it as a traditional information extraction task, and forgo the specific properties of social media, such as the possible noise in user generated posts and the varying impacts; In contrast, we extract aspects by maximizing their representativeness, which is a …
Time-Aware Conversion Prediction, Wendi Ji, Xiaoling Wang, Feida Zhu
Time-Aware Conversion Prediction, Wendi Ji, Xiaoling Wang, Feida Zhu
Research Collection School Of Computing and Information Systems
The importance of product recommendation has been well recognized as a central task in business intelligence for e-commerce websites. Interestingly, what has been less aware of is the fact that different products take different time periods for conversion. The “conversion” here refers to actually a more general set of pre-defined actions, including for example purchases or registrations in recommendation and advertising systems. The mismatch between the product’s actual conversion period and the application’s target conversion period has been the subtle culprit compromising many existing recommendation algorithms.The challenging question: what products should be recommended for a given time period to maximize …
Indexing Metric Uncertain Data For Range Queries And Range Joins, Lu Chen, Yunjun Gao, Aoxiao Zhong, Christian S. Jensen, Gang Chen, Baihua Zheng
Indexing Metric Uncertain Data For Range Queries And Range Joins, Lu Chen, Yunjun Gao, Aoxiao Zhong, Christian S. Jensen, Gang Chen, Baihua Zheng
Research Collection School Of Computing and Information Systems
Range queries and range joins in metric spaces have applications in many areas, including GIS, computational biology, and data integration, where metric uncertain data exist in different forms, resulting from circumstances such as equipment limitations, high-throughput sequencing technologies, and privacy preservation. We represent metric uncertain data by using an object-level model and a bi-level model, respectively. Two novel indexes, the uncertain pivot B+-tree (UPB-tree) and the uncertain pivot B+-forest (UPB-forest), are proposed in order to support probabilistic range queries and range joins for a wide range of uncertain data types and similarity metrics. Both index structures use a small set …
On Efficiently Finding Reverse K-Nearest Neighbors Over Uncertain Graphs, Yunjun Gao, Xiaoye Miao, Gang Chen, Baihua Zheng, Deng Cai, Huiyong Cui
On Efficiently Finding Reverse K-Nearest Neighbors Over Uncertain Graphs, Yunjun Gao, Xiaoye Miao, Gang Chen, Baihua Zheng, Deng Cai, Huiyong Cui
Research Collection School Of Computing and Information Systems
Reverse k-nearest neighbor (RkNN) query on graphs returns the data objects that take a specified query object q as one of their k-nearest neighbors. It has significant influence in many real-life applications including resource allocation and profile-based marketing. However, to the best of our knowledge, there is little previous work on RkNN search over uncertain graph data, even though many complex networks such as traffic networks and protein–protein interaction networks are often modeled as uncertain graphs. In this paper, we systematically study the problem of reversek-nearest neighbor search on uncertain graphs (UG-RkNN search for short), where graph edges contain uncertainty. …
Pivot-Based Metric Indexing, Lu Chen, Yunjun Gao, Baihua Zheng, Christian S. Jensen, Hanyu Yang, Keyu Yang
Pivot-Based Metric Indexing, Lu Chen, Yunjun Gao, Baihua Zheng, Christian S. Jensen, Hanyu Yang, Keyu Yang
Research Collection School Of Computing and Information Systems
The general notion of a metric space encompasses a diverse range of data types and accompanying similarity measures. Hence, metric search plays an important role in a wide range of settings, including multimedia retrieval, data mining, and data integration. With the aim of accelerating metric search, a collection of pivot-based indexing techniques for metric data has been proposed, which reduces the number of potentially expensive similarity comparisons by exploiting the triangle inequality for pruning and validation. However, no comprehensive empirical study of those techniques exists. Existing studies each offers only a narrower coverage, and they use different pivot selection strategies …
Geometric Approaches For Top-K Queries [Tutorial], Kyriakos Mouratidis
Geometric Approaches For Top-K Queries [Tutorial], Kyriakos Mouratidis
Research Collection School Of Computing and Information Systems
Top-k processing is a well-studied problem with numerous applications that is becoming increasingly relevant with the growing availability of recommendation systems and decision-making software. The objective of this tutorial is twofold. First, we will delve into the geometric aspects of top-k processing. Second, we will cover complementary features to top-k queries, with strong practical relevance and important applications, that have a computational geometric nature. The tutorial will close with insights in the effect of dimensionality on the meaningfulness of top-k queries, and interesting similarities to nearest neighbor search.
Semantic Visualization For Short Texts With Word Embeddings, Van Minh Tuan Le, Hady W. Lauw
Semantic Visualization For Short Texts With Word Embeddings, Van Minh Tuan Le, Hady W. Lauw
Research Collection School Of Computing and Information Systems
Semantic visualization integrates topic modeling and visualization, such that every document is associated with a topic distribution as well as visualization coordinates on a low-dimensional Euclidean space. We address the problem of semantic visualization for short texts. Such documents are increasingly common, including tweets, search snippets, news headlines, or status updates. Due to their short lengths, it is difficult to model semantics as the word co-occurrences in such a corpus are very sparse. Our approach is to incorporate auxiliary information, such as word embeddings from a larger corpus, to supplement the lack of co-occurrences. This requires the development of a …