Open Access. Powered by Scholars. Published by Universities.®
Numerical Analysis and Scientific Computing Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- Singapore Management University (671)
- University of Dayton (31)
- Central Bank of Nigeria (20)
- University of Arkansas, Fayetteville (7)
- University of Nebraska - Lincoln (6)
-
- LSU New Orleans (4)
- California Polytechnic State University, San Luis Obispo (3)
- City University of New York (CUNY) (3)
- Montclair State University (3)
- Purdue University (3)
- San Jose State University (3)
- Technological University Dublin (3)
- Bryant University (2)
- Chapman University (2)
- Clemson University (2)
- East Tennessee State University (2)
- Embry-Riddle Aeronautical University (2)
- Georgia Southern University (2)
- Old Dominion University (2)
- Portland State University (2)
- Southern Methodist University (2)
- The University of Akron (2)
- University of Kentucky (2)
- University of Nevada, Las Vegas (2)
- California State University, San Bernardino (1)
- Claremont Colleges (1)
- Columbus State University (1)
- DePaul University (1)
- Eastern Washington University (1)
- Fort Hays State University (1)
- Keyword
-
- Data mining (25)
- Social media (20)
- Query processing (19)
- Classification (17)
- Online learning (16)
-
- Twitter (15)
- Machine learning (13)
- Algorithms (12)
- Machine Learning (12)
- Neural networks (11)
- Algorithm (10)
- Natural language processing (9)
- Sentiment analysis (9)
- Artificial intelligence (8)
- Deep Learning (8)
- Spatial databases (8)
- Data structures (7)
- Database (7)
- Location-based services (7)
- Recommender systems (7)
- Spatial database (7)
- Text mining (7)
- Data models (6)
- Deep learning (6)
- Feature extraction (6)
- Information retrieval (6)
- Online Learning (6)
- Road network (6)
- Semantics (6)
- Social network (6)
- Publication Year
- Publication
-
- Research Collection School Of Computing and Information Systems (661)
- Computer Science Faculty Publications (33)
- CBN Journal of Applied Statistics (JAS) (20)
- Dissertations and Theses Collection (Open Access) (6)
- Graduate Theses and Dissertations (5)
-
- LSU New Orleans Theses and Dissertations (4)
- Theses and Dissertations (4)
- Department of Computer Science Faculty Scholarship and Creative Works (3)
- All Dissertations (2)
- College of Graduate Studies: Theses & Dissertations (2)
- Department of Agricultural and Biological Systems Engineering: Dissertations, Theses, and Student Research (2)
- Electronic Theses and Dissertations (2)
- Honors Projects in Information Systems and Analytics (2)
- Research Collection Lee Kong Chian School Of Business (2)
- SMU Data Science Review (2)
- STAR Program Research Presentations (2)
- SWITCH (2)
- School of Computing: Dissertations, Theses, and Student Research (2)
- The Summer Undergraduate Research Fellowship (SURF) Symposium (2)
- UNLV Theses, Dissertations, Professional Papers, and Capstones (2)
- Williams Honors College, Honors Research Projects (2)
- Asian Management Insights (1)
- Bulletin of TUIT: Management and Communication Technologies (1)
- CMC Senior Theses (1)
- College of Computing and Digital Media Dissertations (1)
- Computer Science and Computer Engineering Faculty Publications and Presentations (1)
- Computer Science and Computer Engineering Undergraduate Honors Theses (1)
- Computer Science and Software Engineering (1)
- Conference papers (1)
- Department of Computer Science Publications (1)
- Publication Type
Articles 331 - 360 of 808
Full-Text Articles in Numerical Analysis and Scientific Computing
Efficient Reverse Top-K Boolean Spatial Keyword Queries On Road Networks, Yunjun Gao, Xu Qin, Baihua Zheng, Gang Chen
Efficient Reverse Top-K Boolean Spatial Keyword Queries On Road Networks, Yunjun Gao, Xu Qin, Baihua Zheng, Gang Chen
Research Collection School Of Computing and Information Systems
Reverse k nearest neighbor (RkNN) queries have a broad application base such as decision support, profile-based marketing, and resource allocation. Previous work on RkNN search does not take textual information into consideration or limits to the Euclidean space. In the real world, however, most spatial objects are associated with textual information and lie on road networks. In this paper, we introduce a new type of queries, namely, reverse top-k Boolean spatial keyword (RkBSK) retrieval, which assumes objects are on the road network and considers both spatial and textual information. Given a data set P on a road network and a …
Best Upgrade Plans For Single And Multiple Source-Destination Pairs, Yimin Lin, Kyriakos Mouratidis
Best Upgrade Plans For Single And Multiple Source-Destination Pairs, Yimin Lin, Kyriakos Mouratidis
Research Collection School Of Computing and Information Systems
In this paper, we study Resource Constrained Best Upgrade Plan (BUP) computation in road network databases. Consider a transportation network (weighted graph) G where a subset of the edges are upgradable, i.e., for each such edge there is a cost, which if spent, the weight of the edge can be reduced to a specific new value. In the single-pair version of BUP, the input includes a source and a destination in G, and a budget B (resource constraint). The goal is to identify which upgradable edges should be upgraded so that the shortest path distance between source and …
Review Selection Using Micro-Reviews, Thanh-Son Nguyen, Hady W. Lauw, Panayiotis Tsaparas
Review Selection Using Micro-Reviews, Thanh-Son Nguyen, Hady W. Lauw, Panayiotis Tsaparas
Research Collection School Of Computing and Information Systems
Given the proliferation of review content, and the fact that reviews are highly diverse and often unnecessarily verbose, users frequently face the problem of selecting the appropriate reviews to consume. Micro-reviews are emerging as a new type of online review content in the social media. Micro-reviews are posted by users of check-in services such as Foursquare. They are concise (up to 200 characters long) and highly focused, in contrast to the comprehensive and verbose reviews. In this paper, we propose a novel mining problem, which brings together these two disparate sources of review content. Specifically, we use coverage of micro-reviews …
Leading Undergraduate Students To Big Data Generation, Jianjun Yang, Ju Shen
Leading Undergraduate Students To Big Data Generation, Jianjun Yang, Ju Shen
Computer Science Faculty Publications
People are facing a flood of data today. Data are being collected at unprecedented scale in many areas, such as networking, image processing, virtualization, scientific computation, and algorithms. The huge data nowadays are called Big Data. Big data is an all encompassing term for any collection of data sets so large and complex that it becomes difficult to process them using traditional data processing applications. In this article, the authors present a unique way which uses network simulator and tools of image processing to train students abilities to learn, analyze, manipulate, and apply Big Data. Thus they develop students hands-on …
On Efficient K-Optimal-Location-Selection Query Processing In Metric Spaces, Yunjun Gao, Shuyao Qi, Lu Chen, Baihua Zheng, Xinhan Li
On Efficient K-Optimal-Location-Selection Query Processing In Metric Spaces, Yunjun Gao, Shuyao Qi, Lu Chen, Baihua Zheng, Xinhan Li
Research Collection School Of Computing and Information Systems
This paper studies the problem of k-optimal-location-selection (kOLS) retrieval in metric spaces. Given a set DA of customers, a set DB of locations, a constrained region R , and a critical distance dc, a metric kOLS (MkOLS) query retrieves k locations in DB that are outside R but have the maximal optimality scores. Here, the optimality score of a location l∈DB located outside R is defined as the number of the customers in DA that are inside R and meanwhile have their distances to l bounded by …
Joint Search By Social And Spatial Proximity, Kyriakos Mouratidis, Jing Li, Yu Tang, Nikos Mamoulis
Joint Search By Social And Spatial Proximity, Kyriakos Mouratidis, Jing Li, Yu Tang, Nikos Mamoulis
Research Collection School Of Computing and Information Systems
The diffusion of social networks introduces new challenges and opportunities for advanced services, especially so with their ongoing addition of location-based features. We show how applications like company and friend recommendation could significantly benefit from incorporating social and spatial proximity, and study a query type that captures these two-fold semantics. We develop highly scalable algorithms for its processing, and enhance them with elaborate optimizations. Finally, we use real social network data to empirically verify the efficiency and efficacy of our solutions.
Beyond Support And Confidence: Exploring Interestingness Measures For Rule-Based Specification Mining, Bui Tien Duy Le, David Lo
Beyond Support And Confidence: Exploring Interestingness Measures For Rule-Based Specification Mining, Bui Tien Duy Le, David Lo
Research Collection School Of Computing and Information Systems
Numerous rule-based specification mining approaches have been proposed in the literature. Many of these approaches analyze a set of execution traces to discover interesting usage rules, e.g., whenever lock() is invoked, eventually unlock() is invoked. These techniques often generate and enumerate a set of candidate rules and compute some interestingness scores. Rules whose interestingness scores are above a certain threshold would then be output. In past studies, two measures, namely support and confidence, which are well-known measures, are often used to compute these scores. However, aside from these two, many other interestingness measures have been proposed. It is thus unclear …
Hole Detection And Shape-Free Representation And Double Landmarks Based Geographic Routing In Wireless Sensor Networks, Jianjun Yang, Zongming Fei, Ju Shen
Hole Detection And Shape-Free Representation And Double Landmarks Based Geographic Routing In Wireless Sensor Networks, Jianjun Yang, Zongming Fei, Ju Shen
Computer Science Faculty Publications
In wireless sensor networks, an important issue of geographic routing is “local minimum” problem, which is caused by a “hole” that blocks the greedy forwarding process. Existing geographic routing algorithms use perimeter routing strategies to find a long detour path when such a situation occurs. To avoid the long detour path, recent research focuses on detecting the hole in advance, then the nodes located on the boundary of the hole advertise the hole information to the nodes near the hole. Hence the long detour path can be avoided in future routing. We propose a heuristic hole detecting algorithm which identifies …
Bridging The Vocabulary Gap Between Health Seekers And Healthcare Knowledge, Liqiang Nie, Yiliang Zhao, Akbari Mohammad, Jialie Shen, Tat-Seng Chua
Bridging The Vocabulary Gap Between Health Seekers And Healthcare Knowledge, Liqiang Nie, Yiliang Zhao, Akbari Mohammad, Jialie Shen, Tat-Seng Chua
Research Collection School Of Computing and Information Systems
The vocabulary gap between health seekers and providers has hindered the cross-system operability and the interuser reusability. To bridge this gap, this paper presents a novel scheme to code the medical records by jointly utilizing local mining and global learning approaches, which are tightly linked and mutually reinforced. Local mining attempts to code the individual medical record by independently extracting the medical concepts from the medical record itself and then mapping them to authenticated terminologies. A corpus-aware terminology vocabulary is naturally constructed as a byproduct, which is used as the terminology space for global learning. Local mining approach, however, may …
Extracting Interest Tags From Twitter User Biographies, Ying Ding, Jing Jiang
Extracting Interest Tags From Twitter User Biographies, Ying Ding, Jing Jiang
Research Collection School Of Computing and Information Systems
Twitter, one of the most popular social media platforms, has been studied from different angles. One of the important sources of information in Twitter is users’ biographies, which are short self-introductions written by users in free form. Biographies often describe users’ background and interests. However, to the best of our knowledge, there has not been much work trying to extract information from Twitter biographies. In this work, we study how to extract information revealing users’ personal interests from Twitter biographies. A sequential labeling model is trained with automatically constructed labeled data. The popular patterns expressing user interests are extracted and …
Data Preparation For Social Network Mining And Analysis, Yazhe Wang
Data Preparation For Social Network Mining And Analysis, Yazhe Wang
Dissertations and Theses Collection (Open Access)
This dissertation studies the problem of preparing good-quality social network data for data analysis and mining. Modern online social networks such as Twitter, Facebook, and LinkedIn have rapidly grown in popularity. The consequent availability of a wealth of social network data provides an unprecedented opportunity for data analysis and mining researchers to determine useful and actionable information in a wide variety of fields such as social sciences, marketing, management, and security. However, raw social network data are vast, noisy, distributed, and sensitive in nature, which challenge data mining and analysis tasks in storage, efficiency, accuracy, etc. Many mining algorithms cannot …
High-Dimensional Data Stream Classification Via Sparse Online Learning, Dayong Wang, Pengcheng Wu, Peilin Zhao, Yue Wu, Chunyan Miao, Steven C. H. Hoi
High-Dimensional Data Stream Classification Via Sparse Online Learning, Dayong Wang, Pengcheng Wu, Peilin Zhao, Yue Wu, Chunyan Miao, Steven C. H. Hoi
Research Collection School Of Computing and Information Systems
The amount of data in our society has been exploding in the era of big data today. In this paper, we address several open challenges of big data stream classification, including high volume, high velocity, high dimensionality, and high sparsity. Many existing studies in data mining literature solve data stream classification tasks in a batch learning setting, which suffers from poor efficiency and scalability when dealing with big data. To overcome the limitations, this paper investigates an online learning framework for big data stream classification tasks. Unlike some existing online data stream classification techniques that are often based on first-order …
Online Passive Aggressive Active Learning And Its Applications, Jing Lu, Peilin Zhao, Steven C. H. Hoi
Online Passive Aggressive Active Learning And Its Applications, Jing Lu, Peilin Zhao, Steven C. H. Hoi
Research Collection School Of Computing and Information Systems
We investigate online active learning techniques for classification tasks in data stream mining applications. Unlike traditional learning approaches (either batch or online learning) that often require to request the class label of each incoming instance, online active learning queries only a subset of informative incoming instances to update the classification model, which aims to maximize classification performance using minimal human labeling effort during the entire online stream data mining task. In this paper, we present a new family of algorithms for online active learning called Passive-Aggressive Active (PAA) learning algorithms by adapting the popular Passive-Aggressive algorithms in an online active …
Generative Modeling Of Entity Comparisons In Text, Maksim Tkachenko, Hady W. Lauw
Generative Modeling Of Entity Comparisons In Text, Maksim Tkachenko, Hady W. Lauw
Research Collection School Of Computing and Information Systems
Users frequently rely on online reviews for decision making. In addition to allowing users to evaluate the quality of individual products, reviews also support comparison shopping. One key user activity is to compare two (or more) products based on a specific aspect. However, making a comparison across two different reviews, written by different authors, is not always equitable due to the different standards and preferences of individual authors. Therefore, we focus instead on comparative sentences, whereby two products are compared directly by a review author within a single sentence. We study the problem of comparative relation mining. Given a set …
Modloc: Localizing Multiple Objects In Dynamic Indoor Environment, Xiaonan Guo, Dian Zhang, Kaishun Wu, Lionel M. Ni
Modloc: Localizing Multiple Objects In Dynamic Indoor Environment, Xiaonan Guo, Dian Zhang, Kaishun Wu, Lionel M. Ni
Research Collection School Of Computing and Information Systems
Radio frequency (RF) based technologies play an important role in indoor localization, since Radio Signal Strength (RSS) can be easily measured by various wireless devices without additional cost. Among these, radio map based technologies (also referred as fingerprinting technologies) are attractive due to high accuracy and easy deployment. However, these technologies have not been extensively applied on real environment for two fatal limitations. First, it is hard to localize multiple objects. When the number of target objects is unknown, constructing a radio map of multiple objects is almost impossible. Second, environment changes will generate different multipath signals and severely disturb …
Dynamic Clustering Of Contextual Multi-Armed Bandits, Trong T. Nguyen, Hady W. Lauw
Dynamic Clustering Of Contextual Multi-Armed Bandits, Trong T. Nguyen, Hady W. Lauw
Research Collection School Of Computing and Information Systems
With the prevalence of the Web and social media, users increasingly express their preferences online. In learning these preferences, recommender systems need to balance the trade-off between exploitation, by providing users with more of the "same", and exploration, by providing users with something "new" so as to expand the systems' knowledge. Multi-armed bandit (MAB) is a framework to balance this trade-off. Most of the previous work in MAB either models a single bandit for the whole population, or one bandit for each user. We propose an algorithm to divide the population of users into multiple clusters, and to customize the …
Networked Employment Discrimination, Tamara Kneese
Networked Employment Discrimination, Tamara Kneese
Media Studies
Employers often struggle to assess qualified applicants, particularly in contexts where they receive hundreds of applications for job openings. In an effort to increase efficiency and improve the process, many have begun employing new tools to sift through these applications, looking for signals that a candidate is “the best fit.” Some companies use tools that offer algorithmic assessments of workforce data to identify the variables that lead to stronger employee performance, or to high employee attrition rates, while others turn to third party ranking services to identify the top applicants in a labor pool. Still others eschew automated systems, but …
Cost-Sensitive Online Classification, Jialei Wang, Peilin Zhao, Steven C. H. Hoi
Cost-Sensitive Online Classification, Jialei Wang, Peilin Zhao, Steven C. H. Hoi
Research Collection School Of Computing and Information Systems
Both cost-sensitive classification and online learning have been extensively studied in data mining and machine learning communities, respectively. However, very limited study addresses an important intersecting problem, that is, “Cost-Sensitive Online Classification". In this paper, we formally study this problem, and propose a new framework for Cost-Sensitive Online Classification by directly optimizing cost-sensitive measures using online gradient descent techniques. Specifically, we propose two novel cost-sensitive online classification algorithms, which are designed to directly optimize two well-known cost-sensitive measures: (i) maximization of weighted sum of sensitivity and specificity, and (ii) minimization of weighted misclassification cost. We analyze the theoretical bounds of …
Sharing Political News: The Balancing Act Of Intimacy And Socialization In Selective Exposure, Jisun An, Daniele Quercia, Meeyoung Cha, Krishna Gummadi, Jon Crowcroft
Sharing Political News: The Balancing Act Of Intimacy And Socialization In Selective Exposure, Jisun An, Daniele Quercia, Meeyoung Cha, Krishna Gummadi, Jon Crowcroft
Research Collection School Of Computing and Information Systems
One might think that, compared to traditional media, social media sites allow people to choose more freely what to read and what to share, especially for politically oriented news. However, reading and sharing habits originate from deeply ingrained behaviors that might be hard to change. To test the extent to which this is true, we propose a Political News Sharing (PoNS) model that holistically captures four key aspects of social psychology: gratification, selective exposure, socialization, and trust & intimacy. Using real instances of political news sharing in Twitter, we study the predictive power of these features. As one might expect, …
Online Probabilistic Learning For Fuzzy Inference System, Richard Jayadi Oentaryo, Meng Joo Er, San Linn, Xiang Li
Online Probabilistic Learning For Fuzzy Inference System, Richard Jayadi Oentaryo, Meng Joo Er, San Linn, Xiang Li
Research Collection School Of Computing and Information Systems
Online learning is a key methodology for expert systems to gracefully cope with dynamic environments. In the context of neuro-fuzzy systems, research efforts have been directed toward developing online learning methods that can update both system structure and parameters on the fly. However, the current online learning approaches often rely on heuristic methods that lack a formal statistical basis and exhibit limited scalability in the face of large data stream. In light of these issues, we develop a new Sequential Probabilistic Learning for Adaptive Fuzzy Inference System (SPLAFIS) that synergizes the Bayesian Adaptive Resonance Theory (BART) and Rule-Wise Decoupled Extended …
Interestingness-Driven Diffussion Process Summarization In Dynamic Networks, Qiang Qu, Siyuan Liu, Christian Jensen, Feida Zhu, Christos Faloutsos
Interestingness-Driven Diffussion Process Summarization In Dynamic Networks, Qiang Qu, Siyuan Liu, Christian Jensen, Feida Zhu, Christos Faloutsos
Research Collection School Of Computing and Information Systems
The widespread use of social networks enables the rapid diffusion of information, e.g., news, among users in very large communities. It is a substantial challenge to be able to observe and understand such diffusion processes, which may be modeled as networks that are both large and dynamic. A key tool in this regard is data summarization. However, few existing studies aim to summarize graphs/networks for dynamics. Dynamic networks raise new challenges not found in static settings, including time sensitivity and the needs for online interestingness evaluation and summary traceability, which render existing techniques inapplicable. We study the topic of dynamic …
Opinion Mining Of Sociopolitical Comments From Social Media, Swapna Gottipati
Opinion Mining Of Sociopolitical Comments From Social Media, Swapna Gottipati
Dissertations and Theses Collection (Open Access)
Opinions are central to almost all human activities by influencing greatly the decision making process. In this thesis, we present the problems of mining issues, extracting entities and suggestive opinions towards the entities, detecting thoughtful comments, and extracting stances and ideological expressions from online comments in the sociopolitical domain. This study is essential for opinion mining applications that are beneficial for policy makers, government sectors and social organizations. Much work has been done to try to uncover consumer sentiments from online comments to help businesses improve their products and services. However, sociopolitical opinion mining poses new challenges due to complex …
Direct Neighbor Search, Jilian Zhang, Kyriakos Mouratidis, Hwee Hwa Pang
Direct Neighbor Search, Jilian Zhang, Kyriakos Mouratidis, Hwee Hwa Pang
Research Collection School Of Computing and Information Systems
In this paper we study a novel query type, called direct neighbor query. Two objects in a dataset are direct neighbors (DNs) if a window selection may exclusively retrieve these two objects. Given a source object, a DN search computes all of its direct neighbors in the dataset. The DNs define a new type of affinity that differs from existing formulations (e.g., nearest neighbors, nearest surrounders, reverse nearest neighbors, etc.) and finds application in domains where user interests are expressed in the form of windows, i.e., multi-attribute range selections. Drawing on key properties of the DN relationship, we develop an …
Semantic Visualization For Spherical Representation, Tuan M. V. Le, Hady W. Lauw
Semantic Visualization For Spherical Representation, Tuan M. V. Le, Hady W. Lauw
Research Collection School Of Computing and Information Systems
Visualization of high-dimensional data such as text documents is widely applicable. The traditional means is to find an appropriate embedding of the high-dimensional representation in a low-dimensional visualizable space. As topic modeling is a useful form of dimensionality reduction that preserves the semantics in documents, recent approaches aim for a visualization that is consistent with both the original word space, as well as the semantic topic space. In this paper, we address the semantic visualization problem. Given a corpus of documents, the objective is to simultaneously learn the topic distributions as well as the visualization coordinates of documents. We propose …
Online Multiple Kernel Regression, Doyen Sahoo, Steven C. H. Hoi, Bin Li
Online Multiple Kernel Regression, Doyen Sahoo, Steven C. H. Hoi, Bin Li
Research Collection School Of Computing and Information Systems
Kernel-based regression represents an important family of learning techniques for solving challenging regression tasks with non-linear patterns. Despite being studied extensively, most of the existing work suffers from two major drawbacks: (i) they are often designed for solving regression tasks in a batch learning setting, making them not only computationally inefficient and but also poorly scalable in real-world applications where data arrives sequentially; and (ii) they usually assume a fixed kernel function is given prior to the learning task, which could result in poor performance if the chosen kernel is inappropriate. To overcome these drawbacks, this paper presents a novel …
Structure Preserving Large Imagery Reconstruction, Ju Shen, Jianjun Yang, Sami Taha Abu Sneineh, Bryson Payne, Markus Hitz
Structure Preserving Large Imagery Reconstruction, Ju Shen, Jianjun Yang, Sami Taha Abu Sneineh, Bryson Payne, Markus Hitz
Computer Science Faculty Publications
With the explosive growth of web-based cameras and mobile devices, billions of photographs are uploaded to the internet. We can trivially collect a huge number of photo streams for various goals, such as image clustering, 3D scene reconstruction, and other big data applications. However, such tasks are not easy due to the fact the retrieved photos can have large variations in their view perspectives, resolutions, lighting, noises, and distortions. Furthermore, with the occlusion of unexpected objects like people, vehicles, it is even more challenging to find feature correspondences and reconstruct realistic scenes. In this paper, we propose a structure-based image …
Manifold Learning For Jointly Modeling Topic And Visualization, Tuan Minh Van Le, Hady W. Lauw
Manifold Learning For Jointly Modeling Topic And Visualization, Tuan Minh Van Le, Hady W. Lauw
Research Collection School Of Computing and Information Systems
Classical approaches to visualization directly reduce a document's high-dimensional representation into visualizable two or three dimensions, using techniques such as multidimensional scaling. More recent approaches consider an intermediate representation in topic space, between word space and visualization space, which preserves the semantics by topic modeling. We call the latter semantic visualization problem, as it seeks to jointly model topic and visualization. While previous approaches aim to preserve the global consistency, they do not consider the local consistency in terms of the intrinsic geometric structure of the document manifold. We therefore propose an unsupervised probabilistic model, called Semafore, which aims to …
Soml: Sparse Online Metric Learning With Application To Image Retrieval, Xingyu Gao, Steven C. H. Hoi, Yongdong Zhang, Ji Wan, Jintao Li
Soml: Sparse Online Metric Learning With Application To Image Retrieval, Xingyu Gao, Steven C. H. Hoi, Yongdong Zhang, Ji Wan, Jintao Li
Research Collection School Of Computing and Information Systems
Image similarity search plays a key role in many multimedia applications, where multimedia data (such as images and videos) are usually represented in high-dimensional feature space. In this paper, we propose a novel Sparse Online Metric Learning (SOML) scheme for learning sparse distance functions from large-scale high-dimensional data and explore its application to image retrieval. In contrast to many existing distance metric learning algorithms that are often designed for low-dimensional data, the proposed algorithms are able to learn sparse distance metrics from high-dimensional data in an efficient and scalable manner. Our experimental results show that the proposed method achieves better …
Learning Relative Similarity By Stochastic Dual Coordinate Ascent, Pengcheng Wu, Ding Yi, Peilin Zhao, Chunyan Miao, Steven C. H. Hoi
Learning Relative Similarity By Stochastic Dual Coordinate Ascent, Pengcheng Wu, Ding Yi, Peilin Zhao, Chunyan Miao, Steven C. H. Hoi
Research Collection School Of Computing and Information Systems
Learning relative similarity from pairwise instances is an important problem in machine learning and has a wide range of applications. Despite being studied for years, some existing methods solved by Stochastic Gradient Descent (SGD) techniques generally suffer from slow convergence. In this paper, we investigate the application of Stochastic Dual Coordinate Ascent (SDCA) technique to tackle the optimization task of relative similarity learning by extending from vector to matrix parameters. Theoretically, we prove the optimal linear convergence rate for the proposed SDCA algorithm, beating the well-known sublinear convergence rate by the previous best metric learning algorithms. Empirically, we conduct extensive …
Predicting The Popularity Of Web 2.0 Items Based On User Comments, Xiangnan He, Ming Gao, Min-Yen Kan, Yiqun Liu, Kazunari Sugiyama
Predicting The Popularity Of Web 2.0 Items Based On User Comments, Xiangnan He, Ming Gao, Min-Yen Kan, Yiqun Liu, Kazunari Sugiyama
Research Collection School Of Computing and Information Systems
In the current Web 2.0 era, the popularity of Web resources fluctuates ephemerally, based on trends and social interest. As a result, content-based relevance signals are insufficient to meet users' constantly evolving information needs in searching for Web 2.0 items. Incorporating future popularity into ranking is one way to counter this. However, predicting popularity as a third party (as in the case of general search engines) is difficult in practice, due to their limited access to item view histories. To enable popularity prediction externally without excessive crawling, we propose an alternative solution by leveraging user comments, which are more accessible …