Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences Commons™

Open Access. Powered by Scholars. Published by Universities.®

Singapore Management University

Discipline
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 6211 - 6240 of 9024

Full-Text Articles in Computer Sciences

Jointly Modeling Aspects, Ratings And Sentiments For Movie Recommendation (Jmars), Qiming Diao, Minghui Qiu, Chao-Yuan Wu, Alexander J. Smola, Jing Jiang, Chong Wang Aug 2014

Jointly Modeling Aspects, Ratings And Sentiments For Movie Recommendation (Jmars), Qiming Diao, Minghui Qiu, Chao-Yuan Wu, Alexander J. Smola, Jing Jiang, Chong Wang

Research Collection School Of Computing and Information Systems

Recommendation and review sites offer a wealth of information beyond ratings. For instance, on IMDb users leave reviews, commenting on different aspects of a movie (e.g. actors, plot, visual effects), and expressing their sentiments (positive or negative) on these aspects in their reviews. This suggests that uncovering aspects and sentiments will allow us to gain a better understanding of users, movies, and the process involved in generating ratings. The ability to answer questions such as “Does this user care more about the plot or about the special effects?” or ”What is the quality of the movie in terms of acting?” …


Automatic Fine-Grained Issue Report Reclassification, Pavneet Singh Kochhar, Ferdian Thung, David Lo Aug 2014

Automatic Fine-Grained Issue Report Reclassification, Pavneet Singh Kochhar, Ferdian Thung, David Lo

Research Collection School Of Computing and Information Systems

Issue tracking systems are valuable resources during software maintenance activities. These systems contain different categories of issue reports such as bug, request for improvement (RFE), documentation, refactoring, task etc. While logging issue reports into a tracking system, reporters can indicate the category of the reports. Herzig et al. Recently reported that more than 40% of issue reports are given wrong categories in issue tracking systems. Among issue reports that are marked as bugs, more than 30% of them are not bug reports. The misclassification of issue reports can adversely affects developers as they then need to manually identify the categories …


Managing Complexity Through Selective Decoupling, C. Jason Woodard, Eric K. Clemons Aug 2014

Managing Complexity Through Selective Decoupling, C. Jason Woodard, Eric K. Clemons

Research Collection School Of Computing and Information Systems

Designers of complex systems are often confounded by the tendency for design changes to increase performance on some dimensions while decreasing it on others. While adopting a more modular architecture may temper these opposing effects, modularization can also deprive designers of opportunities to harness complementarities among system elements. This paper explores this tension using an NK model in which product designers can modify the structure of their fitness landscapes by suppressing or restoring interactions between components. We find that these changes can lead to improved performance by flattening harmful interactions that would otherwise cause search efforts to become trapped on …


The Learning Curves In Open-Source Software (Oss) Development Network, Youngsoo Kim, Lingxiao Jiang Aug 2014

The Learning Curves In Open-Source Software (Oss) Development Network, Youngsoo Kim, Lingxiao Jiang

Research Collection School Of Computing and Information Systems

We examine the learning curves of individual software developers in Open-Source Software (OSS) Development. We collected the dataset of multi-year code change histories from the repositories for five open source software projects involving more than 100 developers. We build and estimate regression models to assess individual developers' learning progress (in reducing the likelihood they may make a bug). Our estimation results show that developer's coding experience does not decrease bug ratios while cumulative bug-fixing experience leads to learning progress. The results may have implications and provoke future research on project management about allocating resources on tasks that add new code …


Automated Prediction Of Glasgow Outcome Scale For Traumatic Brain Injury, Bolan Su, Thien Anh Dinh, A. K. Ambastha, Tianxia Gong, Tomi Silander, Shijian Lu, C. C. Tchoyoson Lim, Boon Chuan Pang, Cheng Kiang Lee, Tze-Yun Leong, Chew Lim Tan Aug 2014

Automated Prediction Of Glasgow Outcome Scale For Traumatic Brain Injury, Bolan Su, Thien Anh Dinh, A. K. Ambastha, Tianxia Gong, Tomi Silander, Shijian Lu, C. C. Tchoyoson Lim, Boon Chuan Pang, Cheng Kiang Lee, Tze-Yun Leong, Chew Lim Tan

Research Collection School Of Computing and Information Systems

Clinical features found in brain CT scan images are widely used in traumatic brain injury (TBI) as indicators for Glasgow Outcome Scale (GOS) prediction. However, due to the lack of automated methods to measure and quantify the CT scan image features, the computerized prediction of GOS in TBI has not been well studied. This paper introduces an automated GOS prediction system for traumatic brain CT images. Different from most existing systems that perform the prognosis based on pre-processed data, our system directly works on brain CT scan images based on the image features. Our system can also be extended to …


Guest Editorial: Special Issue On Brain Inspired Models Of Cognitive Memory, Huajin Tang, Kiruthika Ramanathan, Ning Nign Aug 2014

Guest Editorial: Special Issue On Brain Inspired Models Of Cognitive Memory, Huajin Tang, Kiruthika Ramanathan, Ning Nign

Research Collection School Of Computing and Information Systems

Current memory technologies have experienced significant progress in terms of storage capacity, operation speed, integration capability, etc. However, their functions are highly constrained in storing and transferring data in space and time, prompting the need for improvement. In contrast to physical memories, the biological counterpart – cognitive memory – has versatile functions. For instance, human memory stores data associatively such that different modalities of data could be retrieved simultaneously; it can learn different concepts, categorize and store them in an organized manner; it can process and store data concurrently and in a distributed fashion; it can restore content even if …


Integrating Motivated Learning And K-Winner-Take-All To Coordinate Multi-Agent Reinforcement Learning, Teck-Hou Teng, Ah-Hwee Tan, Janusz Starzyk, Yuan-Sin Tan, Loo-Nin Teow Aug 2014

Integrating Motivated Learning And K-Winner-Take-All To Coordinate Multi-Agent Reinforcement Learning, Teck-Hou Teng, Ah-Hwee Tan, Janusz Starzyk, Yuan-Sin Tan, Loo-Nin Teow

Research Collection School Of Computing and Information Systems

This work addresses the coordination issue in distributed optimization problem (DOP) where multiple distinct and time-critical tasks are performed to satisfy a global objective function. The performance of these tasks has to be coordinated due to the sharing of consumable resources and the dependency on non-consumable resources. Knowing that it can be sub-optimal to predefine the performance of the tasks for large DOPs, the multi-agent reinforcement learning (MARL) framework is adopted wherein an agent is used to learn the performance of each distinct task using reinforcement learning. To coordinate MARL, we propose a novel coordination strategy integrating Motivated Learning (ML) …


On Macro And Micro Exploration Of Hashtag Diffusion In Twitter, Yazhe Wang, Baihua Zheng Aug 2014

On Macro And Micro Exploration Of Hashtag Diffusion In Twitter, Yazhe Wang, Baihua Zheng

Research Collection School Of Computing and Information Systems

This exploratory work studies hashtag diffusion in Twitter. The analysis is conducted from two aspects. From the macro perspective, we study general properties of hashtag diffusion, and classify hashtags into three main classes based on their temporal dynamics referred as 'single spike', 'multi-spikes', and 'fluctuation', and find that each of these classes has some unique characteristics. From the micro perspective, we investigate individual diffusion.We adopt Edelman's 'topology of influence' theory to identify four type of users with different influence levels in diffusion based on their dynamic retweet behaviors. The results of our study are useful for gaining more insights of …


Reducing Carbon Emission Of Ocean Shipments By Optimizing Container Size Selection, Edwin Lik Ming Chong, Nang Laik Ma, Kar Way Tan Aug 2014

Reducing Carbon Emission Of Ocean Shipments By Optimizing Container Size Selection, Edwin Lik Ming Chong, Nang Laik Ma, Kar Way Tan

Research Collection School Of Computing and Information Systems

Human’s impact on earth through global warming is more or less an accepted fact. Ocean freight is estimated to contribute 4-5% of global carbon emissions and manufacturing companies can aid in reducing this amount. Many companies that ship goods through full container loads do not have the capabilities to ensure the containers they are using minimizes their carbon footprint. One of the reasons is the choice of non-ideal container sizes for their shipments. This paper provides a mathematical model to minimize companies’ shipping carbon footprints by selecting the ideal container sizes appropriate for their shipment volumes. Using data from a …


An Empirical Study Of Off-Line Configuration And On-Line Adaptation In Operator Selection, Zhi Yuan, Stephanus Daniel Handoko, Duc Thien Nguyen, Hoong Chuin Lau Aug 2014

An Empirical Study Of Off-Line Configuration And On-Line Adaptation In Operator Selection, Zhi Yuan, Stephanus Daniel Handoko, Duc Thien Nguyen, Hoong Chuin Lau

Research Collection School Of Computing and Information Systems

Automating the process of finding good parameter settings is important in the design of high-performing algorithms. These automatic processes can generally be categorized into off-line and on-line methods. Off-line configuration consists in learning and selecting the best setting in a training phase, and usually fixes it while solving an instance. On-line adaptation methods on the contrary vary the parameter setting adaptively during each algorithm run. In this work, we provide an empirical study of both approaches on the operator selection problem, explore the possibility of varying parameter value by a non-adaptive distribution tuned off-line, and incorporate the off-line with on-line …


A Mathematical Model And Metaheuristics For Time Dependent Orienteering Problem, Aldy Gunawan, Zhi Yuan, Hoong Chuin Lau Aug 2014

A Mathematical Model And Metaheuristics For Time Dependent Orienteering Problem, Aldy Gunawan, Zhi Yuan, Hoong Chuin Lau

Research Collection School Of Computing and Information Systems

This paper presents a generalization of the Orienteering Problem, the Time-Dependent Orienteering Problem (TDOP) which is based on the real-life application of providing automatic tour guidance to a large leisure facility such as a theme park. In this problem, the travel time between two nodes depends on the time when the trip starts. We formulate the problem as an integer linear programming (ILP) model. We then develop various heuristics in a step by step fashion: greedy construction, local search and variable neighborhood descent, and two versions of iterated local search. The proposed metaheuristics were tested on modified benchmark instances, randomly …


Online Multiple Kernel Regression, Doyen Sahoo, Steven C. H. Hoi, Bin Li Aug 2014

Online Multiple Kernel Regression, Doyen Sahoo, Steven C. H. Hoi, Bin Li

Research Collection School Of Computing and Information Systems

Kernel-based regression represents an important family of learning techniques for solving challenging regression tasks with non-linear patterns. Despite being studied extensively, most of the existing work suffers from two major drawbacks: (i) they are often designed for solving regression tasks in a batch learning setting, making them not only computationally inefficient and but also poorly scalable in real-world applications where data arrives sequentially; and (ii) they usually assume a fixed kernel function is given prior to the learning task, which could result in poor performance if the chosen kernel is inappropriate. To overcome these drawbacks, this paper presents a novel …


Ranking Model Selection And Fusion For Effective Microblog Search, Zhongyu Wei, Wei Gao, Tarek El-Ganainy, Walid Magdy, Kam-Fai Wong Jul 2014

Ranking Model Selection And Fusion For Effective Microblog Search, Zhongyu Wei, Wei Gao, Tarek El-Ganainy, Walid Magdy, Kam-Fai Wong

Research Collection School Of Computing and Information Systems

Re-ranking was shown to have positive impact on the effectiveness for microblog search. Yet existing approaches mostly focused on using a single ranker to learn some better ranking function with respect to various relevance features. Given various available rank learners (such as learning to rank algorithms), in this work, we mainly study an orthogonal problem where multiple learned ranking models form an ensemble for re-ranking the retrieved tweets than just using a single ranking model in order to achieve higher search effectiveness. We explore the use of query-sensitive model selection and rank fusion methods based on the result lists produced …


Symbolic Analysis Of An Electric Vehicle Charging Protocol, Li Li, Jun Pang, Yang Liu, Jun Sun, Jin Song Dong Jul 2014

Symbolic Analysis Of An Electric Vehicle Charging Protocol, Li Li, Jun Pang, Yang Liu, Jun Sun, Jin Song Dong

Research Collection School Of Computing and Information Systems

In this paper, we describe our analysis of a recently proposed electric vehicle charing protocol. The protocol builds on complex cryptographic primitives such as commitment, zeroknowledge proofs, BBS+ signature and etc. Moreover, interesting properties such as secrecy, authentication, anonymity, and location privacy are claimed on this protocol. It thus presents a challenge for formal verification, as no single existing tool for security protocol analysis support for all the required features. In our analysis, we employ and combine the strength of two stateof-the-art symbolic verifiers, Tamarin and ProVerif, to check all important properties of the protocol.


Detecting Differences Across Multiple Instances Of Code Clones, Yun Lin, Zhenchang Xing, Yinxing Xue, Yang Liu, Xin Peng, Jun Sun, Wenyun Zhao Jul 2014

Detecting Differences Across Multiple Instances Of Code Clones, Yun Lin, Zhenchang Xing, Yinxing Xue, Yang Liu, Xin Peng, Jun Sun, Wenyun Zhao

Research Collection School Of Computing and Information Systems

Clone detectors find similar code fragments (i.e., instances of code clones) and report large numbers of them for industrial systems. To maintain or manage code clones, developers often have to investigate differences of multiple cloned code fragments. However,existing program differencing techniques compare only two code fragments at a time. Developers then have to manually combine several pairwise differencing results. In this paper, we present an approach to automatically detecting differences across multiple clone instances. We have implemented our approach as an Eclipse plugin and evaluated its accuracy with three Java software systems. Our evaluation shows that our algorithm has precision …


Multi-Cost And Upgradable Spatial Network Databases, Yimin Lin Jul 2014

Multi-Cost And Upgradable Spatial Network Databases, Yimin Lin

Dissertations and Theses Collection (Open Access)

In this dissertation, we first consider data processing problems in multi-cost networks and in upgradable networks. These network types are motivated by real-life situations, which do not fall under the standard spatial network formulation and have not received much attention from database researchers. In a multi-cost network (MCN), each edge is associated with more than one weight type that may affect the user-specific perception of distance. We study two query types on MCNs, namely, the MCN skyline and the MCN top-k query. In an upgradable network, a subset of the edges are amenable to weight reduction, at a cost (e.g., …


Synchronicity: Pushing The Envelope Of Fine-Grained Localization With Distributed Mimo, Jie Xiong, Kyle Jamieson, Karthikeyan Sundaresan Jul 2014

Synchronicity: Pushing The Envelope Of Fine-Grained Localization With Distributed Mimo, Jie Xiong, Kyle Jamieson, Karthikeyan Sundaresan

Research Collection School Of Computing and Information Systems

Indoor localization of mobile devices and tags has received much attention recently, with encouraging fine-grained localization results available with enough line-of-sight coverage and enough hardware infrastructure. Synchronicity is a location system that aims to push the envelope of highly-accurate localization systems further in both dimensions, requiring less line-of-sight and less infrastructure. With Distributed MIMO network of wireless LAN access points (APs) as a starting point, we leverage the time synchronization that such a network affords to localize with time-difference-of-arrival information at the APs. We contribute novel super-resolution signal processing algorithms and reflection path elimination schemes, yielding superior results even in …


Influences Of Influential Users: An Empirical Study Of Music Social Network, Jing Ren, Zhiyong Cheng, Jialie Shen, Feida Zhu Jul 2014

Influences Of Influential Users: An Empirical Study Of Music Social Network, Jing Ren, Zhiyong Cheng, Jialie Shen, Feida Zhu

Research Collection School Of Computing and Information Systems

Influential user can play a crucial role in online social networks. This paper documents an empirical study aiming at exploring the effects of influential users in the context of music social network. To achieve this goal, music diffusion graph is developed to model how information propagates over network. We also propose a heuristic method to measure users' influences. Using the real data from Last. fm, our empirical test demonstrates key effects of influential users and reveals limitations of existing influence identification/characterization schemes.


Predicting The Popularity Of Web 2.0 Items Based On User Comments, Xiangnan He, Ming Gao, Min-Yen Kan, Yiqun Liu, Kazunari Sugiyama Jul 2014

Predicting The Popularity Of Web 2.0 Items Based On User Comments, Xiangnan He, Ming Gao, Min-Yen Kan, Yiqun Liu, Kazunari Sugiyama

Research Collection School Of Computing and Information Systems

In the current Web 2.0 era, the popularity of Web resources fluctuates ephemerally, based on trends and social interest. As a result, content-based relevance signals are insufficient to meet users' constantly evolving information needs in searching for Web 2.0 items. Incorporating future popularity into ranking is one way to counter this. However, predicting popularity as a third party (as in the case of general search engines) is difficult in practice, due to their limited access to item view histories. To enable popularity prediction externally without excessive crawling, we propose an alternative solution by leveraging user comments, which are more accessible …


Fsph: Fitted Spectral Hashing For Efficient Similarity Search, Yong-Dong Zhang, Yu Wang, Sheng Tang, Steven C. H. Hoi, Jin-Tao Li Jul 2014

Fsph: Fitted Spectral Hashing For Efficient Similarity Search, Yong-Dong Zhang, Yu Wang, Sheng Tang, Steven C. H. Hoi, Jin-Tao Li

Research Collection School Of Computing and Information Systems

Spectral hashing (SpH) is an efficient and simple binary hashing method, which assumes that data are sampled from a multidimensional uniform distribution. However, this assumption is too restrictive in practice. In this paper we propose an improved method, fitted spectral hashing (FSpH), to relax this distribution assumption. Our work is based on the fact that one-dimensional data of any distribution could be mapped to a uniform distribution without changing the local neighbor relations among data items. We have found that this mapping on each PCA direction has certain regular pattern, and could be fitted well by S-curve function (Sigmoid function). …


Verifying Monadic Second-Order Properties Of Graph Programs, Christopher M. Poskitt, Detlef Plump Jul 2014

Verifying Monadic Second-Order Properties Of Graph Programs, Christopher M. Poskitt, Detlef Plump

Research Collection School Of Computing and Information Systems

The core challenge in a Hoare- or Dijkstra-style proof system for graph programs is in defining a weakest liberal precondition construction with respect to a rule and a postcondition. Previous work addressing this has focused on assertion languages for first-order properties, which are unable to express important global properties of graphs such as acyclicity, connectedness, or existence of paths. In this paper, we extend the nested graph conditions of Habel, Pennemann, and Rensink to make them equivalently expressive to monadic second-order logic on graphs. We present a weakest liberal precondition construction for these assertions, and demonstrate its use in verifying …


On Predicting Religion Labels In Microblogging Networks, Minh Thap Nguyen, Ee Peng Lim Jul 2014

On Predicting Religion Labels In Microblogging Networks, Minh Thap Nguyen, Ee Peng Lim

Research Collection School Of Computing and Information Systems

Religious belief plays an important role in how people behave, influencing how they form preferences, interpret events around them, and develop relationships with others. Traditionally, the religion labels of user population are obtained by conducting a large scale census study. Such an approach is both high cost and time consuming. In this paper, we study the problem of predicting users' religion labels using their microblogging data. We formulate religion label prediction as a classification task, and identify content, structure and aggregate features considering their self and social variants for representing a user. We introduce the notion of representative user to …


Scalable Detection Of Missed Cross-Function Refactorings, Narcisa Andreea Milea, Lingxiao Jiang, Siau-Cheng Khoo Jul 2014

Scalable Detection Of Missed Cross-Function Refactorings, Narcisa Andreea Milea, Lingxiao Jiang, Siau-Cheng Khoo

Research Collection School Of Computing and Information Systems

Refactoring is an important way to improve the design of existing code. Identifying refactoring opportunities (i.e., code fragments that can be refactored) in large code bases is a challenging task. In this paper, we propose a novel, automated and scalable technique for identifying cross-function refactoring opportunities that span more than one function (e.g., Extract Method and Inline Method). The key of our technique is the design of efficient vector inlining operations that emulate the effect of method inlining among code fragments, so that the problem of identifying cross-function refactoring can be reduced to the problem of finding similar vectors before …


Cloud-Based Query Evaluation For Energy-Efficient Mobile Sensing, Tianli Mo, Sougata Sen, Lipyeow Lim, Archan Misra, Rajesh Krishna Balan, Youngki Lee Jul 2014

Cloud-Based Query Evaluation For Energy-Efficient Mobile Sensing, Tianli Mo, Sougata Sen, Lipyeow Lim, Archan Misra, Rajesh Krishna Balan, Youngki Lee

Research Collection School Of Computing and Information Systems

In this paper, we reduce the energy overheads of continuous mobile sensing for context-aware applications that are interested in collective context or events. We propose a cloud-based query management and optimization framework, called CloQue, which can support concurrent queries, executing over thousands of individual smartphones. CloQue exploits correlation across context of different users to reduce energy overheads via two key innovations: i) Dynamically reordering the order of predicate processing to preferentially select predicates with not just lower sensing cost and higher selectivity, but that maximally reduce the uncertainty about other context predicates, and ii) intelligently propagating the query evaluation results …


Decentralized Multi-Agent Reinforcement Learning In Average-Reward Dynamic Dcops, Duc Thien Nguyen, William Yeoh, Hoong Chuin Lau, Shlomo Zilberstein, Chongjie Zhang Jul 2014

Decentralized Multi-Agent Reinforcement Learning In Average-Reward Dynamic Dcops, Duc Thien Nguyen, William Yeoh, Hoong Chuin Lau, Shlomo Zilberstein, Chongjie Zhang

Research Collection School Of Computing and Information Systems

Researchers have introduced the Dynamic Distributed Constraint Optimization Problem (Dynamic DCOP) formulation to model dynamically changing multi-agent coordination problems, where a dynamic DCOP is a sequence of (static canonical) DCOPs, each partially different from the DCOP preceding it. Existing work typically assumes that the problem in each time step is decoupled from the problems in other time steps, which might not hold in some applications. Therefore, in this paper, we make the following contributions: (i) We introduce a new model, called Markovian Dynamic DCOPs (MD-DCOPs), where the DCOP in the next time step is a function of the value assignments …


Reinforcement Learning For Adaptive Operator Selection In Memetic Search Applied To Quadratic Assignment Problem, Stephanus Daniel Handoko, Duc Thien Nguyen, Zhi Yuan, Hoong Chuin Lau Jul 2014

Reinforcement Learning For Adaptive Operator Selection In Memetic Search Applied To Quadratic Assignment Problem, Stephanus Daniel Handoko, Duc Thien Nguyen, Zhi Yuan, Hoong Chuin Lau

Research Collection School Of Computing and Information Systems

Memetic search is well known as one of the state-of-the-art metaheuristics for finding high-quality solutions to NP-hard problems. Its performance is often attributable to appropriate design, including the choice of its operators. In this paper, we propose a Markov Decision Process model for the selection of crossover operators in the course of the evolutionary search. We solve the proposed model by a Q-learning method. We experimentally verify the efficacy of our proposed approach on the benchmark instances of Quadratic Assignment Problem.


Collaborative Error Reduction For Hierarchical Classification, Shiai Zhu, Xiao-Yong Wei, Chong-Wah Ngo Jul 2014

Collaborative Error Reduction For Hierarchical Classification, Shiai Zhu, Xiao-Yong Wei, Chong-Wah Ngo

Research Collection School Of Computing and Information Systems

Hierarchical classification (HC) is a popular and efficient way for detecting the semantic concepts from the images. The conventional method always selects the branch with the highest classification response. This branch selection strategy has a risk of propagating classification errors from higher levels of the hierarchy to the lower levels. We argue that the local strategy is too arbitrary, because the candidate nodes are considered individually, which ignores the semantic and context relationships among concepts. In this paper, we first propose a novel method for HC, which is able to utilize the semantic relationship among candidate nodes and their children …


User Daily Activity Pattern Learning: A Multi-Memory Modeling Approach, Shan Gao, Ah-Hwee Tan Jul 2014

User Daily Activity Pattern Learning: A Multi-Memory Modeling Approach, Shan Gao, Ah-Hwee Tan

Research Collection School Of Computing and Information Systems

In this paper, we propose a multi-memory model, ADLART model, to discover the daily activity pattern of a sensor monitored user from his/her activities of daily living (ADL). The proposed model mimics the human multiple memory system which comprises a working memory, an episodic memory, and a semantic memory. Through encoding user's daily activities patterns in episodic memory and extracting the regularities of activity routines in semantic memory, the ADLART system is able to learn, recognize, compare, and retrieve daily ADL patterns of the user. Experiments are presented to show the performance of the ADLART model using different parameter settings …


Facilitating Crowd Sourced Software Engineering Via Stack Overflow, Ohad Barzilay, Christoph Treude, Alexey Zagalsky Jul 2014

Facilitating Crowd Sourced Software Engineering Via Stack Overflow, Ohad Barzilay, Christoph Treude, Alexey Zagalsky

Research Collection School Of Computing and Information Systems

The open source community, as well as numerous technical blogs and community web sites, put online vast quantities of free source code, ranging from snippets to full-blown products. This code embodies the software development community’s domain knowledge, and mirrors the structure of the Internet: it is distributed rather than hierarchical; it is chaotic, incomplete, and inconsistent. StackOverflow.com is a Question and Answer (Q&A) website which uses social media to facilitate knowledge exchange between programmers by mitigating the pitfalls involved in using code from the Internet. Its design nurtures a community of developers, and enables crowd sourced software engineering activities ranging …


Fsph: Fitted Spectral Hashing For Efficient Similarity Search, Yong-Dong Zhang, Yu Wang, Sheng Tang, Steven C. H. Hoi, Jin-Tao Li Jul 2014

Fsph: Fitted Spectral Hashing For Efficient Similarity Search, Yong-Dong Zhang, Yu Wang, Sheng Tang, Steven C. H. Hoi, Jin-Tao Li

Research Collection School Of Computing and Information Systems

Spectral hashing (SpH) is an efficient and simple binary hashing method, which assumes that data are sampled from a multidimensional uniform distribution. However, this assumption is too restrictive in practice. In this paper we propose an improved method, fitted spectral hashing (FSpH), to relax this distribution assumption. Our work is based on the fact that one-dimensional data of any distribution could be mapped to a uniform distribution without changing the local neighbor relations among data items. We have found that this mapping on each PCA direction has certain regular pattern, and could be fitted well by S-curve function (Sigmoid function). …