Efficient Processing Of Exact Top-K Queries Over Disk-Resident Sorted Lists,
2010
Singapore Management University
Efficient Processing Of Exact Top-K Queries Over Disk-Resident Sorted Lists, Hwee Hwa Pang, Xuhua Ding, Baihua Zheng
Research Collection School Of Computing and Information Systems
The top-k query is employed in a wide range of applications to generate a ranked list of data that have the highest aggregate scores over certain attributes. As the pool of attributes for selection by individual queries may be large, the data are indexed with per-attribute sorted lists, and a threshold algorithm (TA) is applied on the lists involved in each query. The TA executes in two phases--find a cut-off threshold for the top-k result scores, then evaluate all the records that could score above the threshold. In this paper, we focus on exact top-k queries that involve monotonic linear …
Player Performance Prediction In Massively Multiplayer Online Role-Playing Games (Mmorpgs),
2010
Singapore Management University
Player Performance Prediction In Massively Multiplayer Online Role-Playing Games (Mmorpgs), Kyong Jin Shim, Richa Sharan, Jaideep Srivastava
Research Collection School Of Computing and Information Systems
In this study, we propose a comprehensive performance management tool for measuring and reporting operational activities of game players. This study uses performance data of game players in EverQuest II, a popular MMORPG developed by Sony Online Entertainment, to build performance prediction models forgame players. The prediction models provide a projection of player’s future performance based on his past performance, which is expected to be a useful addition to existing player performance monitoring tools. First, we show that variations of PECOTA [2] and MARCEL [3], two most popular baseball home run prediction methods, can be used for game player performance …
Prediction Of Protein Subcellular Localization: A Machine Learning Approach,
2010
Singapore Management University
Prediction Of Protein Subcellular Localization: A Machine Learning Approach, Kyong Jin Shim
Research Collection School Of Computing and Information Systems
Subcellular localization is a key functional characteristic of proteins. Optimally combining available information is one of the key challenges in today's knowledge-based subcellular localization prediction approaches. This study explores machine learning approaches for the prediction of protein subcellular localization that use resources concerning Gene Ontology and secondary structures. Using the spectrum kernel for feature representation of amino acid sequences and secondary structures, we explore an SVM-based learning method that classifies six subcellular localization sites: endoplasmic reticulum, extracellular, Golgi, membrane, mitochondria, and nucleus.
Two-View Transductive Support Vector Machines,
2010
Nanyang Technological University
Two-View Transductive Support Vector Machines, Guangxia Li, Steven C. H. Hoi, Kuiyu Chang
Research Collection School Of Computing and Information Systems
Obtaining high-quality and up-to-date labeled data can be difficult in many real-world machine learning applications, especially for Internet classification tasks like review spam detection, which changes at a very brisk pace. For some problems, there may exist multiple perspectives, so called views, of each data sample. For example, in text classification, the typical view contains a large number of raw content features such as term frequency, while a second view may contain a small but highly-informative number of domain specific features. We thus propose a novel two-view transductive SVM that takes advantage of both the abundant amount of unlabeled data …
Mixed Operators In Compressed Sensing,
2010
University of California - Los Angeles
Mixed Operators In Compressed Sensing, Matthew A. Herman, Deanna Needell
CMC Faculty Publications and Research
Applications of compressed sensing motivate the possibility of using different operators to encode and decode a signal of interest. Since it is clear that the operators cannot be too different, we can view the discrepancy between the two matrices as a perturbation. The stability of L1-minimization and greedy algorithms to recover the signal in the presence of additive noise is by now well-known. Recently however, work has been done to analyze these methods with noise in the measurement matrix, which generates a multiplicative noise term. This new framework of generalized perturbations (i.e., both additive and multiplicative noise) extends the prior …
Pretty Lights,
2010
California Polytechnic State University - San Luis Obispo
Pretty Lights, Nicholas (Nick) Delmas, Matthew (Matt) Maniaci
Computer Engineering
Digital media players often include a visualization component that allows a user to watch a visualization synchronized to their music or videos. This project uses the visualization plugin API of an existing media playback program (WinAmp) but it displays its visuals using physical LED lights. Instead of outputting visuals to the computer screen, data is sent over USB to a micro controller that runs the LED lights. This project aims to give users a more visceral visual experience than traditional visualizations on the computer screen.
A Genetic Algorithm Approach For Optimized Routing,
2010
Old Dominion University
A Genetic Algorithm Approach For Optimized Routing, Pavithra Gudur
Electrical & Computer Engineering Theses & Dissertations
Genetic Algorithms find several applications in a variety of fields, such as engineering, management, finance, chemistry, scheduling, data mining and so on, where optimization plays a key role. This technique represents a numerical optimization technique that is modeled after the natural process of selection based on the Darwinian principle of evolution. The Genetic Algorithm (GA) is one among several optimization techniques and attempts to obtain the desired solution by generating a set of possible candidate solutions or populations. These populations are then compared and the best solutions from the set are retained. Subsequently, new candidate solutions are produced, and the …
Continuous Spatial Assignment Of Moving Users,
2010
University of Hong Kong
Continuous Spatial Assignment Of Moving Users, Hou U Leong, Kyriakos Mouratidis, Nikos Mamoulis
Research Collection School Of Computing and Information Systems
Consider a set of servers and a set of users, where each server has a coverage region (i.e., an area of service) and a capacity (i.e., a maximum number of users it can serve). Our task is to assign every user to one server subject to the coverage and capacity constraints. To offer the highest quality of service, we wish to minimize the average distance between users and their assigned server. This is an instance of a well-studied problem in operations research, termed optimal assignment. Even though there exist several solutions for the static case (where user locations are fixed), …
Information Search Patterns In E-Commerce Product Comparison Services,
2010
Singapore Management University
Information Search Patterns In E-Commerce Product Comparison Services, Fiona Fui-Hoon Nah, Weiyin Hong, Liqiang Chen, Hong-Hee Lee
Research Collection School Of Computing and Information Systems
To facilitate product selection and purchase decisions on e-commerce Web sites, the presentation of product information is very important. In this research, the authors study how disposition styles influence users’ search patterns in product comparison services of e-commerce Web sites. The results show that people use relatively more feature paths and less product paths in vertical disposition style than horizontal disposition style. The findings also indicate that there are relatively more feature paths and less product paths in the first half than second half of the information search paths. This is consistent with Gensch’s two-stage choice model which suggests that …
Optimal Matching Between Spatial Datasets Under Capacity Constraints,
2010
University of Hong Kong
Optimal Matching Between Spatial Datasets Under Capacity Constraints, Hou U Leong, Kyriakos Mouratidis, Man Lung Yiu, Nikos Mamoulis
Research Collection School Of Computing and Information Systems
Consider a set of customers (e.g., WiFi receivers) and a set of service providers (e.g., wireless access points), where each provider has a capacity and the quality of service offered to its customers is anti-proportional to their distance. The capacity constrained assignment (CCA) is a matching between the two sets such that (i) each customer is assigned to at most one provider, (ii) every provider serves no more customers than its capacity, (iii) the maximum possible number of customers are served, and (iv) the sum of Euclidean distances within the assigned provider-customer pairs is minimized. Although max-flow algorithms are applicable …
What Is Twitter, A Social Network Or A News Media?,
2010
Singapore Management University
What Is Twitter, A Social Network Or A News Media?, Haewoon Kwak, Changhyun Lee, Hosung: Moon Park
Research Collection School Of Computing and Information Systems
Twitter, a microblogging service less than three years old, commands more than 41 million users as of July 2009 and is growing fast. Twitter users tweet about any topic within the 140-character limit and follow others to receive their tweets. The goal of this paper is to study the topological characteristics of Twitter and its power as a new medium of information sharing.We have crawled the entire Twitter site and obtained 41.7 million user profiles, 1.47 billion social relations, 4,262 trending topics, and 106 million tweets. In its follower-following topology analysis we have found a non-power-law follower distribution, a short …
Mining Diversity On Networks,
2010
Tsinghua University
Mining Diversity On Networks, Lu Liu, Feida Zhu, Chen Chen, Xifeng Yan, Jiawei Han, Philip Yu, Shiqiang Yang
Research Collection School Of Computing and Information Systems
Despite the recent emergence of many large-scale networks in different application domains, an important measure that captures a participant’s diversity in the network has been largely neglected in previous studies. Namely, diversity characterizes how diverse a given node connects with its peers. In this paper, we give a comprehensive study of this concept. We first lay out two criteria that capture the semantic meaning of diversity, and then propose a compliant definition which is simple enough to embed the idea. An efficient top-k diversity ranking algorithm is developed for computation on dynamic networks. Experiments on both synthetic and real datasets …
Algorithms For Constrained K-Nearest Neighbor Queries Over Moving Object Trajectories,
2010
Singapore Management University
Algorithms For Constrained K-Nearest Neighbor Queries Over Moving Object Trajectories, Yunjun Gao, Baihua Zheng, Gencai Chen, Qing Li, Chun Chen
Research Collection School Of Computing and Information Systems
An important query for spatio-temporal databases is to find nearest trajectories of moving objects. Existing work on this topic focuses on the closest trajectories in the whole data space. In this paper, we introduce and solve constrained k-nearest neighbor (CkNN) queries and historical continuous CkNN (HCCkNN) queries on R-tree-like structures storing historical information about moving object trajectories. Given a trajectory set D, a query object (point or trajectory) q, a temporal extent T, and a constrained region CR, (i) a CkNN query over trajectories retrieves from D within T, the k (≥ 1) trajectories that lie closest to q and …
Do You Trust To Get Trust? A Study Of Trust Reciprocity Behaviors And Reciprocal Trust Prediction,
2010
Singapore Management University
Do You Trust To Get Trust? A Study Of Trust Reciprocity Behaviors And Reciprocal Trust Prediction, Viet-An Nguyen, Ee Peng Lim, Hwee Hoon Tan, Jing Jiang, Aixin Sun
Research Collection School Of Computing and Information Systems
Trust reciprocity, a special form of link reciprocity, exists in many networks of trust among users. In this paper, we seek to determine the extent to which reciprocity exists in a trust network and develop quantitative models for measuring reciprocity and reciprocity related behaviors. We identify several reciprocity behaviors and their respective measures. These behavior measures can be employed for predicting if a trustee will return trust to her trustor given that the latter initiates a trust link earlier. We develop for this reciprocal trust prediction task a number of ranking method and classification methods, and evaluated them on an …
Policy-Driven Distributed And Collaborative Demand Response In Multi-Domain Commercial Buildings,
2010
Singapore Management University
Policy-Driven Distributed And Collaborative Demand Response In Multi-Domain Commercial Buildings, Archan Misra, Henning Schulzrinne
Research Collection School Of Computing and Information Systems
Enabling a sophisticated Demand Response (DR) framework, whereby individual consumers adapt their electricity consumption in response to price variations, is a major objective of the emerging Smart Grid. We first point out why the current model, of EMS-based centralized control of a static repository of high load appliances, is inappropriate for supporting DR in future commercial buildings and campuses, where the consuming appliances are controlled by multiple users. To enable DR in such multi-domain environments, we envision a more collaborative and autonomous model, where a large set of heterogeneous smart electrical devices autonomously self-organize and negotiate their collective DR. Enabling …
Data Mining Based Predictive Models For Overall Health Indices,
2010
University of Minnesota - Twin Cities
Data Mining Based Predictive Models For Overall Health Indices, Ridhima Rajkumar, Kyong Jin Shim, Jaideep Srivastava
Research Collection School Of Computing and Information Systems
In this study, we infer health care indices of individuals using their pharmacy medical and prescription claims. Specifically, we focus on the widely used Charlson Index. We use data mining techniques to formulate the problem of classifying Charlson Index (CI) and build predictive models to predict individual health index score. First, we present comparative analyses of several classification algorithms. Second, our study shows that certain ensemble algorithms lead to higher prediction accuracy in comparison to base algorithms. Third, we introduce cost-sensitive learning to the classification algorithms and show that the inclusion of cost-sensitive learning leads to improved prediction accuracy. The …
Finding Influentials Based On The Temporal Order Of Information Adoption In Twitter,
2010
Singapore Management University
Finding Influentials Based On The Temporal Order Of Information Adoption In Twitter, Changhyun Lee, Haewoon Kwak, Hosung Park, Sue Moon
Research Collection School Of Computing and Information Systems
Twitter offers an explicit mechanism to facilitate information diffusion and has emerged as a new medium for communication. Many approaches to find influentials have been proposed, but they do not consider the temporal order of information adoption. In this work, we propose a novel method to find influentials by considering both the link structure and the temporal order of information adoption in Twitter. Our method finds distinct influentials who are not discovered by other methods.
Time Dependent Channel Packet Calculation Of Two Nucleon Scattering Matrix Elements,
2010
Air Force Institute of Technology
Time Dependent Channel Packet Calculation Of Two Nucleon Scattering Matrix Elements, Brian S. Davis
Theses and Dissertations
A new approach to calculating nucleon-nucleon scattering matrix elements using a proven atomic time-dependent wave packet technique is investigated. Wave packets containing centripetal barrier information are prepared in close proximity to nuclear well. This is accomplished by first using an analytic equation to determine the wave packets in a suitable intermediate asymptotic state where the centripetal barrier is negligible. Then, the split operator technique is used to propagate the wave packets back to their original positions under the full Hamiltonian. Here, one wave packet is held stationary while the other is allowed to evolve and explore the nuclear well. Scattering …
Software Internationalization: A Framework Validated Against Industry Requirements For Computer Science And Software Engineering Programs,
2010
California Polytechnic State University, San Luis Obispo
Software Internationalization: A Framework Validated Against Industry Requirements For Computer Science And Software Engineering Programs, John Huân Vũ
Master's Theses
View John Huân Vũ's thesis presentation at http://youtu.be/y3bzNmkTr-c.
In 2001, the ACM and IEEE Computing Curriculum stated that it was necessary to address "the need to develop implementation models that are international in scope and could be practiced in universities around the world." With increasing connectivity through the internet, the move towards a global economy and growing use of technology places software internationalization as a more important concern for developers. However, there has been a "clear shortage in terms of numbers of trained persons applying for entry-level positions" in this area. Eric Brechner, Director of Microsoft Development Training, suggested …
Top-K Aggregation Queries Over Large Networks,
2010
University of California at Santa Barbara, USA
Top-K Aggregation Queries Over Large Networks, Xifeng Yan, Bin He, Feida Zhu, Jiawei Han
Research Collection School Of Computing and Information Systems
Searching and mining large graphs today is critical to a variety of application domains, ranging from personalized recommendation in social networks, to searches for functional associations in biological pathways. In these domains, there is a need to perform aggregation operations on large-scale networks. Unfortunately the existing implementation of aggregation operations on relational databases does not guarantee superior performance in network space, especially when it involves edge traversals and joins of gigantic tables. In this paper, we investigate the neighborhood aggregation queries: Find nodes that have top-k highest aggregate values over their h-hop neighbors. While these basic queries are common in …
