Exclusive Lasso For Multi-Task Feature Selection,
2010
Michigan State University
Exclusive Lasso For Multi-Task Feature Selection, Yang Zhou, Rong Jin, Steven C. H. Hoi
Research Collection School Of Computing and Information Systems
We propose a novel group regularization which we call exclusive lasso. Unlike the group lasso regularizer that assumes co-varying variables in groups, the proposed exclusive lasso regularizer models the scenario when variables in the same group compete with each other. Analysis is presented to illustrate the properties of the proposed regularizer. We present a framework of kernel-based multi-task feature selection algorithm based on the proposed exclusive lasso regularizer. An efficient algorithm is derived to solve the related optimization problem. Experiments with document categorization show that our approach outperforms state-of-the-art algorithms for multi-task feature selection.
Two-View Transductive Support Vector Machines,
2010
Nanyang Technological University
Two-View Transductive Support Vector Machines, Guangxia Li, Steven C. H. Hoi, Kuiyu Chang
Research Collection School Of Computing and Information Systems
Obtaining high-quality and up-to-date labeled data can be difficult in many real-world machine learning applications, especially for Internet classification tasks like review spam detection, which changes at a very brisk pace. For some problems, there may exist multiple perspectives, so called views, of each data sample. For example, in text classification, the typical view contains a large number of raw content features such as term frequency, while a second view may contain a small but highly-informative number of domain specific features. We thus propose a novel two-view transductive SVM that takes advantage of both the abundant amount of unlabeled data …
Exploiting Query Logs For Cross-Lingual Query Suggestions.,
2010
Singapore Management University
Exploiting Query Logs For Cross-Lingual Query Suggestions., Wei Gao, Cheng Niu, Jian-Yun Nie, Ming Zhou, Kam-Fai Wong, Hsiao-Wuen Hon
Research Collection School Of Computing and Information Systems
Query suggestion aims to suggest relevant queries for a given query, which helps users better specify their information needs. Previous work on query suggestion has been limited to the same language. In this article, we extend it to cross-lingual query suggestion (CLQS): for a query in one language, we suggest similar or relevant queries in other languages. This is very important to the scenarios of cross-language information retrieval (CLIR) and other related cross-lingual applications. Instead of relying on existing query translation technologies for CLQS, we present an effective means to map the input query of one language to queries of …
Learning User Profiles For Personalized Information Dissemination,
2010
Singapore Management University
Learning User Profiles For Personalized Information Dissemination, Ah-Hwee Tan, Christine Teo
Research Collection School Of Computing and Information Systems
Personalized information systems represent the recent effort of delivering information to users more effectively in the modern electronic age. This paper illustrates how a supervised Adaptive Resonance Theory (ART) system, known as fuzzy ARAM, can be used to learn user profiles for personalized information dissemination. ARAM learning is on-line, fast, and incremental. Acquisition of new knowledge does not require re-training on previously learned cases. ARAM integrates both user-defined and system-learned knowledge in a single framework. Therefore inconsistency between the two knowledge sources will not arise. ARAM has been used to develop a personalized news system known as PIN. Preliminary experiments …
Understanding The Values Of Mobile Technology In Education: A Value-Focused Thinking Approach,
2010
Missouri University of Science and Technology
Understanding The Values Of Mobile Technology In Education: A Value-Focused Thinking Approach, Hong Sheng, Keng Siau, Fiona Nah
Research Collection School Of Computing and Information Systems
Mobile technology has mobilized the human interaction in all dimensions by supporting mobile collaboration. As collaboration is key to learning in today's educational environment, mobile technology has tremendous potential in supporting and improving education and its delivery. Given that mobile technology for education is a new phenomenon that is gaining popularity, the values of using mobile technology to support education need to be further researched and better understood. In this research, we used the Value-Focused Thinking approach to interview students and instructors to identify the values of education that are enabled by mobile technology. These values are represented in the …
A Comparative Study On Text Categorization,
2010
University of Nevada Las Vegas
A Comparative Study On Text Categorization, Aditya Chainulu Karamcheti
UNLV Theses, Dissertations, Professional Papers, and Capstones
Automated text categorization is a supervised learning task, defined as assigning category labels to new documents based on likelihood suggested by a training set of labeled documents. Two examples of methodology for text categorizations are Naive Bayes and K-Nearest Neighbor.
In this thesis, we implement two categorization engines based on Naive Bayes and K-Nearest Neighbor methodology. We then compare the effectiveness of these two engines by calculating standard precision and recall for a collection of documents. We will further report on time efficiency of these two engines.
Rapport: Semantic-Sensitive Namespace Management In Large-Scale File Systems,
2010
Huazhong Univ. of Sci. & Tech.
Rapport: Semantic-Sensitive Namespace Management In Large-Scale File Systems, Yu Hua, Hong Jiang, Yifeng Zhu, Dan Feng
School of Computing: Technical Reports
Explosive growth in volume and complexity of data exacerbates the key challenge to effectively and efficiently manage data in a way that fundamentally improves the ease and efficacy of their use. Existing large-scale file systems rely on hierarchically structured namespace that leads to severe performance bottlenecks and renders it impossible to support real-time queries on multi-dimensional attributes. This paper proposes a novel semantic-sensitive scheme, called Rapport, to provide dynamic and adaptive namespace management and support complex queries. The basic idea is to build files’ namespace by utilizing their semantic correlation and exploiting dynamic evolution of attributes to support namespace management. …
Trust In Social And Sensor Networks,
2010
Wright State University - Main Campus
Trust In Social And Sensor Networks, Pramod Anantharam, Krishnaprasad Thirunarayan, Cory Andrew Henson, Amit P. Sheth
Kno.e.sis Publications
Trust can be defined as the perception of the trustor about the degree to which the trustee would satisfy an expectation about a transaction constituting risk. Trust plays a pivotal role when the risk in believing incorrect information is high. With Web 2.0 where user generated content and real time interactions dominate, the openness of data contribution may hinder the quality of information we can get.
What Goes Around Comes Around - Improving Linked Open Data Through On-Demand Model Creation,
2010
Wright State University - Main Campus
What Goes Around Comes Around - Improving Linked Open Data Through On-Demand Model Creation, Christopher Thomas, Wenbo Wang, Pankaj Mehra, Delroy H. Cameron, Pablo N. Mendes, Amit P. Sheth
Kno.e.sis Publications
Web 2.0 has changed the way we share and keep up with information. We communicate through social media platforms and make the information we exchange to a large extent publicly available. Linked Open Data (LOD) follows the same paradigm of sharing information but also makes it machine accessible. LOD provides an abundance of structured information albeit in a less formally rigorous form than would be desirable for Semantic Web applications. Nevertheless, most of the LOD assertions are community reviewed and we can rely on their accuracy to a large extent. In this work we want to follow the Web 2.0 …
Information Technology Implementation Decisions To Support The Kentucky Mesonet,
2010
Western Kentucky University
Information Technology Implementation Decisions To Support The Kentucky Mesonet, D. Michael Grogan
Masters Theses & Specialist Projects
The Kentucky Mesonet is a high-density, mesoscale network of automated meteorological and climatological sensing platforms being developed across the commonwealth. Data communications, collection, processing, and delivery mechanisms play a critical role in such networks, and the World Meteorological Organization recognizes that “an observing system is not complete unless it is connected to other systems that deliver the data to the users.” This document reviews the implementation steps, decisions, and rationale surrounding communications and computing infrastructure development to support the Mesonet. A general overview of the network and technology-related research is provided followed by a review of pertinent literature related to …
Dynamic Associative Relationships On The Linked Open Data Web,
2010
Wright State University - Main Campus
Dynamic Associative Relationships On The Linked Open Data Web, Pablo N. Mendes, Pavan Kapanipathi, Delroy H. Cameron, Amit P. Sheth
Kno.e.sis Publications
We provide a definition of context based on theme, time and location, and propose a mixed retrieval/extraction model for the dynamic suggestion of trending relationships to LOD resources.
Semantics-Empowered Text Exploration For Knowledge Discovery,
2010
Wright State University - Main Campus
Semantics-Empowered Text Exploration For Knowledge Discovery, Delroy H. Cameron, Pablo N. Mendes, Amit P. Sheth, Victor Chan
Kno.e.sis Publications
The interaction paradigm offered by most contemporary Web Information Systems is a search-and-sift paradigm in which users manually seek information using hyperlinked documents. This paradigm is derived from a document-centric model that gives users minimal support for scanning through high volumes of text. We present a novel information exploration paradigm based on a data-centric view of corpora, along with a prototype implementation that demonstrates the value in content-driven navigation. We leverage semantic metadata to link data in documents by exploiting named relationships between entities. We also present utilities for gathering user generated navigation trails, critical for knowledge discovery. We discuss …
A Social Network Based Study Of Software Team Dynamics,
2010
Singapore Management University
A Social Network Based Study Of Software Team Dynamics, Subhajit Datta, Vikrant S. Kaulgoud, Vibhu Saujanya Sharma, Nishant Kumar
Research Collection School Of Computing and Information Systems
Members of software project teams have specific roles and responsibilities which are formally defined during project inception or at the start of a life cycle activity. Often, the team structure undergoes spontaneous changes as delivery deadlines draw near and critical tasks have to be completed. Some members -- depending on their skill or seniority -- need to take on more responsibilities, while others end up being peripheral to the project's execution. We posit that this kind of ad hoc reorganization of a team's structure can be discerned from the project's bug tracker. In this paper, we extract a social network …
A Pda Intervention To Sustain Smoking Cessation In Clients With Socioeconomic Vulnerability,
2010
University of Nebraska Medical Center
A Pda Intervention To Sustain Smoking Cessation In Clients With Socioeconomic Vulnerability, Lynne Buchanan, Deepak Khazanchi
Information Systems and Quantitative Analysis Faculty Publications
This article describes a pilot study to explore use of a personal digital assistant (PDA) to sustain smoking cessation after discharge in clients with socioeconomic vulnerability. The major aim is to describe technology acceptance (perceived ease of use, usefulness, and attitude), portability, technical difficulty, satisfaction, and use time. The sample includes 31 medical surgical clients with average age of 47.35 (±13.3), average household income of $13,629 (±8,204), average number in the household of 2.67 (±2.22), and average education of 11th grade. The results demonstrate mean use time of 9.28 (±3.23) hr, or about 1 hr over 8 weeks. Technology acceptance …
Optimal Matching Between Spatial Datasets Under Capacity Constraints,
2010
University of Hong Kong
Optimal Matching Between Spatial Datasets Under Capacity Constraints, Hou U Leong, Kyriakos Mouratidis, Man Lung Yiu, Nikos Mamoulis
Research Collection School Of Computing and Information Systems
Consider a set of customers (e.g., WiFi receivers) and a set of service providers (e.g., wireless access points), where each provider has a capacity and the quality of service offered to its customers is anti-proportional to their distance. The capacity constrained assignment (CCA) is a matching between the two sets such that (i) each customer is assigned to at most one provider, (ii) every provider serves no more customers than its capacity, (iii) the maximum possible number of customers are served, and (iv) the sum of Euclidean distances within the assigned provider-customer pairs is minimized. Although max-flow algorithms are applicable …
Generating Synonyms Based On Query Log Data,
2010
Singapore Management University
Generating Synonyms Based On Query Log Data, Stelios Paparizos, Tao Cheng, Hady W. Lauw
Research Collection School Of Computing and Information Systems
An approach is described for generating synonyms to supplement at least one information item, such as, in one case, a set of related items. The approach can involve an expansion phase, a clean-up phase, and a reduction phase. In the expansion phase, the approach identifies, for each related item, a set of initial synonym candidates. In the clean-up phase, the approach removes noise from the set of initial synonym candidates (if such noise exists), to provide a set of filtered synonym candidate items. In the reduction phase, the approach ranks and applies a threshold (or thresholds) to the set of …
Adaptive Ensemble Classification In P2p Networks,
2010
Nanyang Technological University
Adaptive Ensemble Classification In P2p Networks, Hock Hee Ang, Vivekanand Gopalkrishnan, Steven C. H. Hoi, Wee Keong Ng
Research Collection School Of Computing and Information Systems
Classification in P2P networks has become an important research problem in data mining due to the popularity of P2P computing environments. This is still an open difficult research problem due to a variety of challenges, such as non-i.i.d. data distribution, skewed or disjoint class distribution, scalability, peer dynamism and asynchronism. In this paper, we present a novel P2P Adaptive Classification Ensemble (PACE) framework to perform classification in P2P networks. Unlike regular ensemble classification approaches, our new framework adapts to the test data distribution and dynamically adjusts the voting scheme by combining a subset of classifiers/peers according to the test data …
Continuous Spatial Assignment Of Moving Users,
2010
University of Hong Kong
Continuous Spatial Assignment Of Moving Users, Hou U Leong, Kyriakos Mouratidis, Nikos Mamoulis
Research Collection School Of Computing and Information Systems
Consider a set of servers and a set of users, where each server has a coverage region (i.e., an area of service) and a capacity (i.e., a maximum number of users it can serve). Our task is to assign every user to one server subject to the coverage and capacity constraints. To offer the highest quality of service, we wish to minimize the average distance between users and their assigned server. This is an instance of a well-studied problem in operations research, termed optimal assignment. Even though there exist several solutions for the static case (where user locations are fixed), …
Efficient Skyline Maintenance For Streaming Data With Partially-Ordered Domains,
2010
Singapore Management University
Efficient Skyline Maintenance For Streaming Data With Partially-Ordered Domains, Yuan Fang, Chee-Yong Chan
Research Collection School Of Computing and Information Systems
We address the problem of skyline query processing for a count-based window of continuous streaming data that involves both totally- and partially-ordered attribute domains. In this problem, a fixed-size buffer of the N most recent tuples is dynamically maintained and the key challenge is how to efficiently maintain the skyline of the sliding window of N tuples as new tuples arrive and old tuples expire. We identify the limitations of the state-of-the-art approach STARS, and propose two new approaches, STARS+ and SkyGrid, to address its drawbacks. STARS+ is an enhancement of STARS with three new optimization techniques, while SkyGrid is …
Data Mining Based Predictive Models For Overall Health Indices,
2010
University of Minnesota - Twin Cities
Data Mining Based Predictive Models For Overall Health Indices, Ridhima Rajkumar, Kyong Jin Shim, Jaideep Srivastava
Research Collection School Of Computing and Information Systems
In this study, we infer health care indices of individuals using their pharmacy medical and prescription claims. Specifically, we focus on the widely used Charlson Index. We use data mining techniques to formulate the problem of classifying Charlson Index (CI) and build predictive models to predict individual health index score. First, we present comparative analyses of several classification algorithms. Second, our study shows that certain ensemble algorithms lead to higher prediction accuracy in comparison to base algorithms. Third, we introduce cost-sensitive learning to the classification algorithms and show that the inclusion of cost-sensitive learning leads to improved prediction accuracy. The …
