Open Access. Powered by Scholars. Published by Universities.®

2009

Discipline
Institution
Keyword
Publication
Publication Type

Articles 1 - 30 of 73

Full-Text Articles in Numerical Analysis and Scientific Computing

Sos: Searching Help Pages Of R Packages, Spencer Graves, Sundar Dorai-Raj, Romain François Dec 2009

Sos: Searching Help Pages Of R Packages, Spencer Graves, Sundar Dorai-Raj, Romain François

The R Journal

The sos package provides a means to quickly and flexibly search the help pages of contributed packages, finding functions and datasets in seconds or minutes that could not be found in hours or days by any other means we know. Its findFn function accesses Jonathan Baron’s R Site Search database and returns the matches in a data frame of class "findFn", which can be further manipulated by other sos functions to produce, for example, an Excel file that starts with a summary sheet that makes it relatively easy to prioritize alternative packages for further study. As such, it provides a …


Rattle: A Data Mining Gui For R, Graham J. Williams Dec 2009

Rattle: A Data Mining Gui For R, Graham J. Williams

The R Journal

Data mining delivers insights, pat terns, and descriptive and predictive models from the large amounts of data available today in many organisations. The data miner draws heavily on methodologies, techniques and algorithms from statistics, machine learning, and computer science. R increasingly provides a powerful platform for data mining. However, scripting and programming is sometimes a challenge for data analysts moving into data mining. The Rattle package provides a graphical user interface specifically for data mining using R. It also provides a stepping stone toward using R as a programming language for data analysis.


Copas: An R Package For Fitting The Copas Selection Model, J. Carpenter, G. Rücker, G. Schhwarzer Dec 2009

Copas: An R Package For Fitting The Copas Selection Model, J. Carpenter, G. Rücker, G. Schhwarzer

The R Journal

This article describes the R package copas which is an add-on package to the R pack age meta. The R package copas can be used to f it the Copas selection model to adjust for bias in meta-analysis. A clinical example is used to illustrate fitting and interpreting the Copas selection model.


Party On!, Carolin Strobl, Torsten Hothorn, Achim Zeileis Dec 2009

Party On!, Carolin Strobl, Torsten Hothorn, Achim Zeileis

The R Journal

Random forests are one of the most popular statistical learning algorithms, and a variety of methods for fitting random forests and related recursive partitioning approaches is available in R. This paper points out two important features of the random forest implementation cforest available in the party package: The resulting forests are unbiased and thus prefer able to the randomForest implementation avail able in randomForest if predictor variables are of different types. Moreover, a conditional per mutation importance measure has recently been added to the party package, which can help evaluate the importance of correlated predictor variables. The rationale of this …


Aspects Of The Social Organization And Trajectory Of The R Project, John Fox Dec 2009

Aspects Of The Social Organization And Trajectory Of The R Project, John Fox

The R Journal

Based partly on interviews with members of the R Core team, this paper considers the development of the R Project in the context of open-source software development and, more generally, voluntary activities. The paper de scribes aspects of the social organization of the R Project, including the organization of the R Core team; describes the trajectory of the R Project; seeks to identify factors crucial to the success of R; and speculates about the prospects for R.


Asymptest: A Simple R Package For Classical Parametric Statistical Tests And Confidence Intervals In Large Samples, J.-F. Coeurjolly, R. Drouilhet, P. Lafaye De Micheaux, J.-F. Robineau Dec 2009

Asymptest: A Simple R Package For Classical Parametric Statistical Tests And Confidence Intervals In Large Samples, J.-F. Coeurjolly, R. Drouilhet, P. Lafaye De Micheaux, J.-F. Robineau

The R Journal

asympTest is an R package implementing large sample tests and confidence intervals. One and two sample mean and variance tests (differences and ratios) are considered. The test statistics are all expressed in the same form as the Student t-test, which facilitates their presentation in the classroom. This contribution also fills the gap of a robust (to non-normality) alternative to the chi-square single variance test for large samples, since no such procedure is implemented in standard statistical software.


Convergenceconcepts: An R Package To Investigate Various Modes Of Convergence, Pierre Lafaye De Micheaux, Benoit Liquet Dec 2009

Convergenceconcepts: An R Package To Investigate Various Modes Of Convergence, Pierre Lafaye De Micheaux, Benoit Liquet

The R Journal

ConvergenceConcepts is an R pack age, built upon the tkrplot, tcltk and lattice packages, designed to investigate the convergence of simulated sequences of random variables. Four classical modes of convergence may be studied, namely: almost sure convergence (a.s.), convergence in probability (P), convergence in law (L) and convergence in r-th mean (r). This investigation is performed through ac curate graphical representations. This package may be used as a pedagogical tool. It may give students a better understanding of these notions and help them to visualize these difficult theoretical concepts. Moreover, …


The R Journal (December 2009) 1(2): Complete Issue, The R Foundation Dec 2009

The R Journal (December 2009) 1(2): Complete Issue, The R Foundation

The R Journal

Contributed Research Articles

Aspects of the Social Organization and Trajectory of the R Project, John Fox

Party on! Carolin Strobl, Torsten Hothorn and Achim Zeileis

ConvergenceConcepts: An R Package to Investigate Various Modes of Convergence, Pierre Lafaye de Micheaux and Benoit Liquet

asympTest: A Simple R Package for Classical Parametric Statistical Tests and Confidence Intervals in Large Samples, J.-F. Coeurjolly, R. Drouilhet, P. Lafaye de Micheaux and J.-F. Robineau

copas: An R package for Fitting the Copas Selection Model, J. Carpenter, G. Rücker and G. Schwarzer

Transitioning to R: Replicating SAS, Stata, and SUDAAN Analysis Techniques in Health Policy Data, …


On Strategies For Imbalanced Text Classification Using Svm: A Comparative Study, Aixin Sun, Ee Peng Lim, Ying Liu Dec 2009

On Strategies For Imbalanced Text Classification Using Svm: A Comparative Study, Aixin Sun, Ee Peng Lim, Ying Liu

Research Collection School Of Computing and Information Systems

Many real-world text classification tasks involve imbalanced training examples. The strategies proposed to address the imbalanced classification (e.g., resampling, instance weighting), however, have not been systematically evaluated in the text domain. In this paper, we conduct a comparative study on the effectiveness of these strategies in the context of imbalanced text classification using Support Vector Machines (SVM) classifier. SVM is the interest in this study for its good classification accuracy reported in many text classification tasks. We propose a taxonomy to organize all proposed strategies following the training and the test phases in text classification tasks. Based on the taxonomy, …


Adaptive Type-2 Fuzzy Maintenance Advisor For Offshore Power Systems, Zhaoxia Wang, C. S. Chang, Fan Yang, W. W. Tan Dec 2009

Adaptive Type-2 Fuzzy Maintenance Advisor For Offshore Power Systems, Zhaoxia Wang, C. S. Chang, Fan Yang, W. W. Tan

Research Collection School Of Computing and Information Systems

Proper maintenance strategies are very desirable for minimizing the operational and maintenance costs of power systems without sacrificing reliability. Condition-based maintenance has largely replaced time-based maintenance because of the former's potential economic benefits. As offshore substations are often remotely located, they experience more adverse environments, higher failures, and therefore need more powerful analytical tools than their onshore counterpart. As reliability information collected during operation of an offshore substation can rarely avoid uncertainties, it is essential to obtain consistent estimates of reliability measures under changing environmental and operating conditions. Some attempts with type-1 fuzzy logic were made with limited success in …


To Trust Or Not To Trust? Predicting Online Trusts Using Trust Antecedent Framework, Viet-An Nguyen, Ee Peng Lim, Jing Jiang, Aixin Sun Dec 2009

To Trust Or Not To Trust? Predicting Online Trusts Using Trust Antecedent Framework, Viet-An Nguyen, Ee Peng Lim, Jing Jiang, Aixin Sun

Research Collection School Of Computing and Information Systems

This paper analyzes the trustor and trustee factors that lead to inter-personal trust using a well studied Trust Antecedent framework in management science. To apply these factors to trust ranking problem in online rating systems, we derive features that correspond to each factor and develop different trust ranking models. The advantage of this approach is that features relevant to trust can be systematically derived so as to achieve good prediction accuracy. Through a series of experiments on real data from Epinions, we show that even a simple model using the derived features yields good accuracy and outperforms MoleTrust, a trust …


Trust-Oriented Composite Services Selection And Discovery, Lei Li, Yan Wang, Ee Peng Lim Nov 2009

Trust-Oriented Composite Services Selection And Discovery, Lei Li, Yan Wang, Ee Peng Lim

Research Collection School Of Computing and Information Systems

In Service-Oriented Computing (SOC) environments, service clients interact with service providers for consuming services. From the viewpoint of service clients, the trust level of a service or a service provider is a critical issue to consider in service selection and discovery, particularly when a client is looking for a service from a large set of services or service providers. However, a service may invoke other services offered by different providers forming composite services. The complex invocations in composite services greatly increase the complexity of trust-oriented service selection and discovery. In this paper, we propose novel approaches for composite service representation, …


What Makes Categories Difficult To Classify?, Aixin Sun, Ee Peng Lim, Ying Liu Nov 2009

What Makes Categories Difficult To Classify?, Aixin Sun, Ee Peng Lim, Ying Liu

Research Collection School Of Computing and Information Systems

In this paper, we try to predict which category will be less accurately classified compared with other categories in a classification task that involves multiple categories. The categories with poor predicted performance will be identified before any classifiers are trained and additional steps can be taken to address the predicted poor accuracies of these categories. Inspired by the work on query performance prediction in ad-hoc retrieval, we propose to predict classification performance using two measures, namely, category size and category coherence. Our experiments on 20-Newsgroup and Reuters-21578 datasets show that the Spearman rank correlation coefficient between the predicted rank of …


Trust Relationship Prediction Using Online Product Review Data, Nan Ma, Ee Peng Lim, Viet-An Nguyen, Aixin Sun Nov 2009

Trust Relationship Prediction Using Online Product Review Data, Nan Ma, Ee Peng Lim, Viet-An Nguyen, Aixin Sun

Research Collection School Of Computing and Information Systems

Trust between users is an important piece of knowledge that can be exploited in search and recommendation.Given that user-supplied trust relationships are usually very sparse, we study the prediction of trust relationships using user interaction features in an online user generated review application context. We show that trust relationship prediction can achieve better accuracy when one adopts personalized and cluster-based classification methods. The former trains one classifier for each user using user-specific training data. The cluster-based method first constructs user clusters before training one classifier for each user cluster. Our proposed methods have been evaluated in a series of experiments …


Udel/Smu At Trec 2009 Entity Track, Wei Zheng, Swapna Gottipati, Jing Jiang, Hui Fang Nov 2009

Udel/Smu At Trec 2009 Entity Track, Wei Zheng, Swapna Gottipati, Jing Jiang, Hui Fang

Research Collection School Of Computing and Information Systems

We report our methods and experiment results from the collaborative participation of the InfoLab group from University of Delaware and the school of Information Systems from Singapore Management University in the TREC 2009 Entity track. Our general goal is to study how we may apply language modeling approaches and natural language processing techniques to the task. Specically, we proposed to find supporting information based on segment retrieval, to extract entities using Stanford NER tagger, and to rank entities based on a previously proposed probabilistic framework for expert finding.


Mining Communities In Networks: A Solution For Consistency And Its Evaluation, Haewoon Kwak, Yoonchan Choi, Young-Ho Eom, Hawoong Jeong, Sue Moon Nov 2009

Mining Communities In Networks: A Solution For Consistency And Its Evaluation, Haewoon Kwak, Yoonchan Choi, Young-Ho Eom, Hawoong Jeong, Sue Moon

Research Collection School Of Computing and Information Systems

Online social networks pose significant challenges to computer scientists, physicists, and sociologists alike, for their massive size, fast evolution, and uncharted potential for social computing. One particular problem that has interested us is community identification. Many algorithms based on various metrics have been proposed for communities in networks [18, 24], but a few algorithms scale to very large networks. Three recent community identification algorithms, namely CNM [16], Wakita [59], and Louvain [10], stand out for their scalability to a few millions of nodes. All of them use modularity as the metric of optimization. However, all three algorithms produce inconsistent communities …


Parallel Sets In The Real World: Three Case Studies, Robert Kosara, Caroline Ziemkiewicz, F. Joseph Iii Mako, Tin Seong Kam Oct 2009

Parallel Sets In The Real World: Three Case Studies, Robert Kosara, Caroline Ziemkiewicz, F. Joseph Iii Mako, Tin Seong Kam

Research Collection School Of Computing and Information Systems

Parallel Sets are a visualization technique for categorical data. We recently released an implementation to the public in an effort to make our research useful to real users. This paper presents three case studies of Parallel Sets in use with real data.


Continuous Monitoring Of Spatial Queries In Wireless Broadcast Environments, Kyriakos Mouratidis, Spiridon Bakiras, Dimitris Papadias Oct 2009

Continuous Monitoring Of Spatial Queries In Wireless Broadcast Environments, Kyriakos Mouratidis, Spiridon Bakiras, Dimitris Papadias

Research Collection School Of Computing and Information Systems

Wireless data broadcast is a promising technique for information dissemination that leverages the computational capabilities of the mobile devices in order to enhance the scalability of the system. Under this environment, the data are continuously broadcast by the server, interleaved with some indexing information for query processing. Clients may then tune in the broadcast channel and process their queries locally without contacting the server. Previous work on spatial query processing for wireless broadcast systems has only considered snapshot queries over static data. In this paper, we propose an air indexing framework that 1) outperforms the existing (i.e., snapshot) techniques in …


A Surprise Triggered Adaptive And Reactive (Star) Framework For Online Adaptation In Non-Stationary Environments, Truong-Huy Dinh Nguyen, Tze-Yun Leong Oct 2009

A Surprise Triggered Adaptive And Reactive (Star) Framework For Online Adaptation In Non-Stationary Environments, Truong-Huy Dinh Nguyen, Tze-Yun Leong

Research Collection School Of Computing and Information Systems

We consider the task of developing an adaptive autonomous agent that can interact with non-stationary environments. Traditional learning approaches such as Reinforcement Learning assume stationary characteristics over the course of the problem, and are therefore unable to learn the dynamically changing settings correctly. We introduce a novel adaptive framework that can detect dynamic changes due to non-stationary elements. The Surprise Triggered Adaptive and Reactive (STAR) framework is inspired by human adaptability in dealing with daily life changes. An agent adopting the STAR framework consists primarily of two components, Adapter and Reactor. The Reactor chooses suitable actions based on predictions made …


Monte Carlo Model Of Light Propagation In Tissues And The Effects Of Phase Changes On The Light Intensity, Rakesh Choula Oct 2009

Monte Carlo Model Of Light Propagation In Tissues And The Effects Of Phase Changes On The Light Intensity, Rakesh Choula

Electrical & Computer Engineering Theses & Dissertations

Lasers, due to their unique properties, have a wide range of applications in the medical field. For accurate laser treatments that focus on bio-tissues, prior knowledge of the amount of laser power, spot size and its irradiation time are necessary. In order to predict the effects of lasers on tissues and their bio-effects, a first necessary step is the creation of a model that can predict the temperature distributions within the tissue following laser excitation. This involves modeling light propagation through the tissue with inclusion of internal scattering, and assessment of the energy deposited by the incoming photons. The next …


Visible Reverse K-Nearest Neighbor Query Processing In Spatial Databases, Yunjun Gao, Baihua Zheng, Gencai Chen, Wang-Chien Lee, Ken C. K. Lee, Qing Li Sep 2009

Visible Reverse K-Nearest Neighbor Query Processing In Spatial Databases, Yunjun Gao, Baihua Zheng, Gencai Chen, Wang-Chien Lee, Ken C. K. Lee, Qing Li

Research Collection School Of Computing and Information Systems

Reverse nearest neighbor (RNN) queries have a broad application base such as decision support, profile-based marketing, resource allocation, etc. Previous work on RNN search does not take obstacles into consideration. In the real world, however, there are many physical obstacles (e.g., buildings) and their presence may affect the visibility between objects. In this paper, we introduce a novel variant of RNN queries, namely, visible reverse nearest neighbor (VRNN) search, which considers the impact of obstacles on the visibility of objects. Given a data set P, an obstacle set O, and a query point q in a 2D space, a VRNN …


Detecting Automotive Exhaust Gas Based On Fuzzy Inference System, Li. Shujin, Ming Bai, Quan Wang, Bo Chen, Xiaobing Zhao, Ting Yang, Zhaoxia Wang Sep 2009

Detecting Automotive Exhaust Gas Based On Fuzzy Inference System, Li. Shujin, Ming Bai, Quan Wang, Bo Chen, Xiaobing Zhao, Ting Yang, Zhaoxia Wang

Research Collection School Of Computing and Information Systems

This paper proposes a method of detecting automotive exhaust gas based on fuzzy logic inference after analyzing the principle of the infrared automobile exhaust gas analyzer and the influence of the environmental temperature on analyzer. This paper analyses the measurement error caused by environmental temperature, and then makes a non-linear error correction of temperature for the infrared sensor using fuzzy inference. The results of simulation have clearly demonstrated that the proposed fuzzy compensation scheme is better than the non-fuzzy method.


Inferring Player Rating From Performance Data In Massively Multiplayer Online Role-Playing Games (Mmorpgs), Kyong Jin Shim, Muhammad Aurangzeb Ahmad, Nishith Pathak, Jaideep Srivastava Aug 2009

Inferring Player Rating From Performance Data In Massively Multiplayer Online Role-Playing Games (Mmorpgs), Kyong Jin Shim, Muhammad Aurangzeb Ahmad, Nishith Pathak, Jaideep Srivastava

Research Collection School Of Computing and Information Systems

This paper examines online player performance in EverQuest II, a popular massively multiplayer online role-playing game (MMORPG) developed by Sony Online Entertainment. The study uses the game's player performance data to devise performance metrics for online players. We report three major findings. First, we show that the game's point-scaling system overestimates performances of lower level players and underestimates performances of higher level players. We present a novel point-scaling system based on the game's player performance data that addresses the underestimation and overestimation problems. Second, we present a highly accurate predictive model for player performance as a function of past behavior. …


Essential Spreadsheet Modeling Course For Business Students, Thin Yin Leong, Michelle L. F. Cheong Aug 2009

Essential Spreadsheet Modeling Course For Business Students, Thin Yin Leong, Michelle L. F. Cheong

Research Collection School Of Computing and Information Systems

Ask any student at the Singapore Management University (SMU) to name one of the most practical and useful courses offered by the university. The answer would inevitably include CAT. CAT stands for the "Computer as an Analysis Tool" course. Originally based on a course of the same title offered by the Wharton Business School, the focus of CAT was shifted to provide business students the essential practical skills and necessary "real-world" exposure to better use personal computers for resolving business problems. The course is basically centered on using the Excel spreadsheet to work on ambiguous ill-defined problems.


Multi-Task Transfer Learning For Weakly-Supervised Relation Extraction, Jing Jiang Aug 2009

Multi-Task Transfer Learning For Weakly-Supervised Relation Extraction, Jing Jiang

Research Collection School Of Computing and Information Systems

Creating labeled training data for relation extraction is expensive. In this paper, we study relation extraction in a special weakly-supervised setting when we have only a few seed instances of the target relation type we want to extract but we also have a large amount of labeled instances of other relation types. Observing that different relation types can share certain common structures, we propose to use a multi-task learning method coupled with human guidance to address this weakly-supervised relation extraction problem. The proposed framework models the commonality among different relation types through a shared weight vector, enables knowledge learned from …


Ssnetviz: A Visualization Engine For Heterogeneous Semantic Social Networks, Ee Peng Lim, Maureen Maureen, Nelman Lubis Ibrahim, Aixin Sun, Anwitaman Datta, Kuiyu Chang Aug 2009

Ssnetviz: A Visualization Engine For Heterogeneous Semantic Social Networks, Ee Peng Lim, Maureen Maureen, Nelman Lubis Ibrahim, Aixin Sun, Anwitaman Datta, Kuiyu Chang

Research Collection School Of Computing and Information Systems

SSnetViz is an ongoing research to design and implement a visualization engine for heterogeneous semantic social networks. A semantic social network is a multi-modal network that contains nodes representing di®erent types of people or object entities, and edges representing relationships among them. When multiple heterogeneous semantic social networks are to be visualized together, SSnetViz provides a suite of functions to store heterogeneous semantic social networks, to integrate them for searching and analysis. We will illustrate these functions using social networks related to terrorism research, one crafted by domain experts and another from Wikipedia.


Optimal-Location-Selection Query Processing In Spatial Databases, Yunjun Gao, Baihua Zheng, Gencai Chen, Qing Li Aug 2009

Optimal-Location-Selection Query Processing In Spatial Databases, Yunjun Gao, Baihua Zheng, Gencai Chen, Qing Li

Research Collection School Of Computing and Information Systems

This paper introduces and solves a novel type of spatial queries, namely, Optimal-Location-Selection (OLS) search, which has many applications in real life. Given a data object set D_A, a target object set D_B, a spatial region R, and a critical distance d_c in a multidimensional space, an OLS query retrieves those target objects in D_B that are outside R but have maximal optimality. Here, the optimality of a target object b \in D_B located outside R is defined as the number of the data objects from D_A that are inside R and meanwhile have their distances to b not exceeding …


Scalable Verification For Outsourced Dynamic Databases, Hwee Hwa Pang, Jilian Zhang, Kyriakos Mouratidis Aug 2009

Scalable Verification For Outsourced Dynamic Databases, Hwee Hwa Pang, Jilian Zhang, Kyriakos Mouratidis

Research Collection School Of Computing and Information Systems

Query answers from servers operated by third parties need to be verified, as the third parties may not be trusted or their servers may be compromised. Most of the existing authentication methods construct validity proofs based on the Merkle hash tree (MHT). The MHT, however, imposes severe concurrency constraints that slow down data updates. We introduce a protocol, built upon signature aggregation, for checking the authenticity, completeness and freshness of query answers. The protocol offers the important property of allowing new data to be disseminated immediately, while ensuring that outdated values beyond a pre-set age can be detected. We also …


A Distributed Spatial Index For Error-Prone Wireless Data Broadcast, Baihua Zheng, Wang-Chien Lee, Ken C. K. Lee, Dik Lun Lee, Min Shao Aug 2009

A Distributed Spatial Index For Error-Prone Wireless Data Broadcast, Baihua Zheng, Wang-Chien Lee, Ken C. K. Lee, Dik Lun Lee, Min Shao

Research Collection School Of Computing and Information Systems

Information is valuable to users when it is available not only at the right time but also at the right place. To support efficient location-based data access in wireless data broadcast systems, a distributed spatial index (called DSI) is presented in this paper. DSI is highly efficient because it has a linear yet fully distributed structure that naturally shares links in different search paths. DSI is very resilient to the error-prone wireless communication environment because interrupted search operations based on DSI can be resumed easily. It supports search algorithms for classical location-based queries such as window queries and kNN queries …


On Efficient Mutual Nearest Neighbor Query Processing In Spatial Databases, Yunjun Gao, Baihua Zheng, Gencai Chen, Qing Li Aug 2009

On Efficient Mutual Nearest Neighbor Query Processing In Spatial Databases, Yunjun Gao, Baihua Zheng, Gencai Chen, Qing Li

Research Collection School Of Computing and Information Systems

This paper studies a new form of nearest neighbor queries in spatial databases, namely, mutual nearest neighbour (MNN) search. Given a set D of objects and a query object q, an MNN query returns from D, the set of objects that are among the k1 (≥ 1) nearest neighbors (NNs) of q; meanwhile, have q as one of their k2(≥ 1) NNs. Although MNN queries are useful in many applications involving decision making, data mining, and pattern recognition, it cannot be efficiently handled by existing spatial query processing approaches. In this paper, we present …