Generating Templates Of Entity Summaries With An Entity-Aspect Model And Pattern Mining,
2010
Shanghai Jiaotong University
Generating Templates Of Entity Summaries With An Entity-Aspect Model And Pattern Mining, Peng Li, Jing Jiang, Yinglin Wang
Research Collection School Of Computing and Information Systems
In this paper, we propose a novel approach to automatic generation of summary templates from given collections of summary articles. This kind of summary templates can be useful in various applications. We first develop an entity-aspect LDA model to simultaneously cluster both sentences and words into aspects. We then apply frequent subtree pattern mining on the dependency parse trees of the clustered and labeled sentences to discover sentence patterns that well represent the aspects. Key features of our method include automatic grouping of semantically related sentence patterns and automatic identification of template slots that need to be filled in. We …
Extracting Common Emotions From Blogs Based On Fine-Grained Sentiment Clustering,
2010
Northeastern University
Extracting Common Emotions From Blogs Based On Fine-Grained Sentiment Clustering, Shi Feng, Daling Wang, Ge Yu, Wei Gao, Kam-Fai Wong
Research Collection School Of Computing and Information Systems
Recently, blogs have emerged as the major platform for people to express their feelings and sentiments in the age of Web 2.0. The common emotions, which reflect people’s collective and overall sentiments, are becoming the major concern for governments, business companies and individual users. Different from previous literatures on sentiment classification and summarization, the major issue of common emotion extraction is to find out people’s collective sentiments and their corresponding distributions on the Web. Most existing blog clustering methods take into account keywords, stories or timelines but neglect the embedded sentiments, which are considered very important features of blogs. In …
Hybrid Time-Frequency Domain Analysis For Inverter-Fed Induction Motor Fault Detection,
2010
Singapore Management University
Hybrid Time-Frequency Domain Analysis For Inverter-Fed Induction Motor Fault Detection, T. W. Chua, W. W. Tan, Zhaoxia Wang, C. S. Chang
Research Collection School Of Computing and Information Systems
The detection of faults in an induction motor is important as a part of preventive maintenance. Stator current is one of the most popular signals used for utility-supplied induction motor fault detection as a current sensor can be installed nonintrusively. In variable speeds operation, the use of an inverter to drive the induction motor introduces noise into the stator current so stator current based fault detection techniques become less reliable. This paper presents a hybrid algorithm, which combines time and frequency domain analysis, for broken rotor bar and bearing fault detection. Cluster information obtained by using Independent Component Analysis (ICA) …
Effective Music Tagging Through Advanced Statistical Modeling,
2010
Singapore Management University
Effective Music Tagging Through Advanced Statistical Modeling, Jialie Shen, Meng Wang, Shuicheng Yan, Hwee Hwa Pang, Xian-Sheng Hua
Research Collection School Of Computing and Information Systems
Music information retrieval (MIR) holds great promise as a technology for managing large music archives. One of the key components of MIR that has been actively researched into is music tagging. While significant progress has been achieved, most of the existing systems still adopt a simple classification approach, and apply machine learning classifiers directly on low level acoustic features. Consequently, they suffer the shortcomings of (1) poor accuracy, (2) lack of comprehensive evaluation results and the associated analysis based on large scale datasets, and (3) incomplete content representation, arising from the lack of multimodal and temporal information integration. In this …
Hybrid Time-Frequency Domain Analysis For Inverter-Fed Induction Motor Fault Detection,
2010
Singapore Management University
Hybrid Time-Frequency Domain Analysis For Inverter-Fed Induction Motor Fault Detection, T. W. Chua, W. W. Tan, Zhaoxia Wang, C. S. Chang
Research Collection School Of Computing and Information Systems
The detection of faults in an induction motor is important as a part of preventive maintenance. Stator current is one of the most popular signals used for utility-supplied induction motor fault detection as a current sensor can be installed nonintrusively. In variable speeds operation, the use of an inverter to drive the induction motor introduces noise into the stator current so stator current based fault detection techniques become less reliable. This paper presents a hybrid algorithm, which combines time and frequency domain analysis, for broken rotor bar and bearing fault detection. Cluster information obtained by using Independent Component Analysis (ICA) …
Show Me The Numbers: Visual Analytics For Insights,
2010
Singapore Management University
Show Me The Numbers: Visual Analytics For Insights, Tin Seong Kam
Research Collection School Of Computing and Information Systems
In this highly volatile and fast-paced financial market, traders and managers working in banking and financial organizations must struggle to cope with large and complex data from multi-sources, that move throughout the market at increasingly high speed. The cost of making poor business and investment decisions is very high. This places great demands on data analysts, who are responsible for providing process information, to support the activities of traders and managers. Static reports and traditional business intelligence tools simply cannot keep up with a market that is changing on a second-to-second basis. By the time the traders and bankers have …
The R Journal (June 2010) 2(1): Complete Issue,
2010
University of Nebraska - Lincoln
The R Journal (June 2010) 2(1): Complete Issue, The R Foundation
The R Journal
Contributed Research Articles
IsoGene: An R Package for Analyzing Dose-response Studies in Microarray Experiments, Setia Pramana, Dan Lin, Philippe Haldermans, Ziv Shkedy, Tobias Verbeke, Hinrich Göhlmann, An De Bondt, Willem Talloen, and Luc Bijnens
MCMC for Generalized Linear Mixed Models with glmmBUGS, Patrick Brown and Lutong Zhou
Mapping and Measuring Country Shapes, Nils B. Weidmann and Kristian Skrede Gleditsch
tmvtnorm: A Package for the Truncated Multivariate Normal Distribution, Stefan Wilhelm and B. G. Manjunath
neuralnet: Training of Neural Networks, Frauke Günther and Stefan Fritsch
glmperm: A Permutation of Regressor Residuals Test for Inference in Generalized Linear Models, Wiebke Werft and …
Isogene: An R Package For Analyzing Dose-Response Studies In Microarray Experiments,
2010
Universiteit Hasselt
Isogene: An R Package For Analyzing Dose-Response Studies In Microarray Experiments, Setia Pramana, Dan Lin, Philippe Haldermans, Ziv Shkedy, Tobias Verbeke, Hinrich Göhlmann, An De Bondt, Williem Talloen, Luc Bijnens
The R Journal
IsoGene is an R package for the analysis of dose-response microarray experiments to identify gene or subsets of genes with a mono tone relationship between the gene expression and the doses. Several testing procedures (i.e., the likelihood ratio test, Williams, Marcus, the M, and Modified M), that take into account the order restriction of the means with respect to the increasing doses are implemented in the package. The inference is based on resampling methods, both permutations and the Significance Analysis of Microarrays (SAM).
Two-Sided Exact Tests And Matching Confidence Intervals For Discrete Data,
2010
National Institute of Allergy and Infectious Diseases
Two-Sided Exact Tests And Matching Confidence Intervals For Discrete Data, Michael P. Fay
The R Journal
There is an inherent relationship between two-sided hypothesis tests and confidence intervals. A series of two-sided hypothesis tests may be inverted to obtain the matching 100(1-)% confidence interval defined as the smallest interval that contains all point null parameter values that would not be rejected at the α level. Unfortunately, for discrete data there are several different ways of defining two-sided exact tests and the most commonly used two sided exact tests are defined one way, while the most commonly used exact confidence intervals are inversions of tests defined another way. This can lead to inconsistencies where the exact test …
Neuralnet: Training Of Neural Networks,
2010
University of Bremen
Neuralnet: Training Of Neural Networks, Franke Günther, Stefan Fritsch
The R Journal
Artificial neural networks are applied in many situations. neuralnet is built to train multi-layer perceptrons in the context of regression analyses, i.e. to approximate functional relationships between covariates and response variables. Thus, neural networks are used as extensions of generalized linear models. neuralnet is a very flexible package. The back propagation algorithm and three versions of resilient back-propagation are implemented and it provides a custom-choice of activation and error function. An arbitrary number of covariates and response variables as well as of hidden layers can theoretically be included. The paper gives a brief introduction to multi-layer perceptrons and resilient back-propagation …
Mcmc For Generalized Linear Mixed Models With Glmmbugs,
2010
University of Toronto, and Cancer Care Ontario
Mcmc For Generalized Linear Mixed Models With Glmmbugs, Patrick Brown, Lutong Zhou
The R Journal
The glmmBUGS package is a bridging tool between Generalized Linear Mixed Models (GLMMs) in R and the BUGS language. It provides a simple way of performing Bayesian inference using Markov Chain Monte Carlo (MCMC) methods, taking a model formula and data frame in R and writing a BUGS model file, data file, and initial values files. Functions are provided to reformat and summarize the BUGS results. A key aim of the package is to provide files and objects that can be modified prior to calling BUGS, giving users a platform for customizing and extending the models to accommodate a wide …
Glmperm: A Permutation Of Regressor Residuals Test For Inference In Generalized Linear Models,
2010
German Cancer Research Center
Glmperm: A Permutation Of Regressor Residuals Test For Inference In Generalized Linear Models, Wiebke Werft, Axel Benner
The R Journal
We introduce a new R package called glmperm for inference in generalized linear models especially for small and moderate-sized data sets. The inference is based on the per mutation of regressor residuals test introduced by Potter (2005). The implementation of glmperm outperforms currently available permutation test software as glmperm can be applied in situations where more than one covariate is involved.
Tmvtnorm: A Package For The Truncated Multivariate Normal Distribution,
2010
University of Basel
Tmvtnorm: A Package For The Truncated Multivariate Normal Distribution, Stefan Wilhelm, B. G. Manjunath
The R Journal
In this article we present tmvtnorm, an R package implementation for the truncated multivariate normal distribution. We consider random number generation with rejection and Gibbs sampling, computation of marginal densities as well as computation of the mean and co variance of the truncated variables. This contribution brings together latest research in this field and provides useful methods for both scholars and practitioners when working with truncated normal variables.
Convergence Of The Sinc Method Applied To Volterra Integral Equations,
2010
University of Mohaghegh Ardabili
Convergence Of The Sinc Method Applied To Volterra Integral Equations, M. Zarebnia, J. Rashidinia
Applications and Applied Mathematics: An International Journal (AAM)
A collocation procedure is developed for the linear and nonlinear Volterra integral equations, using the globally defined Sinc and auxiliary basis functions. We analytically show the exponential convergence of the Sinc collocation method for approximate solution of Volterra integral equations. Numerical examples are included to confirm applicability and justify rapid convergence of our method.
Visualizing And Exploring Evolving Information Networks In Wikipedia,
2010
Singapore Management University
Visualizing And Exploring Evolving Information Networks In Wikipedia, Ee Peng Lim, Agus Trisnajaya Kwee, Nelman Lubis Ibrahim, Aixin Sun, Anwitaman Datta, Kuiyu Chang, Maureen Maureen
Research Collection School Of Computing and Information Systems
Information networks in Wikipedia evolve as users collaboratively edit articles that embed the networks. These information networks represent both the structure and content of community’s knowledge and the networks evolve as the knowledge gets updated. By observing the networks evolve and finding their evolving patterns, one can gain higher order knowledge about the networks and conduct longitudinal network analysis to detect events and summarize trends. In this paper, we present SSNetViz+, a visual analytic tool to support visualization and exploration of Wikipedia’s information networks. SSNetViz+ supports time-based network browsing, content browsing and search. Using a terrorism information network as an …
Using Hadoop And Cassandra For Taxi Data Analytics: A Feasibility Study,
2010
Singapore Management University
Using Hadoop And Cassandra For Taxi Data Analytics: A Feasibility Study, Alvin Jun Yong Koh, Xuan Khoa Nguyen, C. Jason Woodard
Research Collection School Of Computing and Information Systems
This paper reports on a preliminary study to assess the feasibility of using the Open Cirrus Cloud Computing Research testbed to provide offline and online analytical support for taxi fleet operations. In the study, we benchmarked the performance gains from distributing the offline analysis of GPS location traces over multiple virtual machines using the Apache Hadoop implementation of the MapReduce paradigm. We also explored the use of the Apache Cassandra distributed database system for online retrieval of vehicle trace data. While configuring the testbed infrastructure was straightforward, we encountered severe I/O bottlenecks in running the benchmarks due to the lack …
Z-Sky: An Efficient Skyline Query Processing Framework Based On Z-Order,
2010
Pennsylvania State University
Z-Sky: An Efficient Skyline Query Processing Framework Based On Z-Order, Ken C. K. Lee, Wang-Chien Lee, Baihua Zheng, Huajing Li, Yuan Tian
Research Collection School Of Computing and Information Systems
Given a set of data points in a multidimensional space, a skyline query retrieves those data points that are not dominated by any other point in the same dataset. Observing that the properties of Z-order space filling curves (or Z-order curves) perfectly match with the dominance relationships among data points in a geometrical data space, we, in this paper, develop and present a novel and efficient processing framework to evaluate skyline queries and their variants, and to support skyline result updates based on Z-order curves. This framework consists of ZBtree, i.e., an index structure to organize a source dataset and …
Efficient Mutual Nearest Neighbor Query Processing For Moving Object Trajectories,
2010
Zhejiang University
Efficient Mutual Nearest Neighbor Query Processing For Moving Object Trajectories, Yunjun Gao, Baihua Zheng, Gencai Chen, Qing Li, Chun Chen, Gang Chen
Research Collection School Of Computing and Information Systems
Given a set D of trajectories, a query object q, and a query time extent Γ, a mutual (i.e., symmetric) nearest neighbor (MNN) query over trajectories finds from D, the set of trajectories that are among the k1 nearest neighbors (NNs) of q within Γ, and meanwhile, have q as one of their k2 NNs. This type of queries is useful in many applications such as decision making, data mining, and pattern recognition, as it considers both the proximity of the trajectories to q and the proximity of q to the trajectories. In this paper, we first formalize MNN search …
Do Wikipedians Follow Domain Experts? A Domain-Specific Study On Wikipedia Contribution,
2010
Nanyang Technological University
Do Wikipedians Follow Domain Experts? A Domain-Specific Study On Wikipedia Contribution, Yi Zhang, Aixin Sun, Anwitaman Datta, Kuiyu Chang, Ee Peng Lim
Research Collection School Of Computing and Information Systems
Wikipedia is one of the most successful online knowledge bases, attracting millions of visits daily. Not surprisingly, its huge success has in turn led to immense research interest for a better understanding of the collaborative knowledge building process. In this paper, we performed a (terrorism) domain-specific case study, comparing and contrasting the knowledge evolution in Wikipedia with a knowledge base created by domain experts. Specifically, we used the Terrorism Knowledge Base (TKB) developed by experts at MIPT. We identified 409 Wikipedia articles matching TKB records, and went ahead to study them from three aspects: creation, revision, and link evolution. We …
Stevent: Spatio-Temporal Event Model For Social Network Discovery,
2010
Singapore Management University
Stevent: Spatio-Temporal Event Model For Social Network Discovery, Hady W. Lauw, Ee Peng Lim, Hwee Hwa Pang, Teck-Tim Tan
Research Collection School Of Computing and Information Systems
Spatio-temporal data concerning the movement of individuals over space and time contains latent information on the associations among these individuals. Sources of spatio-temporal data include usage logs of mobile and Internet technologies. This article defines a spatio-temporal event by the co-occurrences among individuals that indicate potential associations among them. Each spatio-temporal event is assigned a weight based on the precision and uniqueness of the event. By aggregating the weights of events relating two individuals, we can determine the strength of association between them. We conduct extensive experimentation to investigate both the efficacy of the proposed model as well as the …
