Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences Commons

Open Access. Powered by Scholars. Published by Universities.®

Numerical Analysis and Scientific Computing

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 6271 - 6300 of 6663

Full-Text Articles in Computer Sciences

Estimating The Quality Of Postings In The Real-Time Web, Hady W. Lauw, Alexandros Ntoulas, Krishnaram Kenthapadi Feb 2010

Estimating The Quality Of Postings In The Real-Time Web, Hady W. Lauw, Alexandros Ntoulas, Krishnaram Kenthapadi

Research Collection School Of Computing and Information Systems

Millions of users are posting their status updates, interesting findings, news, ideas and observations in real-time on microblogging services such as Twitter, Jaiku and Plurk. This real-time Web can be a great resource of valuable timely information. Since the real-time Web is completely open and decentralized and anyone may post information at whim, distinguishing interesting and popular postings from the mundane ones is a challenging task. In this paper we study the problem of estimating the quality (or “interestingness”) of postings in the real-time Web. We identify several important factors that are indicative of the quality of postings, and present …


Player Performance Prediction In Massively Multiplayer Online Role-Playing Games (Mmorpgs), Kyong Jin Shim, Richa Sharan, Jaideep Srivastava Feb 2010

Player Performance Prediction In Massively Multiplayer Online Role-Playing Games (Mmorpgs), Kyong Jin Shim, Richa Sharan, Jaideep Srivastava

Research Collection School Of Computing and Information Systems

Recent years have seen an ever increasing number of people interacting in the online space. Massively multiplayer online role-playing games (MMORPGs) are personal computer or console-based digital games where thousands of players can simultaneously sign on to the same online, persistent virtual world to interact and collaborate with each other through their in-game characters. In recent years, researchers have found virtual environments to be a sound venue for studying learning, collaboration, social participation, literacy in online space, and learning trajectory at the individual level as well as at the group level. While many games today provide web and GUI-based reports …


Privacy-Preserving Similarity-Based Text Retrieval, Hwee Hwa Pang, Jialie Shen, Ramayya Krishnan Feb 2010

Privacy-Preserving Similarity-Based Text Retrieval, Hwee Hwa Pang, Jialie Shen, Ramayya Krishnan

Research Collection School Of Computing and Information Systems

Users of online services are increasingly wary that their activities could disclose confidential information on their business or personal activities. It would be desirable for an online document service to perform text retrieval for users, while protecting the privacy of their activities. In this article, we introduce a privacy-preserving, similarity-based text retrieval scheme that (a) prevents the server from accurately reconstructing the term composition of queries and documents, and (b) anonymizes the search results from unauthorized observers. At the same time, our scheme preserves the relevance-ranking of the search server, and enables accounting of the number of documents that each …


Using Transitivity With Nearest Neighbor To Reduce Error In Sample-Based Pearson Correlation Coefficients, Taylor Phillips Jan 2010

Using Transitivity With Nearest Neighbor To Reduce Error In Sample-Based Pearson Correlation Coefficients, Taylor Phillips

Electronic Theses and Dissertations

Pearson product-moment correlation coefficients are a well-practiced quantification of linear dependence seen across many fields. When calculating a sample-based correlation coefficient, the accuracy of the estimation is dependent on the quality and quantity of the sample. Like all statistical models, these correlation coefficients can suffer from overfitting, which results in the representation of random error instead of an underlying trend.

In this paper, we discuss how Pearson's product-moment correlation coefficients can utilize information outside of the two items for which the correlation is being computed. By introducing a relationship with one or more additional items that meet specified criterion, our …


Supporting Multiple Paths To Objects In Information Hierarchies: Faceted Classification, Faceted Search, And Symbolic Links, Saverio Perugini Jan 2010

Supporting Multiple Paths To Objects In Information Hierarchies: Faceted Classification, Faceted Search, And Symbolic Links, Saverio Perugini

Computer Science Faculty Publications

We present three fundamental, interrelated approaches to support multiple access paths to each terminal object in information hierarchies: faceted classification, faceted search, and web directories with embedded symbolic links. This survey aims to demonstrate how each approach supports users who seek information from multiple perspectives. We achieve this by exploring each approach, the relationships between these approaches, including tradeoffs, and how they can be used in concert, while focusing on a core set of hypermedia elements common to all. This approach provides a foundation from which to study, understand, and synthesize applications which employ these techniques. This survey does not …


An Intelligent Data-Centric Approach Toward Identification Of Conserved Motifs In Protein Sequences, Kathryn Dempsey Cooper, Benjamin Currall, Richard Hallworth, Hesham Ali Jan 2010

An Intelligent Data-Centric Approach Toward Identification Of Conserved Motifs In Protein Sequences, Kathryn Dempsey Cooper, Benjamin Currall, Richard Hallworth, Hesham Ali

Interdisciplinary Informatics Faculty Proceedings & Presentations

The continued integration of the computational and biological sciences has revolutionized genomic and proteomic studies. However, efficient collaboration between these fields requires the creation of shared standards. A common problem arises when biological input does not properly fit the expectations of the algorithm, which can result in misinterpretation of the output. This potential confounding of input/output is a drawback especially when regarding motif finding software. Here we propose a method for improving output by selecting input based upon evolutionary distance, domain architecture, and known function. This method improved detection of both known and unknown motifs in two separate case studies. …


Parallel And Distributed Simulation Of Parabolic And Telegraphic Equations., Ewedafe Simon Uzezi Jan 2010

Parallel And Distributed Simulation Of Parabolic And Telegraphic Equations., Ewedafe Simon Uzezi

Student Works (2010-2019)

In this thesis, a parallel implementation of explicit/implicit parallel algorithms such as the stationary iterative methods and the class of iterating alternating methods which includes: Alternating Direction Implicit (ADI), Iterative Alternating Direction Explicit (IADE), for D’Yakonov (IADE-DY), Double sweep Mitchell and Fairweather (MF-DS) and Alternating Group Explicit (AGE) method for solving 1-Dimensional (1-D), 2-Dimensional (2-D) Parabolic (special examples including 1-D, 2-D Bio-Heat Equation) and 1-D, 2-D and 3-D Telegraphic Equations on a distributed environment of Message Passing Interface (MPI) and Parallel Virtual Machine (PVM) platform is presented. To correlate the communication activity with computation, we counted events between significant MPI/PVM …


Cbtv: Visualising Case Bases For Similarity Measure Design And Selection, Brian Mac Namee, Sarah Jane Delany Jan 2010

Cbtv: Visualising Case Bases For Similarity Measure Design And Selection, Brian Mac Namee, Sarah Jane Delany

Conference papers

In CBR the design and selection of similarity measures is paramount. Selection can benefit from the use of exploratory visualisation- based techniques in parallel with techniques such as cross-validation ac- curacy comparison. In this paper we present the Case Base Topology Viewer (CBTV) which allows the application of different similarity mea- sures to a case base to be visualised so that system designers can explore the case base and the associated decision boundary space. We show, using a range of datasets and similarity measure types, how the idiosyncrasies of particular similarity measures can be illustrated and compared in CBTV allowing …


Modeling Anticipatory Event Transitions, He Qi, Kuiyu Chang, Ee Peng Lim Jan 2010

Modeling Anticipatory Event Transitions, He Qi, Kuiyu Chang, Ee Peng Lim

Research Collection School Of Computing and Information Systems

Major world events such as terrorist attacks, natural disasters, wars, etc. typically progress through various representative stages/states in time. For example, a volcano eruption could lead to earthquakes, tsunamis, aftershocks, evacuation, rescue efforts, international relief support, rebuilding, and resettlement, etc. By analyzing various types of catastrophical and historical events, we can derive corresponding event transition models to embed useful information at each state. The knowledge embedded in these models can be extremely valuable. For instance, a transition model of the 1918-1920 flu pandemic could be used for the planning and allocation of resources to decisively respond to future occurrences of …


Information Integration For Graph Databases, Ee Peng Lim, Aixin Sun, Anwitaman Datta, Chang Kuiyu Jan 2010

Information Integration For Graph Databases, Ee Peng Lim, Aixin Sun, Anwitaman Datta, Chang Kuiyu

Research Collection School Of Computing and Information Systems

With increasing interest in querying and analyzing graph data from multiple sources, algorithms and tools to integrate different graphs become very important. Integration of graphs can take place at the schema and instance levels. While links among graph nodes pose additional challenges to graph information integration, they can also serve as useful features for matching nodes representing real-world entities. This chapter introduces a general framework to perform graph information integration. It then gives an overview of the state-of-the-art research and tools in graph information integration.


Trust-Oriented Composite Service Selection With Qos Constraints, Lei Li, Yang Wang, Ee Peng Lim Jan 2010

Trust-Oriented Composite Service Selection With Qos Constraints, Lei Li, Yang Wang, Ee Peng Lim

Research Collection School Of Computing and Information Systems

In Service-Oriented Computing (SOC) environments, service clients interact with service providers for consuming services. From the viewpoint of service clients, the trust level of a service or a service provider is a critical factor to consider in service selection, particularlywhen a client is looking for a service from a large set of services or service providers. However, a invoked service may be composed of other services. The complex invocations in composite services greatly increase the complexity of trust-oriented service selection. In this paper, we propose novel approaches for composite service representation, trust evaluation and trust-oriented com-posite service selection (with QoS …


Anonymous Query Processing In Road Networks, Kyriakos Mouratidis, Man Lung Yiu Jan 2010

Anonymous Query Processing In Road Networks, Kyriakos Mouratidis, Man Lung Yiu

Research Collection School Of Computing and Information Systems

The increasing availability of location-aware mobile devices has given rise to a flurry of location-based services (LBSs). Due to the nature of spatial queries, an LBS needs the user position in order to process her requests. On the other hand, revealing exact user locations to a (potentially untrusted) LBS may pinpoint their identities and breach their privacy. To address this issue, spatial anonymity techniques obfuscate user locations, forwarding to the LBS a sufficiently large region instead. Existing methods explicitly target processing in the euclidean space and do not apply when proximity to the users is defined according to network distance …


An Optical Machine Vision System For Applications In Cytopathology, Jonathan Blackledge, Dmitry Dubovitskiy Jan 2010

An Optical Machine Vision System For Applications In Cytopathology, Jonathan Blackledge, Dmitry Dubovitskiy

Articles

This paper discusses a new approach to the processes of object detection, recognition and classification in a digital image focusing on problem in Cytopathology. A unique self learning procedure is presented in order to incorporate expert knowledge. The classification method is based on the application of a set of features which includes fractal parameters such as the Lacunarity and Fourier dimension. Thus, the approach includes the characterisation of an object in terms of its fractal properties and texture characteristics. The principal issues associated with object recognition are presented which include the basic model and segmentation algorithms. The self-learning procedure for …


Dual Phase Learning For Large Scale Video Gait Recognition, Jialie Shen, Hwee Hwa Pang, Dacheng Tao, Xuelong Li Jan 2010

Dual Phase Learning For Large Scale Video Gait Recognition, Jialie Shen, Hwee Hwa Pang, Dacheng Tao, Xuelong Li

Research Collection School Of Computing and Information Systems

Accurate gait recognition from video is a complex process involving heterogenous features, and is still being developed actively. This article introduces a novel framework, called GC2F, for effective and efficient gait recognition and classification. Adopting a ”refinement-and-classification” principle, the framework comprises two components: 1) a classifier to generate advanced probabilistic features from low level gait parameters; and 2) a hidden classifier layer (based on multilayer perceptron neural network) to model the statistical properties of different subject classes. To validate our framework, we have conducted comprehensive experiments with a large test collection, and observed significant improvements in identification accuracy relative to …


Sos: Searching Help Pages Of R Packages, Spencer Graves, Sundar Dorai-Raj, Romain François Dec 2009

Sos: Searching Help Pages Of R Packages, Spencer Graves, Sundar Dorai-Raj, Romain François

The R Journal

The sos package provides a means to quickly and flexibly search the help pages of contributed packages, finding functions and datasets in seconds or minutes that could not be found in hours or days by any other means we know. Its findFn function accesses Jonathan Baron’s R Site Search database and returns the matches in a data frame of class "findFn", which can be further manipulated by other sos functions to produce, for example, an Excel file that starts with a summary sheet that makes it relatively easy to prioritize alternative packages for further study. As such, it provides a …


Rattle: A Data Mining Gui For R, Graham J. Williams Dec 2009

Rattle: A Data Mining Gui For R, Graham J. Williams

The R Journal

Data mining delivers insights, pat terns, and descriptive and predictive models from the large amounts of data available today in many organisations. The data miner draws heavily on methodologies, techniques and algorithms from statistics, machine learning, and computer science. R increasingly provides a powerful platform for data mining. However, scripting and programming is sometimes a challenge for data analysts moving into data mining. The Rattle package provides a graphical user interface specifically for data mining using R. It also provides a stepping stone toward using R as a programming language for data analysis.


Copas: An R Package For Fitting The Copas Selection Model, J. Carpenter, G. Rücker, G. Schhwarzer Dec 2009

Copas: An R Package For Fitting The Copas Selection Model, J. Carpenter, G. Rücker, G. Schhwarzer

The R Journal

This article describes the R package copas which is an add-on package to the R pack age meta. The R package copas can be used to f it the Copas selection model to adjust for bias in meta-analysis. A clinical example is used to illustrate fitting and interpreting the Copas selection model.


Party On!, Carolin Strobl, Torsten Hothorn, Achim Zeileis Dec 2009

Party On!, Carolin Strobl, Torsten Hothorn, Achim Zeileis

The R Journal

Random forests are one of the most popular statistical learning algorithms, and a variety of methods for fitting random forests and related recursive partitioning approaches is available in R. This paper points out two important features of the random forest implementation cforest available in the party package: The resulting forests are unbiased and thus prefer able to the randomForest implementation avail able in randomForest if predictor variables are of different types. Moreover, a conditional per mutation importance measure has recently been added to the party package, which can help evaluate the importance of correlated predictor variables. The rationale of this …


Aspects Of The Social Organization And Trajectory Of The R Project, John Fox Dec 2009

Aspects Of The Social Organization And Trajectory Of The R Project, John Fox

The R Journal

Based partly on interviews with members of the R Core team, this paper considers the development of the R Project in the context of open-source software development and, more generally, voluntary activities. The paper de scribes aspects of the social organization of the R Project, including the organization of the R Core team; describes the trajectory of the R Project; seeks to identify factors crucial to the success of R; and speculates about the prospects for R.


Asymptest: A Simple R Package For Classical Parametric Statistical Tests And Confidence Intervals In Large Samples, J.-F. Coeurjolly, R. Drouilhet, P. Lafaye De Micheaux, J.-F. Robineau Dec 2009

Asymptest: A Simple R Package For Classical Parametric Statistical Tests And Confidence Intervals In Large Samples, J.-F. Coeurjolly, R. Drouilhet, P. Lafaye De Micheaux, J.-F. Robineau

The R Journal

asympTest is an R package implementing large sample tests and confidence intervals. One and two sample mean and variance tests (differences and ratios) are considered. The test statistics are all expressed in the same form as the Student t-test, which facilitates their presentation in the classroom. This contribution also fills the gap of a robust (to non-normality) alternative to the chi-square single variance test for large samples, since no such procedure is implemented in standard statistical software.


Convergenceconcepts: An R Package To Investigate Various Modes Of Convergence, Pierre Lafaye De Micheaux, Benoit Liquet Dec 2009

Convergenceconcepts: An R Package To Investigate Various Modes Of Convergence, Pierre Lafaye De Micheaux, Benoit Liquet

The R Journal

ConvergenceConcepts is an R pack age, built upon the tkrplot, tcltk and lattice packages, designed to investigate the convergence of simulated sequences of random variables. Four classical modes of convergence may be studied, namely: almost sure convergence (a.s.), convergence in probability (P), convergence in law (L) and convergence in r-th mean (r). This investigation is performed through ac curate graphical representations. This package may be used as a pedagogical tool. It may give students a better understanding of these notions and help them to visualize these difficult theoretical concepts. Moreover, …


Adaptive Type-2 Fuzzy Maintenance Advisor For Offshore Power Systems, Zhaoxia Wang, C. S. Chang, Fan Yang, W. W. Tan Dec 2009

Adaptive Type-2 Fuzzy Maintenance Advisor For Offshore Power Systems, Zhaoxia Wang, C. S. Chang, Fan Yang, W. W. Tan

Research Collection School Of Computing and Information Systems

Proper maintenance strategies are very desirable for minimizing the operational and maintenance costs of power systems without sacrificing reliability. Condition-based maintenance has largely replaced time-based maintenance because of the former's potential economic benefits. As offshore substations are often remotely located, they experience more adverse environments, higher failures, and therefore need more powerful analytical tools than their onshore counterpart. As reliability information collected during operation of an offshore substation can rarely avoid uncertainties, it is essential to obtain consistent estimates of reliability measures under changing environmental and operating conditions. Some attempts with type-1 fuzzy logic were made with limited success in …


The R Journal (December 2009) 1(2): Complete Issue, The R Foundation Dec 2009

The R Journal (December 2009) 1(2): Complete Issue, The R Foundation

The R Journal

Contributed Research Articles

Aspects of the Social Organization and Trajectory of the R Project, John Fox

Party on! Carolin Strobl, Torsten Hothorn and Achim Zeileis

ConvergenceConcepts: An R Package to Investigate Various Modes of Convergence, Pierre Lafaye de Micheaux and Benoit Liquet

asympTest: A Simple R Package for Classical Parametric Statistical Tests and Confidence Intervals in Large Samples, J.-F. Coeurjolly, R. Drouilhet, P. Lafaye de Micheaux and J.-F. Robineau

copas: An R package for Fitting the Copas Selection Model, J. Carpenter, G. Rücker and G. Schwarzer

Transitioning to R: Replicating SAS, Stata, and SUDAAN Analysis Techniques in Health Policy Data, …


On Strategies For Imbalanced Text Classification Using Svm: A Comparative Study, Aixin Sun, Ee Peng Lim, Ying Liu Dec 2009

On Strategies For Imbalanced Text Classification Using Svm: A Comparative Study, Aixin Sun, Ee Peng Lim, Ying Liu

Research Collection School Of Computing and Information Systems

Many real-world text classification tasks involve imbalanced training examples. The strategies proposed to address the imbalanced classification (e.g., resampling, instance weighting), however, have not been systematically evaluated in the text domain. In this paper, we conduct a comparative study on the effectiveness of these strategies in the context of imbalanced text classification using Support Vector Machines (SVM) classifier. SVM is the interest in this study for its good classification accuracy reported in many text classification tasks. We propose a taxonomy to organize all proposed strategies following the training and the test phases in text classification tasks. Based on the taxonomy, …


To Trust Or Not To Trust? Predicting Online Trusts Using Trust Antecedent Framework, Viet-An Nguyen, Ee Peng Lim, Jing Jiang, Aixin Sun Dec 2009

To Trust Or Not To Trust? Predicting Online Trusts Using Trust Antecedent Framework, Viet-An Nguyen, Ee Peng Lim, Jing Jiang, Aixin Sun

Research Collection School Of Computing and Information Systems

This paper analyzes the trustor and trustee factors that lead to inter-personal trust using a well studied Trust Antecedent framework in management science. To apply these factors to trust ranking problem in online rating systems, we derive features that correspond to each factor and develop different trust ranking models. The advantage of this approach is that features relevant to trust can be systematically derived so as to achieve good prediction accuracy. Through a series of experiments on real data from Epinions, we show that even a simple model using the derived features yields good accuracy and outperforms MoleTrust, a trust …


Trust-Oriented Composite Services Selection And Discovery, Lei Li, Yan Wang, Ee Peng Lim Nov 2009

Trust-Oriented Composite Services Selection And Discovery, Lei Li, Yan Wang, Ee Peng Lim

Research Collection School Of Computing and Information Systems

In Service-Oriented Computing (SOC) environments, service clients interact with service providers for consuming services. From the viewpoint of service clients, the trust level of a service or a service provider is a critical issue to consider in service selection and discovery, particularly when a client is looking for a service from a large set of services or service providers. However, a service may invoke other services offered by different providers forming composite services. The complex invocations in composite services greatly increase the complexity of trust-oriented service selection and discovery. In this paper, we propose novel approaches for composite service representation, …


What Makes Categories Difficult To Classify?, Aixin Sun, Ee Peng Lim, Ying Liu Nov 2009

What Makes Categories Difficult To Classify?, Aixin Sun, Ee Peng Lim, Ying Liu

Research Collection School Of Computing and Information Systems

In this paper, we try to predict which category will be less accurately classified compared with other categories in a classification task that involves multiple categories. The categories with poor predicted performance will be identified before any classifiers are trained and additional steps can be taken to address the predicted poor accuracies of these categories. Inspired by the work on query performance prediction in ad-hoc retrieval, we propose to predict classification performance using two measures, namely, category size and category coherence. Our experiments on 20-Newsgroup and Reuters-21578 datasets show that the Spearman rank correlation coefficient between the predicted rank of …


Trust Relationship Prediction Using Online Product Review Data, Nan Ma, Ee Peng Lim, Viet-An Nguyen, Aixin Sun Nov 2009

Trust Relationship Prediction Using Online Product Review Data, Nan Ma, Ee Peng Lim, Viet-An Nguyen, Aixin Sun

Research Collection School Of Computing and Information Systems

Trust between users is an important piece of knowledge that can be exploited in search and recommendation.Given that user-supplied trust relationships are usually very sparse, we study the prediction of trust relationships using user interaction features in an online user generated review application context. We show that trust relationship prediction can achieve better accuracy when one adopts personalized and cluster-based classification methods. The former trains one classifier for each user using user-specific training data. The cluster-based method first constructs user clusters before training one classifier for each user cluster. Our proposed methods have been evaluated in a series of experiments …


Mining Communities In Networks: A Solution For Consistency And Its Evaluation, Haewoon Kwak, Yoonchan Choi, Young-Ho Eom, Hawoong Jeong, Sue Moon Nov 2009

Mining Communities In Networks: A Solution For Consistency And Its Evaluation, Haewoon Kwak, Yoonchan Choi, Young-Ho Eom, Hawoong Jeong, Sue Moon

Research Collection School Of Computing and Information Systems

Online social networks pose significant challenges to computer scientists, physicists, and sociologists alike, for their massive size, fast evolution, and uncharted potential for social computing. One particular problem that has interested us is community identification. Many algorithms based on various metrics have been proposed for communities in networks [18, 24], but a few algorithms scale to very large networks. Three recent community identification algorithms, namely CNM [16], Wakita [59], and Louvain [10], stand out for their scalability to a few millions of nodes. All of them use modularity as the metric of optimization. However, all three algorithms produce inconsistent communities …


Udel/Smu At Trec 2009 Entity Track, Wei Zheng, Swapna Gottipati, Jing Jiang, Hui Fang Nov 2009

Udel/Smu At Trec 2009 Entity Track, Wei Zheng, Swapna Gottipati, Jing Jiang, Hui Fang

Research Collection School Of Computing and Information Systems

We report our methods and experiment results from the collaborative participation of the InfoLab group from University of Delaware and the school of Information Systems from Singapore Management University in the TREC 2009 Entity track. Our general goal is to study how we may apply language modeling approaches and natural language processing techniques to the task. Specically, we proposed to find supporting information based on segment retrieval, to extract entities using Stanford NER tagger, and to rank entities based on a previously proposed probabilistic framework for expert finding.