Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences Commons™

Open Access. Powered by Scholars. Published by Universities.®

Numerical Analysis and Scientific Computing

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 6151 - 6180 of 6662

Full-Text Articles in Computer Sciences

The Digitize Package: Extracting Numerical Data From Scatterplots, Timothée Poisot Jun 2011

The Digitize Package: Extracting Numerical Data From Scatterplots, Timothée Poisot

The R Journal

I present the small R package digitize, designed to extract data from scatter plots with a simple method and suited to small datasets. I present an application of this method to the ex traction of data from a graph whose source is not available.


Analyzing An Electronic Limit Order Book, David Kane, Andrew Liu, Khanh Nguyen Jun 2011

Analyzing An Electronic Limit Order Book, David Kane, Andrew Liu, Khanh Nguyen

The R Journal

The orderbook package provides facilities for exploring and visualizing the data associated with an order book: the electronic collection of the outstanding limit orders for a financial instrument. This article provides an overview of the orderbook package and examples of its use.


Testthat: Get Started With Testing, Hadley Wickham Jun 2011

Testthat: Get Started With Testing, Hadley Wickham

The R Journal

Software testing is important, but many of us don’t do it because it is frustrating and boring. testthat is a new testing framework for R that is easy learn and use, and integrates with your existing workflow. This paper shows how, with illustrations from existing packages.


Rworldmap: A New R Package For Mapping Global Data, Andy South Jun 2011

Rworldmap: A New R Package For Mapping Global Data, Andy South

The R Journal

rworldmap is a relatively new package available on CRAN for the mapping and visualisation of global data. The vision is to make the display of global data easier, to facilitate understanding and communication. The initial focus is on data referenced by country or grid due to the frequency of use of such data in global assessments. Tools to link data referenced by country (either name or code) to a map, and then to display the map are provided as are functions to map global gridded data. Country and gridded functions accept the same arguments to specify the nature of categories …


Differential Evolution With Deoptim, David Ardia, Kris Boudt, Peter Carl, Katharine M. Mullen, Brian G. Peterson Jun 2011

Differential Evolution With Deoptim, David Ardia, Kris Boudt, Peter Carl, Katharine M. Mullen, Brian G. Peterson

The R Journal

The R package DEoptim implements the Differential Evolution algorithm. This algorithm is an evolutionary technique similar to classic genetic algorithms that is useful for the solution of global optimization problems. In this note we provide an introduction to the package and demonstrate its utility for financial applications by solving a non-convex portfolio optimization problem.


Probabilistic Weather Forecasting In R, Chris Fraley, Adrian Raftery, Tilmann Gneiting, Mclean Sloughter, Veronica Berrocol Jun 2011

Probabilistic Weather Forecasting In R, Chris Fraley, Adrian Raftery, Tilmann Gneiting, Mclean Sloughter, Veronica Berrocol

The R Journal

This article describes two R packages for probabilistic weather forecasting, ensembleBMA, which offers ensemble post-processing via Bayesian model averaging (BMA), and Prob ForecastGOP, which implements the geostatistical output perturbation (GOP) method. BMA forecasting models use mixture distributions, in which each component corresponds to an ensemble member, and the form of the component distribution depends on the weather parameter (temperature, quantitative precipitation or wind speed). The model parameters are estimated from training data. The GOP technique uses geostatistical methods to produce probabilistic fore casts of entire weather fields for temperature or pressure, based on a single numerical forecast on …


Continuous Visible Nearest Neighbor Query Processing In Spatial Databases, Yunjun Gao, Baihua Zheng, Gencai Chen, Qing Li, Xiaofa Guo Jun 2011

Continuous Visible Nearest Neighbor Query Processing In Spatial Databases, Yunjun Gao, Baihua Zheng, Gencai Chen, Qing Li, Xiaofa Guo

Research Collection School Of Computing and Information Systems

In this paper, we identify and solve a new type of spatial queries, called continuous visible nearest neighbor (CVNN) search. Given a data set P, an obstacle set O, and a query line segment q in a two-dimensional space, a CVNN query returns a set of $${\langle p, R\rangle}$$ tuples such that $${p \in P}$$ is the nearest neighbor to every point r along the interval $${R \subseteq q}$$ as well as pis visible to r. Note that p may be NULL, meaning that all points in P are invisible to all points in R due to the obstruction of …


Link Type Based Pre-Cluster Pair Model For Coreference Resolution, Yang Song, Houfeng Wang, Jing Jiang Jun 2011

Link Type Based Pre-Cluster Pair Model For Coreference Resolution, Yang Song, Houfeng Wang, Jing Jiang

Research Collection School Of Computing and Information Systems

This paper presents our participation in the CoNLL-2011 shared task, Modeling Unrestricted Coreference in OntoNotes. Coreference resolution, as a difficult and challenging problem in NLP, has attracted a lot of attention in the research community for a long time. Its objective is to determine whether two mentions in a piece of text refer to the same entity. In our system, we implement mention detection and coreference resolution seperately. For mention detection, a simple classification based method combined with several effective features is developed. For coreference resolution, we propose a link type based pre-cluster pair model. In this model, pre-clustering of …


Topical Keyphrase Extraction From Twitter, Xin Zhao, Jing Jiang, Jing He, Yang Song, Palakorn Achananuparp, Ee Peng Lim, Xiaoming Li Jun 2011

Topical Keyphrase Extraction From Twitter, Xin Zhao, Jing Jiang, Jing He, Yang Song, Palakorn Achananuparp, Ee Peng Lim, Xiaoming Li

Research Collection School Of Computing and Information Systems

Summarizing and analyzing Twitter content is an important and challenging task. In this paper, we propose to extract topical keyphrases as one way to summarize Twitter. We propose a context-sensitive topical PageRank method for keyword ranking and a probabilistic scoring function that considers both relevance and interestingness of keyphrases for keyphrase ranking. We evaluate our proposed methods on a large Twitter data set. Experiments show that these methods are very effective for topical keyphrase extraction.


Online Fault Detection Of Induction Motors Using Frequency Domain Independent Components Analysis, Zhaoxia Wang, C. S. Chang Jun 2011

Online Fault Detection Of Induction Motors Using Frequency Domain Independent Components Analysis, Zhaoxia Wang, C. S. Chang

Research Collection School Of Computing and Information Systems

This paper proposes an online fault detection method for induction motors using frequency-domain independent component analysis. Frequency-domain results, which are obtained by applying Fast Fourier Transform (FFT) to measured stator current time-domain waveforms, are analyzed with the aim of extracting frequency signatures of healthy and faulty motors with broken rotor-bar or bearing problem. Independent components analysis (ICA) is applied for such an aim to the FFT results. The obtained independent components as well as the FFT results are then used to obtain the combined fault signatures. The proposed method overcomes problems occurring in many existing FFT-based methods. Results using laboratory-collected …


Supervisory Evolutionary Optimization Strategy For Adaptive Maintenance Schedules, Zhaoxia Wang, C. S. Chang Jun 2011

Supervisory Evolutionary Optimization Strategy For Adaptive Maintenance Schedules, Zhaoxia Wang, C. S. Chang

Research Collection School Of Computing and Information Systems

No abstract provided.


Using Service Responsibility Tables To Supplement Uml In Analyzing E-Service Systems, Xin Tan, Steven Alter, Keng Siau Jun 2011

Using Service Responsibility Tables To Supplement Uml In Analyzing E-Service Systems, Xin Tan, Steven Alter, Keng Siau

Research Collection School Of Computing and Information Systems

This paper proposes using Service Responsibility Tables (SRTs) as a tool in analyzing e-service systems. First it discusses difficulties and deficiencies of using formal modeling languages such as UML in analyzing e-service systems. It proposes using SRTs as an informal language and lightweight analytical tool to be used by business professionals in analyzing e-service systems. SRTs are based on a service value chain framework but do not rely on abstract concepts and constructs, and therefore can be used by business professionals to supplement UML. We suggest a set of heuristics for transforming SRTs into two key UML diagrams, thereby illustrating …


Parallelizing Scale Invariant Feature Transform On A Distributed Memory Cluster, Stanislav Bobovych May 2011

Parallelizing Scale Invariant Feature Transform On A Distributed Memory Cluster, Stanislav Bobovych

Computer Science and Computer Engineering Undergraduate Honors Theses

Scale Invariant Feature Transform (SIFT) is a computer vision algorithm that is widely-used to extract features from images. We explored accelerating an existing implementation of this algorithm with message passing in order to analyze large data sets. We successfully tested two approaches to data decomposition in order to parallelize SIFT on a distributed memory cluster.


Continuous Nearest Neighbor Search In The Presence Of Obstacles, Yunjun Gao, Baihua Zheng, Gang Chen, Chun Chen, Qing Li May 2011

Continuous Nearest Neighbor Search In The Presence Of Obstacles, Yunjun Gao, Baihua Zheng, Gang Chen, Chun Chen, Qing Li

Research Collection School Of Computing and Information Systems

Despite the ubiquity of physical obstacles (e.g., buildings, hills, and blindages, etc.) in the real world, most of spatial queries ignore the obstacles. In this article, we study a novel form of continuous nearest-neighbor queries in the presence of obstacles, namely continuous obstructed nearest-neighbor (CONN) search, which considers the impact of obstacles on the distance between objects. Given a data setP, an obstacle set O, and a query line segment q, in a two-dimensional space, a CONN query retrieves the nearest neighbor p ∈ P of each point p′ on q according to the obstructed distance, the shortest path between …


Methods For Multilevel Parallelism On Gpu Clusters: Application To A Multigrid Accelerated Navier-Stokes Solver, Dana A. Jacobsen May 2011

Methods For Multilevel Parallelism On Gpu Clusters: Application To A Multigrid Accelerated Navier-Stokes Solver, Dana A. Jacobsen

Boise State University Theses and Dissertations

Computational Fluid Dynamics (CFD) is an important field in high performance computing with numerous applications. Solving problems in thermal and fluid sciences demands enormous computing resources and has been one of the primary applications used on supercomputers and large clusters. Modern graphics processing units (GPUs) with many-core architectures have emerged as general-purpose parallel computing platforms that can accelerate simulation science applications substantially. While significant speedups have been obtained with single and multiple GPUs on a single workstation, large problems require more resources. Conventional clusters of central processing units (CPUs) are now being augmented with GPUs in each compute-node to tackle …


Fragile Online Relationship: A First Look At Unfollow Dynamics In Twitter, Haewoon Kwak, Hyunwoo Chun, Sue. Moon May 2011

Fragile Online Relationship: A First Look At Unfollow Dynamics In Twitter, Haewoon Kwak, Hyunwoo Chun, Sue. Moon

Research Collection School Of Computing and Information Systems

We analyze the dynamics of the behavior known as 'unfollow' in Twitter. We collected daily snapshots of the online relationships of 1.2 million Korean-speaking users for 51 days as well as all of their tweets. We found that Twitter users frequently unfollow. We then discover the major factors, including the reciprocity of the relationships, the duration of a relationship, the followees' informativeness, and the overlap of the relationships, which affect the decision to unfollow. We conduct interview with 22 Korean respondents to supplement the quantitative results.They unfollowed those who left many tweets within a short time, created tweets about uninteresting …


Comparing Twitter And Traditional Media Using Topic Models, Wayne Xin Zhao, Jing Jiang, Jianshu Weng, Jing He, Ee Peng Lim, Hongfei Yan, Xiaoming Li Apr 2011

Comparing Twitter And Traditional Media Using Topic Models, Wayne Xin Zhao, Jing Jiang, Jianshu Weng, Jing He, Ee Peng Lim, Hongfei Yan, Xiaoming Li

Research Collection School Of Computing and Information Systems

Twitter as a new form of social media can potentially contain much useful information, but content analysis on Twitter has not been well studied. In particular, it is not clear whether as an information source Twitter can be simply regarded as a faster news feed that covers mostly the same information as traditional news media. In This paper we empirically compare the content of Twitter with a traditional news medium, New York Times, using unsupervised topic modeling. We use a Twitter-LDA model to discover topics from a representative sample of the entire Twitter. We then use text mining techniques to …


A Probabilistic Analysis Of Misparking In Reservation Based Parking Garages, Vikas G. Ashok Apr 2011

A Probabilistic Analysis Of Misparking In Reservation Based Parking Garages, Vikas G. Ashok

Computer Science Theses & Dissertations

Parking in major cities is an expensive and annoying affair, the reason ascribed to the limited availability of parking space. Modern parking garages provide parking reservation facility, thereby ensuring availability to prospective customers. Misparking in such reservation based parking garages creates confusion and aggravates driver frustration. The general conception about misparking is that it tends to completely cripple the normal functioning of the system leading to chaos and confusion. A single mispark tends to have a ripple effect and therefore spawns a chain of misparks. The chain terminates when the last mispark occurs at the parking slot reserved by the …


Efficient Topological Olap On Information Networks, Qiang Qu, Feida Zhu, Xifeng Yan, Jiawei Han, Philip Yu, Hongyan Li Apr 2011

Efficient Topological Olap On Information Networks, Qiang Qu, Feida Zhu, Xifeng Yan, Jiawei Han, Philip Yu, Hongyan Li

Research Collection School Of Computing and Information Systems

We propose a framework for efficient OLAP on information networks with a focus on the most interesting kind, the topological OLAP (called “T-OLAP”), which incurs topological changes in the underlying networks. T-OLAP operations generate new networks from the original ones by rolling up a subset of nodes chosen by certain constraint criteria. The key challenge is to efficiently compute measures for the newly generated networks and handle user queries with varied constraints. Two effective computational techniques, T-Distributiveness and T-Monotonicity are proposed to achieve efficient query processing and cube materialization. We also provide a T-OLAP query processing framework into which these …


Predicting Item Adoption Using Social Correlation, Freddy Chong-Tat Chua, Hady W. Lauw, Ee Peng Lim Apr 2011

Predicting Item Adoption Using Social Correlation, Freddy Chong-Tat Chua, Hady W. Lauw, Ee Peng Lim

Research Collection School Of Computing and Information Systems

Users face a dazzling array of choices on the Web when it comes to choosing which product to buy, which video to watch, etc. The trend of social information processing means users increasingly rely not only on their own preferences, but also on friends when making various adoption decisions. In this paper, we investigate the effects of social correlation on users’ adoption of items. Given a user-user social graph and an item-user adoption graph, we seek to answer the following questions: 1) whether the items adopted by a user correlate to items adopted by her friends, and 2) how to …


Mkboost: A Framework Of Multiple Kernel Boosting, Hao Xia, Steven C. H. Hoi Apr 2011

Mkboost: A Framework Of Multiple Kernel Boosting, Hao Xia, Steven C. H. Hoi

Research Collection School Of Computing and Information Systems

Multiple kernel learning (MKL) has been shown as a promising machine learning technique for data mining tasks by integrating with multiple diverse kernel functions. Traditional MKL methods often formulate the problem as an optimization task of learning both optimal combination of kernels and classifiers, and attempt to resolve the challenging optimization task by various techniques. Unlike the existing MKL methods, in this paper, we investigate a boosting framework of exploring multiple kernel learning for classification tasks. In particular, we present a novel framework of Multiple Kernel Boosting (MKBoost), which applies boosting techniques for learning kernel-based classifiers with multiple kernels. Based …


Multi-Objective Zone Mapping In Large-Scale Distributed Virtual Environments, Nguyen Binh Duong Ta, Suiping Zhou, Wentong Cai, Xueyan Tang, Rassul Avani Mar 2011

Multi-Objective Zone Mapping In Large-Scale Distributed Virtual Environments, Nguyen Binh Duong Ta, Suiping Zhou, Wentong Cai, Xueyan Tang, Rassul Avani

Research Collection School Of Computing and Information Systems

In large-scale distributed virtual environments (DVEs), the NP-hard zone mapping problem concerns how to assign distinct zones of the virtual world to a number of distributed servers to improve overall interactivity. Previously, this problem has been formulated as a single-objective optimization problem, in which the objective is to minimize the total number of clients that are without QoS. This approach may cause considerable network traffic and processing overhead, as a large number of zones may need to be migrated across servers. In this paper, we introduce a multi-objective approach to the zone mapping problem, in which both the total number …


Mining Social Images With Distance Metric Learning For Automated Image Tagging, Pengcheng Wu, Steven C. H. Hoi, Peilin Zhao, Ying He Feb 2011

Mining Social Images With Distance Metric Learning For Automated Image Tagging, Pengcheng Wu, Steven C. H. Hoi, Peilin Zhao, Ying He

Research Collection School Of Computing and Information Systems

With the popularity of various social media applications, massive social images associated with high quality tags have been made available in many social media web sites nowadays. Mining social images on the web has become an emerging important research topic in web search and data mining. In this paper, we propose a machine learning framework for mining social images and investigate its application to automated image tagging. To effectively discover knowledge from social images that are often associated with multimodal contents (including visual images and textual tags), we propose a novel Unified Distance Metric Learning (UDML) scheme, which not only …


Interdisciplinary Computational Projects Utilizing The Hpc Cluster, Demos Athanasopoulos Jan 2011

Interdisciplinary Computational Projects Utilizing The Hpc Cluster, Demos Athanasopoulos

Cornerstone 3 Reports : Interdisciplinary Informatics

No abstract provided.


A Parallel Graph Sampling Algorithm For Analyzing Gene Correlation Networks, Kathryn Dempsey Cooper, Kanimathi Duraisamy, Hesham Ali, Sanjukta Bhowmick Jan 2011

A Parallel Graph Sampling Algorithm For Analyzing Gene Correlation Networks, Kathryn Dempsey Cooper, Kanimathi Duraisamy, Hesham Ali, Sanjukta Bhowmick

Interdisciplinary Informatics Faculty Publications

Effcient analysis of complex networks is often a challenging task due to its large size and the noise inherent in the system. One popular method of overcoming this problem is through graph sampling, that is extracting a representative subgraph from the larger network. The accuracy of the sample is validated by comparing the combinatorial properties of the subgraph and the original network. However, there has been little study in comparing networks based on the applications that they represent. Furthermore, sampling methods are generally applied agnostically, without mapping to the requirements of the underlying analysis. In this paper,we introduce a parallel …


In-Degree Dynamics Of Large-Scale P2p Systems, Zhongmei Yao, Daren B. H. Cline, Dmitri Loguinov Jan 2011

In-Degree Dynamics Of Large-Scale P2p Systems, Zhongmei Yao, Daren B. H. Cline, Dmitri Loguinov

Computer Science Faculty Publications

This paper builds a complete modeling framework for understanding user churn and in-degree dynamics in unstructured P2P systems in which each user can be viewed as a stationary alternating renewal process. While the classical Poisson result on the superposition of n stationary renewal processes for n→∞ requires that each point process become sparser as n increases, it is often difficult to rigorously show this condition in practice. In this paper, we first prove that despite user heterogeneity and non-Poisson arrival dynamics, a superposition of edge-arrival processes to a live user under uniform selection converges to a Poisson process when …


Evaluation Of Essential Genes In Correlation Networks Using Measures Of Centrality, Kathryn Dempsey Cooper, Hesham Ali Jan 2011

Evaluation Of Essential Genes In Correlation Networks Using Measures Of Centrality, Kathryn Dempsey Cooper, Hesham Ali

Interdisciplinary Informatics Faculty Proceedings & Presentations

Correlation networks are emerging as powerful tools for modeling relationships in high-throughput data such as gene expression. Other types of biological networks, such as protein-protein interaction networks, are popular targets of study in network theory, and previous analysis has revealed that network structures identified using graph theoretic techniques often relate to certain biological functions. Structures such as highly connected nodes and groups of nodes have been found to correspond to essential genes and protein complexes, respectively. The correlation network, which measures the level of co-variation of gene expression levels, shares some structural properties with other types of biological networks. We …


A Novel Correlation Networks Approach For The Identification Of Gene Targets, Kathryn Dempsey Cooper, Stephen Bonasera, Dhundy Raj Bastola, Hesham Ali Jan 2011

A Novel Correlation Networks Approach For The Identification Of Gene Targets, Kathryn Dempsey Cooper, Stephen Bonasera, Dhundy Raj Bastola, Hesham Ali

Interdisciplinary Informatics Faculty Proceedings & Presentations

Correlation networks are emerging as a powerful tool for modeling temporal mechanisms within the cell. Particularly useful in examining coexpression within microarray data, studies have determined that correlation networks follow a power law degree distribution and thus manifest properties such as the existence of “hub” nodes and semicliques that potentially correspond to critical cellular structures. Difficulty lies in filtering coincidental relationships from causative structures in these large, noise-heavy networks. As such, computational expenses and algorithm availability limit accurate comparison, making it difficult to identify changes between networks. In this vein, we present our work identifying temporal relationships from microarray data …


A Noise Reducing Sampling Approach For Uncovering Critical Properties In Large Scale Biological Networks, Karthik Duraisamy, Kathryn Dempsey Cooper, Hesham Ali, Sanjukta Bhowmick Jan 2011

A Noise Reducing Sampling Approach For Uncovering Critical Properties In Large Scale Biological Networks, Karthik Duraisamy, Kathryn Dempsey Cooper, Hesham Ali, Sanjukta Bhowmick

Interdisciplinary Informatics Faculty Proceedings & Presentations

A correlation network is a graph-based representation of relationships among genes or gene products, such as proteins. The advent of high-throughput bioinformatics has resulted in the generation of volumes of data that require sophisticated in silico models, such as the correlation network, for in-depth analysis. Each element in our network represents expression levels of multiple samples of one gene and an edge connecting two nodes reflects the correlation level between the two corresponding genes in the network according to the Pearson correlation coefficient. Biological networks made in this manner are generally found to adhere to a scale-free structural nature, that …


Real-Time Road Traffic Prediction With Spatio-Temporal Correlations, Wanli Min, Laura Wynter Jan 2011

Real-Time Road Traffic Prediction With Spatio-Temporal Correlations, Wanli Min, Laura Wynter

Research Collection School Of Computing and Information Systems

Real-time road traffic prediction is a fundamental capability needed to make use of advanced, smart transportation technologies. Both from the point of view of network operators as well as from the point of view of travelers wishing real-time route guidance, accurate short-term traffic prediction is a necessary first step. While techniques for short-term traffic prediction have existed for some time, emerging smart transportation technologies require the traffic prediction capability to be both fast and scalable to full urban networks. We present a method that has proven to be able to meet this challenge. The method presented provides predictions of speed …