Open Access. Powered by Scholars. Published by Universities.®

2012

Discipline
Institution
Keyword
Publication
Publication Type

Articles 1 - 30 of 75

Full-Text Articles in Numerical Analysis and Scientific Computing

The R Journal (December 2012) 4(2): Complete Issue, The R Foundation Dec 2012

The R Journal (December 2012) 4(2): Complete Issue, The R Foundation

The R Journal

Contributing Articles

What's in a Name? Paul Murrell

It's Not What You Draw, It's What You Don't Draw, Paul Murrell

Debugging grid Graphics, Paul Murrell and Velvet Ly

frailtyHL: A Package for Fitting Frailty Models with H-likelihood, Il Do Ha, Maengseok Noh, and Youngjo Lee

influence.ME: Tools for Detecting Influential Data in Mixed Effects Models, Rense Nieuwenhuis, Manfred te Grotenhuis and Ben Pelzer

The crs Package: Nonparametric Regression Splines for Continuous and Categorical Predictors, Zhenghua Nie and Jeffrey S. Racine

Rfit: Rank-based Estimation for Linear Models, John D. Kloke and Joseph W. McKean

Graphical Markov Models with Mixed Graphs in …


The Crs Package: Nonparametric Regression Splines For Continuous And Categorical Predictors, Zhenghua Nie, Jeffery S. Racine Dec 2012

The Crs Package: Nonparametric Regression Splines For Continuous And Categorical Predictors, Zhenghua Nie, Jeffery S. Racine

The R Journal

A new package crs is introduced for computing nonparametric regression (and quantile) splines in the presence of both continuous and categorical predictors. B-splines are employed in the regression model for the continuous predictors and kernel weighting is employed for the categorical predictors. We also de velop a simple R interface to NOMAD, which is a mixed integer optimization solver used to compute optimal regression spline solutions.


Influence.Me: Tools For Detecting Influential Data In Mixed Effects Models, Rense Nieuwenhuis, Manfred Te Grotenhuis, Ben Pelzer Dec 2012

Influence.Me: Tools For Detecting Influential Data In Mixed Effects Models, Rense Nieuwenhuis, Manfred Te Grotenhuis, Ben Pelzer

The R Journal

influence.ME provides tools for detecting influential data in mixed effects models. The application of these models has become common practice, but the development of diagnostic tools has lagged behind. influence.ME calculates standardized measures of influential data for the point estimates of generalized mixed effects models, such as DFBETAS, Cook’s distance, as well as percentile change and a test for changing levels of significance. influence.ME calculates these measures of influence while ac counting for the nesting structure of the data. The package and measures of influential data are introduced, a practical example is given, and strategies for dealing with influential data …


Graphical Markov Models With Mixed Graphs In R, Kayvan Sadeghi, Giovanni M. Marchetti Dec 2012

Graphical Markov Models With Mixed Graphs In R, Kayvan Sadeghi, Giovanni M. Marchetti

The R Journal

In this paper we provide a short tuto rial illustrating the new functions in the package ggm that deal with ancestral, summary and ribbonless graphs. These are mixed graphs (containing three types of edges) that are important because they capture the modified independence structure after marginalisation over, and conditioning on, nodes of directed acyclic graphs. We provide functions to verify whether a mixed graph implies that A is independent of B given C for any disjoint sets of nodes and to generate maximal graphs inducing the same independence structure of non-maximal graphs. Finally, we provide functions to decide on the …


What's In A Name?, Paul Murrell Dec 2012

What's In A Name?, Paul Murrell

The R Journal

Any shape that is drawn using the grid graphics package can have a name associated with it. If a name is provided, it is possible to access, query, and modify the shape after it has been drawn. These facilities allow for very detailed customisations of plots and also for very general transformations of plots that are drawn by packages based on grid.


Frailtyhl: A Package For Fitting Frailty Models With H-Likelihood, Il Do Ha, Maengseok Noh, Youngjo Lee Dec 2012

Frailtyhl: A Package For Fitting Frailty Models With H-Likelihood, Il Do Ha, Maengseok Noh, Youngjo Lee

The R Journal

We present the frailtyHL package for fitting semi-parametric frailty models using h likelihood. This package allows lognormal or gamma frailties for random-effect distribution, and it fits shared or multilevel frailty models for correlated survival data. Functions are provided to format and summarize the frailtyHL results. The estimates of fixed effects and frailty parameters and their standard errors are calculated. We illustrate the use of our package with three well known data sets and compare our results with various alternative R-procedures.


Rfit: Rank-Based Estimation For Linear Models, John D. Kloke, Joseph W. Mckeen Dec 2012

Rfit: Rank-Based Estimation For Linear Models, John D. Kloke, Joseph W. Mckeen

The R Journal

In the nineteen seventies, Jureĉková and Jaeckel proposed rank estimation for linear models. Since that time, several authors have developed inference and diagnostic methods for these estimators. These rank-based estimators and their associated inference are highly efficient and are robust to outliers in response space. The methods include estimation of standard errors, tests of general linear hypotheses, confidence intervals, diagnostic procedures including studentized residuals, and measures of influential cases. We have developed an R package, Rfit, for computing of these robust procedures. In this paper we highlight the main features of the pack age. The package uses standard linear …


It's Not What You Draw, It's What You Don't Draw, Paul Murrell Dec 2012

It's Not What You Draw, It's What You Don't Draw, Paul Murrell

The R Journal

The R graphics engine has new support for drawing complex paths via the functions polypath() and grid.path(). This article explains what is meant by a complex path and demonstrates the usefulness of complex paths in drawing non-trivial shapes, logos, customised data symbols, and maps.


Debugging Grid Graphics, Paul Murrell, Velvet Ly Dec 2012

Debugging Grid Graphics, Paul Murrell, Velvet Ly

The R Journal

A graphical scene that has been produced using the grid graphics package consists of grobs (graphical objects) and viewports. This article describes functions that allow the exploration and inspection of the grobs and viewports in a grid scene, including several functions that are available in a new package called gridDe bug. The ability to explore the grobs and view ports in a grid scene is useful for adding more drawing to a scene that was produced using grid and for understanding and debugging the grid code that produced a scene.


The State Of Naming Conventions In R, Rasmus Bååth Dec 2012

The State Of Naming Conventions In R, Rasmus Bååth

The R Journal

Most programming language communities have naming conventions that are generally agreed upon, that is, a set of rules that governs how functions and variables are named. This is not the case with R, and a review of unofficial style guides and naming convention us age on CRAN shows that a number of different naming conventions are currently in use. Some naming conventions are, however, more popular than others and as a newcomer to the R community or as a developer of a new package this could be useful to consider when choosing what naming convention to adopt.


Design And Optimization Of Four-Dimensional Cone-Beam Computed To Mography In Image-Guided Radiation Therapy, Moiz Ahmad Dec 2012

Design And Optimization Of Four-Dimensional Cone-Beam Computed To Mography In Image-Guided Radiation Therapy, Moiz Ahmad

Dissertations and Theses (Open Access)

The influence of respiratory motion on patient anatomy poses a challenge to accurate radiation therapy, especially in lung cancer treatment. Modern radiation therapy planning uses models of tumor respiratory motion to account for target motion in targeting. The tumor motion model can be verified on a per-treatment session basis with four-dimensional cone-beam computed tomography (4D-CBCT), which acquires an image set of the dynamic target throughout the respiratory cycle during the therapy session. 4D-CBCT is undersampled if the scan time is too short. However, short scan time is desirable in clinical practice to reduce patient setup time. This dissertation presents the …


A Survey Of Recommender Systems In Twitter, Su Mon Kywe, Ee Peng Lim, Feida Zhu Dec 2012

A Survey Of Recommender Systems In Twitter, Su Mon Kywe, Ee Peng Lim, Feida Zhu

Research Collection School Of Computing and Information Systems

Twitter is a social information network where short messages or tweets are shared among a large number of users through a very simple messaging mechanism. With a population of more than 100M users generating more than 300M tweets each day, Twitter users can be easily overwhelmed by the massive amount of information available and the huge number of people they can interact with. To overcome the above information overload problem, recommender systems can be introduced to help users make the appropriate selection. Researchers have began to study recommendation problems in Twitter but their works usually address individual recommendation tasks. There …


Detecting Anomalies In Bipartite Graphs With Mutual Dependency Principles, Hanbo Dai, Feida Zhu, Ee Peng Lim, Hwee Hwa Pang Dec 2012

Detecting Anomalies In Bipartite Graphs With Mutual Dependency Principles, Hanbo Dai, Feida Zhu, Ee Peng Lim, Hwee Hwa Pang

Research Collection School Of Computing and Information Systems

Bipartite graphs can model many real life applications including users-rating-products in online marketplaces, users-clicking-webpages on the World Wide Web and users referring users in social networks. In these graphs, the anomalousness of nodes in one partite often depends on that of their connected nodes in the other partite. Previous studies have shown that this dependency can be positive (the anomalousness of a node in one partite increases or decreases along with that of its connected nodes in the other partite) or negative (the anomalousness of a node in one partite rises or falls in opposite direction to that of its …


Cost-Sensitive Online Classification, Jialei Wang, Peilin Zhao, Steven C. H. Hoi Dec 2012

Cost-Sensitive Online Classification, Jialei Wang, Peilin Zhao, Steven C. H. Hoi

Research Collection School Of Computing and Information Systems

Both cost-sensitive classification and online learning have been extensively studied in data mining and machine learning communities, respectively. However, very limited study addresses an important intersecting problem, that is, “Cost-Sensitive Online Classification". In this paper, we formally study this problem, and propose a new framework for Cost-Sensitive Online Classification by directly optimizing cost-sensitive measures using online gradient descent techniques. Specifically, we propose two novel cost-sensitive online classification algorithms, which are designed to directly optimize two well-known cost-sensitive measures: (i) maximization of weighted sum of sensitivity and specificity, and (ii) minimization of weighted misclassification cost. We analyze the theoretical bounds of …


Visualization For Anomaly Detection And Data Management By Leveraging Network, Sensor And Gis Techniques, Zhaoxia Wang, Chee Seng Chong, Rick S. M. Goh, Wanqing Zhou, Dan Peng, Hoong Chor Chin Dec 2012

Visualization For Anomaly Detection And Data Management By Leveraging Network, Sensor And Gis Techniques, Zhaoxia Wang, Chee Seng Chong, Rick S. M. Goh, Wanqing Zhou, Dan Peng, Hoong Chor Chin

Research Collection School Of Computing and Information Systems

This paper studies the importance of visualization for discerning and interpreting patterns of data and its application for solving real problems, such as anomaly detection and data management. There are various ways to realize visualization to cater to the needs of numerous real life applications. Depending on needs, a combination of some of these ways may be required for presenting an effective visualization. The authors present visualization schemes for anomaly detection/condition monitoring and data management by leveraging network techniques and combining them with modern techniques such as sensor, database, mobile communication, GPS and GIS techniques. Two case studies are presented …


Fast And Accurate Psd Matrix Estimation By Row Reduction, Hiroshi Kuwajima, Takashi Washio, Ee Peng Lim Nov 2012

Fast And Accurate Psd Matrix Estimation By Row Reduction, Hiroshi Kuwajima, Takashi Washio, Ee Peng Lim

Research Collection School Of Computing and Information Systems

Fast and accurate estimation of missing relations, e.g., similarity, distance and kernel, among objects is now one of the most important techniques required by major data mining tasks, because the missing information of the relations is needed in many applications such as economics, psychology, and social network communities. Though some approaches have been proposed in the last several years, the practical balance between their required computation amount and obtained accuracy are insufficient for some class of the relation estimation. The objective of this paper is to formalize a problem to quickly and efficiently estimate missing relations among objects from the …


Impact Of Multimedia In Sina Weibo: Popularity And Life Span, Xun Zhao, Feida Zhu, Weining Qian, Aoying Zhou Nov 2012

Impact Of Multimedia In Sina Weibo: Popularity And Life Span, Xun Zhao, Feida Zhu, Weining Qian, Aoying Zhou

Research Collection School Of Computing and Information Systems

Multimedia contents such as images and videos are widely used in social network sites nowadays. Sina Weibo, a Chinese microblogging service, is one of the first microblog platforms to incorporate multimedia content sharing features. This work provides statistical analysis on how multimedia contents are produced, consumed, and propagated in Sina Weibo. Based on 230 million tweets and 1.8 million user profiles in Sina Weibo, we study the impact of multimedia contents on the popularity of both users and tweets as well as tweet life span. Our preliminary study shows that multimedia tweets dominant pure text ones in SinaWeibo. Multimedia contents …


Divad: A Dynamic And Interactive Visual Analytical Dashboard For Exploring And Analyzing Transport Data, Tin Seong Kam, Ketan Barshikar, Shaun Jun Hua Tan Nov 2012

Divad: A Dynamic And Interactive Visual Analytical Dashboard For Exploring And Analyzing Transport Data, Tin Seong Kam, Ketan Barshikar, Shaun Jun Hua Tan

Research Collection School Of Computing and Information Systems

The advances in location-based data collection technologies such as GPS, RFID etc. and the rapid reduction of their costs provide us with a huge and continuously increasing amount of data about movement of vehicles, people and goods in an urban area. This explosive growth of geospatially-referenced data has far outpaced the planner’s ability to utilize and transform the data into insightful information thus creating an adverse impact on the return on the investment made to collect and manage this data. Addressing this pressing need, we designed and developed DIVAD, a dynamic and interactive visual analytics dashboard to allow city planners …


A Unified Learning Framework For Auto Face Annotation By Mining Web Facial Images, Dayong Wang, Steven C. H. Hoi, Ying He Nov 2012

A Unified Learning Framework For Auto Face Annotation By Mining Web Facial Images, Dayong Wang, Steven C. H. Hoi, Ying He

Research Collection School Of Computing and Information Systems

Auto face annotation plays an important role in many real-world multimedia information and knowledge management systems. Recently there is a surge of research interests in mining weakly-labeled facial images on the internet to tackle this long-standing research challenge in computer vision and image understanding. In this paper, we present a novel unified learning framework for face annotation by mining weakly labeled web facial images through interdisciplinary efforts of combining sparse feature representation, content-based image retrieval, transductive learning and inductive learning techniques. In particular, we first introduce a new search-based face annotation paradigm using transductive learning, and then propose an effective …


Self-Regulating Action Exploration In Reinforcement Learning, Teck-Hou Teng, Ah-Hwee Tan Oct 2012

Self-Regulating Action Exploration In Reinforcement Learning, Teck-Hou Teng, Ah-Hwee Tan

Research Collection School Of Computing and Information Systems

The basic tenet of a learning process is for an agent to learn for only as much and as long as it is necessary. With reinforcement learning, the learning process is divided between exploration and exploitation. Given the complexity of the problem domain and the randomness of the learning process, the exact duration of the reinforcement learning process can never be known with certainty. Using an inaccurate number of training iterations leads either to the non-convergence or the over-training of the learning agent. This work addresses such issues by proposing a technique to self-regulate the exploration rate and training duration …


Entity Synonyms For Structured Web Search, Tao Cheng, Hady W. Lauw, Stelios Paparizos Oct 2012

Entity Synonyms For Structured Web Search, Tao Cheng, Hady W. Lauw, Stelios Paparizos

Research Collection School Of Computing and Information Systems

Nowadays, there are many queries issued to search engines targeting at finding values from structured data (e.g., movie showtime of a specific location). In such scenarios, there is often a mismatch between the values of structured data (how content creators describe entities) and the web queries (how different users try to retrieve them). Therefore, recognizing the alternative ways people use to reference an entity, is crucial for structured web search. In this paper, we study the problem of automatic generation of entity synonyms over structured data toward closing the gap between users and structured data. We propose an offline, data-driven …


Influentials, Novelty, And Social Contagion: The Viral Power Of Average Friends, Close Communities, And Old News, Nicholas Harrigan, Palakorn Achananuparp, Ee Peng Lim Oct 2012

Influentials, Novelty, And Social Contagion: The Viral Power Of Average Friends, Close Communities, And Old News, Nicholas Harrigan, Palakorn Achananuparp, Ee Peng Lim

Research Collection School Of Computing and Information Systems

What is the effect of (1) popular individuals, and (2) community structures on the retransmission of socially contagious behavior? We examine a community of Twitter users over a five month period, operationalizing social contagion as ‘retweeting’, and social structure as the count of subgraphs (small patterns of ties and nodes) between users in the follower/following network. We find that popular individuals act as ‘inefficient hubs’ for social contagion: they have limited attention, are overloaded with inputs, and therefore display limited responsiveness to viral messages. We argue this contradicts the ‘law of the few’ and ‘influentials hypothesis’. We find that community …


Semi-Automatic Simulation Initialization By Mining Structured And Unstructured Data Formats From Local And Web Data Sources, Olcay Sahin Oct 2012

Semi-Automatic Simulation Initialization By Mining Structured And Unstructured Data Formats From Local And Web Data Sources, Olcay Sahin

Computational Modeling & Simulation Engineering Theses & Dissertations

Initialization is one of the most important processes for obtaining successful results from a simulation. However, initialization is a challenge when 1) a simulation requires hundreds or even thousands of input parameters or 2) re-initializing the simulation due to different initial conditions or runtime errors. These challenges lead to the modeler spending more time initializing a simulation and may lead to errors due to poor input data.

This thesis proposes two semi-automatic simulation initialization approaches that provide initialization using data mining from structured and unstructured data formats from local and web data sources. First, the System Initialization with Retrieval (SIR) …


Retrieval Of Sub-Pixel-Based Fire Intensity And Its Application For Characterizing Smoke Injection Heights And Fire Weather In North America, David Peterson Sep 2012

Retrieval Of Sub-Pixel-Based Fire Intensity And Its Application For Characterizing Smoke Injection Heights And Fire Weather In North America, David Peterson

Department of Earth and Atmospheric Sciences: Dissertations, Theses, and Student Research

For over two decades, satellite sensors have provided the locations of global fire activity with ever-increasing accuracy. However, the ability to measure fire intensity, know as fire radiative power (FRP), and its potential relationships to meteorology and smoke plume injection heights, are currently limited by the pixel resolution. This dissertation describes the development of a new, sub-pixel-based FRP calculation (FRPf) for fire pixels detected by the MODerate Resolution Imaging Spectroradiometer (MODIS) fire detection algorithm (Collection 5), which is subsequently applied to several large wildfire events in North America. The methodology inherits an earlier bi-spectral algorithm for retrieving sub-pixel …


The Implementation Of The Shear Correlation Function And The Matter Power Spectrum In R, Allison A. Scheppelmann, Deborah J. Bard Aug 2012

The Implementation Of The Shear Correlation Function And The Matter Power Spectrum In R, Allison A. Scheppelmann, Deborah J. Bard

STAR Program Research Presentations

Weak gravitational lensing is an important tool in understanding the large-scale structure of the universe. One component in understanding the effect of weak gravitational lensing is the shear correlation function and matter power spectrum. The calculation of these values is often complicated and time consuming. In order to decrease the cost of these calculations they were implemented in R using parallelization. This resulted in the calculations completing faster and the process to be easily changed in order to fit the need of each researcher using the algorithms created in R.


A Domain Specific Model For Generating Etl Workflows From Business Intents, Wesley Deneke Aug 2012

A Domain Specific Model For Generating Etl Workflows From Business Intents, Wesley Deneke

Graduate Theses and Dissertations

Extract-Transform-Load (ETL) tools have provided organizations with the ability to build and maintain workflows (consisting of graphs of data transformation tasks) that can process the flood of digital data. Currently, however, the specification of ETL workflows is largely manual, human time intensive, and error prone. As these workflows become increasingly complex, the users that build and maintain them must retain an increasing amount of knowledge specific to how to produce solutions to business objectives using their domain's ETL workflow system. A program that can reduce the human time and expertise required to define such workflows, producing accurate ETL solutions with …


On The K-Mer Frequency Spectra Of Organism Genome And Proteome Sequences With A Preliminary Machine Learning Assessment Of Prime Predictability, Nathan O. Schmidt Aug 2012

On The K-Mer Frequency Spectra Of Organism Genome And Proteome Sequences With A Preliminary Machine Learning Assessment Of Prime Predictability, Nathan O. Schmidt

Boise State University Theses and Dissertations

A regular expression and region-specific filtering system for biological records at the National Center for Biotechnology database is integrated into an object oriented sequence counting application, and a statistical software suite is designed and deployed to interpret the resulting k-mer frequencies|with a priority focus on nullomers. The proteome k-mer frequency spectra of ten model organisms and the genome k-mer frequency spectra of two bacteria and virus strains for the coding and non-coding regions are comparatively scrutinized. We observe that the naturally-evolved (NCBI/organism) and the artificially-biased (randomly-generated) sequences exhibit a clear deviation from the artificially-unbiased (randomly-generated) histogram distributions. …


Shortest Path Computation With No Information Leakage, Kyriakos Mouratidis, Man Lung Yiu Aug 2012

Shortest Path Computation With No Information Leakage, Kyriakos Mouratidis, Man Lung Yiu

Research Collection School Of Computing and Information Systems

Shortest path computation is one of the most common queries in location-based services (LBSs). Although particularly useful, such queries raise serious privacy concerns. Exposing to a (potentially untrusted) LBS the client’s position and her destination may reveal personal information, such as social habits, health condition, shopping preferences, lifestyle choices, etc. The only existing method for privacy-preserving shortest path computation follows the obfuscation paradigm; it prevents the LBS from inferring the source and destination of the query with a probability higher than a threshold. This implies, however, that the LBS still deduces some information (albeit not exact) about the client’s location …


Modeling Concept Dynamics For Large Scale Music Search, Jialie Shen, Hwee Hwa Pang, Meng Wang, Shuicheng Yan Aug 2012

Modeling Concept Dynamics For Large Scale Music Search, Jialie Shen, Hwee Hwa Pang, Meng Wang, Shuicheng Yan

Research Collection School Of Computing and Information Systems

Continuing advances in data storage and communication technologies have led to an explosive growth in digital music collections. To cope with their increasing scale, we need effective Music Information Retrieval (MIR) capabilities like tagging, concept search and clustering. Integral to MIR is a framework for modelling music documents and generating discriminative signatures for them. In this paper, we introduce a multimodal, layered learning framework called DMCM. Distinguished from the existing approaches that encode music as an ensemble of order-less feature vectors, our framework extracts from each music document a variety of acoustic features, and translates them into low-level encodings over …


A Non-Parametric Visual-Sense Model Of Images: Extending The Cluster Hypothesis Beyond Text, Kong-Wah Wan, Ah-Hwee Tan, Joo-Hwee Lim, Liang-Tien Chia Aug 2012

A Non-Parametric Visual-Sense Model Of Images: Extending The Cluster Hypothesis Beyond Text, Kong-Wah Wan, Ah-Hwee Tan, Joo-Hwee Lim, Liang-Tien Chia

Research Collection School Of Computing and Information Systems

The main challenge of a search engine is to find information that are relevant and appropriate. However, this can become difficult when queries are issued using ambiguous words. Rijsbergen first hypothesized a clustering approach for web pages wherein closely associated pages are treated as a semantic group with the same relevance to the query (Rijsbergen 1979). In this paper, we extend Rijsbergen’s cluster hypothesis to multimedia content such as images. Given a user query, the polysemy in the return image set is related to the many possible meanings of the query. We develop a method to cluster the polysemous images …