Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 121 - 150 of 204

Full-Text Articles in Numerical Analysis and Computation

Multiple Subject Barycentric Discriminant Analysis (Musubada): How To Assign Scans To Categories Without Using Spatial Normalization, Hervé Abdi, Lynne J. Williams, Andrew C. Connolly, M. Ida Gobbini Dec 2012

Multiple Subject Barycentric Discriminant Analysis (Musubada): How To Assign Scans To Categories Without Using Spatial Normalization, Hervé Abdi, Lynne J. Williams, Andrew C. Connolly, M. Ida Gobbini

Dartmouth Scholarship

We present a new discriminant analysis (DA) method called Multiple Subject Barycentric Discriminant Analysis (MUSUBADA) suited for analyzing fMRI data because it handles datasets with multiple participants that each provides different number of variables (i.e., voxels) that are themselves grouped into regions of interest (ROIs). Like DA, MUSUBADA (1) assigns observations to predefined categories, (2) gives factorial maps displaying observations and categories, and (3) optimally assigns observations to categories. MUSUBADA handles cases with more variables than observations and can project portions of the data table (e.g., subtables, which can represent participants or ROIs) on the factorial maps. Therefore MUSUBADA can …


Modeling And Mathematical Analysis Of Plant Models In Ecology, Eric A. Eager Aug 2012

Modeling And Mathematical Analysis Of Plant Models In Ecology, Eric A. Eager

Department of Mathematics: Dissertations, Theses, and Student Research

Population dynamics tries to explain in a simple mechanistic way the variations of the size and structure of biological populations. In this dissertation we use mathematical modeling and analysis to study the various aspects of the dynamics of plant populations and their seed banks.

In Chapter 2 we investigate the impact of structural model uncertainty by considering different nonlinear recruitment functions in an integral projection model for Cirsium canescens. We show that, while having identical equilibrium populations, these two models can elicit drastically different transient dynamics. We then derive a formula for the sensitivity of the equilibrium population to …


Half-Life Learning Curves In The Defense Acquisition Life Cycle, Adedeji B. Badiru Jul 2012

Half-Life Learning Curves In The Defense Acquisition Life Cycle, Adedeji B. Badiru

Faculty Publications

Learning curves are useful for assessing performance improvement due to the positive impact of learning. In recent years, the deleterious effects of forgetting have also been recognized. Workers experience forgetting or decline in performance over time. Consequently, contemporary learning curves have attempted to incorporate forgetting components into learning curves. An area of increasing interest is the study of how fast and how far the forgetting impact can influence overall performance. This article introduces the concept of half-life analysis of learning curves using the concept of growth and decay, with particular emphasis on applications in the defense acquisition process. The computational …


Random Number Generation: Types And Techniques, David F. Dicarlo Apr 2012

Random Number Generation: Types And Techniques, David F. Dicarlo

Senior Honors Theses

What does it mean to have random numbers? Without understanding where a group of numbers came from, it is impossible to know if they were randomly generated. However, common sense claims that if the process to generate these numbers is truly understood, then the numbers could not be random. Methods that are able to let their internal workings be known without sacrificing random results are what this paper sets out to describe. Beginning with a study of what it really means for something to be random, this paper dives into the topic of random number generators and summarizes the key …


Flexible Distributed Lag Models Using Random Functions With Application To Estimating Mortality Displacement From Heat-Related Deaths, Roger D. Peng Dec 2011

Flexible Distributed Lag Models Using Random Functions With Application To Estimating Mortality Displacement From Heat-Related Deaths, Roger D. Peng

Johns Hopkins University, Dept. of Biostatistics Working Papers

No abstract provided.


Applying Gmdh-Type Neural Network And Genetic Algorithm For Stock Price Prediction Of Iranian Cement Sector, Saeed Fallahi, Meysam Shaverdi, Vahab Bashiri Dec 2011

Applying Gmdh-Type Neural Network And Genetic Algorithm For Stock Price Prediction Of Iranian Cement Sector, Saeed Fallahi, Meysam Shaverdi, Vahab Bashiri

Applications and Applied Mathematics: An International Journal (AAM)

The cement industry is one of the most important and profitable industries in Iran and great content of financial resources are investing in this sector yearly. In this paper a GMDH-type neural network and genetic algorithm is developed for stock price prediction of cement sector. For stocks price prediction by GMDH type-neural network, we are using earnings per share (EPS), Prediction Earnings Per Share (PEPS), Dividend per share (DPS), Price-earnings ratio (P/E), Earnings-price ratio (E/P) as input data and stock price as output data. For this work, data of ten cement companies is gathering from Tehran stock exchange (TSE) in …


A Comparison Of Spatio-Temporal Prediction Methods Of Cancer Incidence In The U.S, Michelle Hamlyn Aug 2011

A Comparison Of Spatio-Temporal Prediction Methods Of Cancer Incidence In The U.S, Michelle Hamlyn

UNLV Theses, Dissertations, Professional Papers, and Capstones

Cancer is the cause of one out of four deaths in the United States, and in 2009, researchers expected over 1.5 million new patients to be diagnosed with some form of cancer. People diagnosed with cancer, whether a common or rare type, need to undergo treatments, the amount and kind of which will depend on the severity of the cancer. So how do healthcare providers know how much funding is needed for treatment? What would better enable a pharmaceutical company to determine how much to allocate for research and development of drugs, the amount of each drug to manufacture, or …


Variable Importance Analysis With The Multipim R Package, Stephan J. Ritter, Nicholas P. Jewell, Alan E. Hubbard Jul 2011

Variable Importance Analysis With The Multipim R Package, Stephan J. Ritter, Nicholas P. Jewell, Alan E. Hubbard

U.C. Berkeley Division of Biostatistics Working Paper Series

We describe the R package multiPIM, including statistical background, functionality and user options. The package is for variable importance analysis, and is meant primarily for analyzing data from exploratory epidemiological studies, though it could certainly be applied in other areas as well. The approach taken to variable importance comes from the causal inference field, and is different from approaches taken in other R packages. By default, multiPIM uses a double robust targeted maximum likelihood estimator (TMLE) of a parameter akin to the attributable risk. Several regression methods/machine learning algorithms are available for estimating the nuisance parameters of the models, including …


Analysis Of Roms Estimated Posterior Error Utilizing 4dvar Data Assimilation, Joseph Patrick Horton Jun 2011

Analysis Of Roms Estimated Posterior Error Utilizing 4dvar Data Assimilation, Joseph Patrick Horton

Mathematics

The appropriateness of the approximate error calculated by the Regional Ocean Modeling System (ROMS) is analyzed using Four-Dimensional Data Assimilation (4DVAR) performed on a numerical model of the San Luis Obispo Bay. An effective method of sampling data to minimize the actual error associated with the assimilated numerical model is explored by using different data sampling methods. An idealized state of the SLO bay region ("Real Run") is created to be used as the real ocean, then a numerical model of this region is created approximating this Real Run; this is known as the "Simulated State". By taking samples from …


Generalized Bathtub Hazard Models For Binary-Transformed Climate Data, James Polcer May 2011

Generalized Bathtub Hazard Models For Binary-Transformed Climate Data, James Polcer

Masters Theses & Specialist Projects

In this study, we use a hazard-based modeling as an alternative statistical framework to time series methods as applied to climate data. Data collected from the Kentucky Mesonet will be used to study the distributional properties of the duration of high and low-energy wind events relative to an arbitrary threshold. Our objectiveswere to fit bathtub models proposed in literature, propose a generalized bathtub model, apply these models to Kentucky Mesonet data, and make recommendations as to feasibility of wind power generation. Using two different thresholds (1.8 and 10 mph respectively), results show that the Hjorth bathtub model consistently performed better …


Unlv Enrollment Forecasting, Sabrina Beckman, Stefan Cline, Monika Neda Apr 2011

Unlv Enrollment Forecasting, Sabrina Beckman, Stefan Cline, Monika Neda

Festival of Communities: UG Symposium (Posters)

Our project investigates the future enrollment of undergraduates at UNLV in the entire university, the College of Science, and the Department of Mathematical Sciences. The method used for the forecast, is the well-known least-squares method, for which a mathematical description will be presented. Studies for the numerical error are pursued too. The study will include graphs that describe the past and future behavior for different parameter settings. Mathematical results obtained show that the university will continue to grow given the current trends of enrollment.


Reliability Measures Of A Three-State Complex System: A Copula Approach, Mangey Ram Dec 2010

Reliability Measures Of A Three-State Complex System: A Copula Approach, Mangey Ram

Applications and Applied Mathematics: An International Journal (AAM)

Improvement in reliability and production play a very important role in system design. The two key factors, considered in predicting system reliability, are failure distribution of the component and system configuration. This research discusses the mathematical modeling of a highly reliable complex system, which is in three states i.e. normal, partial failed (degraded state) and complete failed state. The system, partial failed is due to the partial failure of internal components or redundancies and completely failed is due to catastrophic failure of the system. Repair rates are general functions of the time spent. All the transition rates are constant except …


Random Walks With Elastic And Reflective Lower Boundaries, Lucas Clay Devore Dec 2009

Random Walks With Elastic And Reflective Lower Boundaries, Lucas Clay Devore

Masters Theses & Specialist Projects

No abstract provided.


Targeted Maximum Likelihood Estimation: A Gentle Introduction, Susan Gruber, Mark J. Van Der Laan Aug 2009

Targeted Maximum Likelihood Estimation: A Gentle Introduction, Susan Gruber, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

This paper provides a concise introduction to targeted maximum likelihood estimation (TMLE) of causal effect parameters. The interested analyst should gain sufficient understanding of TMLE from this introductory tutorial to be able to apply the method in practice. A program written in R is provided. This program implements a basic version of TMLE that can be used to estimate the effect of a binary point treatment on a continuous or binary outcome.


Shrinkage Estimation Of Expression Fold Change As An Alternative To Testing Hypotheses Of Equivalent Expression, Zahra Montazeri, Corey M. Yanofsky, David R. Bickel Aug 2009

Shrinkage Estimation Of Expression Fold Change As An Alternative To Testing Hypotheses Of Equivalent Expression, Zahra Montazeri, Corey M. Yanofsky, David R. Bickel

COBRA Preprint Series

Research on analyzing microarray data has focused on the problem of identifying differentially expressed genes to the neglect of the problem of how to integrate evidence that a gene is differentially expressed with information on the extent of its differential expression. Consequently, researchers currently prioritize genes for further study either on the basis of volcano plots or, more commonly, according to simple estimates of the fold change after filtering the genes with an arbitrary statistical significance threshold. While the subjective and informal nature of the former practice precludes quantification of its reliability, the latter practice is equivalent to using a …


Model-Based Clustering Of Methylation Array Data: A Recursive-Partitioning Algorithm For High-Dimensional Data Arising As A Mixture Of Beta Distributions, E. Andres Houseman, Brock C. Christensen, Ru-Fang Yeh, Carmen J. Marsit, Margaret R. Karagas, Margaret Wrensch, Heather H. Nelson, Joseph Wiemels, Shichun Zheng, John K. Wiencke, Karl T. Kelsey Jun 2008

Model-Based Clustering Of Methylation Array Data: A Recursive-Partitioning Algorithm For High-Dimensional Data Arising As A Mixture Of Beta Distributions, E. Andres Houseman, Brock C. Christensen, Ru-Fang Yeh, Carmen J. Marsit, Margaret R. Karagas, Margaret Wrensch, Heather H. Nelson, Joseph Wiemels, Shichun Zheng, John K. Wiencke, Karl T. Kelsey

Harvard University Biostatistics Working Paper Series

No abstract provided.


Methods Of Assessing And Ranking Probable Sources Of Error, Nataniel Greene May 2008

Methods Of Assessing And Ranking Probable Sources Of Error, Nataniel Greene

Publications and Research

A classical method for ranking n potential events as sources of error is Bayes' theorem. However, a ranking based on Bayes' theorem lacks a fundamental symmetry: the ranking in terms of blame for error will not be the reverse of the ranking in terms of credit for lack of error. While this is not a flaw in Bayes' theorem, it does lead one to inquire whether there are related methods which have such symmetry. Related methods explored here include the logical version of Bayes' theorem based on probabilities of conditionals, probabilities of biconditionals, and ratios or differences of credit to …


An Overview Of Conditionals And Biconditionals In Probability, Nataniel Greene Mar 2008

An Overview Of Conditionals And Biconditionals In Probability, Nataniel Greene

Publications and Research

Conditional and biconditional statements are a standard part of symbolic logic but they have only recently begun to be explored in probability for applications in artificial intelligence. Here we give a brief overview of the major theorems involved and illustrate them using two standard model problems from conditional probability.


A Method For Visualizing Multivariate Time Series Data, Roger D. Peng Feb 2008

A Method For Visualizing Multivariate Time Series Data, Roger D. Peng

Johns Hopkins University, Dept. of Biostatistics Working Papers

Visualization and exploratory analysis is an important part of any data analysis and is made more challenging when the data are voluminous and high-dimensional. One such example is environmental monitoring data, which are often collected over time and at multiple locations, resulting in a geographically indexed multivariate time series. Financial data, although not necessarily containing a geographic component, present another source of high-volume multivariate time series data. We present the mvtsplot function which provides a method for visualizing multivariate time series data. We outline the basic design concepts and provide some examples of its usage by applying it to a …


Bayesian Analysis For Penalized Spline Regression Using Win Bugs, Ciprian M. Crainiceanu, David Ruppert, M.P. Wand Dec 2007

Bayesian Analysis For Penalized Spline Regression Using Win Bugs, Ciprian M. Crainiceanu, David Ruppert, M.P. Wand

Johns Hopkins University, Dept. of Biostatistics Working Papers

Penalized splines can be viewed as BLUPs in a mixed model framework, which allows the use of mixed model software for smoothing. Thus, software originally developed for Bayesian analysis of mixed models can be used for penalized spline regression. Bayesian inference for nonparametric models enjoys the flexibility of nonparametric models and the exact inference provided by the Bayesian inferential machinery. This paper provides a simple, yet comprehensive, set of programs for the implementation of nonparametric Bayesian analysis in WinBUGS. MCMC mixing is substantially improved from the previous versions by using low{rank thin{plate splines instead of truncated polynomial basis. Simulation time …


Loss-Based Estimation With Evolutionary Algorithms And Cross-Validation, David Shilane, Richard H. Liang, Sandrine Dudoit Nov 2007

Loss-Based Estimation With Evolutionary Algorithms And Cross-Validation, David Shilane, Richard H. Liang, Sandrine Dudoit

U.C. Berkeley Division of Biostatistics Working Paper Series

Many statistical inference methods rely upon selection procedures to estimate a parameter of the joint distribution of explanatory and outcome data, such as the regression function. Within the general framework for loss-based estimation of Dudoit and van der Laan, this project proposes an evolutionary algorithm (EA) as a procedure for risk optimization. We also analyze the size of the parameter space for polynomial regression under an interaction constraints along with constraints on either the polynomial or variable degree.


Survival Analysis With Large Dimensional Covariates: An Application In Microarray Studies, David A. Engler, Yi Li Jul 2007

Survival Analysis With Large Dimensional Covariates: An Application In Microarray Studies, David A. Engler, Yi Li

Harvard University Biostatistics Working Paper Series

Use of microarray technology often leads to high-dimensional and low- sample size data settings. Over the past several years, a variety of novel approaches have been proposed for variable selection in this context. However, only a small number of these have been adapted for time-to-event data where censoring is present. Among standard variable selection methods shown both to have good predictive accuracy and to be computationally efficient is the elastic net penalization approach. In this paper, adaptation of the elastic net approach is presented for variable selection both under the Cox proportional hazards model and under an accelerated failure time …


Simultaneous Confidence Intervals Based On The Percentile Bootstrap Approach, Micha Mandel, Rebecca A. Betensky Jun 2007

Simultaneous Confidence Intervals Based On The Percentile Bootstrap Approach, Micha Mandel, Rebecca A. Betensky

Harvard University Biostatistics Working Paper Series

No abstract provided.


A Distributed Parabolic Control With Mixed Boundary Conditions, Jose-Luis Menaldi, Domingo Alberto Tarzia Jan 2007

A Distributed Parabolic Control With Mixed Boundary Conditions, Jose-Luis Menaldi, Domingo Alberto Tarzia

Mathematics Faculty Research Publications

We study the asymptotic behavior of an optimal distributed control problem where the state is given by the heat equation with mixed boundary conditions. The parameter α intervenes in the Robin boundary condition and it represents the heat transfer coefficient on a portion Γ1 of the boundary of a given regular n-dimensional domain. For each α, the distributed parabolic control problem optimizes the internal energy g. It is proven that the optimal control ĝα with optimal state uĝαα and optimal adjoint state pĝαα are convergent as α → 1 …


Diffusion And Fractional Diffusion Based Models For Multiple Light Scattering And Image Analysis, Jonathan Blackledge Jan 2007

Diffusion And Fractional Diffusion Based Models For Multiple Light Scattering And Image Analysis, Jonathan Blackledge

Articles

This paper considers a fractional light diffusion model as an approach to characterizing the case when intermediate scattering processes are present, i.e. the scattering regime is neither strong nor weak. In order to introduce the basis for this approach, we revisit the elements of formal scattering theory and the classical diffusion problem in terms of solutions to the inhomogeneous wave and diffusion equations respectively. We then address the significance of these equations in terms of a random walk model for multiple scattering. This leads to the proposition of a fractional diffusion equation for modelling intermediate strength scattering that is based …


Spatio-Temporal Analysis Of Areal Data And Discovery Of Neighborhood Relationships In Conditionally Autoregressive Models, Subharup Guha, Louise Ryan Nov 2006

Spatio-Temporal Analysis Of Areal Data And Discovery Of Neighborhood Relationships In Conditionally Autoregressive Models, Subharup Guha, Louise Ryan

Harvard University Biostatistics Working Paper Series

No abstract provided.


Bayesian Smoothing Of Irregularly-Spaced Data Using Fourier Basis Functions, Christopher J. Paciorek Aug 2006

Bayesian Smoothing Of Irregularly-Spaced Data Using Fourier Basis Functions, Christopher J. Paciorek

Harvard University Biostatistics Working Paper Series

No abstract provided.


Posterior Simulation In The Generalized Linear Model With Semiparmetric Random Effects, Subharup Guha May 2006

Posterior Simulation In The Generalized Linear Model With Semiparmetric Random Effects, Subharup Guha

Harvard University Biostatistics Working Paper Series

Generalized linear mixed models with semiparametric random effects are useful in a wide variety of Bayesian applications. When the random effects arise from a mixture of Dirichlet process (MDP) model, normal base measures and Gibbs sampling procedures based on the Pólya urn scheme are often used to simulate posterior draws. These algorithms are applicable in the conjugate case when (for a normal base measure) the likelihood is normal. In the non-conjugate case, the algorithms proposed by MacEachern and Müller (1998) and Neal (2000) are often applied to generate posterior samples. Some common problems associated with simulation algorithms for non-conjugate MDP …


Gauss-Seidel Estimation Of Generalized Linear Mixed Models With Application To Poisson Modeling Of Spatially Varying Disease Rates, Subharup Guha, Louise Ryan Oct 2005

Gauss-Seidel Estimation Of Generalized Linear Mixed Models With Application To Poisson Modeling Of Spatially Varying Disease Rates, Subharup Guha, Louise Ryan

Harvard University Biostatistics Working Paper Series

Generalized linear mixed models (GLMMs) provide an elegant framework for the analysis of correlated data. Due to the non-closed form of the likelihood, GLMMs are often fit by computational procedures like penalized quasi-likelihood (PQL). Special cases of these models are generalized linear models (GLMs), which are often fit using algorithms like iterative weighted least squares (IWLS). High computational costs and memory space constraints often make it difficult to apply these iterative procedures to data sets with very large number of cases.

This paper proposes a computationally efficient strategy based on the Gauss-Seidel algorithm that iteratively fits sub-models of the GLMM …


Computational Techniques For Spatial Logistic Regression With Large Datasets, Christopher J. Paciorek, Louise Ryan Oct 2005

Computational Techniques For Spatial Logistic Regression With Large Datasets, Christopher J. Paciorek, Louise Ryan

Harvard University Biostatistics Working Paper Series

In epidemiological work, outcomes are frequently non-normal, sample sizes may be large, and effects are often small. To relate health outcomes to geographic risk factors, fast and powerful methods for fitting spatial models, particularly for non-normal data, are required. We focus on binary outcomes, with the risk surface a smooth function of space. We compare penalized likelihood models, including the penalized quasi-likelihood (PQL) approach, and Bayesian models based on fit, speed, and ease of implementation.

A Bayesian model using a spectral basis representation of the spatial surface provides the best tradeoff of sensitivity and specificity in simulations, detecting real spatial …