Open Access. Powered by Scholars. Published by Universities.®

Applied Statistics Commons™

Open Access. Powered by Scholars. Published by Universities.®

2012

Discipline
Institution
Keyword
Publication
Publication Type

Articles 31 - 60 of 81

Full-Text Articles in Applied Statistics

Comparative Analysis Of Dispersion Parameter Estimates In Loglinear Modeling: Applied To E-Commerce Sales And Customer Data, Scott Davis Sep 2012

Comparative Analysis Of Dispersion Parameter Estimates In Loglinear Modeling: Applied To E-Commerce Sales And Customer Data, Scott Davis

Statistics

When loglinear models are applied to count data the issue of over-dispersion often arises. Moment and maximum likelihood estimation methods in accounting for over-dispersion are widely used because they allow for model checking tools such as Chi-square, F, and likelihood ratio tests. Here is a comparison between R functions that each uses one method; glm.nb uses MLE, and glm.poisson.disp uses MME. The Index of Dissimilarity and visual model selection (ECDF plots) are also incorporated. These are applied to sales data using product and customer information compiled over the last five years that was generously provided by an e-commerce company.


The Implementation Of The Shear Correlation Function And The Matter Power Spectrum In R, Allison A. Scheppelmann, Deborah J. Bard Aug 2012

The Implementation Of The Shear Correlation Function And The Matter Power Spectrum In R, Allison A. Scheppelmann, Deborah J. Bard

STAR Program Research Presentations

Weak gravitational lensing is an important tool in understanding the large-scale structure of the universe. One component in understanding the effect of weak gravitational lensing is the shear correlation function and matter power spectrum. The calculation of these values is often complicated and time consuming. In order to decrease the cost of these calculations they were implemented in R using parallelization. This resulted in the calculations completing faster and the process to be easily changed in order to fit the need of each researcher using the algorithms created in R.


From Unbiased Numerical Estimates To Unbiased Interval Estimates, Baokun Li, Gang Xiang, Vladik Kreinovich, Panagios Moscopoulos Aug 2012

From Unbiased Numerical Estimates To Unbiased Interval Estimates, Baokun Li, Gang Xiang, Vladik Kreinovich, Panagios Moscopoulos

Departmental Technical Reports (CS)

One of the main objectives of statistics is to estimate the parameters of a probability distribution based on a sample taken from this distribution. Of course, since the sample is finite, the estimate X is, in general, different from the actual value x of the corresponding parameter. What we can require is that the corresponding estimate is unbiased, i.e., that the mean value of the difference X - x is equal to 0: E[X] = x. In some problems, unbiased estimates are not possible. We show that in some such problems, it is possible to have interval unbiased estimates, i.e., …


Significant Themes In 19th-Century Literature, Matthew L. Jockers, David Mimno Aug 2012

Significant Themes In 19th-Century Literature, Matthew L. Jockers, David Mimno

Department of English: Faculty Publications

External factors such as author gender, author nationality, and date of publication affect both the choice of literary themes in novels and the expression of those themes, but the extent of this association is difficult to quantify. In this work, we apply statistical methods to identify and extract hundreds of "topics" from a corpus of 3,346 works of 19th-century British, Irish, and American fiction. We use these topics as a measurable, data-driven proxy for literary themes. External factors may predict fluctuations in the use of themes and the individual word choices within themes. We use topics to measure the evidence …


Analysis Of Bank Failure And Size Of Assets, Guancun Zhong Aug 2012

Analysis Of Bank Failure And Size Of Assets, Guancun Zhong

UNLV Theses, Dissertations, Professional Papers, and Capstones

The financial health of the banking industry is an important prerequisite for economic stability and growth. Bank failures in the United States have run in cycles largely associated with the collapse of economic bubbles. The number of bank failures has increased dramatically over the last thirty years (Halling and Hayden, 2007). In this thesis, we try to address the following two questions: 1) What is the relationship, if any, between a bank's asset size and its likelihood of failures? 2) How can we use statistical tools to predict the numbers of bank failures in the future? Various modeling techniques are …


Reliability Models For Hpc Applications And A Cloud Economic Model, Thanadech Thanakornworakij Jul 2012

Reliability Models For Hpc Applications And A Cloud Economic Model, Thanadech Thanakornworakij

Doctoral Dissertations

With the enormous number of computing resources in HPC and Cloud systems, failures become a major concern. Therefore, failure behaviors such as reliability, failure rate, and mean time to failure need to be understood to manage such a large system efficiently.

This dissertation makes three major contributions in HPC and Cloud studies. First, a reliability model with correlated failures in a k-node system for HPC applications is studied. This model is extended to improve accuracy by accounting for failure correlation. Marshall-Olkin Multivariate Weibull distribution is improved by excess life, conditional Weibull, to better estimate system reliability. Also, the univariate …


Meta-Heuristics Analysis For Technologically Complex Programs: Understanding The Impact Of Total Constraints For Schedule, Quality And Cost, Henry Darrel Webb Jul 2012

Meta-Heuristics Analysis For Technologically Complex Programs: Understanding The Impact Of Total Constraints For Schedule, Quality And Cost, Henry Darrel Webb

EMSE Doctoral Projects

Program management data associated with a technically complex radio frequency electronics base communication system has been collected and analyzed to identify heuristics which may be utilized in addition to existing processes and procedures to provide indicators that a program is trending to failure. Analysis of the collected data includes detailed schedule analysis, detailed earned value management analysis and defect analysis within the framework of a Firm Fixed Price (FFP) incentive fee contract.

This project develops heuristics and provides recommendations for analysis of complex project management efforts such as those discussed herein. The analysis of the effects of the constraints on …


Response Surface Optimization Of Electron Beam Freeform Fabrication Depositions Using Design Of Experiments, Patricia A. Quigley Jul 2012

Response Surface Optimization Of Electron Beam Freeform Fabrication Depositions Using Design Of Experiments, Patricia A. Quigley

Engineering Management & Systems Engineering Theses & Dissertations

The Electron Beam Freeform Fabrication (EBF3 ) System is a material depositing, layer additive technique that produces three dimensional (3D) parts out of a wide range of metals in high vacuum, using an electron beam and wire feedstock. Screening deposition trials on a titanium alloy, Ti-6Al-4V, at the National Aeronautics Space Administration (NASA) revealed selective vaporization of the aluminum content of linear prototypes when subjected to chemical analysis. In this study, the aluminum content, bead height and bead width output responses were analyzed from a systematic study of the effects that the interactions of the EBF3 processing parameters …


A Statistical Model To Determine Multiple Binding Sites Of A Transcription Factor On Dna Using Chip-Seq Data, Rasika Jayatillake Jul 2012

A Statistical Model To Determine Multiple Binding Sites Of A Transcription Factor On Dna Using Chip-Seq Data, Rasika Jayatillake

Mathematics & Statistics Theses & Dissertations

Protein-DNA interaction is vital to many biological processes in cells such as cell division, embryo development and regulating gene expression. Chromatin Immunoprecipitation followed by massively parallel sequencing (ChIP-seq) is a new technology that can reveal protein binding sites in genome with superior accuracy. Although many methods have been proposed to find binding sites for ChIP-seq data, they can find only one binding site within a short region of the genome. In this study we introduce a statistical model to identify multiple binding sites of a transcription factor within a short region of the genome using the ChIP-seq data. Mapped sequence …


Investigation Of Trends And Predictive Effectiveness Of Crash Severity Models, James E. Mooradian Jun 2012

Investigation Of Trends And Predictive Effectiveness Of Crash Severity Models, James E. Mooradian

Master's Theses

This thesis describes analysis using ordinal logistic regression to uncover temporal patterns in the severity level (fatal, serious injury, minor injury, slight injury or no injury) for persons involved in highway crashes in Connecticut, focusing on the demographic split between senior travelers (65 years and over) and non-senior travelers. Existing state sources provide data describing the time and weather conditions for each crash and the vehicles and persons involved over the time period from 1995 to 2009 as well as the traffic volumes and the characteristics of the roads on which these crashes occurred. Findings indicate an overall increase in …


Analysing Domestic Electricity Smart Metering Data Using Self Organising Maps, Fintan Mcloughlin, Aidan Duffy, Michael Conlon Jun 2012

Analysing Domestic Electricity Smart Metering Data Using Self Organising Maps, Fintan Mcloughlin, Aidan Duffy, Michael Conlon

Conference Papers

This paper investigates a method of classifying domestic electricity load profiles through Self Organising Maps (SOMs). Approximately four thousand customers are divided into groups based on their electricity demand patterns. Dwelling and occupant characteristics are then investigated for each group. The results show that SOMs are an effective way of classifying customers into groups in terms of their electrical load profile and that certain dwelling and occupant characteristics are significant factors in determining which group they end up in.


Analysis Of Dietary Patterns Over Freshman Year Of College, Chelsea Lofland Jun 2012

Analysis Of Dietary Patterns Over Freshman Year Of College, Chelsea Lofland

Statistics

This analysis is an investigation of changes in Cal Poly students’ eating habits over freshman year. The motivation behind this was an interest in college students’ lifestyles; college is the first time most students live on their own and it can be an important maturation period. College is stressful, exciting, liberating, and terrifying all at the same time. This distinctive life experience, along with my desire to handle big and messy data, led me to this research question.

The response variable analyzed was food consumption and the explanatory variables were: sex, race, quarter, food group, stress, exercise, BMI, sleep quality …


Improvement Of Statistical Process Control At St. Jude Medical's Cardiac Manufacturing Facility, Christopher Lance Edwards Jun 2012

Improvement Of Statistical Process Control At St. Jude Medical's Cardiac Manufacturing Facility, Christopher Lance Edwards

Master's Theses

Sig sigma is a methodology where companies strive to reproduce results ending up having a 99.9996% chance their product will be void of defects. In order for companies to reach six sigma, statistical process control (SPC) needs to be introduced. SPC has many different tools associated with it, control charts being one of them. Control charts play a vital role in managing how a process is behaving. Control charts allow users to identify special causes, or shifts, and can therefore change the process to keep producing good products, free of defects.

There are many factories and manufacturing facilities having implemented …


Analyzing Multiple Independent Spatial Point Processes, Neal Grantham May 2012

Analyzing Multiple Independent Spatial Point Processes, Neal Grantham

Statistics

No abstract provided.


The Impact Of Violating Factor Scaling Method Assumptions On Latent Mean Difference Testing In Structured Means Models, Dandan Wang, Tiffany A. Whittaker, S. Natasha Beretvas May 2012

The Impact Of Violating Factor Scaling Method Assumptions On Latent Mean Difference Testing In Structured Means Models, Dandan Wang, Tiffany A. Whittaker, S. Natasha Beretvas

Journal of Modern Applied Statistical Methods

Type I error rates and power of the likelihood ratio test and bias of the standardized effect size measure associated with the latent mean difference in structured means modeling are examined when violating the assumptions underlying the two available factor scaling methods under various conditions. Implications and recommendations are discussed.


New Approximate Bayesian Confidence Intervals For The Coefficient Of Variation Of A Gaussian Distribution, Vincent A. R. Camara May 2012

New Approximate Bayesian Confidence Intervals For The Coefficient Of Variation Of A Gaussian Distribution, Vincent A. R. Camara

Journal of Modern Applied Statistical Methods

Confidence intervals are constructed for the coefficient of variation of a Gaussian distribution. Considering the square error and the Higgins-Tsokos loss functions, approximate Bayesian models are derived and compared to a published classical model. The models are shown to have great coverage accuracy. The classical model does not always yield the best confidence intervals; the proposed models often perform better.


A Poisson Regression Model For Female Radium Dial Workers, Tze-San Lee May 2012

A Poisson Regression Model For Female Radium Dial Workers, Tze-San Lee

Journal of Modern Applied Statistical Methods

A Poisson regression model with interaction terms was applied to study the dose response relationship for radium-induced skeletal cancers. The model showed that the expected frequency count of bone tumors depended not only on the logarithmic dose and the time since first exposure, but also on the interaction between the logarithmic dose and the time since first exposure, whereas the dose-response model for head tumors depended only on the logarithmic dose.


Jmasm 32: Sas Template For Single-Subject Experimental Designs, Hyewon Chung, Jiseon Kim, Ryoungsun Park May 2012

Jmasm 32: Sas Template For Single-Subject Experimental Designs, Hyewon Chung, Jiseon Kim, Ryoungsun Park

Journal of Modern Applied Statistical Methods

Meta-analysis has been used to synthesize research findings and to evaluate the effectiveness of treatments or the accuracy of diagnostic tools. Although meta-analytic techniques were developed to synthesize the results of several studies, controversy exists as to how to quantify the results from singlesubject experimental designs (SSEDs). The most commonly used metrics are reviewed, including nonregression and regression based methods. The application of the SAS template is demonstrated through simulated data sets. The SAS templates can be modified to accommodate a more complex data structure.


The Length-Biased Lognormal Distribution And Its Application In The Analysis Of Data From Oil Field Exploration Studies, Makarand V. Ratnaparkhi, Uttara V. Naik-Nimbalkar May 2012

The Length-Biased Lognormal Distribution And Its Application In The Analysis Of Data From Oil Field Exploration Studies, Makarand V. Ratnaparkhi, Uttara V. Naik-Nimbalkar

Journal of Modern Applied Statistical Methods

The length-biased version of the lognormal distribution and related estimation problems are considered and sized-biased data arising in the exploration of oil fields is analyzed. The properties of the estimators are studied using simulations and the use of sample mode as an estimate of the lognormal parameter is discussed.


Four Period Crossover Designs, James F. Reed Iii May 2012

Four Period Crossover Designs, James F. Reed Iii

Journal of Modern Applied Statistical Methods

In higher-order four period crossover designs with two treatments, sixteen possible treatment sequences can result: AAAA, AAAB, AABA, AABB, ABAA, ABAB, ABBA, ABBB and their duals. Higher-order crossover designs are useful for several reasons: they allow estimation of a treatment effect even in the presence of a carry-over effect, they provide estimates of intra-subject variability and they draw inference on the carry-over effect. The real question related to a two-treatment four-period crossover design is the real world application of these designs. This article considers four designs: Design I: ABBA and its dual; Design II: ABBA, AABB and their duals, Design …


Gamma-Pareto Distribution And Its Applications, Ayman Alzaatreh, Felix Famoye, Carl Lee May 2012

Gamma-Pareto Distribution And Its Applications, Ayman Alzaatreh, Felix Famoye, Carl Lee

Journal of Modern Applied Statistical Methods

A new distribution, the gamma-Pareto, is defined and studied and various properties of the distribution are obtained. Results for moments, limiting behavior and entropies are provided. The method of maximum likelihood is proposed for estimating the parameters and the distribution is applied to fit three real data sets.


A Weighted Exponential Detection Function Model For Line Transect Data, Faisal Ababneh, Omar M. Eidous May 2012

A Weighted Exponential Detection Function Model For Line Transect Data, Faisal Ababneh, Omar M. Eidous

Journal of Modern Applied Statistical Methods

A new parametric model is proposed for modeling the density function of perpendicular distances in line transects sampling. The model can be considered a weighted exponential model in the sense that it combines two exponential models with different weights. The proposed model is appealing because it is monotone decreasing with distance from transect line; in contrast to the classical exponential model, it satisfies the shoulder condition at the origin. Simulation results for a wide range of target densities show reasonable and good performances of the weighted exponential model in most considered cases compared to the classical exponential and the half-normal …


Ordinal Regression Analysis: Using Generalized Ordinal Logistic Regression Models To Estimate Educational Data, Xing Liu, Hari Koirala May 2012

Ordinal Regression Analysis: Using Generalized Ordinal Logistic Regression Models To Estimate Educational Data, Xing Liu, Hari Koirala

Journal of Modern Applied Statistical Methods

The proportional odds (PO) assumption for ordinal regression analysis is often violated because it is strongly affected by sample size and the number of covariate patterns. To address this issue, the partial proportional odds (PPO) model and the generalized ordinal logit model were developed. However, these models are not typically used in research. One likely reason for this is the restriction of current statistical software packages: SPSS cannot perform the generalized ordinal logit model analysis and SAS requires data restructuring. This article illustrates the use of generalized ordinal logistic regression models to predict mathematics proficiency levels using Stata and compares …


Robust Regression Estimates In The Prediction Of Latent Variables In Structural Equation Models, Marcelo Angelo Cirillo, Lúcia Pereira Barroso May 2012

Robust Regression Estimates In The Prediction Of Latent Variables In Structural Equation Models, Marcelo Angelo Cirillo, Lúcia Pereira Barroso

Journal of Modern Applied Statistical Methods

The incorporation of the robust regression methods Least Median Square (LMS) and Least Trimmed Squares (LTS) is proposed in structural equation modeling. Results show that, in situations of high deviations of symmetry, the evaluated methods would be recommended for applications including smaller sample sizes.


Comparison Of Re-Sampling Methods To Generalized Linear Models And Transformations In Factorial And Fractional Factorial Designs, Maher Qumsiyeh, Gerald Shaughnessy May 2012

Comparison Of Re-Sampling Methods To Generalized Linear Models And Transformations In Factorial And Fractional Factorial Designs, Maher Qumsiyeh, Gerald Shaughnessy

Journal of Modern Applied Statistical Methods

Experimental situations in which observations are not normally distributed frequently occur in practice. A common situation occurs when responses are discrete in nature, for example counts. One way to analyze such experimental data is to use a transformation for the responses; another is to use a link function based on a generalized linear model (GLM) approach. Re-sampling is employed as an alternative method to analyze non-normal, discrete data. Results are compared to those obtained by the previous two methods.


Statistical Inferences For Lomax Distribution Based On Record Values (Bayesian And Classical), Parviz Nasiri, Saman Hosseini May 2012

Statistical Inferences For Lomax Distribution Based On Record Values (Bayesian And Classical), Parviz Nasiri, Saman Hosseini

Journal of Modern Applied Statistical Methods

A maximum likelihood estimation (MLE) based on records is obtained and a proper prior distribution to attain a Bayes estimation (both informative and non-informative) based on records for quadratic loss and squared error loss functions is also calculated. The study considers the shortest confidence interval and Highest Posterior Distribution confidence interval based on records, and using Mean Square Error MSE criteria for point estimation and length criteria for interval estimation, their appropriateness to each other is examined.


Empirical Sampling From Permutation Space With Unique Patterns, Justice I. Odiase May 2012

Empirical Sampling From Permutation Space With Unique Patterns, Justice I. Odiase

Journal of Modern Applied Statistical Methods

The exact distribution of a test statistic ultimately guarantees that the probability of a Type I error is exactly α. Several methods for estimating the exact distribution of a test statistic have evolved over the years with inherent computational problems and varying degrees of accuracy. The unique pattern of permutations resulting from using experimental data to sample within the permutation space without the risk of repeating permutations is identified. The method presented circumvents the theoretical requirements of asymptotic procedures and the computational difficulties associated with an exhaustive enumeration of permutations. Results show that time and space complexities are drastically reduced …


Weight: Does It Really Matter?, Jennifer L. Brown, Gerald Halpin, Glennelle Halpin May 2012

Weight: Does It Really Matter?, Jennifer L. Brown, Gerald Halpin, Glennelle Halpin

Journal of Modern Applied Statistical Methods

Differential weighting promises to improve the validity of a measure. This study examines whether similar results would be found using weighted, unweighted and standardized z scores from the All Stars Core survey. It was concluded that the weighted systems were developed to equate the questions within the scales and to ease the process for customers without access to data analysis programs; however, the standardized scores were the more appropriate method for equating the test items.


Ratio Type Estimator Of Ratio Of Two Population Means In Stratified Random Sampling, Rajesh Tailor, Sunil Chouhan May 2012

Ratio Type Estimator Of Ratio Of Two Population Means In Stratified Random Sampling, Rajesh Tailor, Sunil Chouhan

Journal of Modern Applied Statistical Methods

A ratio estimator is proposed for the ratio of two population means using auxiliary information in stratified random sampling. Bias and mean squared error expressions are obtained under large sample approximation, and the proposed estimator is compared both theoretically and empirically with the conventional estimator of ratio for two population means in stratified random sampling.


Parameter Estimation With Mixture Item Response Theory Models: A Monte Carlo Comparison Of Maximum Likelihood And Bayesian Methods, W. Holmes Finch, Brian F. French May 2012

Parameter Estimation With Mixture Item Response Theory Models: A Monte Carlo Comparison Of Maximum Likelihood And Bayesian Methods, W. Holmes Finch, Brian F. French

Journal of Modern Applied Statistical Methods

The Mixture Item Response Theory (MixIRT) can be used to identify latent classes of examinees in data as well as to estimate item parameters such as difficulty and discrimination for each of the groups. Parameter estimation via maximum likelihood (MLE) and Bayesian estimation based on the Markov Chain Monte Carlo (MCMC) are compared for classification accuracy and parameter estimation bias for difficulty and discrimination. Standard error magnitude and coverage rates were compared across number of items, number of latent groups, group size ratio, total sample size and underlying item response model. Results show that MCMC provides more accurate group membership …