Open Access. Powered by Scholars. Published by Universities.®

Applied Statistics Commons™

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 2011 - 2040 of 2922

Full-Text Articles in Applied Statistics

Electoral Voting And Population Distribution In The United States, Paul Kvam Feb 2010

Electoral Voting And Population Distribution In The United States, Paul Kvam

Department of Math & Statistics Faculty Publications

In the United States, the electoral system for determining the president is controversial and sometimes confusing to voters keeping track of election outcomes. Instead of directly counting votes to decide the winner of a presidential election, individual states send a representative number of electors to the Electoral College, and they are trusted to cast their collective vote for the candidate who won the popular vote in their state.

Under the current rules, the value of a vote differs from state to state. A large state such as California has an immense effect on the national election, but, compared to a …


Mathematical Themes In Economics, Machine Learning, And Bioinformatics, Matt Bogard Jan 2010

Mathematical Themes In Economics, Machine Learning, And Bioinformatics, Matt Bogard

Economics Faculty Publications

Graduate students in economics are often introduced to some very useful mathematical tools that many outside the discipline may not associate with training in economics. This essay looks at some of these tools and concepts, including constrained optimization, separating hyperplanes, supporting hyperplanes, and ‘duality.’ Applications of these tools are explored including topics from machine learning and bioinformatics.


Sustainable Agriculture Bibliography, Matt Bogard Jan 2010

Sustainable Agriculture Bibliography, Matt Bogard

Agriculture Department Seminar Series

An annotated bibliography related to the sustainability of biotechnology and pharmaceutical technologies used in modern agriculture.


Using R, Matt Bogard Jan 2010

Using R, Matt Bogard

Economics Faculty Publications

R is a statistical programming language with a command line interface that is becoming more and more popular every day. I have used R for data visualization, data mining/machine learning, as well as social network analysis. Initially embraced largely in academia, R is becoming the software of choice in various corporate settings.


Using Twitter To Demonstrate Basic Concepts From Network Analysis, Matt Bogard Jan 2010

Using Twitter To Demonstrate Basic Concepts From Network Analysis, Matt Bogard

Economics Faculty Publications

Social network analysis focuses on finding patterns in interactions between people or entities. These patterns may be described in the form of a network. Network analysis in general has many applications including models of student integration and persistence, business to business supply chains, terrorist cells, or analysis of social media such as Facebook and Twitter. This presentation provides a reference for basic concepts from social network analysis with examples using tweets from Twitter.


Traveling Wave Solutions For A Nonlocal Reaction-Diffusion Model Of Influenza A Drift, Joaquin Riviera, Yi Li Jan 2010

Traveling Wave Solutions For A Nonlocal Reaction-Diffusion Model Of Influenza A Drift, Joaquin Riviera, Yi Li

Mathematics and Statistics Faculty Publications

In this paper we discuss the existence of traveling wave solutions for a nonlocal reaction-diffusion model of Influenza A proposed in Lin et. al. (2003). The proof for the existence of the traveling wave takes advantage of the different time scales between the evolution of the disease and the progress of the disease in the population. Under this framework we are able to use the techniques from geometric singular perturbation theory to prove the existence of the traveling wave.


Confidence Intervals In Survival Analysis, Tan Shay Kee Jan 2010

Confidence Intervals In Survival Analysis, Tan Shay Kee

Student Works (2010-2019)

In obtaining confidence interval for the survivor function using Greenwood’s formula and in performing log-rank test for comparing the survivor functions of two groups of individuals, only the information given by the first two moments of the relevant statistics are used. Presently we show that by incorporating the information given by the third and fourth moments of the statistics, the performance of the confidence interval and statistical test can be improved. When the survivor function can be described by the Weibull distribution, the knowledge regarding the survivor function can be obtained through the estimation of the Weibull scale and shape …


Economic Risk Assessment Using The Fractal Market Hypothesis, Jonathan Blackledge, Marek Rebow Jan 2010

Economic Risk Assessment Using The Fractal Market Hypothesis, Jonathan Blackledge, Marek Rebow

Conference papers

This paper considers the Fractal Market Hypothesi (FMH) for assessing the risk(s) in developing a financial portfolio based on data that is available through the Internet from an increasing number of sources. Most financial risk management systems are still based on the Efficient Market Hypothesis which often fails due to the inaccuracies of the statistical models that underpin the hypothesis, in particular, that financial data are based on stationary Gaussian processes. The FMH considered in this paper assumes that financial data are non-stationary and statistically self-affine so that a risk analysis can, in principal, be applied at any time scale …


Some Problems Of Outliers In Circular Data., Ali H.M. Abuzaid Jan 2010

Some Problems Of Outliers In Circular Data., Ali H.M. Abuzaid

Student Works (2010-2019)

This study considers three problems of outliers in circular statistics. The first problem is an attempt to use the standard outlier detection procedures for linear data set by approximating circular variables by linear variables. This is possible for large values of concentration parameter. Series of simulation studies are carried out to specify the accepted value of the concentration parameter so that the von Mises distribution can be approximated by normal distribution. The second is the problem of outliers in circular samples. Two numerical tests of discordancy are proposed to identify outliers. The test statistics are based on the summation of …


Impact Of Higher Education Workplace Relations Requirements On Nteu Membership Density, Tulsi Laxman Panchani Jan 2010

Impact Of Higher Education Workplace Relations Requirements On Nteu Membership Density, Tulsi Laxman Panchani

Theses : Honours

The National Tertiary Education Union (NTEU) is the only union working entirely in tertiary education around Australia. The union has over twenty four thousand members comprising of academic and general staff. NTEU maintains membership records at three levels, national, state and branch. The information collected includes gender, age group, employment type and work classification. In late April 2005, Higher Education Workplace Relations Requirements (HEWRRs) legislation was introduced by the Australian government. This legislation imposed restrictions on the interaction between the universities and the union and also curtailed an automatic involvement of the union in the resolution of workplace issues. The …


Encryption Using Deterministic Chaos, Jonathan Blackledge, Nikolai Ptitsyn Jan 2010

Encryption Using Deterministic Chaos, Jonathan Blackledge, Nikolai Ptitsyn

Articles

The concepts of randomness, unpredictability, complexity and entropy form the basis of modern cryptography and a cryptosystem can be interpreted as the design of a key-dependent bijective transformation that is unpredictable to an observer for a given computational resource. For any cryptosystem, including a Pseudo-Random Number Generator (PRNG), encryption algorithm or a key exchange scheme, for example, a cryptanalyst has access to the time series of a dynamic system and knows the PRNG function (the algorithm that is assumed to be based on some iterative process) which is taken to be in the public domain by virtue of the Kerchhoff-Shannon …


Canonical Correlation Analysis For Longitudinal Data, Raymond Mccollum Jan 2010

Canonical Correlation Analysis For Longitudinal Data, Raymond Mccollum

Mathematics & Statistics Theses & Dissertations

Data (multivariate data) on two sets of vectors commonly occur in applications. Statistical analysis of these data is usually done using a canonical correlation analysis (CCA). Occurrence of these data at multiple occasions or conditions leads to longitudinal multivariate data for a CCA. We address the problem of canonical correlation analysis on longitudinal data when the data have a Kronecker product covariance structure. Using structured correlation matrices we model the dependency of repeatedly observed data. Recent work of Srivastava, Nahtman, and von Rosen (2008) developed an iterative algorithm to determine the maximum likelihood estimate of the Kronecker product covariance structure …


Analysis Of Models For Longitudinal And Clustered Binary Data, Weiming Yang Jan 2010

Analysis Of Models For Longitudinal And Clustered Binary Data, Weiming Yang

Mathematics & Statistics Theses & Dissertations

This dissertation deals with modeling and statistical analysis of longitudinal and clustered binary data. Such data consists of observations on a dichotomous response variable generated from multiple time or cluster points, that exhibit either decaying correlation or equi-correlated dependence. The current literature addresses modeling the dependence using an appropriate correlation structure, but ignores the feasible bounds on the correlation parameter imposed by the marginal means.

The first part of this dissertation deals with two multivariate probability models, the first order Markov chain model and the multivariate probit model, that adhere to the feasible bounds on the correlation. For both the …


Modeling Residential Foreclosures In Kent County, Kaitlyn Ratkowiak Dec 2009

Modeling Residential Foreclosures In Kent County, Kaitlyn Ratkowiak

Student Summer Scholars Manuscripts

Residential Foreclosures in Kent County have become commonplace in the past few years. In this project, we hope to analyze data on foreclosures since 2004 to learn more about the mounting crisis, with the hope that we can identify neighborhoods at risk of foreclosures and its associated consequences.


Random Walks With Elastic And Reflective Lower Boundaries, Lucas Clay Devore Dec 2009

Random Walks With Elastic And Reflective Lower Boundaries, Lucas Clay Devore

Masters Theses & Specialist Projects

No abstract provided.


U.S. Chamber Of Commerce Liability Survey: Inaccurate, Unfair, And Bad For Business, Theodore Eisenberg Dec 2009

U.S. Chamber Of Commerce Liability Survey: Inaccurate, Unfair, And Bad For Business, Theodore Eisenberg

Cornell Law Faculty Publications

The U.S. Chamber of Commerce uses its Survey of State Liability to criticize judiciaries and seek legal change but no detailed evaluation of the survey’s quality exists. This article presents evidence that the survey is substantively inaccurate and methodologically flawed. It incorrectly characterizes state law; respondents provide less than 10 percent correct answers for objectively verifiable responses. It is internally inconsistent; a state threatened with judicial hellhole status ranked first in the survey while venues not on the list ranked lower. The absence of correlation between survey rankings and observable activity suggests that other factors drive the rankings. Two factors …


A Comparison Of Frequentist And Bayesian Approaches To The Estimation Of Long-Stay Per-Diems, Jeff Hatcher, Jason M. Sutherland Nov 2009

A Comparison Of Frequentist And Bayesian Approaches To The Estimation Of Long-Stay Per-Diems, Jeff Hatcher, Jason M. Sutherland

Dartmouth Scholarship

Within many diagnosis related group (DRG) systems, there is recognition that a single cost weight per DRG is not suitable, and that cost weights should take into account extremely lengthy hospital stays. Long lengths of stay are considered to be due to factors largely beyond the control of the hospital, and a single weight per DRG would potentially place hospitals under financial risk.

Within Canada's acute-care, inpatient grouping methodology - Case Mix Groups (CMG+) - long-stay episodes represent approximately 4.5% of all discharges. Within a CMG (analogous to DRG), the cost weight assigned to long-stay cases consists of the typical …


Application Of The Truncated Skew Laplace Probability Distribution In Maintenance System, Gokarna R. Aryal, Chris P. Tsokos Nov 2009

Application Of The Truncated Skew Laplace Probability Distribution In Maintenance System, Gokarna R. Aryal, Chris P. Tsokos

Journal of Modern Applied Statistical Methods

A random variable X is said to have the skew-Laplace probability distribution if its pdf is given by f(x) = 2g(x)G(λx), where g (.) and G (.), respectively, denote the pdf and the cdf of the Laplace distribution. When the skew Laplace distribution is truncated on the left at 0 it is called it the truncated skew Laplace (TSL) distribution. This article provides a comparison of TSL distribution with twoparameter gamma model and the hypoexponential model, and an application of the subject model in maintenance system is studied.


Examples Of Computing Power For Zero-Inflated And Overdispersed Count Data, Suzanne R. Doyle Nov 2009

Examples Of Computing Power For Zero-Inflated And Overdispersed Count Data, Suzanne R. Doyle

Journal of Modern Applied Statistical Methods

Examples of zero-inflated Poisson and negative binomial regression models were used to demonstrate conditional power estimation, utilizing the method of an expanded data set derived from probability weights based on assumed regression parameter values. SAS code is provided to calculate power for models with a binary or continuous covariate associated with zero-inflation.


An Inductive Approach To Calculate The Mle For The Double Exponential Distribution, W. J. Hurley Nov 2009

An Inductive Approach To Calculate The Mle For The Double Exponential Distribution, W. J. Hurley

Journal of Modern Applied Statistical Methods

Norton (1984) presented a calculation of the MLE for the parameter of the double exponential distribution based on the calculus. An inductive approach is presented here.


New Effect Size Rules Of Thumb, Shlomo S. Sawilowsky Nov 2009

New Effect Size Rules Of Thumb, Shlomo S. Sawilowsky

Journal of Modern Applied Statistical Methods

Recommendations to expand Cohen’s (1988) rules of thumb for interpreting effect sizes are given to include very small, very large, and huge effect sizes. The reasons for the expansion, and implications for designing Monte Carlo studies, are discussed.


Generating And Comparing Aggregate Variables For Use Across Datasets In Multilevel Analysis, James Chowhan, Laura Duncan Nov 2009

Generating And Comparing Aggregate Variables For Use Across Datasets In Multilevel Analysis, James Chowhan, Laura Duncan

Journal of Modern Applied Statistical Methods

This article examines the creation of contextual aggregate variables from one dataset for use with another dataset in multilevel analysis. The process of generating aggregate variables and methods of assessing the validity of the constructed aggregates are presented, together with the difficulties that this approach presents.


Detecting Lag-One Autocorrelation In Interrupted Time Series Experiments With Small Datasets, Clare Riviello, S. Natasha Beretvas Nov 2009

Detecting Lag-One Autocorrelation In Interrupted Time Series Experiments With Small Datasets, Clare Riviello, S. Natasha Beretvas

Journal of Modern Applied Statistical Methods

The power and type I error rates of eight indices for lag-one autocorrelation detection were assessed for interrupted time series experiments (ITSEs) with small numbers of data points. Performance of Huitema and McKean’s (2000) zHM statistic was modified and compared with the zHM, five information criteria and the Durbin-Watson statistic.


Relationship Between Internal Consistency And Goodness Of Fit Maximum Likelihood Factor Analysis With Varimax Rotation, Gibbs Y. Kanyongo, James B. Schreiber Nov 2009

Relationship Between Internal Consistency And Goodness Of Fit Maximum Likelihood Factor Analysis With Varimax Rotation, Gibbs Y. Kanyongo, James B. Schreiber

Journal of Modern Applied Statistical Methods

This study investigates how reliability (internal consistency) affects model-fitting in maximum likelihood exploratory factor analysis (EFA). This was accomplished through an examination of goodness of fit indices between the population and the sample matrices. Monte Carlo simulations were performed to create pseudo-populations with known parameters. Results indicated that the higher the internal consistency the worse the fit. It is postulated that the observations are similar to those from structural equation modeling where a good fit with low correlations can be observed and also the reverse with higher item correlations.


Estimating Model Complexity Of Feed-Forward Neural Networks, Douglas Landsittel Nov 2009

Estimating Model Complexity Of Feed-Forward Neural Networks, Douglas Landsittel

Journal of Modern Applied Statistical Methods

In a previous simulation study, the complexity of neural networks for limited cases of binary and normally-distributed variables based the null distribution of the likelihood ratio statistic and the corresponding chi-square distribution was characterized. This study expands on those results and presents a more general formulation for calculating degrees of freedom.


Level Robust Methods Based On The Least Squares Regression Estimator, Marie Ng, Rand R. Wilcox Nov 2009

Level Robust Methods Based On The Least Squares Regression Estimator, Marie Ng, Rand R. Wilcox

Journal of Modern Applied Statistical Methods

Heteroscedastic consistent covariance matrix (HCCM) estimators provide ways for testing hypotheses about regression coefficients under heteroscedasticity. Recent studies have found that methods combining the HCCM-based test statistic with the wild bootstrap consistently perform better than non-bootstrap HCCM-based methods (Davidson & Flachaire, 2008; Flachaire, 2005; Godfrey, 2006). This finding is more closely examined by considering a broader range of situations which were not included in any of the previous studies. In addition, the latest version of HCCM, HC5 (Cribari-Neto, et al., 2007), is evaluated.


Least Error Sample Distribution Function, Vassili F. Pastushenko Nov 2009

Least Error Sample Distribution Function, Vassili F. Pastushenko

Journal of Modern Applied Statistical Methods

Email: The empirical distribution function (ecdf) is unbiased in the usual sense, but shows certain order bias. Pyke suggested discrete ecdf using expectations of order statistics. Piecewise constant optimal ecdf saves 200%/N of sample size N. Results are compared with linear interpolation for U(0, 1), which require up to sixfold shorter samples at the same accuracy.


Confidence Interval Estimation For Intraclass Correlation Coefficient Under Unequal Family Sizes, Madhusudan Bhandary, Koji Fujiwara Nov 2009

Confidence Interval Estimation For Intraclass Correlation Coefficient Under Unequal Family Sizes, Madhusudan Bhandary, Koji Fujiwara

Journal of Modern Applied Statistical Methods

Confidence intervals (based on the χ2 -distribution and (Z) standard normal distribution) for the intraclass correlation coefficient under unequal family sizes based on a single multinormal sample have been proposed. It has been found that the confidence interval based on the χ2 -distribution consistently and reliably produces better results in terms of shorter average interval length than the confidence interval based on the standard normal distribution: especially for larger sample sizes for various intraclass correlation coefficient values. The coverage probability of the interval based on the χ2 -distribution is competitive with the coverage probability of the interval …


On Some Discrete Distributions And Their Applications With Real Life Data, Shipra Banik, B. M. Golam Kibria Nov 2009

On Some Discrete Distributions And Their Applications With Real Life Data, Shipra Banik, B. M. Golam Kibria

Journal of Modern Applied Statistical Methods

This article reviews some useful discrete models and compares their performance in terms of the high frequency of zeroes, which is observed in many discrete data (e.g., motor crash, earthquake, strike data, etc.). A simulation study is conducted to determine how commonly used discrete models (such as the binomial, Poisson, negative binomial, zero-inflated and zero-truncated models) behave if excess zeroes are present in the data. Results indicate that the negative binomial model and the ZIP model are better able to capture the effect of excess zeroes. Some real-life environmental data are used to illustrate the performance of the proposed models.


Closed Form Confidence Intervals For Small Sample Matched Proportions, James F. Reed Iii Nov 2009

Closed Form Confidence Intervals For Small Sample Matched Proportions, James F. Reed Iii

Journal of Modern Applied Statistical Methods

The behavior of the Wald-z, Wald-c, Quesenberry-Hurst, Wald-m and Agresti-Min methods was investigated for matched proportions confidence intervals. It was concluded that given the widespread use of the repeated-measure design, pretest-posttest design, matched-pairs design, and cross-over design, the textbook Wald-z method should be abandoned in favor of the Agresti-Min alternative.