Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 5611 - 5640 of 12832

Full-Text Articles in Statistics and Probability

Location Optimization Of A Coal Power Plant To Balance Coal Supply And Electric Transmission Costs Against Plant’S Emission Exposure, Najam Khan Jan 2018

Location Optimization Of A Coal Power Plant To Balance Coal Supply And Electric Transmission Costs Against Plant’S Emission Exposure, Najam Khan

Electronic Theses and Dissertations

This research is focused on developing a location analysis methodology that can minimize the pollutant exposure to the public while ensuring that the combined costs of electric transmission losses and coal logistics are minimized. Coal power plants will provide a critical contribution towards meeting electricity demands for various nations in the foreseeable future. The site selection for a new coal power plant is extremely important from an investment point of view. The operational costs for running a coal power plant can be minimized by a combined emphasis on placing a coal power plant near coal mines as well as customers. …


Mechanism Design, Matching Theory And The Stable Roommates Problem, Yashaswi Mohanty Jan 2018

Mechanism Design, Matching Theory And The Stable Roommates Problem, Yashaswi Mohanty

Honors Theses

This thesis consists of two independent albeit related chapters. The first chapter introduces concepts from mechanism design and matching theory, and discusses potential applications of this theory, particularly in relation to dorm allocations in colleges. The second chapter investigates a subset of the dorm allocation problem, namely that of matching roommates. In particular, the paper looks at the probability of solvability of random instances of the stable roommates game under the condition that preferences are not completely random and exogenous but endogenously determined through a dependence on room choice. These probabilities are estimated using Monte-Carlo simulations and then compared with …


Analysis Of Residential And Auto Break-In Records In Taipei City, Afnan Althoupety, Aishwarya Joy, Juchun Cheng, Priyanka Patil, Tejas Deshpande Jan 2018

Analysis Of Residential And Auto Break-In Records In Taipei City, Afnan Althoupety, Aishwarya Joy, Juchun Cheng, Priyanka Patil, Tejas Deshpande

Engineering and Technology Management Student Projects

Taipei City is the capital of Taiwan. It has population of 2.7 million living in the city area of 271 km2 (104 mi2). There are totally 12 administrative districts in this city. To maintain the safety of the city, Taipei City Police Bureau has arranged regular patrol routes with focus on the high-risk area where residential and auto break-in occurs. Due to limited police resource, resident neighborhood also organized volunteered patrol teams to enhance the security in residential area. Based on past file record history, the Bureau would like to understand the high-risk districts and time schedule to improve their …


A Quantitative Analysis Of Intermediate Forms Within Astarte From The Atlantic Coastal Plain, Philip Roberson Jan 2018

A Quantitative Analysis Of Intermediate Forms Within Astarte From The Atlantic Coastal Plain, Philip Roberson

Murray State Theses and Dissertations

The Atlantic Coastal Plain has long been recognized as a natural laboratory useful for testing hypotheses about various environmental and ecological effects on marine fauna. For studies such as these to continue being conducted in a rigorous and easily repeatable manner, a reliable taxonomy must be established for genera within this physiographic province. The bivalve genus, Astarte, is a cosmopolitan genus that is commonly found within the Atlantic Coastal Plain. This genus has many formally recognized species, even though it lacks many features that would encourage diversification, marking it as a taxonomic group in need of potential revision. The …


Students’ Interpretations Of Categorical Data Using Dynamic Graphical Representations, Adam Eide Jan 2018

Students’ Interpretations Of Categorical Data Using Dynamic Graphical Representations, Adam Eide

Master's Theses and Doctoral Dissertations

Statistical association is an important concept in statistics. An exploratory study examined how students reason about statistical association utilizing graphical representations constructed with CODAP, a dynamic statistical graphing software. Task-based interviews were conducted with three 6th grade students prior to formal instruction. Students’ conceptions of a statistical relationship, proportional reasoning skill level, ability to interpret bivariate categorical graphs (particularly segmented bar graphs and two-way binned plots), and ability to identify association of two categorical variables were all investigated through interview tasks and responses to inquiry. Students were found to have developing proportional reasoning skills and struggled to correctly define and …


Type I General Exponential Class Of Distributions, Gholamhossein G. Hamedani, Haitham M. Yousof, Mahdi Rasekhi, Morad Alizadeh, Seyed Morteza Najibi Jan 2018

Type I General Exponential Class Of Distributions, Gholamhossein G. Hamedani, Haitham M. Yousof, Mahdi Rasekhi, Morad Alizadeh, Seyed Morteza Najibi

Mathematics, Statistics and Computer Science Faculty Research and Publications

We introduce a new family of continuous distributions and study the mathematical properties of the new family. Some useful characterizations based on the ratio of two truncated moments and hazard function are also presented. We estimate the model parameters by the maximum likelihood method and assess its performance based on biases and mean squared errors in a simulation study framework.


Prediction Intervals For Functional Data, Nicholas Rios Jan 2018

Prediction Intervals For Functional Data, Nicholas Rios

Theses, Dissertations and Culminating Projects

The prediction of functional data samples has been the focus of several functional data analysis endeavors. This work describes the use of dynamic function-on-function regression for dynamic prediction of the future trajectory as well as the construction of dynamic prediction intervals for functional data. The overall goals of this thesis are to assess the efficacy of Dynamic Penalized Function-on-Function Regression (DPFFR) and to compare DPFFR prediction intervals with those of other dynamic prediction methods. To make these comparisons, metrics are used that measure prediction error, prediction interval width, and prediction interval coverage. Simulations and applications to financial stock data from …


Spatio-Temporal Frequency Separation With Application Of Kolmogorov-Zurbenko Filters To The Multivariate Analysis Of Melanoma Prevalence, Edward Valachovic Jan 2018

Spatio-Temporal Frequency Separation With Application Of Kolmogorov-Zurbenko Filters To The Multivariate Analysis Of Melanoma Prevalence, Edward Valachovic

Legacy Theses & Dissertations (2009 - 2024)

Time Series Analysis is the observation of variables recorded across time. Observations are visualized and analysis often performed in the native time domain. It is common for a time series to be the dependent variable of more than one factor. Several factors can have concurrent and combined effects. The time domain presents an obstacle due to constructive and destructive interference of factors at each time point. Unless effects are clearly pronounced and separable, the entanglement of factors along with the presence and intensity of random variation can obscure true relationships.


Stress-Strength Estimation And Its Applications In Clinical Trials, Dinesh Kumar Jan 2018

Stress-Strength Estimation And Its Applications In Clinical Trials, Dinesh Kumar

Legacy Theses & Dissertations (2009 - 2024)

Stress Strength model P(X


Race, Ethnicity, And The Great Recession : A National Evaluation Of Mortgages And Subprime Lending, 2004-2010, Meghan M. O'Neil Jan 2018

Race, Ethnicity, And The Great Recession : A National Evaluation Of Mortgages And Subprime Lending, 2004-2010, Meghan M. O'Neil

Legacy Theses & Dissertations (2009 - 2024)

The dissertation analyzes multilevel models to predict mortgage origination and the allocation of subprime credit pre-and-post Great Recession. With representative samples from two full years of mortgage applications filed in the top 100 U.S. metropolitan areas, the dissertation uncovers evidence of persistent disparities by race and neighborhood minority concentration despite controls for socioeconomic, demographic, assimilation and housing variables. Mortgage outcomes varied by applicant race, neighborhood racial composition and neighborhood racial change. Findings suggest evidence of Fair Housing Act violations and disparate impacts towards minority homebuyers and minority neighborhoods. Results lend support for spatial assimilation theories in explaining much of the …


Non-Stationary Counts With Mixture Distributions, Ziqiang Lin Jan 2018

Non-Stationary Counts With Mixture Distributions, Ziqiang Lin

Legacy Theses & Dissertations (2009 - 2024)

We study a new non--stationary mixture Pengram and thinning model for time series of counts that include the effect of covariate variables on the outcome variable. Properties of the model and performance are discussed. It has a simpler likelihood function than the non--stationary INAR(1) model and therefore MLE estimators for the model's parameters are easier to find. Therefore the model offers an alternative to non--stationary INAR(1).


Accounting For Spatial Autocorrelation In Modeling The Distribution Of Water Quality Variables, Lorrayne Miralha Jan 2018

Accounting For Spatial Autocorrelation In Modeling The Distribution Of Water Quality Variables, Lorrayne Miralha

Theses and Dissertations--Geography

Several studies in hydrology have reported differences in outcomes between models in which spatial autocorrelation (SAC) is accounted for and those in which SAC is not. However, the capacity to predict the magnitude of such differences is still ambiguous. In this thesis, I hypothesized that SAC, inherently possessed by a response variable, influences spatial modeling outcomes. I selected ten watersheds in the USA and analyzed them to determine whether water quality variables with higher Moran’s I values undergo greater increases in the coefficient of determination (R²) and greater decreases in residual SAC (rSAC) after spatial modeling. I compared non-spatial ordinary …


Improved Methods And Selecting Classification Types For Time-Dependent Covariates In The Marginal Analysis Of Longitudinal Data, I-Chen Chen Jan 2018

Improved Methods And Selecting Classification Types For Time-Dependent Covariates In The Marginal Analysis Of Longitudinal Data, I-Chen Chen

Theses and Dissertations--Epidemiology and Biostatistics

Generalized estimating equations (GEE) are popularly utilized for the marginal analysis of longitudinal data. In order to obtain consistent regression parameter estimates, these estimating equations must be unbiased. However, when certain types of time-dependent covariates are presented, these equations can be biased unless an independence working correlation structure is employed. Moreover, in this case regression parameter estimation can be very inefficient because not all valid moment conditions are incorporated within the corresponding estimating equations. Therefore, approaches using the generalized method of moments or quadratic inference functions have been proposed for utilizing all valid moment conditions. However, we have found that …


An Investigation Of Atomic Structures Derived From X-Ray Crystallography And Cryo-Electron Microscopy Using Distal Blocks Of Side-Chains, Lin Chen, Jing He, Salim Sazzed, Rayshawn Walker Jan 2018

An Investigation Of Atomic Structures Derived From X-Ray Crystallography And Cryo-Electron Microscopy Using Distal Blocks Of Side-Chains, Lin Chen, Jing He, Salim Sazzed, Rayshawn Walker

Computer Science Faculty Publications

Cryo-electron microscopy (cryo-EM) is a structure determination method for large molecular complexes. As more and more atomic structures are determined using this technique, it is becoming possible to perform statistical characterization of side-chain conformations. Two data sets were involved to characterize block lengths for each of the 18 types of amino acids. One set contains 9131 structures resolved using X-ray crystallography from density maps with better than or equal to 1.5 Å resolutions, and the other contains 237 protein structures derived from cryo-EM density maps with 2-4 Å resolutions. The results show that the normalized probability density function of block …


Statistical Algorithms And Bioinformatics Tools Development For Computational Analysis Of High-Throughput Transcriptomic Data, Adam Mcdermaid Jan 2018

Statistical Algorithms And Bioinformatics Tools Development For Computational Analysis Of High-Throughput Transcriptomic Data, Adam Mcdermaid

Electronic Theses and Dissertations

Next-Generation Sequencing technologies allow for a substantial increase in the amount of data available for various biological studies. In order to effectively and efficiently analyze this data, computational approaches combining mathematics, statistics, computer science, and biology are implemented. Even with the substantial efforts devoted to development of these approaches, numerous issues and pitfalls remain. One of these issues is mapping uncertainty, in which read alignment results are biased due to the inherent difficulties associated with accurately aligning RNA-Sequencing reads. GeneQC is an alignment quality control tool that provides insight into the severity of mapping uncertainty in each annotated gene from …


Variable Selection Techniques For Clustering On The Unit Hypersphere, Damon Bayer Jan 2018

Variable Selection Techniques For Clustering On The Unit Hypersphere, Damon Bayer

Electronic Theses and Dissertations

Mixtures of von Mises-Fisher distributions have been shown to be an effective model for clustering data on a unit hypersphere, but variable selection for these models remains an important and challenging problem. In this paper, we derive two variants of the expectation-maximization framework, which are each used to identify a specific type of irrelevant variables for these models. The first type are noise variables, which are not useful for separating any pairs of clusters. The second type are redundant variables, which may be useful for separating pairs of clusters, but do not enable any additional separation beyond the separability provided …


New Developments Of Dimension Reduction, Lei Huo Jan 2018

New Developments Of Dimension Reduction, Lei Huo

Doctoral Dissertations

"Variable selection becomes more crucial than before, since high dimensional data are frequently seen in many research areas. Many model-based variable selection methods have been developed. However, the performance might be poor when the model is mis-specified. Sufficient dimension reduction (SDR, Li 1991; Cook 1998) provides a general framework for model-free variable selection methods.

In this thesis, we first propose a novel model-free variable selection method to deal with multi-population data by incorporating the grouping information. Theoretical properties of our proposed method are also presented. Simulation studies show that our new method significantly improves the selection performance compared with those …


A Bayesian Model For Spectral Density Estimation, Yi Xie Jan 2018

A Bayesian Model For Spectral Density Estimation, Yi Xie

Open Access Theses & Dissertations

When we analyze a stationary time series, one of the questions we often meet is how to estimate its spectral density. Many approaches have been proposed to this end. In this paper we estimate the spectral density of a stationary time series nonparametrically. We fit a nonparametric regression model to the log periodogram and use third-degree B-spline functions as basis functions. Since the the number of basis functions is relatively large, we place priors such as random-walk and regularized horseshoe on the coefficients of the basis functions to avoid over-fitting and smooth the log periodogram.


Integrated Statistical And Machine Learning Algorithms For Predicting And Classifying G Protein-Coupled Receptors, Fredrick Ayivor Jan 2018

Integrated Statistical And Machine Learning Algorithms For Predicting And Classifying G Protein-Coupled Receptors, Fredrick Ayivor

Open Access Theses & Dissertations

G protein-coupled receptors (GPCRs) are transmembrane proteins with important functions in signal transduction and often serve as drug targets. With increasing availability of protein sequence information, there is much interest in computationally predicting GPCRs and classifying them according to their biological roles. Such predictions are cost-efficient and can be valuable guides for designing wet lab experiments to help elucidate signaling pathways and expedite drug discovery. There are existing computational tools of GPCR prediction that involve principal component analysis (PCA), intimate sorting (IS), support vector machine, and random forest (RF) techniques using various sequence derived features. While accuracies of over 90\% …


Statistics And Biomechanics: An Interdisciplinary Evaluation Of The Mathematical, Practical, And Athletic Applications Of Principal Component Analysis, Sydney Grace Davis Jan 2018

Statistics And Biomechanics: An Interdisciplinary Evaluation Of The Mathematical, Practical, And Athletic Applications Of Principal Component Analysis, Sydney Grace Davis

Undergraduate Honors Theses

Excerpt from Introduction

Coaches and athletes around the world are in constant pursuit of improving their athletic performance. For some, a routine amount of weight lifting, cardiovascular exercise, agilities and flexibility training with gradual advancement may be enough to see growth. However, many are not satisfied and turn to in-depth analyses of their techniques in order to measure their progress. After an intense workout or competition, coaches spend time breaking down athletic performances based on major movements. By compartmentalizing these activities, they can identify which motions are efficient and which ones hinder fluid motion. From this evaluation and discernment, athletes …


Categorizing A Continuous Predictor Subject To Measurement Error, Betsabé G. Blas Achic, Tianying Wang, Ya Su, Victor Kipnis, Kevin Dodd, Raymond J. Carroll Jan 2018

Categorizing A Continuous Predictor Subject To Measurement Error, Betsabé G. Blas Achic, Tianying Wang, Ya Su, Victor Kipnis, Kevin Dodd, Raymond J. Carroll

Statistics Faculty Publications

Epidemiologists often categorize a continuous risk predictor, even when the true risk model is not a categorical one. Nonetheless, such categorization is thought to be more robust and interpretable, and thus their goal is to fit the categorical model and interpret the categorical parameters. We address the question: with measurement error and categorization, how can we do what epidemiologists want, namely to estimate the parameters of the categorical model that would have been estimated if the true predictor was observed? We develop a general methodology for such an analysis, and illustrate it in linear and logistic regression. Simulation studies are …


Time Dependent Attribute-Level Best Worst Discrete Choice Modelling, Amanda Working, Mohammed Alqawba, Norou Diawara, Ling Li Jan 2018

Time Dependent Attribute-Level Best Worst Discrete Choice Modelling, Amanda Working, Mohammed Alqawba, Norou Diawara, Ling Li

Mathematics & Statistics Faculty Publications

Discrete choice models (DCMs) are applied in statistical modelling of consumer behavior. Such models are used in many areas including social sciences, health economics, transportation research, and health systems research and they are time dependent. In this manuscript, we review references on the study of such models, develop DCMs with emphasis on time dependent best-worst choice and discrimination between choice attributes. Referenced measurements of the dynamic DCMs are simulated. Expected utilities over time are derived using Markov decision processes. We study attributes and attribute-levels associated with the quality of life of seniors, report the estimation results, and discuss our findings.


Microrna Expression Patterns In Human Anterior Cingulate And Motor Cortex: A Study Of Dementia With Lewy Bodies Cases And Controls, Peter T. Nelson, Wang-Xia Wang, Sarah A. Janse, Katherine L. Thompson Jan 2018

Microrna Expression Patterns In Human Anterior Cingulate And Motor Cortex: A Study Of Dementia With Lewy Bodies Cases And Controls, Peter T. Nelson, Wang-Xia Wang, Sarah A. Janse, Katherine L. Thompson

Sanders-Brown Center on Aging Faculty Publications

Overview

MicroRNAs (miRNAs) have been implicated in neurodegenerative diseases including Parkinson’s disease and Alzheimer’s disease (AD). Here, we evaluated the expression of miRNAs in anterior cingulate (AC; Brodmann area [BA] 24) and primary motor (MO; BA 4) cortical tissue from aged human brains in the University of Kentucky AD Center autopsy cohort, with a focus on dementia with Lewy bodies (DLB).

Methods

RNA was isolated from gray matter of brain samples with pathology-defined DLB, AD, AD+DLB, and low-pathology controls, with n=52 cases initially included (n=23 with DLB), all with low (<4hrs) postmortem intervals. RNA was profiled using Exiqon miRNA microarrays. Quantitative PCR for post-hoc replication was performed on separate cases (n=6 controls) and included RNA isolated from gray matter of MO, AC, primary somatosensory (BA 3), and dorsolateral prefrontal (BA 9) cortical regions.

Results

The miRNA expression patterns differed substantially according to …


Automated Tree-Level Forest Quantification Using Airborne Lidar, Hamid Hamraz Jan 2018

Automated Tree-Level Forest Quantification Using Airborne Lidar, Hamid Hamraz

Theses and Dissertations--Computer Science

Traditional forest management relies on a small field sample and interpretation of aerial photography that not only are costly to execute but also yield inaccurate estimates of the entire forest in question. Airborne light detection and ranging (LiDAR) is a remote sensing technology that records point clouds representing the 3D structure of a forest canopy and the terrain underneath. We present a method for segmenting individual trees from the LiDAR point clouds without making prior assumptions about tree crown shapes and sizes. We then present a method that vertically stratifies the point cloud to an overstory and multiple understory tree …


Modeling And Mapping Location-Dependent Human Appearance, Zachary Bessinger Jan 2018

Modeling And Mapping Location-Dependent Human Appearance, Zachary Bessinger

Theses and Dissertations--Computer Science

Human appearance is highly variable and depends on individual preferences, such as fashion, facial expression, and makeup. These preferences depend on many factors including a person's sense of style, what they are doing, and the weather. These factors, in turn, are dependent upon geographic location and time. In our work, we build computational models to learn the relationship between human appearance, geographic location, and time. The primary contributions are a framework for collecting and processing geotagged imagery of people, a large dataset collected by our framework, and several generative and discriminative models that use our dataset to learn the relationship …


Occurrence And Attributes Of Two Echinoderm-Bearing Faunas From The Upper Mississippian (Chesterian; Lower Serpukhovian) Ramey Creek Member, Slade Formation, Eastern Kentucky, U.S.A., Ann Well Harris Jan 2018

Occurrence And Attributes Of Two Echinoderm-Bearing Faunas From The Upper Mississippian (Chesterian; Lower Serpukhovian) Ramey Creek Member, Slade Formation, Eastern Kentucky, U.S.A., Ann Well Harris

Theses and Dissertations--Earth and Environmental Sciences

Well-preserved echinoderm faunas are rare in the fossil record, and when uncovered, understanding their occurrence can be useful in interpreting other faunas. In this study, two such faunas of the same age from separate localities in the shallow-marine Ramey Creek Member of the Slade Formation in the Upper Mississippian (Chesterian) rocks of eastern Kentucky are examined. Of the more than 5,000 fossil specimens from both localities, only 9–34 percent were echinoderms from 3–5 classes. Nine non-echinoderm (8 invertebrate and one vertebrate) classes occurred at both localities, but of these, bryozoans, brachiopods and sponges dominated. To understand the attributes of both …


High Dimensional Multivariate Inference Under General Conditions, Xiaoli Kong Jan 2018

High Dimensional Multivariate Inference Under General Conditions, Xiaoli Kong

Theses and Dissertations--Statistics

In this dissertation, we investigate four distinct and interrelated problems for high-dimensional inference of mean vectors in multi-groups.

The first problem concerned is the profile analysis of high dimensional repeated measures. We introduce new test statistics and derive its asymptotic distribution under normality for equal as well as unequal covariance cases. Our derivations of the asymptotic distributions mimic that of Central Limit Theorem with some important peculiarities addressed with sufficient rigor. We also derive consistent and unbiased estimators of the asymptotic variances for equal and unequal covariance cases respectively.

The second problem considered is the accurate inference for high-dimensional repeated …


Accounting For Matching Uncertainty In Photographic Identification Studies Of Wild Animals, Amanda R. Ellis Jan 2018

Accounting For Matching Uncertainty In Photographic Identification Studies Of Wild Animals, Amanda R. Ellis

Theses and Dissertations--Statistics

I consider statistical modelling of data gathered by photographic identification in mark-recapture studies and propose a new method that incorporates the inherent uncertainty of photographic identification in the estimation of abundance, survival and recruitment. A hierarchical model is proposed which accepts scores assigned to pairs of photographs by pattern recognition algorithms as data and allows for uncertainty in matching photographs based on these scores. The new models incorporate latent capture histories that are treated as unknown random variables informed by the data, contrasting past models having the capture histories being fixed. The methods properly account for uncertainty in the matching …


The Family Of Conditional Penalized Methods With Their Application In Sufficient Variable Selection, Jin Xie Jan 2018

The Family Of Conditional Penalized Methods With Their Application In Sufficient Variable Selection, Jin Xie

Theses and Dissertations--Statistics

When scientists know in advance that some features (variables) are important in modeling a data, then these important features should be kept in the model. How can we utilize this prior information to effectively find other important features? This dissertation is to provide a solution, using such prior information. We propose the Conditional Adaptive Lasso (CAL) estimates to exploit this knowledge. By choosing a meaningful conditioning set, namely the prior information, CAL shows better performance in both variable selection and model estimation. We also propose Sufficient Conditional Adaptive Lasso Variable Screening (SCAL-VS) and Conditioning Set Sufficient Conditional Adaptive Lasso Variable …


Mixtures-Of-Regressions With Measurement Error, Xiaoqiong Fang Jan 2018

Mixtures-Of-Regressions With Measurement Error, Xiaoqiong Fang

Theses and Dissertations--Statistics

Finite Mixture model has been studied for a long time, however, traditional methods assume that the variables are measured without error. Mixtures-of-regression model with measurement error imposes challenges to the statisticians, since both the mixture structure and the existence of measurement error can lead to inconsistent estimate for the regression coefficients. In order to solve the inconsistency, We propose series of methods to estimate the mixture likelihood of the mixtures-of-regressions model when there is measurement error, both in the responses and predictors. Different estimators of the parameters are derived and compared with respect to their relative efficiencies. The simulation results …