Statistically Analyzing Assembly Line Processing Times Through Incorporation Of Product Variation,
2017
Murray State University
Statistically Analyzing Assembly Line Processing Times Through Incorporation Of Product Variation, Kyle Rehr, Matthew Farr
Scholars Week
Timing methods and performance metrics are important in the heavily industrialized world we live in. Industrial plants use metrics to measure quality of production, help make decisions, and drive the strategy of the organization. However, there are many factors to be considered when measuring performance based on a metric; of which we will be analyzing the importance of product variation. We will be analyzing assembly line timings, whilst controlling for product variance, to show the importance differences between products makes in one’s ability to predict performance. In addition, we will be analyzing the current “statistical” methods used by an industrial …
Assessing The Impact Of Retreat Mechanisms In A Simple Antarctic Ice Sheet Model Using Bayesian Calibration,
2017
The Pennsylvania State University
Assessing The Impact Of Retreat Mechanisms In A Simple Antarctic Ice Sheet Model Using Bayesian Calibration, Kelsey L. Ruckert, Gary Shaffer, David Pollard, Yawen Guan, Tony E. Wong, Chris E. Forest, Klaus Keller
Department of Statistics: Faculty Publications
The response of the Antarctic ice sheet (AIS) to changing climate forcings is an important driver of sea-level changes. Anthropogenic climate change may drive a sizeable AIS tipping point response with subsequent increases in coastal flooding risks. Many studies analyzing flood risks use simple models to project the future responses of AIS and its sea-level contributions. These analyses have provided important new insights, but they are often silent on the effects of potentially important processes such as Marine Ice Sheet Instability (MISI) or Marine Ice Cliff Instability (MICI). These approximations can be well justified and result in more parsimonious and …
Perennial Warm-Season Grasses For Producing Biofuel And Enhancing Soil Properties: An Alternative To Corn Residue Removal,
2017
University of Nebraska-Lincoln
Perennial Warm-Season Grasses For Producing Biofuel And Enhancing Soil Properties: An Alternative To Corn Residue Removal, Humberto Blanco-Canqui, Robert B. Mitchell, Virginia L. Jin, Marty R. Schmer, Kent M. Eskridge
Department of Statistics: Faculty Publications
Removal of corn (Zea mays L.) residues at high rates for biofuel and other off-farm uses may negatively impact soil and the environment in the long term. Biomass removal from perennial warm-season grasses (WSGs) grown in marginally-productive lands could be an alternative to corn residue removal as biofuel feedstocks while controlling water and wind erosion, sequestering carbon (C), cycling water and nutrients, and enhancing other soil ecosystem services. We compared wind and water erosion potential, soil compaction, soil hydraulic properties, soil organic C (SOC), and soil fertility between biomass removal from WSGs and corn residue removal from rainfed no-till …
Impact Of Menthol Smoking On Nicotine Dependence
For Diverse Racial/Ethnic Groups Of Daily Smokers,
2017
University of Central Florida
Impact Of Menthol Smoking On Nicotine Dependence For Diverse Racial/Ethnic Groups Of Daily Smokers, Julia N. Soulakova, Ryan R. Danczak
Department of Statistics: Faculty Publications
Introduction: The aims of this study were to evaluate whether menthol smoking and race/ethnicity are associated with nicotine dependence in daily smokers. Methods: The study used two subsamples of U.S. daily smokers who responded to the 2010–2011 Tobacco Use Supplement to the Current Population Survey. The larger subsample consisted of 18,849 non-Hispanic White (NHW), non-Hispanic Black (NHB), and Hispanic (HISP) smokers. The smaller subsample consisted of 1112 non-Hispanic American Indian/Alaska Native (AIAN), non-Hispanic Asian (ASIAN), non-Hispanic Hawaiian/Pacific Islander (HPI), and non-Hispanic Multiracial (MULT) smokers. Results: For larger (smaller) groups the rates were 45% (33%) for heavy smoking (16+ cig/day), 59% …
Evaluating Current Practices In Shelf Life Estimation,
2017
Merck & Co. Inc.
Evaluating Current Practices In Shelf Life Estimation, Robert Capen, David Christopher, Patrick Forenzo, Kim Huynh-Ba, David Leblond, Oscar Liu, John O'Neill, Nate Patterson, Michelle Quinlan, Radhika Rajagopalan, James Schwenke, Walter W. Stroup
Department of Statistics: Faculty Publications
The current International Council for Harmonisation of Technical Requirements for Pharmaceuticals for Human Use (ICH) methods for determining the supported shelf life of a drug product, described in ICH guidance documents Q1A and Q1E, are evaluated in this paper. To support this evaluation, an industry data set is used which is comprised of 26 individual stability batches of a common drug product where most batches are measured over a 24 month storage period. Using randomly sampled sets of 3 or 6 batches from the industry data set, the current ICH methods are assessed from three perspectives. First, the distributional properties …
Generalized Confidence Intervals Compatible With The Min Test For Simultaneous Comparisons Of One Subpopulation To Several Other Subpopulations,
2017
University of Nebraska-Lincoln
Generalized Confidence Intervals Compatible With The Min Test For Simultaneous Comparisons Of One Subpopulation To Several Other Subpopulations, Julia N. Soulakova
Department of Statistics: Faculty Publications
A problem where one subpopulation is compared to several other subpopulations in terms of means with the goal of estimating the smallest difference between the means commonly arises in biology, medicine, and many other scientific fields. A generalization of Strassburger, Bretz and Hochberg (2004) approach for two comparisons is presented for cases with three and more comparisons. The method allows constructing an interval-estimator for the smallest mean difference, which is compatible with the Min test. An application to a fluency-disorder study is illustrated. Simulations confirmed adequate probability coverage for normally distributed outcomes for a number of designs.
Increasing Genomic-Enabled Prediction Accuracy
By Modeling Genotype X Environment Interactions
In Kansas Wheat,
2017
University of Nebraska-Lincoln
Increasing Genomic-Enabled Prediction Accuracy By Modeling Genotype X Environment Interactions In Kansas Wheat, Diego Jarquin, Cristiano Lemas Da Silva, R. Chris Gaynor, Jesse Poland, Allan Fritz, Reka Howard, Sarah Battenfield, José Crossa
Department of Statistics: Faculty Publications
Wheat (Triticum aestivum L.) breeding programs test experimental lines in multiple locations over multiple years to get an accurate assessment of grain yield and yield stability. Selections in early generations of the breeding pipeline are based on information from only one or few locations and thus materials are advanced with little knowledge of the genotype × environment interaction (G × E) effects. Later, large trials are conducted in several locations to assess the performance of more advanced lines across environments. Genomic selection (GS) models that include G × E covariates allow us to borrow information not only from related …
Application Of Response Surface Methods To
Determine Conditions For Optimal
Genomic Prediction,
2017
University of Nebraska-Lincoln
Application Of Response Surface Methods To Determine Conditions For Optimal Genomic Prediction, Reka Howard, Alicia L. Carriquiry, William D. Beavis
Department of Statistics: Faculty Publications
An epistatic genetic architecture can have a significant impact on prediction accuracies of genomic prediction (GP) methods. Machine learning methods predict traits comprised of epistatic genetic architectures more accurately than statistical methods based on additive mixed linear models. The differences between these types of GP methods suggest a diagnostic for revealing genetic architectures underlying traits of interest. In addition to genetic architecture, the performance of GP methods may be influenced by the sample size of the training population, the number of QTL, and the proportion of phenotypic variability due to genotypic variability (heritability). Possible values for these factors and the …
Trans-Ancestry Fine Mapping And Molecular Assays
Identify Regulatory Variants At The Angptl8
Hdl-C Gwas Locus,
2017
University of North Carolina at Chapel Hill
Trans-Ancestry Fine Mapping And Molecular Assays Identify Regulatory Variants At The Angptl8 Hdl-C Gwas Locus, Maren E. Cannon, Qing Duan, Ying Wu, Monica Zeynalzadeh, Zheng Xu, Antti J. Kangas, Pasi Soininen, Mika Ala-Korpela, Mete Civelek, Aldons J. Lusis, Johanna Kuusisto, Francis S. Collins, Michael Boehnke, Hua Tang, Markku Laakso, Yun Li, Karen L. Mohlke
Department of Statistics: Faculty Publications
Recent genome-wide association studies (GWAS) have identified variants associated with highdensity lipoprotein cholesterol (HDL-C) located in or near the ANGPTL8 gene. Given the extensive sharing of GWAS loci across populations, we hypothesized that at least one shared variant at this locus affects HDL-C. The HDL-C–associated variants are coincident with expression quantitative trait loci for ANGPTL8 and DOCK6 in subcutaneous adipose tissue; however, only ANGPTL8 expression levels are associated with HDL-C levels. We identified a 400-bp promoter region of ANGPTL8 and enhancer regions within 5 kb that contribute to regulating expression in liver and adipose. To identify variants functionally responsible for …
A Bayes Interpretation Of Stacking For M-Complete And M-Open Settings,
2017
University of Nebraska-Lincoln
A Bayes Interpretation Of Stacking For M-Complete And M-Open Settings, Tri Le, Bertrand S. Clarke
Department of Statistics: Faculty Publications
In M-open problems where no true model can be conceptualized, it is common to back off from modeling and merely seek good prediction. Even in M-complete problems, taking a predictive approach can be very useful. Stacking is a model averaging procedure that gives a composite predictor by combining individual predictors from a list of models using weights that optimize a cross validation criterion. We show that the stacking weights also asymptotically minimize a posterior expected loss. Hence we formally provide a Bayesian justification for cross-validation. Often the weights are constrained to be positive and sum to one. For greater generality, …
Detecting Discordance Enrichment Among A Series Of Two-Sample Genome-Wide Expression Data Sets,
2017
George Washington University
Detecting Discordance Enrichment Among A Series Of Two-Sample Genome-Wide Expression Data Sets, Yinglei Lai, Fanni Zhang, Tapan Nayak, Reza Modarres, Norman H. Lee, Timothy A. Mccaffrey
Epidemiology Faculty Publications
Background
With the current microarray and RNA-seq technologies, two-sample genome-wide expression data have been widely collected in biological and medical studies. The related differential expression analysis and gene set enrichment analysis have been frequently conducted. Integrative analysis can be conducted when multiple data sets are available. In practice, discordant molecular behaviors among a series of data sets can be of biological and clinical interest.
Methods
In this study, a statistical method is proposed for detecting discordance gene set enrichment. Our method is based on a two-level multivariate normal mixture model. It is statistically efficient with linearly increased parameter space when …
What’S Brewing? A Statistics Education Discovery Project,
2017
CUNY Guttman Community College
What’S Brewing? A Statistics Education Discovery Project, Marla A. Sole, Sharon L. Weinberg
Publications and Research
We believe that students learn best, are actively engaged, and are genuinely interested when working on real-world problems. This can be done by giving students the opportunity to work collaboratively on projects that investigate authentic, familiar problems. This article shares one such project that was used in an introductory statistics course. We describe the steps taken to investigate why customers are charged more for iced coffee than hot coffee, which included collecting data and using descriptive and inferential statistical analysis. Interspersed throughout the article, we describe strategies that can help teachers implement the project and scaffold material to assist students …
Selection Portfolio: Applying Modern Portfolio Theory To Personnel Selection,
2017
Minnesota State University, Mankato
Selection Portfolio: Applying Modern Portfolio Theory To Personnel Selection, Eric Leingang
All Graduate Theses, Dissertations, and Other Capstone Projects
Modern Portfolio Theory (MPT) is a framework for building a portfolio of risky assets such that the ratio of risk to return is minimized. While this theory has been used in the field of financial economics for over sixty years, the method has not yet been applied to compensatory personnel selection. A common method for personnel selection is multiple regression to maximize the predicted performance of the selected group given a cut-off score on the predictor(s). Recognizing that maximizing the performance of the selected group is not the only consideration, and that, for many jobs and organizations, the outcomes of …
A Traders Guide To The Predictive Universe- A Model For Predicting Oil Price Targets And Trading On Them,
2016
Washington University in St. Louis
A Traders Guide To The Predictive Universe- A Model For Predicting Oil Price Targets And Trading On Them, Jimmie Harold Lenz
Doctor of Business Administration Dissertations
At heart every trader loves volatility; this is where return on investment comes from, this is what drives the proverbial “positive alpha.” As a trader, understanding the probabilities related to the volatility of prices is key, however if you could also predict future prices with reliability the world would be your oyster. To this end, I have achieved three goals with this dissertation, to develop a model to predict future short term prices (direction and magnitude), to effectively test this by generating consistent profits utilizing a trading model developed for this purpose, and to write a paper that anyone with …
Tutorial For Using The Center For High Performance Computing At The University Of Utah And An Example Using Random Forest,
2016
Utah State University
Tutorial For Using The Center For High Performance Computing At The University Of Utah And An Example Using Random Forest, Stephen Barton
All Graduate Plan B and other Reports, Spring 1920 to Spring 2023
Random Forests are very memory intensive machine learning algorithms and most computers would fail at building models from datasets with millions of observations. Using the Center for High Performance Computing (CHPC) at the University of Utah and an airline on-time arrival dataset with 7 million observations from the U.S. Department of Transportation Bureau of Transportation Statistics we built 316 models by adjusting the depth of the trees and randomness of each forest and compared the accuracy and time each took. Using this dataset we discovered that substantial restrictions to the size of trees, observations allowed for each tree, and variables …
Performance-Constrained Binary Classification Using Ensemble Learning: An Application To Cost-Efficient Targeted Prep Strategies,
2016
Division of Biostatistics, School of Public Health, University of California, Berkeley
Performance-Constrained Binary Classification Using Ensemble Learning: An Application To Cost-Efficient Targeted Prep Strategies, Wenjing Zheng, Laura Balzer, Maya L. Petersen, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Binary classifications problems are ubiquitous in health and social science applications. In many cases, one wishes to balance two conflicting criteria for an optimal binary classifier. For instance, in resource-limited settings, an HIV prevention program based on offering Pre-Exposure Prophylaxis (PrEP) to select high-risk individuals must balance the sensitivity of the binary classifier in detecting future seroconverters (and hence offering them PrEP regimens) with the total number of PrEP regimens that is financially and logistically feasible for the program to deliver. In this article, we consider a general class of performance-constrained binary classification problems wherein the objective function and the …
Optimal Design Of Low-Density Snp Arrays
For Genomic Prediction: Algorithm And
Applications,
2016
GeneSeek (a Neogen Company)
Optimal Design Of Low-Density Snp Arrays For Genomic Prediction: Algorithm And Applications, Xiao-Lin Wu, Jiaqi Xu, Guofei Feng, George R. Wiggans, Jeremy F. Taylor, Jun He, Changsong Qian, Jiansheng Qiu, Barry Simpson, Jeremy Walker, Stewart Bauck
Department of Statistics: Faculty Publications
Low-density (LD) single nucleotide polymorphism (SNP) arrays provide a cost-effective solution for genomic prediction and selection, but algorithms and computational tools are needed for the optimal design of LD SNP chips. A multiple-objective, local optimization (MOLO) algorithm was developed for design of optimal LD SNP chips that can be imputed accurately to medium-density (MD) or high-density (HD) SNP genotypes for genomic prediction. The objective function facilitates maximization of non-gap map length and system information for the SNP chip, and the latter is computed either as locus-averaged (LASE) or haplotype-averaged Shannon entropy (HASE) and adjusted for uniformity of the SNP distribution. …
Probabilistic Methods In Information Theory,
2016
Cal State University-San Bernardino
Probabilistic Methods In Information Theory, Erik W. Pachas
Electronic Theses, Projects, and Dissertations
Given a probability space, we analyze the uncertainty, that is, the amount of information of a finite system, by studying the entropy of the system. We also extend the concept of entropy to a dynamical system by introducing a measure preserving transformation on a probability space. After showing some theorems and applications of entropy theory, we study the concept of ergodicity, which helps us to further analyze the information of the system.
Passive Visual Analytics Of Social Media Data For Detection Of Unusual Events,
2016
Purdue University
Passive Visual Analytics Of Social Media Data For Detection Of Unusual Events, Kush Rustagi, Junghoon Chae
The Summer Undergraduate Research Fellowship (SURF) Symposium
Now that social media sites have gained substantial traction, huge amounts of un-analyzed valuable data are being generated. Posts containing images and text have spatiotemporal data attached as well, having immense value for increasing situational awareness of local events, providing insights for investigations and understanding the extent of incidents, their severity, and consequences, as well as their time-evolving nature. However, the large volume of unstructured social media data hinders exploration and examination. To analyze such social media data, the S.M.A.R.T system provides the analyst with an interactive visual spatiotemporal analysis and spatial decision support environment that assists in evacuation planning …
Well I'Ll Be Damned - Insights Into Predictive Value Of Pedigree Information In Horse Racing,
2016
University of Southampton
Well I'Ll Be Damned - Insights Into Predictive Value Of Pedigree Information In Horse Racing, Timothy Baker Mr, Ming-Chien Sung, Johnnie Johnson Professor, Tiejun Ma
International Conference on Gambling & Risk Taking
Fundamental form characteristics like how fast a horse ran at its last start, are widely used to help predict the outcome of horse racing events. The exception being in races where horses haven’t previously competed, such as Maiden races, where there is little or no publicly available past performance information. In these types of events bettors need only consider a simplified suite of factors however this is offset by a higher level of uncertainty. This paper examines the inherent information content embedded within a horse’s ancestry and the extent to which this information is discounted in the United Kingdom bookmaker …
