Open Access. Powered by Scholars. Published by Universities.®

Applied Statistics Commons™

Open Access. Powered by Scholars. Published by Universities.®

University of Nebraska - Lincoln

Discipline
Keyword
Publication Year
Publication
Publication Type

Articles 1 - 30 of 43

Full-Text Articles in Applied Statistics

Changepoint Detection As Model Selection: A General Framework, Michael A. Grantham Dec 2025

Changepoint Detection As Model Selection: A General Framework, Michael A. Grantham

Dissertations and Doctoral Documents, University of Nebraska-Lincoln, 2023–

This dissertation presents a general framework for changepoint detection based on ℓ0 model selection. The core method, Iteratively Reweighted Fused Lasso (IRFL), improves upon the generalized lasso by adaptively reweighting penalties to enhance support recovery and minimize criteria such as the Bayesian Information Criterion (BIC). The approach allows for flexible modeling of seasonal patterns, linear and quadratic trends, and autoregressive dependence in the presence of changepoints.

Simulation studies demonstrate that IRFL achieves accurate changepoint detection across a wide range of challenging scenarios, including those involving nuisance factors such as trends, seasonal patterns, and serially correlated errors. The framework is …


On Bayesian Empirical Likelihood-Based Method For Complex Survey Data With Application To Non-Probability Sampling, Md Hasibur Rahman Dec 2025

On Bayesian Empirical Likelihood-Based Method For Complex Survey Data With Application To Non-Probability Sampling, Md Hasibur Rahman

Department of Statistics: Dissertations, Theses, and Student Research

This thesis develops a Bayesian empirical likelihood (BEL) framework for inference under complex survey designs and extends it to non-probability sampling. Parametric likelihood based methods are difficult to apply to complex survey data because the likelihood is rarely available in closed form. EL provides a flexible alternative by replacing the parametric likelihood with an empirical likelihood constructed from moment conditions. The proposed method first integrates empirical likelihood constraints with survey design features then extends BEL to non-probability sampling through selection models and design consistent restrictions. Posterior inference is carried out using a Metropolis–Hastings MCMC algorithm. A real-data analysis further illustrates …


Detection Of Activity Cliffs Produced By Anti-Cancer Drugs And An Algorithm For Reliable Predictions In Affected Areas, Sarah Josephine Aurit Aug 2025

Detection Of Activity Cliffs Produced By Anti-Cancer Drugs And An Algorithm For Reliable Predictions In Affected Areas, Sarah Josephine Aurit

Department of Statistics: Dissertations, Theses, and Student Research

An activity cliff (AC) occurs when drugs close in chemical space produce dissimilar biological results. We focus on developing an inferential procedure to detect the presence of ACs in a chemical landscape. If detected, we provide a distance-based procedure that can be used to identify regions of stability in the chemical landscape of interest and generate prediction with higher precision in those areas of stability. We conceptualize the chemical landscape as a spatial random field and use spatial models for prediction of efficacy for new drugs based on “distance” in chemical space. We argue that an AC manifests itself by …


Online Prediction Of Streaming Data, Aleena Chanda Aug 2025

Online Prediction Of Streaming Data, Aleena Chanda

Dissertations and Doctoral Documents, University of Nebraska-Lincoln, 2023–

We present two new approaches for point prediction with streaming data based on a) the Count-Min sketch and b) Gaussian Process Priors with random bias. The methods are intended for the most general case where no true model can be usefully formulated for the data stream. In statistical contexts, this is often called the M open problem class. For the Count Min Sketch method we show that the predicted distribution function ^F converges to F under the assumption that the data consists of i.i.d samples from a fixed distribution function F. To implement the Gaussian Process Prior methods, we used …


Multivariate Mixture Regression Models With Known Group Membership And Informative Priors, Pahalapathirage Dona Kalani Hasanthika Aug 2025

Multivariate Mixture Regression Models With Known Group Membership And Informative Priors, Pahalapathirage Dona Kalani Hasanthika

Dissertations and Doctoral Documents, University of Nebraska-Lincoln, 2023–

We introduced couple different novel approaches to incorporate latent variable information to multivariate mixture regression models with both Gaussian and count data. We also evaluated the performance of these models with existing best approaches with simulated data from various sampling structures and also evaluated one of the models performance with rice metabolite data that provided some novel insights as well as validating existing literature about performance and behavior of these metabolites. We validated the method using extensive simulations and a real-world application. In both quantitative covariate designs and complex treatment design simulations, our method consistently outperformed established tools like limma, …


Leveraging Historical Data For Estimating Genetic Gain And Implementing Genomic Selection In A Student Led Barley Breeding Program, Sydney Graham May 2025

Leveraging Historical Data For Estimating Genetic Gain And Implementing Genomic Selection In A Student Led Barley Breeding Program, Sydney Graham

Department of Statistics: Dissertations, Theses, and Student Research

In Nebraska, winter feed barley presents an emerging market for producers and an opportunity to diversify cropping systems. The University of Nebraska Barley Breeding Program aims to develop high-yielding, winter-hardy varieties. A unique aspect of this program is that doctoral students serve as barley breeders and are responsible for crossing, data collection, and advancement decisions. While this provides hands-on experience for the students, the impact of student leadership has not been examined.

This study used a historical data set to evaluate the realized genetic gain of the breeding program, and as a training population for genomic selection. The dataset consisted …


Morphometric Analysis And Taxonomic Re-Evaluation Of Pepsis Cerberus Lucas And P. Elegans Lepeletier (Hymenoptera: Pompilidae: Pepsinae: Pepsini), Frank E. Kurczewski, Akira Shimizu, Diane H. Kiernan May 2024

Morphometric Analysis And Taxonomic Re-Evaluation Of Pepsis Cerberus Lucas And P. Elegans Lepeletier (Hymenoptera: Pompilidae: Pepsinae: Pepsini), Frank E. Kurczewski, Akira Shimizu, Diane H. Kiernan

Insecta Mundi

Hurd (1952) separated Pepsis cerberus Lucas from P. elegans Lepeletier (Hymenoptera: Pompilidae: Pepsinae: Pepsini) based on external morphology and biogeography. Vardy (2005) synonymized the familiar and historically well-documented P. cerberus and P. elegans, combining these Nearctic taxa with several Neotropical variants in an extremely broad definition of P. menechma Lepeletier. In doing so, Vardy (2005) breached the principle of nomenclatural stability. He ignored the prevailing usage and clearly violated articles 23.2, 23.3 and 23.9.1.2 of the ICZN (1999). Morphological differences, ecological divergence, and narrow sympatric geographic distribution of P. cerberus and P. elegans …


Statistical And Machine Learning Approaches To Describe Factors Affecting Preweaning Mortality Of Piglets, Md Towfiqur Rahman, Tami M. Brown-Brandl, Gary A. Rohrer, Sudhendu R. Sharma, Vamsi Manthena, Yeyin Shi Oct 2023

Statistical And Machine Learning Approaches To Describe Factors Affecting Preweaning Mortality Of Piglets, Md Towfiqur Rahman, Tami M. Brown-Brandl, Gary A. Rohrer, Sudhendu R. Sharma, Vamsi Manthena, Yeyin Shi

Department of Agricultural and Biological Systems Engineering: Faculty Publications

High preweaning mortality (PWM) rates for piglets are a significant concern for the worldwide pork industries, causing economic loss and well-being issues. This study focused on identifying the factors affecting PWM, overlays, and predicting PWM using historical production data with statistical and machine learning models. Data were collected from 1,982 litters from the United States Meat Animal Research Center, Nebraska, over the years 2016 to 2021. Sows were housed in a farrowing building with three rooms, each with 20 farrowing crates, and taken care of by well-trained animal caretakers. A generalized linear model was used to analyze the various sow, …


A Classical Fall Statistics Problem, Timothy Meyer Oct 2023

A Classical Fall Statistics Problem, Timothy Meyer

Cornhusker Economics

An evaluation of traditional baseball measures and suggestions for alternatives, centering on statistics related to the offensive quality of a player.


Exploring Experimental Design And Multivariate Analysis Techniques For Evaluating Community Structure Of Bacteria In Microbiome Data, Kelsey Karnik Aug 2023

Exploring Experimental Design And Multivariate Analysis Techniques For Evaluating Community Structure Of Bacteria In Microbiome Data, Kelsey Karnik

Department of Statistics: Dissertations, Theses, and Student Research

The gut microbiome plays a crucial role in human health, and by working collaboratively with microbiologists, we aim to further our understanding of the human gut and its impact on human health. Promoting a diverse microbiome is emphasized throughout microbiology literature, and involving a statistician in designing experiments to relate gut bacteria and some measured health outcome is crucial for ensuring valid and accurate results. By adopting new experimental design and analysis methods, researchers can begin to gain a deeper understanding of how the genetics of our food affect the composition of taxa within the gut microbiome. This dissertation is …


Examining The Effect Of Word Embeddings And Preprocessing Methods On Fake News Detection, Jessica Hauschild May 2023

Examining The Effect Of Word Embeddings And Preprocessing Methods On Fake News Detection, Jessica Hauschild

Department of Statistics: Dissertations, Theses, and Student Research

The words people choose to use hold a lot of power, whether that be in spreading truth or deception. As listeners and readers, we do our best to understand how words are being used. There are many current methods in computer science literature attempting to embed words into numerical information for statistical analyses. Some of these embedding methods, such as Bag of Words, treat words as independent, while others, such as Word2Vec, attempt to gain information about the context of words. It is of interest to compare how well these various methods of translating text into numerical data work specifically …


Risk-Based Machine Learning Approaches For Probabilistic Transient Stability, Umair Shahzad Dec 2021

Risk-Based Machine Learning Approaches For Probabilistic Transient Stability, Umair Shahzad

Department of Electrical and Computer Engineering: Dissertations, Theses, and Student Research

Power systems are getting more complex than ever and are consequently operating close to their limit of stability. Moreover, with the increasing demand of renewable wind generation, and the requirement to maintain a secure power system, the importance of transient stability cannot be overestimated. Considering its significance in power system security, it is important to propose a different approach for enhancing the transient stability, considering uncertainties. Current deterministic industry practices of transient stability assessment ignore the probabilistic nature of variables (fault type, fault location, fault clearing time, etc.). These approaches typically provide a conservative criterion and can result in expensive …


Interval Estimation Of Proportion Of Second-Level Variance In Multi-Level Modeling, Steven Svoboda Oct 2020

Interval Estimation Of Proportion Of Second-Level Variance In Multi-Level Modeling, Steven Svoboda

The Nebraska Educator: A Student-Led Journal

Physical, behavioral and psychological research questions often relate to hierarchical data systems. Examples of hierarchical data systems include repeated measures of students nested within classrooms, nested within schools and employees nested within supervisors, nested within organizations. Applied researchers studying hierarchical data structures should have an estimate of the intraclass correlation coefficient (ICC) for every nested level in their analyses because ignoring even relatively small amounts of interdependence is known to inflate Type I error rate in single-level models. Traditionally, researchers rely upon the ICC as a point estimate of the amount of interdependency in their data. Recent methods utilizing an …


Statistical Methodology To Establish A Benchmark For Evaluating Antimicrobial Resistance Genes Through Real Time Pcr Assay, Enakshy Dutta Jul 2020

Statistical Methodology To Establish A Benchmark For Evaluating Antimicrobial Resistance Genes Through Real Time Pcr Assay, Enakshy Dutta

Department of Statistics: Dissertations, Theses, and Student Research

Novel diagnostic tests are usually compared with gold standard tests for evaluating diagnostic accuracy. For assessing antimicrobial resistance (AMR) to bovine respiratory disease (BRD) pathogens, phenotypic broth microdilution method is used as gold standard (GS). The objective of the thesis is to evaluate the optimal cycle threshold (Ct) generated by real-time polymerase chain reaction (rtPCR) to genes that confer resistance that will translate to the phenotypic classification of AMR. Data from two different methodologies are assessed to identify Ct that will discriminate between resistance (R) and susceptibility (S). First, the receiver operating characteristic (ROC) curve was used to determine the …


Using Stability To Select A Shrinkage Method, Dean Dustin May 2020

Using Stability To Select A Shrinkage Method, Dean Dustin

Department of Statistics: Dissertations, Theses, and Student Research

Shrinkage methods are estimation techniques based on optimizing expressions to find which variables to include in an analysis, typically a linear regression. The general form of these expressions is the sum of an empirical risk plus a complexity penalty based on the number of parameters. Many shrinkage methods are known to satisfy an ‘oracle’ property meaning that asymptotically they select the correct variables and estimate their coefficients efficiently. In Section 1.2, we show oracle properties in two general settings. The first uses a log likelihood in place of the empirical risk and allows a general class of penalties. The second …


Group Testing Identification: Objective Functions, Implementation, And Multiplex Assays, Brianna D. Hitt Apr 2020

Group Testing Identification: Objective Functions, Implementation, And Multiplex Assays, Brianna D. Hitt

Department of Statistics: Dissertations, Theses, and Student Research

Group testing is the process of combining items into groups to test for a binary characteristic. One of its most widely used applications is infectious disease testing. In this context, specimens (e.g., blood, urine) are amalgamated into groups and tested. For groups that test positive, there are many algorithmic retesting procedures available to identify positive individuals. The appeal of group testing is that the overall number of tests needed is significantly less than for individual testing when disease prevalence is small and an appropriate algorithm is chosen. Group testing has a number of applications beyond infectious disease testing, such as …


The Role Of Topography, Soil, And Remotely Sensed Vegetation Condition Towards Predicting Crop Yield, Trenton E. Franz, Sayli Pokal, Justin P. Gibson, Yuzhen Zhou, Hamed Gholizadeh, Fatima Amor Tenorio, Daran Rudnick, Derek M. Heeren, Matthew F. Mccabe, Matteo Ziliani, Zhenong Jin, Kaiyu Guan, Ming Pan, John Gates, Brian Wardlow Jan 2020

The Role Of Topography, Soil, And Remotely Sensed Vegetation Condition Towards Predicting Crop Yield, Trenton E. Franz, Sayli Pokal, Justin P. Gibson, Yuzhen Zhou, Hamed Gholizadeh, Fatima Amor Tenorio, Daran Rudnick, Derek M. Heeren, Matthew F. Mccabe, Matteo Ziliani, Zhenong Jin, Kaiyu Guan, Ming Pan, John Gates, Brian Wardlow

School of Natural Resources: Faculty Publications

Foreknowledge of the spatiotemporal drivers of crop yield would provide a valuable source of information to optimize on-farm inputs and maximize profitability. In recent years, an abundance of spatial data providing information on soils, topography, and vegetation condition have become available from both proximal and remote sensing platforms. Given the wide range of data costs (between USD $0−50/ha), it is important to understand where often limited financial resources should be directed to optimize field production. Two key questions arise. First, will these data actually aid in better fine-resolution yield prediction to help optimize crop management and farm economics? Second, what …


Optimal Design For A Causal Structure, Zaher Kmail Aug 2019

Optimal Design For A Causal Structure, Zaher Kmail

Department of Statistics: Dissertations, Theses, and Student Research

Linear models and mixed models are important statistical tools. But in many natural phenomena, there is more than one endogenous variable involved and these variables are related in a sophisticated way. Structural Equation Modeling (SEM) is often used to model the complex relationships between the endogenous and exogenous variables. It was first implemented in research to estimate the strength and direction of direct and indirect effects among variables and to measure the relative magnitude of each causal factor.

Historically, traditional optimal design theory focuses on univariate linear, nonlinear, and mixed models. There is no current literature on the subject of …


Statistical Investigation Of Road And Railway Hazardous Materials Transportation Safety, Amirfarrokh Iranitalab Nov 2018

Statistical Investigation Of Road And Railway Hazardous Materials Transportation Safety, Amirfarrokh Iranitalab

Department of Civil and Environmental Engineering: Dissertations, Theses, and Student Research

Transportation of hazardous materials (hazmat) in the United States (U.S.) constituted 22.8% of the total tonnage transported in 2012 with an estimated value of more than 2.3 billion dollars. As such, hazmat transportation is a significant economic activity in the U.S. However, hazmat transportation exposes people and environment to the infrequent but potentially severe consequences of incidents resulting in hazmat release. Trucks and trains carried 63.7% of the hazmat in the U.S. in 2012 and are the major foci of this dissertation. The main research objectives were 1) identification and quantification of the effects of different factors on occurrence and …


Quantitative Appraisal Of Non-Irrigated Cropland In South Dakota, Shelby Riggs Oct 2018

Quantitative Appraisal Of Non-Irrigated Cropland In South Dakota, Shelby Riggs

Honors Program: Senior Projects (Public)

This appraisal attempts to remove subjectivity from the appraisal process and replace it with quantitative analysis of known data to generate a fair market value of the subject property. Two methods of appraisal were used, the income approach and the comparable sales approach. For the income approach, I used the average cash rent for the region, the current property taxes for the subject property, and a capitalization rate based on Stokes' (2018) capitalization rate formula to arrive at my income-based valuation. For the comparable sales approach, I utilized Stokes' (2018) research in optimization modeling to estimate a market value for …


Essentials Of Structural Equation Modeling, Mustafa Emre Civelek Mar 2018

Essentials Of Structural Equation Modeling, Mustafa Emre Civelek

Zea E-Books Collection

Structural Equation Modeling is a statistical method increasingly used in scientific studies in the fields of Social Sciences. It is currently a preferred analysis method, especially in doctoral dissertations and academic researches. However, since many universities do not include this method in the curriculum of undergraduate and graduate courses, students and scholars try to solve the problems they encounter by using various books and internet resources.

This book aims to guide the researcher who wants to use this method in a way that is free from math expressions. It teaches the steps of a research program using structured equality modeling …


The Impact Of Truncating Data On The Predictive Ability For Single-Step Genomic Best Linear Unbiased Prediction, Jeremy T. Howard, Thomas A. Rathje, Caitlyn E. Bruns, Danielle F. Wilson-Wells, Stephen D. Kachman, Matthew L. Spangler Jan 2018

The Impact Of Truncating Data On The Predictive Ability For Single-Step Genomic Best Linear Unbiased Prediction, Jeremy T. Howard, Thomas A. Rathje, Caitlyn E. Bruns, Danielle F. Wilson-Wells, Stephen D. Kachman, Matthew L. Spangler

Department of Animal Science: Faculty Publications

Simulated and swine industry data sets were utilized to assess the impact of removing older data on the predictive ability of selection candidate estimated breeding values (EBV) when using single-step genomic best linear unbiased prediction (ssGBLUP). Simulated data included thirty replicates designed to mimic the structure of swine data sets. For the simulated data, varying amounts of data were truncated based on the number of ancestral generations back from the selection candidates. The swine data sets consisted of phenotypic and genotypic records for three traits across two breeds on animals born from 2003 to 2017. Phenotypes and genotypes were iteratively …


Methods To Account For Breed Composition In A Bayesian Gwas Method Which Utilizes Haplotype Clusters, Danielle F. Wilson-Wells Aug 2016

Methods To Account For Breed Composition In A Bayesian Gwas Method Which Utilizes Haplotype Clusters, Danielle F. Wilson-Wells

Department of Statistics: Dissertations, Theses, and Student Research

In livestock, prediction of an animal’s genetic merit using genomic information is becoming increasingly common. The models used to make these predictions typically assume that we are sampling from a homogeneous population. However, in both commercial and experimental populations the sire and dam of an individual may be a mixture of different breeds. Haplotype models can capture this population structure.

Two models based on breed specific haplotype clusters where developed to account for differences across multiple breeds. The first model utilizes the breed composition of the individual, while the second utilizes the breed composition from the sire and dam. Haplotype …


Beta-Binomial Kriging: A New Approach To Modeling Spatially Correlated Proportions, Aimee Schwab Aug 2015

Beta-Binomial Kriging: A New Approach To Modeling Spatially Correlated Proportions, Aimee Schwab

Department of Statistics: Dissertations, Theses, and Student Research

Spatially correlated count data sets appear often in applied data analysis problems, but there is little consensus in the literature about how best to analyze the data. The two prevailing approaches provide accurate parameter estimates and predictions, at the cost of model interpretability and simplicity. This dissertation will present a new approach to modeling spatially correlated binomial observations: beta-binomial kriging. The model proposed here is a modified form of spatial kriging which assumes the data are generated from a correlated beta-binomial distribution. Given this assumption, the spatial parameters and predicted values can be estimated using simple matrix algebra. Beta-binomial kriging …


A Comparison Of Population-Averaged And Cluster-Specific Approaches In The Context Of Unequal Probabilities Of Selection, Natalie A. Koziol May 2015

A Comparison Of Population-Averaged And Cluster-Specific Approaches In The Context Of Unequal Probabilities Of Selection, Natalie A. Koziol

College of Education and Human Sciences: Dissertations, Theses, and Student Research

Sampling designs of large-scale, federally funded studies are typically complex, involving multiple design features (e.g., clustering, unequal probabilities of selection). Researchers must account for these features in order to obtain unbiased point estimators and make valid inferences about population parameters. Single-level (i.e., population-averaged) and multilevel (i.e., cluster-specific) methods provide two alternatives for modeling clustered data. Single-level methods rely on the use of adjusted variance estimators to account for dependency due to clustering, whereas multilevel methods incorporate the dependency into the specification of the model.

Although the literature comparing single-level and multilevel approaches is vast, comparisons have been limited to the …


A New Approach To Modeling Multivariate Time Series On Multiple Temporal Scales, Tucker Zeleny May 2015

A New Approach To Modeling Multivariate Time Series On Multiple Temporal Scales, Tucker Zeleny

Department of Statistics: Dissertations, Theses, and Student Research

In certain situations, observations are collected on a multivariate time series at a certain temporal scale. However, there may also exist underlying time series behavior on a larger temporal scale that is of interest. Often times, identifying the behavior of the data over the course of the larger scale is the key objective. Because this large scale trend is not being directly observed, describing the trends of the data on this scale can be more difficult. To further complicate matters, the observed data on the smaller time scale may be unevenly spaced from one larger scale time point to the …


Best Practice Recommendations For Data Screening, Justin A. Desimone, Peter D. Harms, Alice J. Desimone Feb 2015

Best Practice Recommendations For Data Screening, Justin A. Desimone, Peter D. Harms, Alice J. Desimone

Department of Management: Faculty Publications

Survey respondents differ in their levels of attention and effort when responding to items. There are a number of methods researchers may use to identify respondents who fail to exert sufficient effort in order to increase the rigor of analysis and enhance the trustworthiness of study results. Screening techniques are organized into three general categories, which differ in impact on survey design and potential respondent awareness. Assumptions and considerations regarding appropriate use of screening techniques are discussed along with descriptions of each technique. The utility of each screening technique is a function of survey design and administration. Each technique has …


Equate: Observed-Score Linking And Equating In R, Anthony D. Albano Jan 2015

Equate: Observed-Score Linking And Equating In R, Anthony D. Albano

Department of Educational Psychology: Faculty Publications

Linking and equating are statistical procedures used to convert scores from one measurement scale to another. These procedures are most often used in testing programs that involve multiple test forms, where adjustments are made for form difficulty differences when creating a measurement scale that is common across forms. Linking and equating methods are traditionally distinguished by the type of scores they are applied to, whether observed scores or scores from an item response theory model. Methods are also distinguished by the study design under which measurements are taken. The R package equate (Albano, 2014) is free, open-source software for conducting …


New Statistical Methods For Analysis Of Historical Data From Wildlife Populations, Trevor Hefley Mar 2014

New Statistical Methods For Analysis Of Historical Data From Wildlife Populations, Trevor Hefley

Department of Statistics: Dissertations, Theses, and Student Research

Wildlife biologists, many times with the help of ordinary citizens, have developed and maintained long-term datasets for monitoring the status of wildlife populations. These datasets can range from a collection of citizen-reported sightings of a rare species, to datasets collected by biologists using standardized methods. The commonality is that these datasets span a temporal and spatial scale that is beyond the scope of most scientific studies. Ensuring the continued persistence of wildlife populations requires predictions of the impact of human actions. Regardless if the predictions are quantitative or qualitative, the best we can do is use the past data to …


A Test For Detecting Changes In Closed Networks Based On The Number Of Communications Between Nodes, Christopher S. Wichman Jul 2013

A Test For Detecting Changes In Closed Networks Based On The Number Of Communications Between Nodes, Christopher S. Wichman

Department of Statistics: Dissertations, Theses, and Student Research

This dissertation presents a formal method for detecting changes in a closed communications network based on an “abnormal” shift in the number of communications between some of the nodes. The method relies on the analyst’s ability to define the network of interest; capture the number of communications between nodes; and to establish a history of normal communications flow between nodes over fixed intervals of time. A metric multi-dimensional scaling technique is then used to represent the network at each time interval with a k-dimensional (k = 1, 2, …) configuration. The affine bi-dimensional regression coefficient of determination (aR2) …