Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Biostatistics (61)
- Applied Statistics (58)
- Statistical Models (48)
- Statistical Methodology (41)
- Social and Behavioral Sciences (39)
-
- Mathematics (31)
- Other Statistics and Probability (30)
- Life Sciences (26)
- Statistical Theory (26)
- Education (25)
- Multivariate Analysis (23)
- Applied Mathematics (20)
- Data Science (18)
- Computer Sciences (17)
- Longitudinal Data Analysis and Time Series (17)
- Sociology (16)
- Bioinformatics (15)
- Medicine and Health Sciences (15)
- Educational Assessment, Evaluation, and Research (12)
- Probability (12)
- Categorical Data Analysis (11)
- Quantitative, Qualitative, Comparative, and Historical Methodologies (11)
- Business (10)
- Psychology (9)
- Survival Analysis (9)
- Engineering (8)
- Geography (8)
- Artificial Intelligence and Robotics (7)
- Institution
- Keyword
-
- Morgridge College of Education (46)
- Research Methods and Information Science (43)
- Research Methods and Statistics (42)
- Statistics (14)
- Causal inference (8)
-
- Machine learning (8)
- Bayesian (6)
- Simulation (6)
- Propensity score (5)
- Deep learning (4)
- Education (4)
- Machine Learning (4)
- Missing data (4)
- Mixed data (4)
- Regression (4)
- Variable selection (4)
- Average treatment effect (3)
- Bioinformatics (3)
- Bootstrap (3)
- Breast cancer (3)
- Confirmatory factor analysis (3)
- Depression (3)
- Diabetes (3)
- Feature Selection (3)
- Grounded theory (3)
- Mathematics (3)
- Meta-analysis (3)
- Metabolomics (3)
- Mixed methods (3)
- Negative binomial distribution (3)
Articles 151 - 180 of 261
Full-Text Articles in Statistics and Probability
Evaluation Of Using The Bootstrap Procedure To Estimate The Population Variance, Nghia Trong Nguyen
Evaluation Of Using The Bootstrap Procedure To Estimate The Population Variance, Nghia Trong Nguyen
Electronic Theses and Dissertations
The bootstrap procedure is widely used in nonparametric statistics to generate an empirical sampling distribution from a given sample data set for a statistic of interest. Generally, the results are good for location parameters such as population mean, median, and even for estimating a population correlation. However, the results for a population variance, which is a spread parameter, are not as good due to the resampling nature of the bootstrap method. Bootstrap samples are constructed using sampling with replacement; consequently, groups of observations with zero variance manifest in these samples. As a result, a bootstrap variance estimator will carry a …
Longitudinal Tracking Of Physiological State With Electromyographic Signals., Robert Warren Stallard
Longitudinal Tracking Of Physiological State With Electromyographic Signals., Robert Warren Stallard
Electronic Theses and Dissertations
Electrophysiological measurements have been used in recent history to classify instantaneous physiological configurations, e.g., hand gestures. This work investigates the feasibility of working with changes in physiological configurations over time (i.e., longitudinally) using a variety of algorithms from the machine learning domain. We demonstrate a high degree of classification accuracy for a binary classification problem derived from electromyography measurements before and after a 35-day bedrest. The problem difficulty is increased with a more dynamic experiment testing for changes in astronaut sensorimotor performance by taking electromyography and force plate measurements before, during, and after a jump from a small platform. A …
Multi Self-Adapting Particle Swarm Optimization Algorithm (Msapso)., Gerhard Koch
Multi Self-Adapting Particle Swarm Optimization Algorithm (Msapso)., Gerhard Koch
Electronic Theses and Dissertations
The performance and stability of the Particle Swarm Optimization algorithm depends on parameters that are typically tuned manually or adapted based on knowledge from empirical parameter studies. Such parameter selection is ineffectual when faced with a broad range of problem types, which often hinders the adoption of PSO to real world problems. This dissertation develops a dynamic self-optimization approach for the respective parameters (inertia weight, social and cognition). The effects of self-adaption for the optimal balance between superior performance (convergence) and the robustness (divergence) of the algorithm with regard to both simple and complex benchmark functions is investigated. This work …
Geostatistical Analysis Of Potential Sinkhole Risk: Examining Spatial And Temporal Climate Relationships In Tennessee And Florida, Kimberly Blazzard
Geostatistical Analysis Of Potential Sinkhole Risk: Examining Spatial And Temporal Climate Relationships In Tennessee And Florida, Kimberly Blazzard
Electronic Theses and Dissertations
Sinkholes are a significant hazard for the southeastern United States. Although differences in climate are known to affect karst environments differently, quantitative analyses correlating sinkhole formation with climate variables is lacking. A temporal linear regression for Florida sinkholes and two modeled regressions for Tennessee sinkholes were produced: a general linearized logistic regression and a MaxEnt derived species distribution model. Temporal results showed highly significant correlations with precipitation, teleconnection patterns, temperature, and CO2, while spatial results showed highly significant correlations with precipitation, wind speed, solar radiation, and maximum temperature. Regression results indicated that some sinkhole formation variability could be …
Improving The Detection Limit Of Tau Aggregates For Use With Biological Samples, Emily Rickman Hager
Improving The Detection Limit Of Tau Aggregates For Use With Biological Samples, Emily Rickman Hager
Electronic Theses and Dissertations
The protein Tau is found in neurofibrillary tangles in Alzheimer's disease and over 20 other neurodegenerative diseases. An assay has been developed to detect minute amounts of fibrils from human brain tissue. This assay subjects brain tissue extract and recombinant Tau to several rounds of sonication and incubation. Incubation allows recombinant Tau to add itself to the ends of the existing fibrils in brain tissue extract. Sonication breaks the existing fibrils in the brain tissue extract offering more ends for Tau to add onto. Cycles of sonication and incubation have been shown to allow for amplification of Tau fibrils from …
The Impact Of Data Sovereignty On American Indian Self-Determination: A Framework Proof Of Concept Using Data Science, Joseph Carver Robertson
The Impact Of Data Sovereignty On American Indian Self-Determination: A Framework Proof Of Concept Using Data Science, Joseph Carver Robertson
Electronic Theses and Dissertations
The Data Sovereignty Initiative is a collection of ideas that was designed to create SMART solutions for tribal communities. This concept was to develop a horizontal governance framework to create a strategic act of sovereignty using data science. The core concept of this idea was to present data sovereignty as a way for tribal communities to take ownership of data in order to affect policy and strategic decisions that are data driven in nature. The case studies in this manuscript were developed around statistical theories of spatial statistics, exploratory data analysis, and machine learning. And although these case studies are …
Location Optimization Of A Coal Power Plant To Balance Coal Supply And Electric Transmission Costs Against Plant’S Emission Exposure, Najam Khan
Electronic Theses and Dissertations
This research is focused on developing a location analysis methodology that can minimize the pollutant exposure to the public while ensuring that the combined costs of electric transmission losses and coal logistics are minimized. Coal power plants will provide a critical contribution towards meeting electricity demands for various nations in the foreseeable future. The site selection for a new coal power plant is extremely important from an investment point of view. The operational costs for running a coal power plant can be minimized by a combined emphasis on placing a coal power plant near coal mines as well as customers. …
Statistical Algorithms And Bioinformatics Tools Development For Computational Analysis Of High-Throughput Transcriptomic Data, Adam Mcdermaid
Statistical Algorithms And Bioinformatics Tools Development For Computational Analysis Of High-Throughput Transcriptomic Data, Adam Mcdermaid
Electronic Theses and Dissertations
Next-Generation Sequencing technologies allow for a substantial increase in the amount of data available for various biological studies. In order to effectively and efficiently analyze this data, computational approaches combining mathematics, statistics, computer science, and biology are implemented. Even with the substantial efforts devoted to development of these approaches, numerous issues and pitfalls remain. One of these issues is mapping uncertainty, in which read alignment results are biased due to the inherent difficulties associated with accurately aligning RNA-Sequencing reads. GeneQC is an alignment quality control tool that provides insight into the severity of mapping uncertainty in each annotated gene from …
Variable Selection Techniques For Clustering On The Unit Hypersphere, Damon Bayer
Variable Selection Techniques For Clustering On The Unit Hypersphere, Damon Bayer
Electronic Theses and Dissertations
Mixtures of von Mises-Fisher distributions have been shown to be an effective model for clustering data on a unit hypersphere, but variable selection for these models remains an important and challenging problem. In this paper, we derive two variants of the expectation-maximization framework, which are each used to identify a specific type of irrelevant variables for these models. The first type are noise variables, which are not useful for separating any pairs of clusters. The second type are redundant variables, which may be useful for separating pairs of clusters, but do not enable any additional separation beyond the separability provided …
Development Of Biclustering Techniques For Gene Expression Data Modeling And Mining, Juan Xie
Development Of Biclustering Techniques For Gene Expression Data Modeling And Mining, Juan Xie
Electronic Theses and Dissertations
The next-generation sequencing technologies can generate large-scale biological data with higher resolution, better accuracy, and lower technical variation than the arraybased counterparts. RNA sequencing (RNA-Seq) can generate genome-scale gene expression data in biological samples at a given moment, facilitating a better understanding of cell functions at genetic and cellular levels. The abundance of gene expression datasets provides an opportunity to identify genes with similar expression patterns across multiple conditions, i.e., co-expression gene modules (CEMs). Genomescale identification of CEMs can be modeled and solved by biclustering, a twodimensional data mining technique that allows clustering of rows and columns in a gene …
Functional Data Analysis Methods For Predicting Disease Status., Sarah Kendrick
Functional Data Analysis Methods For Predicting Disease Status., Sarah Kendrick
Electronic Theses and Dissertations
Introduction: Differential scanning calorimetry (DSC) is used to determine thermally-induced conformational changes of biomolecules within a blood plasma sample. Recent research has indicated that DSC curves (or thermograms) may have different characteristics based on disease status and, thus, may be useful as a monitoring and diagnostic tool for some diseases. Since thermograms are curves measured over a range of temperature values, they are often considered as functional data. In this dissertation we propose and apply functional data analysis (FDA) techniques to analyze DSC data from the Lupus Family Registry and Repository (LFRR). The aim is to develop FDA methods to …
Sample Size Calculations And Normalization Methods For Rna-Seq Data., Xiaohong Li
Sample Size Calculations And Normalization Methods For Rna-Seq Data., Xiaohong Li
Electronic Theses and Dissertations
High-throughput RNA sequencing (RNA-seq) has become the preferred choice for transcriptomics and gene expression studies. With the rapid growth of RNA-seq applications, sample size calculation methods for RNA-seq experiment design and data normalization methods for DEG analysis are important issues to be explored and discussed. The underlying theme of this dissertation is to develop novel sample size calculation methods in RNA-seq experiment design using test statistics. I have also proposed two novel normalization methods for analysis of RNA-seq data. In chapter one, I present the test statistical methods including Wald’s test, log-transformed Wald’s test and likelihood ratio test statistics for …
Examination And Comparison Of The Performance Of Common Non-Parametric And Robust Regression Models, Gregory F. Malek
Examination And Comparison Of The Performance Of Common Non-Parametric And Robust Regression Models, Gregory F. Malek
Electronic Theses and Dissertations
ABSTRACT
Examination and Comparison of the Performance of Common Non-Parametric and Robust Regression Models
By
Gregory Frank Malek
Stephen F. Austin State University, Masters in Statistics Program,
Nacogdoches, Texas, U.S.A.
This work investigated common alternatives to the least-squares regression method in the presence of non-normally distributed errors. An initial literature review identified a variety of alternative methods, including Theil Regression, Wilcoxon Regression, Iteratively Re-Weighted Least Squares, Bounded-Influence Regression, and Bootstrapping methods. These methods were evaluated using a simple simulated example data set, as well as various real data sets, including math proficiency data, Belgian telephone call data, and faculty …
Uses Of The Hypergeometric Distribution For Determining Survival Or Complete Representation Of Subpopulations In Sequential Sampling, Brooke Busbee
Uses Of The Hypergeometric Distribution For Determining Survival Or Complete Representation Of Subpopulations In Sequential Sampling, Brooke Busbee
Electronic Theses and Dissertations
This thesis will explore the hypergeometric probability distribution by looking at many different aspects of the distribution. These include, and are not limited to: history and origin, derivation and elementary applications, properties, relationships to other probability models, kindred hypergeometric distributions and elements of statistical inference associated with the hypergeometric distribution. Once the above are established, an investigation into and furthering of work done by Walton (1986) and Charlambides (2005) will be done. Here, we apply the hypergeometric distribution to sequential sampling in order to determine a surviving subcategory as well as study the problem of and complete representation of the …
Bayesian Approach On Short Time-Course Data Of Protein Phosphorylation, Casual Inference For Ordinal Outcome And Causal Analysis Of Dietary And Physical Activity In T2dm Using Nhanes Data., You Wu
Electronic Theses and Dissertations
This dissertation contains three different projects in proteomics and causal inferences. In the first project, I apply a Bayesian hierarchical model to assess the stability of phosphorylated proteins under short-time cold ischemia. This study provides inference on the stability of these phosphorylated proteins, which is valuable when using these proteins as biomarkers for a disease. in the second project, I perform a comparative study of different confounding-adjusted to estimate the treatment effect when the outcome variable is ordinal using observational data. The adjusted U-statistics method is compared with other methods such as ordinal logistic regression, propensity score based stratification and …
A Cross-Sectional Exploration Of Household Financial Reactions And Homebuyer Awareness Of Registered Sex Offenders In A Rural, Suburban, And Urban County., John Charles Navarro
A Cross-Sectional Exploration Of Household Financial Reactions And Homebuyer Awareness Of Registered Sex Offenders In A Rural, Suburban, And Urban County., John Charles Navarro
Electronic Theses and Dissertations
As stigmatized persons, registered sex offenders betoken instability in communities. Depressed home sale values are associated with the presence of registered sex offenders even though the public is largely unaware of the presence of registered sex offenders. Using a spatial multilevel approach, the current study examines the role registered sex offenders influence sale values of homes sold in 2015 for three U.S. counties (rural, suburban, and urban) located in Illinois and Kentucky within the social disorganization framework. Homebuyers were surveyed to examine whether awareness of local registered sex offenders and the homebuyer’s community type operate as moderators between home selling …
Likelihood-Based Methods For Analysis Of Copy Number Variation Using Next Generation Sequencing Data., Udika Iroshini Bandara
Likelihood-Based Methods For Analysis Of Copy Number Variation Using Next Generation Sequencing Data., Udika Iroshini Bandara
Electronic Theses and Dissertations
A Copy Number Variation (CNV) detection problem is considered using Circular Binary Segmentation (CBS) procedures, including newly developed procedures based on likelihood ratio tests with the parametric bootstrap for models based on discrete distributions for count data (Poisson and negative binomial) and a widely-used DNAcopy package. Results from the literature concerning maximum likelihood estimation for the negative binomial distribution are reviewed. The Newton-Raphson method is used to find the root of the derivative of the profile log likelihood function when applicable, and it is proven that this method converges to the true Maximum Likeihood Estimate (MLE), if the starting point …
Estimation Of The Three Key Parameters And The Lead Time Distribution In Lung Cancer Screening., Ruiqi Liu
Estimation Of The Three Key Parameters And The Lead Time Distribution In Lung Cancer Screening., Ruiqi Liu
Electronic Theses and Dissertations
This dissertation contains three research projects on cancer screening probability modeling. Cancer screening is the primary technique for early detection. The goal of screening is to catch the disease early before clinical symptoms appear. In these projects, the three key parameters and lead time distribution were estimated to provide a statistical point of view on the effectiveness of cancer screening programs. In the first project, cancer screening probability model was used to analyze the computed tomography (CT) scan group in the National Lung Screening Trial (NLST) data. Three key parameters were estimated using Bayesian approach and Markov Chain Monte Carlo …
Novel Statistical Approaches For Missing Values In Truncated High-Dimensional Metabolomics Data With A Detection Threshold., Jasmit Sureshkumar Shah
Novel Statistical Approaches For Missing Values In Truncated High-Dimensional Metabolomics Data With A Detection Threshold., Jasmit Sureshkumar Shah
Electronic Theses and Dissertations
Despite considerable advances in high throughput technology over the last decade, new challenges have emerged related to the analysis, interpretation, and integration of high-dimensional data. The arrival of omics datasets has contributed to the rapid improvement of systems biology, which seeks the understanding of complex biological systems. Metabolomics is an emerging omics field, where mass spectrometry technologies generate high dimensional datasets. As advances in this area are progressing, the need for better analysis methods to provide correct and adequate results are required. While in other omics sectors such as genomics or proteomics there has and continues to be critical understanding …
Spatiotemporal Analyses Of Recycled Water Production, Jana E. Archer
Spatiotemporal Analyses Of Recycled Water Production, Jana E. Archer
Electronic Theses and Dissertations
Increased demands on water supplies caused by population expansion, saltwater intrusion, and drought have led to water shortages which may be addressed by use of recycled water as recycled water products. Study I investigated recycled water production in Florida and California during 2009 to detect gaps in distribution and identify areas for expansion. Gaps were detected along the panhandle and Miami, Florida, as well as the northern and southwestern regions in California. Study II examined gaps in distribution, identified temporal change, and located areas for expansion for Florida in 2009 and 2015. Production increased in the northern and southern regions …
Peptide Identification: Refining A Bayesian Stochastic Model, Theophilus Barnabas Kobina Acquah
Peptide Identification: Refining A Bayesian Stochastic Model, Theophilus Barnabas Kobina Acquah
Electronic Theses and Dissertations
Notwithstanding the challenges associated with different methods of peptide identification, other methods have been explored over the years. The complexity, size and computational challenges of peptide-based data sets calls for more intrusion into this sphere. By relying on the prior information about the average relative abundances of bond cleavages and the prior probability of any specific amino acid sequence, we refine an already developed Bayesian approach in identifying peptides. The likelihood function is improved by adding additional ions to the model and its size is driven by two overall goodness of fit measures. In the face of the complexities associated …
Performance Of Imputation Algorithms On Artificially Produced Missing At Random Data, Tobias O. Oketch
Performance Of Imputation Algorithms On Artificially Produced Missing At Random Data, Tobias O. Oketch
Electronic Theses and Dissertations
Missing data is one of the challenges we are facing today in modeling valid statistical models. It reduces the representativeness of the data samples. Hence, population estimates, and model parameters estimated from such data are likely to be biased.
However, the missing data problem is an area under study, and alternative better statistical procedures have been presented to mitigate its shortcomings. In this paper, we review causes of missing data, and various methods of handling missing data. Our main focus is evaluating various multiple imputation (MI) methods from the multiple imputation of chained equation (MICE) package in the statistical software …
Denoising Tandem Mass Spectrometry Data, Felix Offei
Denoising Tandem Mass Spectrometry Data, Felix Offei
Electronic Theses and Dissertations
Protein identification using tandem mass spectrometry (MS/MS) has proven to be an effective way to identify proteins in a biological sample. An observed spectrum is constructed from the data produced by the tandem mass spectrometer. A protein can be identified if the observed spectrum aligns with the theoretical spectrum. However, data generated by the tandem mass spectrometer are affected by errors thus making protein identification challenging in the field of proteomics. Some of these errors include wrong calibration of the instrument, instrument distortion and noise. In this thesis, we present a pre-processing method, which focuses on the removal of noisy …
A Distribution Of The First Order Statistic When The Sample Size Is Random, Vincent Z. Forgo Mr
A Distribution Of The First Order Statistic When The Sample Size Is Random, Vincent Z. Forgo Mr
Electronic Theses and Dissertations
Statistical distributions also known as probability distributions are used to model a random experiment. Probability distributions consist of probability density functions (pdf) and cumulative density functions (cdf). Probability distributions are widely used in the area of engineering, actuarial science, computer science, biological science, physics, and other applicable areas of study. Statistics are used to draw conclusions about the population through probability models. Sample statistics such as the minimum, first quartile, median, third quartile, and maximum, referred to as the five-number summary, are examples of order statistics. The minimum and maximum observations are important in extreme value theory. This paper will …
Analyzing Electricity Use Of Low Income Weatherization Program Participants Using Propensity Score Analysis And A Hierarchical Linear Growth Model, Ksenia Polson
Electronic Theses and Dissertations
This evaluation utilized propensity score matching methods and a longitudinal hierarchical linear growth model to determine the effect of residential energy efficiency upgrade(s) on household electricity use for the low-income community over the course of a year in the City and County of Denver, Colorado. Propensity score analysis with risk set matching was performed at each month under analysis applying nearest neighbor and nearest neighbor with caliper approaches by balancing covariates across the treatment and control groups. Following the completion of propensity score analysis, the data were aggregated to form a data set that was used in a hierarchical linear …
An Evaluation Of Critical Realignment Theory: Comparing Bayesian And Frequentist Approaches, Tara A. Rhodes
An Evaluation Of Critical Realignment Theory: Comparing Bayesian And Frequentist Approaches, Tara A. Rhodes
Electronic Theses and Dissertations
Prior to this study, critical realignment theory, which presupposes eras of substantial and sustained swings in American political party dominance, had only been evaluated using the classical, frequentist approach to modeling. However, potential for more information concerning these electoral phenomena exists given a shift in the design and approach to realigning elections. This study sought to explore those options through one particular alternative to the classical approach to statistics--in this particular case, the Bayesian approach to statistics. Bayesian methods differ from the frequentist approach in three main ways: the treatment of probability, the treatment of parameters, and the treatment of …
Consumers' Perceptions Of Voluntary And Involuntary Deconsumption: An Exploratory Sequential Scale Development Study, Kranti K. Dugar
Consumers' Perceptions Of Voluntary And Involuntary Deconsumption: An Exploratory Sequential Scale Development Study, Kranti K. Dugar
Electronic Theses and Dissertations
This exploratory sequential mixed methods study of scale development was conducted among baby boomers in the United States to render conceptual clarity to the concepts of voluntary and involuntary deconsumption, to explore deconsumption behavior under the tenets of the attribution theory of motivation, and to examine the components, structures, uses, and measurement properties of scales of voluntary and involuntary deconsumption. It was also an attempt to reiterate the importance of the baby boomer segment(s) for marketing practitioners based on growth, economic viability, and the power of influence, and to establish a deep understanding of the deconsumption processes, which could enable …
A Meta-Analysis Of The Effects Of Incentives On Response Rate In Online Survey Studies, Amal Muhammad Asire
A Meta-Analysis Of The Effects Of Incentives On Response Rate In Online Survey Studies, Amal Muhammad Asire
Electronic Theses and Dissertations
Meta-analysis was used to investigate the effect of incentives on response rates of web-based survey studies. Whereas numerous meta-analyses that address the effect of incentives on increasing response rates in survey studies are available in the literature, these analyses are based on mail surveys, so there is a need for an applied meta-analysis to examine the effect of incentives on response rates in online survey studies. A meta-analysis of an online method of survey administration was used because the use of online surveys has greatly increased, making web-based survey administration an important form of data collection in multiple fields of …
Meta-Analyses Of The Relationship Between Depression And Nine Dimensions Of Perfectionism, Gabriel Lynn Hottinger
Meta-Analyses Of The Relationship Between Depression And Nine Dimensions Of Perfectionism, Gabriel Lynn Hottinger
Electronic Theses and Dissertations
Perfectionism has been shown to be related to depression, but perfectionism is multidimensional. Some dimensions are related to positive psychological characteristics and outcomes and other dimensions are related to negative psychological characteristics and outcomes. This study reports results of nine meta-analyses performed to investigate the association between each of nine subscales of perfectionism and depression to determine which dimensions of perfectionism are most strongly associated with depression. The two subscales that were used from the Hewitt and Flett (1991b) Multidimensional Perfectionism scale were Self-Oriented Perfectionism (SOP) and Socially-Prescribed Perfectionism (SPP). The five subscales that were used from the Frost et …
Identifying Predictors Of Weight Loss And Drop-Out Using Joint Modeling, Valerie Bares
Identifying Predictors Of Weight Loss And Drop-Out Using Joint Modeling, Valerie Bares
Electronic Theses and Dissertations
Profile by Sanford is a membership based weight loss program that helps its members make lifestyle changes with diet, exercise, and one-on-one interactions with a weight loss coach. Discovery of characteristics and behaviors influencing weight loss will benefit current and future members of Profile. This research utilizes massive data from Profile by Sanford to analyze member behavior. Fourteen data sets are evaluated, some containing millions of observations. All data is combined into one comprehensive table of 33,487 members. Members of Profile by Sanford are 77% female and two-thirds of all members start the program classified as obese. Attending meetings with …