Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Applied Statistics (2926)
- Social and Behavioral Sciences (2697)
- Medicine and Health Sciences (2516)
- Biostatistics (2513)
- Mathematics (2041)
-
- Statistical Theory (1633)
- Public Health (1564)
- Statistical Methodology (1564)
- Statistical Models (1310)
- Life Sciences (1286)
- Engineering (1182)
- Epidemiology (1097)
- Computer Sciences (1063)
- Applied Mathematics (989)
- Civil and Environmental Engineering (667)
- Materials Science and Engineering (598)
- Transportation Engineering (588)
- Other Civil and Environmental Engineering (579)
- Construction Engineering and Management (575)
- Structural Materials (575)
- Public Affairs, Public Policy and Public Administration (557)
- Medical Specialties (555)
- Data Science (551)
- Multivariate Analysis (525)
- Education (524)
- Health Services Research (502)
- Design of Experiments and Sample Surveys (501)
- Probability (493)
- Institution
-
- Wayne State University (1162)
- COBRA (1108)
- Changsha University of Science and Technology (570)
- Missouri University of Science and Technology (538)
- University of Kentucky (401)
-
- Marquette University (385)
- University of South Carolina (376)
- Universitas Indonesia (369)
- Loma Linda University (326)
- Utah State University (317)
- University of Nebraska - Lincoln (254)
- University of Nevada, Las Vegas (246)
- Old Dominion University (229)
- Wright State University (206)
- University of New Mexico (202)
- Air Force Institute of Technology (182)
- California Polytechnic State University, San Luis Obispo (180)
- University of South Florida (176)
- Virginia Commonwealth University (176)
- Prairie View A&M University (161)
- Himmelfarb Health Sciences Library, The George Washington University (159)
- Roseman University of Health Sciences (155)
- Brigham Young University (144)
- Georgia Southern University (140)
- University of Texas at El Paso (139)
- City University of New York (CUNY) (138)
- Claremont Colleges (120)
- Western Michigan University (120)
- Southern Methodist University (119)
- University of Arkansas, Fayetteville (116)
- Keyword
-
- Statistics (412)
- Machine learning (153)
- Humans (146)
- Simulation (110)
- Machine Learning (102)
-
- Bayesian (96)
- Regression (92)
- Female (91)
- COVID-19 (81)
- Male (81)
- Classification (76)
- Probability (70)
- Logistic regression (69)
- Mathematics (66)
- Reliability (64)
- Students (63)
- Survival analysis (63)
- Prediction (62)
- Missing data (61)
- Bootstrap (59)
- Epidemiology (59)
- Estimation (56)
- Empirical legal studies (55)
- Education (53)
- Road engineering (53)
- Teachers (53)
- Longitudinal data (52)
- Time series (51)
- Bias (50)
- Forecasting (50)
- Publication Year
- Publication
-
- Journal of Modern Applied Statistical Methods (1093)
- Journal of China & Foreign Highway (570)
- Theses and Dissertations (565)
- Mathematics and Statistics Faculty Research & Creative Works (434)
- Kesmas (352)
-
- Loma Linda University Electronic Theses, Dissertations & Projects (326)
- Mathematics, Statistics and Computer Science Faculty Research and Publications (319)
- Faculty Publications (280)
- Electronic Theses and Dissertations (261)
- U.C. Berkeley Division of Biostatistics Working Paper Series (242)
- UW Biostatistics Working Paper Series (215)
- Harvard University Biostatistics Working Paper Series (212)
- Johns Hopkins University, Dept. of Biostatistics Working Papers (178)
- Department of Statistics: Faculty Publications (162)
- Applications and Applied Mathematics: An International Journal (AAM) (161)
- Mathematics & Statistics ETDs (159)
- Annual Research Symposium (155)
- Mathematics and Statistics Faculty Publications (152)
- USF Tampa Graduate Theses and Dissertations (136)
- Open Access Theses & Dissertations (130)
- All Graduate Theses and Dissertations, Spring 1920 to Summer 2023 (123)
- Dissertations (122)
- The University of Michigan Department of Biostatistics Working Paper Series (111)
- All Graduate Plan B and other Reports, Spring 1920 to Spring 2023 (110)
- Statistics (107)
- Epidemiology Faculty Publications (105)
- Chulalongkorn University Theses and Dissertations (Chula ETD) (99)
- International Conference on Gambling & Risk Taking (94)
- COBRA Preprint Series (91)
- Graduate Theses and Dissertations (86)
- Publication Type
Articles 5281 - 5310 of 12832
Full-Text Articles in Statistics and Probability
Of Typicality And Predictive Distributions In Discriminant Function Analysis, Lyle W. Konigsberg, Susan R. Frankenberg
Of Typicality And Predictive Distributions In Discriminant Function Analysis, Lyle W. Konigsberg, Susan R. Frankenberg
Human Biology Open Access Pre-Prints
While discriminant function analysis is an inherently Bayesian method, researchers attempting to estimate ancestry in human skeletal samples often follow discriminant function analysis with the calculation of frequentist-based typicalities for assigning group membership. Such an approach is problematic in that it fails to account for admixture and for variation in why individuals may be classified as outliers, or non-members of particular groups. This paper presents an argument and methodology for employing a fully Bayesian approach in discriminant function analysis applied to cases of ancestry estimation. The approach requires adding the calculation, or estimation, of predictive distributions as the final step …
Secondary Data Analysis Project, Jonathan M. Gallimore
Secondary Data Analysis Project, Jonathan M. Gallimore
SF 420 PR - Gallimore - Fall 2018
This activity is designed to give students an opportunity to apply what they have learned in statistics to a real dataset.
This activity will help students apply what they have learned in statistics to real world data and answer their own research questions. Students will also practice reporting their results in a paper using APA format.
Quantitative Jeopardy Feud, Jonathan M. Gallimore
Quantitative Jeopardy Feud, Jonathan M. Gallimore
MSF 600 PR - Gallimore - Fall 2018
This activity - Quantitative Jeopardy Feud - is a method for using a game as a final exam.
The Transmuted Geometric-Quadratic Hazard Rate Distribution: Development, Properties, Characterizations And Applications, Fiaz Ahmad Bhatti, Gholamhossein Hamedani, Mustafa Ç. Korkmaz, Munir Ahmad
The Transmuted Geometric-Quadratic Hazard Rate Distribution: Development, Properties, Characterizations And Applications, Fiaz Ahmad Bhatti, Gholamhossein Hamedani, Mustafa Ç. Korkmaz, Munir Ahmad
Mathematics, Statistics and Computer Science Faculty Research and Publications
We propose a five parameter transmuted geometric quadratic hazard rate (TG-QHR) distribution derived from mixture of quadratic hazard rate (QHR), geometric and transmuted distributions via the application of transmuted geometric-G (TG-G) family of Afify et al.(Pak J Statist 32(2), 139-160, 2016). Some of its structural properties are studied. Moments, incomplete moments, inequality measures, residual life functions and some other properties are theoretically taken up. The TG-QHR distribution is characterized via different techniques. Estimates of the parameters for TG-QHR distribution are obtained using maximum likelihood method. The simulation studies are performed on the basis of graphical results to illustrate the performance …
Automatic Knowledge Extraction From Ocr Documents Using Hierarchical Document Analysis, Mohammad Masum, Sai Kosaraju, Tanju Bayramoglu, Girish Modgil, Mingon Kang
Automatic Knowledge Extraction From Ocr Documents Using Hierarchical Document Analysis, Mohammad Masum, Sai Kosaraju, Tanju Bayramoglu, Girish Modgil, Mingon Kang
Published and Grey Literature from PhD Candidates
Industries can improve their business efficiency by analyzing and extracting relevant knowledge from large numbers of documents. Knowledge extraction manually from large volume of documents is labor intensive, unscalable and challenging. Consequently, there have been a number of attempts to develop intelligent systems to automatically extract relevant knowledge from OCR documents. Moreover, the automatic system can improve the capability of search engine by providing application-specific domain knowledge. However, extracting the efficient information from OCR documents is challenging due to highly unstructured format. In this paper, we propose an efficient framework for a knowledge extraction system that takes keywords based queries …
Confidence Intervals For The Area Under The Receiver Operating Characteristic Curve In The Presence Of Ignorable Missing Data, Hunyong Cho, Gregory J. Matthews, Ofer Harel
Confidence Intervals For The Area Under The Receiver Operating Characteristic Curve In The Presence Of Ignorable Missing Data, Hunyong Cho, Gregory J. Matthews, Ofer Harel
Mathematics and Statistics: Faculty Publications and Other Works
Receiver operating characteristic curves are widely used as a measure of accuracy of diagnostic tests and can be summarised using the area under the receiver operating characteristic curve (AUC). Often, it is useful to construct a confidence interval for the AUC; however, because there are a number of different proposed methods to measure variance of the AUC, there are thus many different resulting methods for constructing these intervals. In this article, we compare different methods of constructing Wald‐type confidence interval in the presence of missing data where the missingness mechanism is ignorable. We find that constructing confidence intervals using multiple …
Dietary Inflammatory Index And Biomarkers Of Lipoprotein Metabolism, Inflammation And Glucose Homeostasis In Adults, Catherine Phillips, Nitin Shivappa, James R. Hébert, Ivan Perry
Dietary Inflammatory Index And Biomarkers Of Lipoprotein Metabolism, Inflammation And Glucose Homeostasis In Adults, Catherine Phillips, Nitin Shivappa, James R. Hébert, Ivan Perry
Faculty Publications
Accumulating evidence identifies diet and inflammation as potential mechanisms contributing to cardiometabolic risk. However, inconsistent reports regarding dietary inflammatory potential, biomarkers of cardiometabolic health and metabolic syndrome (MetS) risk exist. Our objective was to examine the relationships between a food frequency questionnaire (FFQ)-derived dietary inflammatory index (DII®), biomarkers of lipoprotein metabolism, inflammation and glucose homeostasis and MetS risk in a cross-sectional sample of 1992 adults. Energy-adjusted DII (E-DII) scores derived from an FFQ were calculated. Lipoprotein particle size and subclass concentrations were measured using nuclear magnetic resonance (NMR) spectroscopy. Serum acute-phase reactants, adipocytokines, pro-inflammatory cytokines and white blood cell (WBC) …
Robust Inference For The Stepped Wedge Design, James P. Hughes, Patrick J. Heagerty, Fan Xia, Yuqi Ren
Robust Inference For The Stepped Wedge Design, James P. Hughes, Patrick J. Heagerty, Fan Xia, Yuqi Ren
UW Biostatistics Working Paper Series
Based on a permutation argument, we derive a closed form expression for an estimate of the treatment effect, along with its standard error, in a stepped wedge design. We show that these estimates are robust to misspecification of both the mean and covariance structure of the underlying data-generating mechanism, thereby providing a robust approach to inference for the treatment effect in stepped wedge designs. We use simulations to evaluate the type I error and power of the proposed estimate and to compare the performance of the proposed estimate to the optimal estimate when the correct model specification is known. The …
Generalizing Multistage Partition Procedures For Two-Parameter Exponential Populations, Rui Wang
Generalizing Multistage Partition Procedures For Two-Parameter Exponential Populations, Rui Wang
LSU New Orleans Theses and Dissertations
ANOVA analysis is a classic tool for multiple comparisons and has been widely used in numerous disciplines due to its simplicity and convenience. The ANOVA procedure is designed to test if a number of different populations are all different. This is followed by usual multiple comparison tests to rank the populations. However, the probability of selecting the best population via ANOVA procedure does not guarantee the probability to be larger than some desired prespecified level. This lack of desirability of the ANOVA procedure was overcome by researchers in early 1950's by designing experiments with the goal of selecting the best …
Bubble Stream Production By Belugas (Delphinapterus Leucas), Megan Slack
Bubble Stream Production By Belugas (Delphinapterus Leucas), Megan Slack
Theses
Bubble stream production in belugas has been poorly characterized and its function is not well understood. I examined behavioral states when producing bubble streams (“bubbling”), and when bubbling calls, to determine whether bubbling was significantly associated with a particular call category or behavioral state. Using 19 hours of video and audio recordings collected over a two-day period, I quantified bubble streams of a 4-month old calf and an unrelated adult female housed together. Based on the overall activity budgets and pool of vocalizations for both animals, I calculated the expected counts of bubble streams with and without vocalizations, assuming that …
Deep Machine Learning For Mechanical Performance And Failure Prediction, Elijah Reber, Nickolas D. Winovich, Guang Lin
Deep Machine Learning For Mechanical Performance And Failure Prediction, Elijah Reber, Nickolas D. Winovich, Guang Lin
The Summer Undergraduate Research Fellowship (SURF) Symposium
Deep learning has provided opportunities for advancement in many fields. One such opportunity is being able to accurately predict real world events. Ensuring proper motor function and being able to predict energy output is a valuable asset for owners of wind turbines. In this paper, we look at how effective a deep neural network is at predicting the failure or energy output of a wind turbine. A data set was obtained that contained sensor data from 17 wind turbines over 13 months, measuring numerous variables, such as spindle speed and blade position and whether or not the wind turbine experienced …
Efvs Effects On Pilot Performance, Michael Campbell, Nsikak Udo-Imeh, Steven J. Landry
Efvs Effects On Pilot Performance, Michael Campbell, Nsikak Udo-Imeh, Steven J. Landry
The Summer Undergraduate Research Fellowship (SURF) Symposium
Flight tests have been conducted at Purdue University using a computer-based flying simulator in an attempt to determine and measure the effects of Enhanced Flight Vision Systems (EFVS) on the performance of pilots during landing. Knowledge of these effects could help guide future design and implementation of EFVS in modern commercial aircraft, and further increase pilots’ ability to control the aircraft in low-visibility conditions. The problem that has faced researchers in the past has revolved around the difficulty in interpreting the data which is generated by these tests. The difficulty in making a generalized conclusion based on the large amount …
Scale-Invariant Geometric Data Analysis (Sigda), Marina Girgis, Max Robinson
Scale-Invariant Geometric Data Analysis (Sigda), Marina Girgis, Max Robinson
STAR Program Research Presentations
The purpose of this research is to introduce a new data analysis method called Scale Invariant Geometric Data Analysis (SIGDA). SIGDA has been shown to be more informative than more common data analysis methods, such as Principal Component Analysis (PCA). SIGDA is used to visualize complex data sets in a way that accurately preserves data patterns and behavior. SIGDA is designed to preserve relative ratios in a numerical matrix, and the number of entries has to be more than the total number of rows and columns. Our research involved providing a simple explanation of SIGDA's mathematical process—simple enough for the …
Development Of A Statistical Model For Discrimination Of Rupture Status In Posterior Communicating Artery Aneurysms, Felicitas J. Detmer, Bong Jae Chung, Fernando Mut, Michael Pritz, Martin Slawski, Farid Hamzei-Sichani, David Kallmes, Christopher Putman, Carlos Jimenez, Juan R. Cebral
Development Of A Statistical Model For Discrimination Of Rupture Status In Posterior Communicating Artery Aneurysms, Felicitas J. Detmer, Bong Jae Chung, Fernando Mut, Michael Pritz, Martin Slawski, Farid Hamzei-Sichani, David Kallmes, Christopher Putman, Carlos Jimenez, Juan R. Cebral
Department of Applied Mathematics and Statistics Faculty Scholarship and Creative Works
Background: Intracranial aneurysms at the posterior communicating artery (PCOM) are known to have high rupture rates compared to other locations. We developed and internally validated a statistical model discriminating between ruptured and unruptured PCOM aneurysms based on hemodynamic and geometric parameters, angio-architectures, and patient age with the objective of its future use for aneurysm risk assessment. Methods: A total of 289 PCOM aneurysms in 272 patients modeled with image-based computational fluid dynamics (CFD) were used to construct statistical models using logistic group lasso regression. These models were evaluated with respect to discrimination power and goodness of fit using tenfold nested …
Feature Screening Of Ultrahigh Dimensional Feature Spaces With Applications In Interaction Screening, Randall D. Reese
Feature Screening Of Ultrahigh Dimensional Feature Spaces With Applications In Interaction Screening, Randall D. Reese
All Graduate Theses and Dissertations, Spring 1920 to Summer 2023
Data for which the number of predictors exponentially exceeds the number of observations is becoming increasingly prevalent in fields such as bioinformatics, medical imaging, computer vision, And social network analysis. One of the leading questions statisticians must answer when confronted with such “big data” is how to reduce a set of exponentially many predictors down to a set of a mere few predictors which have a truly causative effect on the response being modelled. This process is often referred to as feature screening. In this work we propose three new methods for feature screening. The first method we propose (TC-SIS) …
Dynamics Of Paramagnetic And Ferromagnetic Ellipsoidal Particles In Shear Flow Under A Uniform Magnetic Field, Christopher A. Sobecki, Jie Zhang, Yanzhi Zhang, Cheng Wang
Dynamics Of Paramagnetic And Ferromagnetic Ellipsoidal Particles In Shear Flow Under A Uniform Magnetic Field, Christopher A. Sobecki, Jie Zhang, Yanzhi Zhang, Cheng Wang
Mathematics and Statistics Faculty Research & Creative Works
We investigate the two-dimensional dynamic motion of magnetic particles of ellipsoidal shapes in shear flow under the influence of a uniform magnetic field. In the first part, we present a theoretical analysis of the rotational dynamics of the particles in simple shear flow. By considering paramagnetic and ferromagnetic particles, we study the effects of the direction and strength of the magnetic field on the particle rotation. The critical magnetic-field strength, at which particle rotation is impeded, is determined. In a weak-field regime (i.e., below the critical strength) where the particles execute complete rotations, the symmetry property of the rotational velocity …
Wald Confidence Intervals For A Single Poisson Parameter And Binomial Misclassification Parameter When The Data Is Subject To Misclassification, Nishantha Janith Chandrasena Poddiwala Hewage
Wald Confidence Intervals For A Single Poisson Parameter And Binomial Misclassification Parameter When The Data Is Subject To Misclassification, Nishantha Janith Chandrasena Poddiwala Hewage
Electronic Theses and Dissertations
This thesis is based on a Poisson model that uses both error-free data and error-prone data subject to misclassification in the form of false-negative and false-positive counts. We present maximum likelihood estimators (MLEs), Fisher's Information, and Wald statistics for Poisson rate parameter and the two misclassification parameters. Next, we invert the Wald statistics to get asymptotic confidence intervals for Poisson rate parameter and false-negative rate parameter. The coverage and width properties for various sample size and parameter configurations are studied via a simulation study. Finally, we apply the MLEs and confidence intervals to one real data set and another realistic …
A Comparison Of R, Sas, And Python Implementations Of Random Forests, Breckell Soifua
A Comparison Of R, Sas, And Python Implementations Of Random Forests, Breckell Soifua
All Graduate Plan B and other Reports, Spring 1920 to Spring 2023
The Random Forest method is a useful machine learning tool developed by Leo Breiman. There are many existing implementations across different programming languages; the most popular of which exist in R, SAS, and Python. In this paper, we conduct a comprehensive comparison of these implementations with regards to the accuracy, variable importance measurements, and timing. This comparison was done on a variety of real and simulated data with different classification difficulty levels, number of predictors, and sample sizes. The comparison shows unexpectedly different results between the three implementations.
A Programme For Risk Assessment And Minimisation Of Progressive Multifocal Leukoencephalopathy Developed For Vedolizumab Clinical Trials, Asit Parikh, Kristin Stephens, Eugene Major, Irving Fox, Catherine Milch, Serap Sankoh, Michael H. Lev, James M. Provenzale, Jesse Shick, Mark Patti, Megan Mcauliffe, Joseph R. Berger, David B. Clifford
A Programme For Risk Assessment And Minimisation Of Progressive Multifocal Leukoencephalopathy Developed For Vedolizumab Clinical Trials, Asit Parikh, Kristin Stephens, Eugene Major, Irving Fox, Catherine Milch, Serap Sankoh, Michael H. Lev, James M. Provenzale, Jesse Shick, Mark Patti, Megan Mcauliffe, Joseph R. Berger, David B. Clifford
Neurology Faculty Publications
Introduction Over the past decade, the potential for drug-associated progressive multifocal leukoencephalopathy (PML) has become an increasingly important consideration in certain drug development programmes, particularly those of immunomodulatory biologics. Whether the risk of PML with an investigational agent is proven (e.g. extrapolated from relevant experience, such as a class effect) or merely theoretical, the serious consequences of acquiring PML require careful risk minimisation and assessment. No single standard for such risk minimisation exists. Vedolizumab is a recently developed monoclonal antibody to α4β7 integrin. Its clinical development necessitated a dedicated PML risk minimisation assessment as part of a global preapproval regulatory …
Bayesian Analytical Approaches For Metabolomics : A Novel Method For Molecular Structure-Informed Metabolite Interaction Modeling, A Novel Diagnostic Model For Differentiating Myocardial Infarction Type, And Approaches For Compound Identification Given Mass Spectrometry Data., Patrick J. Trainor
Electronic Theses and Dissertations
Metabolomics, the study of small molecules in biological systems, has enjoyed great success in enabling researchers to examine disease-associated metabolic dysregulation and has been utilized for the discovery biomarkers of disease and phenotypic states. In spite of recent technological advances in the analytical platforms utilized in metabolomics and the proliferation of tools for the analysis of metabolomics data, significant challenges in metabolomics data analyses remain. In this dissertation, we present three of these challenges and Bayesian methodological solutions for each. In the first part we develop a new methodology to serve a basis for making higher order inferences in metabolomics, …
Generalized Spatiotemporal Modeling And Causal Inference For Assessing Treatment Effects For Multiple Groups For Ordinal Outcome., Soutik Ghosal
Generalized Spatiotemporal Modeling And Causal Inference For Assessing Treatment Effects For Multiple Groups For Ordinal Outcome., Soutik Ghosal
Electronic Theses and Dissertations
This dissertation consists of three projects and can be categorized in two broad research areas: generalized spatiotemporal modeling and causal inference based on observational data. In the first project, I introduce a Bayesian hierarchical mixed effect hurdle model with a nested random effect structure to model the count for primary care providers and understand their spatial and temporal variation. This study further enables us to identify the health professional shortage areas and the possible impacting factors. In the second project, I have unified popular parametric and nonparametric propensity score-based methods to assess the treatment effect of multiple groups for ordinal …
Adapting To Sparsity And Heavy Tailed Data, Mohamed Abdelkader Abba
Adapting To Sparsity And Heavy Tailed Data, Mohamed Abdelkader Abba
Graduate Theses and Dissertations
The Lasso and the Horseshoe, gold-standards in the frequentist and Bayesian paradigms, critically depend on learning the error variance. This causes a lack of scale invariance and adaptability to heavy-tailed data. The √ Lasso [Belloni et al., 2011] attempt to correct this by using the `1 norm on both the likelihood and the penalty for the objective function. In contrast, there is essentially no methods for uncertainty quantification or automatic parameter tuning via a formal Bayesian treatment of an unknown error distribution. On the other hand, Bayesian shrinkage priors lacking a local shrinkage term fails to adapt to the large …
Empirical Bayesian Approach To Testing Multiple Hypotheses With Separate Priors For Left And Right Alternatives, Naveen K. Bansal, Mehdi Maadooliat, Steven J. Schrodi
Empirical Bayesian Approach To Testing Multiple Hypotheses With Separate Priors For Left And Right Alternatives, Naveen K. Bansal, Mehdi Maadooliat, Steven J. Schrodi
Mathematics, Statistics and Computer Science Faculty Research and Publications
We consider a multiple hypotheses problem with directional alternatives in a decision theoretic framework. We obtain an empirical Bayes rule subject to a constraint on mixed directional false discovery rate (mdFDR≤α) under the semiparametric setting where the distribution of the test statistic is parametric, but the prior distribution is nonparametric. We proposed separate priors for the left tail and right tail alternatives as it may be required for many applications. The proposed Bayes rule is compared through simulation against rules proposed by Benjamini and Yekutieli and Efron. We illustrate the proposed methodology for two sets of …
Surprise Vs. Probability As A Metric For Proof, Edward K. Cheng, Matthew Ginther
Surprise Vs. Probability As A Metric For Proof, Edward K. Cheng, Matthew Ginther
Vanderbilt Law School Faculty Publications
In this Symposium issue celebrating his career, Professor Michael Risinger in Leveraging Surprise proposes using "the fundamental emotion of surprise" as a way of measuring belief for purposes of legal proof. More specifically, Professor Risinger argues that we should not conceive of the burden of proof in terms of probabilities such as 51%, 95%, or even "beyond a reasonable doubt." Rather, the legal system should reference the threshold using "words of estimative surprise" -asking jurors how surprised they would be if the fact in question were not true. Toward this goal (and being averse to cardinality), he suggests categories such …
Distribution Of A Sum Of Random Variables When The Sample Size Is A Poisson Distribution, Mark Pfister
Distribution Of A Sum Of Random Variables When The Sample Size Is A Poisson Distribution, Mark Pfister
Electronic Theses and Dissertations
A probability distribution is a statistical function that describes the probability of possible outcomes in an experiment or occurrence. There are many different probability distributions that give the probability of an event happening, given some sample size n. An important question in statistics is to determine the distribution of the sum of independent random variables when the sample size n is fixed. For example, it is known that the sum of n independent Bernoulli random variables with success probability p is a Binomial distribution with parameters n and p: However, this is not true when the sample size …
The Expected Number Of Patterns In A Random Generated Permutation On [N] = {1,2,...,N}, Evelyn Fokuoh
The Expected Number Of Patterns In A Random Generated Permutation On [N] = {1,2,...,N}, Evelyn Fokuoh
Electronic Theses and Dissertations
Previous work by Flaxman (2004) and Biers-Ariel et al. (2018) focused on the number of distinct words embedded in a string of words of length n. In this thesis, we will extend this work to permutations, focusing on the maximum number of distinct permutations contained in a permutation on [n] = {1,2,...,n} and on the expected number of distinct permutations contained in a random permutation on [n]. We further considered the problem where repetition of subsequences are as a result of the occurrence of (Type A and/or Type B) replications. Our method of enumerating the Type A replications causes double …
Clustering Mixed Data: An Extension Of The Gower Coefficient With Weighted L2 Distance, Augustine Oppong
Clustering Mixed Data: An Extension Of The Gower Coefficient With Weighted L2 Distance, Augustine Oppong
Electronic Theses and Dissertations
Sorting out data into partitions is increasing becoming complex as the constituents of data is growing outward everyday. Mixed data comprises continuous, categorical, directional functional and other types of variables. Clustering mixed data is based on special dissimilarities of the variables. Some data types may influence the clustering solution. Assigning appropriate weight to the functional data may improve the performance of the clustering algorithm. In this paper we use the extension of the Gower coefficient with judciously chosen weight for the L2 to cluster mixed data.The benefits of weighting are demonstrated both in in applications to the Buoy data set …
Generalized Non-Inferential Approach To Modeling Restricted Discrete Choice For The Case Of The Spatial Random Utility, Elena Labzina
Generalized Non-Inferential Approach To Modeling Restricted Discrete Choice For The Case Of The Spatial Random Utility, Elena Labzina
Arts & Sciences Graduate Student Theses and Dissertations
Multinomial logistic regression model (MNL) is a powerful and easily tractable way for measuring the probabilistic impact of input variables on individual categorical choices. Crucially, the standard MNL assumes that all subjects of the study have the same choice sets. In the meanwhile, especially in political science and economics, this condition is frequently violated. Probably, the most graphical example of varying choice sets (VCS) is partially contested elections. Furthermore, the MNL implicitly implies the Independence of the Irregular Alternatives (IIA) assumption by requiring i.i.d errors that contrasts the MNL and the multinomial probit (MNP) and mixed logit (MXL) models. In …
Implementing The Use Of Personal Activity Data In An Introductory Statistics Course, Lacy Christensen
Implementing The Use Of Personal Activity Data In An Introductory Statistics Course, Lacy Christensen
All Graduate Theses and Dissertations, Spring 1920 to Summer 2023
Integrating real data into a classroom is one of the recommendations in the Guidelines for Assessment and Instruction in Statistics Education (GAISE) college report which lays out guidelines for an introductory statistics course (Committee, GAISE College Report ASA Revision, 2016). In order to assess the effect of using real data in a classroom, the students received physical activity trackers to wear during an undergraduate introductory statistics course taught in the summer. This tracker, a Fitbit, enabled students to monitor and record their steps, calories, and active time throughout the class. Collecting personal activity data (PAD) creates a large database which …
Comparison Of Correlation, Partial Correlation, And Conditional Mutual Information For Interaction Effects Screening In Generalized Linear Models, Ji Li
Graduate Theses and Dissertations
Numerous screening techniques have been developed in recent years for genome-wide association studies (GWASs) (Moore et al., 2010). In this thesis, a novel model-free screening method was developed and validated by an extensive simulation study. Many screening methods were mainly focused on main effects, while very few studies considered the models containing both main effects and interaction effects. In this work, the interaction effects were fully considered and three different methods (Pearson’s Correlation Coefficient, Partial Correlation, and Conditional Mutual Information) were tested and their prediction accuracies were compared.
Pearson’s Correlation Coefficient method, which is a direct interaction screening (DIS) procedure, …