Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Applied Statistics (49)
- Social and Behavioral Sciences (38)
- Statistical Theory (32)
- Statistical Models (18)
- Statistical Methodology (15)
-
- Biostatistics (11)
- Mathematics (11)
- Computer Sciences (9)
- Engineering (9)
- Probability (9)
- Applied Mathematics (8)
- Design of Experiments and Sample Surveys (8)
- Other Statistics and Probability (7)
- Business (6)
- Education (6)
- Data Science (5)
- Aerospace Engineering (4)
- Artificial Intelligence and Robotics (4)
- Economics (4)
- Longitudinal Data Analysis and Time Series (4)
- Psychology (4)
- Aviation (3)
- Categorical Data Analysis (3)
- Econometrics (3)
- Finance (3)
- Numerical Analysis and Computation (3)
- Operations Research, Systems Engineering and Industrial Engineering (3)
- Other Applied Mathematics (3)
- Institution
-
- Wayne State University (26)
- Air Force Institute of Technology (7)
- Prairie View A&M University (7)
- California Polytechnic State University, San Luis Obispo (5)
- Utah State University (4)
-
- Marshall University (3)
- Old Dominion University (3)
- University of Nevada, Las Vegas (3)
- Virginia Commonwealth University (3)
- Brigham Young University (2)
- COBRA (2)
- Central Bank of Nigeria (2)
- Dordt University (2)
- East Tennessee State University (2)
- James Madison University (2)
- Stephen F. Austin State University (2)
- University of Kentucky (2)
- University of Nebraska - Lincoln (2)
- University of New Mexico (2)
- Arcadia University (1)
- Arkansas State University (1)
- City University of New York (CUNY) (1)
- Embry-Riddle Aeronautical University (1)
- Georgia Southern University (1)
- LSU Health New Orleans (1)
- Louisiana Tech University (1)
- Murray State University (1)
- Northern Illinois University (1)
- Purdue University (1)
- Smith College (1)
- Publication Year
- Publication
-
- Journal of Modern Applied Statistical Methods (25)
- Theses and Dissertations (11)
- Applications and Applied Mathematics: An International Journal (AAM) (7)
- Electronic Theses and Dissertations (6)
- Master's Theses (3)
-
- Theses, Dissertations and Capstones (3)
- All Graduate Plan B and other Reports, Spring 1920 to Spring 2023 (2)
- CBN Journal of Applied Statistics (JAS) (2)
- Faculty Publications (2)
- Faculty Work Comprehensive List (2)
- Statistics (2)
- All Graduate Theses and Dissertations, Fall 2023 to Present (1)
- All Graduate Theses and Dissertations, Spring 1920 to Summer 2023 (1)
- Applications (1)
- Biostatistics Faculty Publications (1)
- Branch Mathematics and Statistics Faculty and Staff Publications (1)
- Capstone Showcase (1)
- College of Graduate Studies: Theses & Dissertations (1)
- Conference papers (1)
- Department of Sociology: Faculty Publications (1)
- Dissertations (1)
- Dissertations, 2020-current (1)
- Doctor of Education (Ed.D) (1)
- Doctoral Dissertations (1)
- Graduate Research Theses & Dissertations (1)
- Graduate Student Theses, Dissertations, & Professional Papers (1)
- Honors Scholar Theses (1)
- Hospitality Faculty Research (1)
- Industrial Engineering Undergraduate Honors Theses (1)
- International Journal of Aviation, Aeronautics, and Aerospace (1)
- Publication Type
Articles 31 - 60 of 110
Full-Text Articles in Statistics and Probability
The Wargaming Commodity Course Of Action Automated Analysis Method, William T. Deberry
The Wargaming Commodity Course Of Action Automated Analysis Method, William T. Deberry
Theses and Dissertations
This research presents the Wargaming Commodity Course of Action Automated Analysis Method (WCCAAM), a novel approach to assist wargame commanders in developing and analyzing courses of action (COAs) through semi-automation of the Military Decision Making Process (MDMP). MDMP is a seven-step iterative method that commanders and mission partners follow to build an operational course of action to achieve strategic objectives. MDMP requires time, resources, and coordination – all competing items the commander weighs to make the optimal decision. WCCAAM receives the MDMP's Mission Analysis phase as input, converts the wargame into a directed graph, processes a multi-commodity flow algorithm on …
The Simulation Extrapolation Method With Differential Measurement Error, Dominic Partipilo
The Simulation Extrapolation Method With Differential Measurement Error, Dominic Partipilo
Graduate Research Theses & Dissertations
Most of statistical theory operates under the assumption that the true values of covariates have been measured correctly, but it is not always possible to obtain the true values of these covariates. A common issue, specifically in regression models, is that predictors are misclassified or measured with systematic measurement error. There have been many methods developed for handling measurement error, specifically in the case where measurement error is nondifferential, where the measurement error can be treated as independent from the covariates. The frequentist method known as simulation extrapolation (SIMEX) is one of these methods that specifically handles the case for …
The Odd Inverse Rayleigh Family Of Distributions: Simulation & Application To Real Data, Saeed E. Hemeda, Muhammad A. Ul Haq
The Odd Inverse Rayleigh Family Of Distributions: Simulation & Application To Real Data, Saeed E. Hemeda, Muhammad A. Ul Haq
Applications and Applied Mathematics: An International Journal (AAM)
A new family of inverse probability distributions named inverse Rayleigh family is introduced to generate many continuous distributions. The shapes of probability density and hazard rate functions are investigated. Some Statistical measures of the new generator including moments, quantile and generating functions, entropy measures and order statistics are derived. The Estimation of the model parameters is performed by the maximum likelihood estimation method. Furthermore, a simulation study is used to estimate the parameters of one of the members of the new family. The data application shows that the new family models can be useful to provide better fits than other …
Interval Estimation Of Proportion Of Second-Level Variance In Multi-Level Modeling, Steven Svoboda
Interval Estimation Of Proportion Of Second-Level Variance In Multi-Level Modeling, Steven Svoboda
The Nebraska Educator: A Student-Led Journal
Physical, behavioral and psychological research questions often relate to hierarchical data systems. Examples of hierarchical data systems include repeated measures of students nested within classrooms, nested within schools and employees nested within supervisors, nested within organizations. Applied researchers studying hierarchical data structures should have an estimate of the intraclass correlation coefficient (ICC) for every nested level in their analyses because ignoring even relatively small amounts of interdependence is known to inflate Type I error rate in single-level models. Traditionally, researchers rely upon the ICC as a point estimate of the amount of interdependency in their data. Recent methods utilizing an …
Can Auxiliary Information Improve Rasch Estimation At Small Sample Sizes?, Derek Sauder
Can Auxiliary Information Improve Rasch Estimation At Small Sample Sizes?, Derek Sauder
Dissertations, 2020-current
The Rasch model is commonly used to calibrate multiple choice items. However, the sample sizes needed to estimate the Rasch model can be difficult to attain (e.g., consider a small testing company trying to pretest new items). With small sample sizes, auxiliary information besides the item responses may improve estimation of the item parameters. The purpose of this study was to determine if incorporating item property information (i.e., characteristics of the items related to item difficulty) in a random effects linear logistic test model (RE-LLTM) would improve estimation of item difficulty. A simulation study was conducted that varied sample size, …
Propensity Score Matching And Generalized Boosted Modeling In The Context Of Model Misspecification: A Simulation Study, Briana G. Craig
Propensity Score Matching And Generalized Boosted Modeling In The Context Of Model Misspecification: A Simulation Study, Briana G. Craig
Masters Theses, 2020-current
In the absence of random assignment, researchers must consider the impact of selection bias – pre-existing covariate differences between groups due to differences among those entering into treatment and those otherwise unable to participate. Propensity score matching (PSM) and generalized boosted modeling (GBM) are two quasi-experimental pre-processing methods that strive to reduce the impact of selection bias before analyzing a treatment effect. PSM and GBM both examine a treatment and comparison group and either match or weight members of those groups to create new, balanced groups. The new, balanced groups theoretically can then be used as a proxy for the …
Residual-Matching: An Efficient Alternative To Random Sampling In Human Subjects Research Recruitment, Andrew Hooyman, Matthew J. Huentelman, Sydney Y. Schaefer
Residual-Matching: An Efficient Alternative To Random Sampling In Human Subjects Research Recruitment, Andrew Hooyman, Matthew J. Huentelman, Sydney Y. Schaefer
Kinesiology and Nutrition Sciences Faculty Research
Given the time- and resource-intense nature of human subjects research, we have developed a more intelligent approach to participant recruitment above and beyond random sampling that leverages pilot or preliminary results to reduce the overall number of participants needed for recruitment from an existing electronic cohort or database. Using open-access data from the General Social Survey (GSS) of the National Opinion Research Center, we generated pilot and validation datasets through a simulation to establish moderate and weak relationships based on linear regression. We then compared the performance of our residual-matching method against random sampling in their probabilities of achieving a …
Dot: Gene-Set Analysis By Combining Decorrelated Association Statistics, Olga A. Vsevolozhskaya, Min Shi, Fengjiao Hu, Dmitri V. Zaykin
Dot: Gene-Set Analysis By Combining Decorrelated Association Statistics, Olga A. Vsevolozhskaya, Min Shi, Fengjiao Hu, Dmitri V. Zaykin
Biostatistics Faculty Publications
Historically, the majority of statistical association methods have been designed assuming availability of SNP-level information. However, modern genetic and sequencing data present new challenges to access and sharing of genotype-phenotype datasets, including cost of management, difficulties in consolidation of records across research groups, etc. These issues make methods based on SNP-level summary statistics particularly appealing. The most common form of combining statistics is a sum of SNP-level squared scores, possibly weighted, as in burden tests for rare variants. The overall significance of the resulting statistic is evaluated using its distribution under the null hypothesis. Here, we demonstrate that this basic …
Using Hac Estimators For Intervention Analysis, Ashok K. Singh, Rohan J. Dalpatadu
Using Hac Estimators For Intervention Analysis, Ashok K. Singh, Rohan J. Dalpatadu
Hospitality Faculty Research
The purpose of this article is to present an alternative method for intervention analysis of time series data that is simpler to use than the traditional method of fitting an explanatory Autoregressive Integrated Moving Average (ARIMA) model. Time series regression analysis is commonly used to test the effect of an event on a time series. An econometric modeling method, which uses a heteroskedasticity and autocorrelation consistent (HAC) estimator of the covariance matrix instead of fitting an ARIMA model, is proposed as an alternative. The method of parametric bootstrap is used to compare the two approaches for intervention analysis. The results …
Assessing Robustness Of The Rasch Mixture Model To Detect Differential Item Functioning - A Monte Carlo Simulation Study, Jinjin Huang
Assessing Robustness Of The Rasch Mixture Model To Detect Differential Item Functioning - A Monte Carlo Simulation Study, Jinjin Huang
Electronic Theses and Dissertations
Measurement invariance is crucial for an effective and valid measure of a construct. Invariance holds when the latent trait varies consistently across subgroups; in other words, the mean differences among subgroups are only due to true latent ability differences. Differential item functioning (DIF) occurs when measurement invariance is violated. There are two kinds of traditional tools for DIF detection: non-parametric methods and parametric methods. Mantel Haenszel (MH), SIBTEST, and standardization are examples of non-parametric DIF detection methods. The majority of parametric DIF detection methods are item response theory (IRT) based. Both non-parametric methods and parametric methods compare differences among subgroups …
Spatio-Temporal Cluster Detection And Local Moran Statistics Of Point Processes, Jennifer L. Matthews
Spatio-Temporal Cluster Detection And Local Moran Statistics Of Point Processes, Jennifer L. Matthews
Mathematics & Statistics Theses & Dissertations
Moran's index is a statistic that measures spatial dependence, quantifying the degree of dispersion or clustering of point processes and events in some location/area. Recognizing that a single Moran's index may not give a sufficient summary of the spatial autocorrelation measure, a local indicator of spatial association (LISA) has gained popularity. Accordingly, we propose extending LISAs to time after partitioning the area and computing a Moran-type statistic for each subarea. Patterns between the local neighbors are unveiled that would not otherwise be apparent. We consider the measures of Moran statistics while incorporating a time factor under simulated multilevel Palm distribution, …
Jmasm 51: Bayesian Reliability Analysis Of Binomial Model – Application To Success/Failure Data, M. Tanwir Akhtar, Athar Ali Khan
Jmasm 51: Bayesian Reliability Analysis Of Binomial Model – Application To Success/Failure Data, M. Tanwir Akhtar, Athar Ali Khan
Journal of Modern Applied Statistical Methods
Reliability data are generated in the form of success/failure. An attempt was made to model such type of data using binomial distribution in the Bayesian paradigm. For fitting the Bayesian model both analytic and simulation techniques are used. Laplace approximation was implemented for approximating posterior densities of the model parameters. Parallel simulation tools were implemented with an extensive use of R and JAGS. R and JAGS code are developed and provided. Real data sets are used for the purpose of illustration.
Exploring The Variance Of The Sample Variance Through Estimation And Simulation, Christina Stradwick
Exploring The Variance Of The Sample Variance Through Estimation And Simulation, Christina Stradwick
Theses, Dissertations and Capstones
In this thesis, we examine properties of the variance of the sample variance, which we will denote V (S 2 ). We derive a formula for this variance and show that it only depends on the sample size, variance, and kurtosis of the underlying distribution. We also derive the maximum likelihood estimators for this parameter, Vˆ (S 2 ), under the normal, exponential, Bernoulli, and Poisson distributions and end the thesis with simulations demonstrating the distributions of these estimators.
Comparing Performance Of Gene Set Test Methods Using Biologically Relevant Simulated Data, Richard M. Lambert
Comparing Performance Of Gene Set Test Methods Using Biologically Relevant Simulated Data, Richard M. Lambert
All Graduate Theses and Dissertations, Spring 1920 to Summer 2023
Today we know that there are many genetically driven diseases and health conditions. These problems often manifest only when a set of genes are either active or inactive. Recent technology allows us to measure the activity level of genes in cells, which we call gene expression. It is of great interest to society to be able to statistically compare the gene expression of a large number of genes between two or more groups. For example, we may want to compare the gene expression of a group of cancer patients with a group of non-cancer patients to better understand the genetic …
Predictions Generated From A Simulation Engine For Gene Expression Micro-Arrays For Use In Research Laboratories, Gopinath R. Mavankal, John Blevins, Dominique Edwards, Monnie Mcgee, Andrew Hardin
Predictions Generated From A Simulation Engine For Gene Expression Micro-Arrays For Use In Research Laboratories, Gopinath R. Mavankal, John Blevins, Dominique Edwards, Monnie Mcgee, Andrew Hardin
SMU Data Science Review
In this paper we introduce the technical components, the biology and data science involved in the use of microarray technology in biological and clinical research. We discuss how laborious experimental protocols involved in obtaining this data used in laboratories could benefit from using simulations of the data. We discuss the approach used in the simulation engine from [7]. We use this simulation engine to generate a prediction tool in Power BI, a Microsoft, business intelligence tool for analytics and data visualization [22]. This tool could be used in any laboratory using micro-arrays to improve experimental design by comparing how predicted …
Estimation Of The Burr Xii-Exponential Distribution Parameters, Gholamhossein Yari, Zahra Tondpour
Estimation Of The Burr Xii-Exponential Distribution Parameters, Gholamhossein Yari, Zahra Tondpour
Applications and Applied Mathematics: An International Journal (AAM)
The Burr XII distribution is one of the most important distributions in Survival analysis. In this article, we introduce the new wider Burr XII-G family of distributions. A special model in the new family called Burr XII-exponential distribution that has constant, decreasing and unimodal hazard rate functions is investigated. We discuss the estimation of this distribution parameters by maximum likelihood, three modifications of maximum likelihood and Bayes methods. In Bayes method, we use the uniform, triangular and Burr XII-uniform priors for posterior analysis and obtain Bayes estimations under two different loss functions. We obtain two approximations of the Bayes estimations, …
A Study Of Flight Simulation Training Time, Aircraft Training Time, And Pilot Competence As Measured By The Naval Standard Score, Aaron D. Judy
A Study Of Flight Simulation Training Time, Aircraft Training Time, And Pilot Competence As Measured By The Naval Standard Score, Aaron D. Judy
Doctor of Education (Ed.D)
The purpose of the study was to investigate the relationships between US Navy T-45C flight simulation training time, actual aircraft training time, and intermediate and advanced jet pilot competence as measured by the Naval Standard Score (NSS). Examining the relationships between US Navy T-45C flight simulation time and actual aircraft flight time may provide further information on flight simulation training versus actual aircraft training to aviation authorities, flight instructors, the military aviation community, the commercial aviation community, and academia. The study was non-experimental, correlational, causal-comparative with an emphasis upon the establishment of mathematic and predictive relationships using archival data from …
Comparison Of The Performance Of Simple Linear Regression And Quantile Regression With Non-Normal Data: A Simulation Study, Marjorie Howard
Comparison Of The Performance Of Simple Linear Regression And Quantile Regression With Non-Normal Data: A Simulation Study, Marjorie Howard
Theses and Dissertations
Linear regression is a widely used method for analysis that is well understood across a wide variety of disciplines. In order to use linear regression, a number of assumptions must be met. These assumptions, specifically normality and homoscedasticity of the error distribution can at best be met only approximately with real data. Quantile regression requires fewer assumptions, which offers a potential advantage over linear regression. In this simulation study, we compare the performance of linear (least squares) regression to quantile regression when these assumptions are violated, in order to investigate under what conditions quantile regression becomes the more advantageous method …
Experimental Design And Data Analysis In Computer Simulation Studies In The Behavioral Sciences, Michael Harwell, Nidhi Kohli, Yadira Peralta
Experimental Design And Data Analysis In Computer Simulation Studies In The Behavioral Sciences, Michael Harwell, Nidhi Kohli, Yadira Peralta
Journal of Modern Applied Statistical Methods
Treating computer simulation studies as statistical sampling experiments subject to established principles of experimental design and data analysis should further enhance their ability to inform statistical practice and a program of statistical research. Latin hypercube designs to enhance generalizability and meta-analytic methods to analyze simulation results are presented.
Performance Evaluation Of Confidence Intervals For Ordinal Coefficient Alpha, Heather J. Turner, Prathiba Natesan, Robin K. Henson
Performance Evaluation Of Confidence Intervals For Ordinal Coefficient Alpha, Heather J. Turner, Prathiba Natesan, Robin K. Henson
Journal of Modern Applied Statistical Methods
The aim of this study was to investigate the performance of the Fisher, Feldt, Bonner, and Hakstian and Whalen (HW) confidence intervals methods for the non-parametric reliability estimate, ordinal alpha. All methods yielded unacceptably low coverage rates and potentially increased Type-I error rates.
A Semiparametric Estimation For The Nonlinear Vector Autoregressive Time Series Model, Rahman Farnoosh, Mahtab Hajebi, Seyed J. Mortazavi
A Semiparametric Estimation For The Nonlinear Vector Autoregressive Time Series Model, Rahman Farnoosh, Mahtab Hajebi, Seyed J. Mortazavi
Applications and Applied Mathematics: An International Journal (AAM)
In this paper, the nonlinear vector autoregressive model is considered and a semiparametric method is proposed to estimate the nonlinear vector regression function. We use Taylor series expansion up to the second order which has a parametric framework as a representation of the nonlinear vector regression function. After the parameters are estimated through the least squares method, the obtained nonlinear vector regression function is adjusted by a nonparametric diagonal matrix, and the proposed diagonal matrix is also estimated through the nonparametric smooth-kernel approach. Estimating the parameters can yield the desired estimate of the vector regression function based on the data. …
A Simulation Of Anthropogenic Mammoth Extinction, Matthew Klapman
A Simulation Of Anthropogenic Mammoth Extinction, Matthew Klapman
Undergraduate Honors Papers
There are multiple hypotheses as to why the Columbian Mammoth (Mammuthus columbi) and other megafauna in North America went extinct relatively recently and relatively quickly. The most popular of which are disease, climate change, meteorite strikes, and over hunting by humans [2, 9]. There is evidence to show that a combination of factors contributed to the megafaunal extinction, but ”overkill” explores the idea that early humans migrated onto the continent and then hunted the mammoths and other megafauna to extinction. The overkill hypothesis was first proposed by anthropologist Paul Martin in 1973 [8]. Evidence from radiocarbon dating shows that the …
Neural Network Predictions Of A Simulation-Based Statistical And Graph Theoretic Study Of The Board Game Risk, Jacob Munson
Neural Network Predictions Of A Simulation-Based Statistical And Graph Theoretic Study Of The Board Game Risk, Jacob Munson
Murray State Theses and Dissertations
We translate the RISK board into a graph which undergoes updates as the game advances. The dissection of the game into a network model in discrete time is a novel approach to examining RISK. A review of the existing statistical findings of skirmishes in RISK is provided. The graphical changes are accompanied by an examination of the statistical properties of RISK. The game is modeled as a discrete time dynamic network graph, with the various features of the game modeled as properties of the network at a given time. As the network is computationally intensive to implement, results are produced …
Technology Design: The Movement Of Means, Yu Gu
Technology Design: The Movement Of Means, Yu Gu
Open Educational Resources
In order to promote students’ conceptual understanding and learning experience in introductory statistics, a technology task, which focuses on the probability distribution in which means are defined, was created using TinkerPlots, an exploratory dataanalysis and modeling software. The targeted audiences range from senior high school grade levels to college freshmen who are starting their introductory course in statistics. Students will be guided to explore and discover the movement behaviors of means of a set of numbers randomly generated from a fixed range of values characterized by a predetermined probability distribution. The cognitive, mathematical, technological and pedagogical natures of the task, …
A Statistical Approach To Characterize And Detect Degradation Within The Barabasi-Albert Network, Mohd-Fairul Mohd-Zaid
A Statistical Approach To Characterize And Detect Degradation Within The Barabasi-Albert Network, Mohd-Fairul Mohd-Zaid
Theses and Dissertations
Social Network Analysis (SNA) is widely used by the intelligence community when analyzing the relationships between individuals within groups of interest. Hence, any tools that can be quantitatively shown to help improve the analyses are advantageous for the intelligence community. To date, there have been no methods developed to characterize a real world network as a Barabasi-Albert network which is a type of network with properties contained in many real-world networks. In this research, two newly developed statistical tests using the degree distribution and the L-moments of the degree distribution are proposed with application to classifying networks and detecting degradation …
Jmasm35: A Percentile-Based Power Method: Simulating Multivariate Non-Normal Continuous Distributions (Sas), Jennifer Koran, Todd C. Headrick
Jmasm35: A Percentile-Based Power Method: Simulating Multivariate Non-Normal Continuous Distributions (Sas), Jennifer Koran, Todd C. Headrick
Journal of Modern Applied Statistical Methods
The conventional power method transformation is a moment-matching technique that simulates non-normal distributions with controlled measures of skew and kurtosis. The percentile-based power method is an alternative that uses the percentiles of a distribution in lieu of moments. This article presents a SAS/IML macro that implements the percentile-based power method.
Implementation And Validation Of A Probabilistic Open Source Baseball Engine (Posbe): Modeling Hitters And Pitchers, Rhett Tracy Schaefer
Implementation And Validation Of A Probabilistic Open Source Baseball Engine (Posbe): Modeling Hitters And Pitchers, Rhett Tracy Schaefer
Open Access Theses
This manuscript details the implementation and validation of an open source probabilistic baseball engine (POSBE) that focuses on the hitter and pitcher model of the simulation. The simulation produced outcomes that parallel those observed in actual professional Major League Baseball games. The observed data were taken from the nineteen games played between the New York Yankees (NYY) and Boston Red Sox (BOS) during the 2015 season. The potential hitter/pitcher outcomes of interest were singles, doubles, triples, homeruns, walks, hit-by-pitch, and strikeouts. The nineteen game series was simulated 1000 times, resulting in a total of 19,000 simulations. The eighteen hitters and …
Determining The Optimal Work Breakdown Structure For Government Acquisition Contracts, Brian J. Fitzpatrick
Determining The Optimal Work Breakdown Structure For Government Acquisition Contracts, Brian J. Fitzpatrick
Theses and Dissertations
The optimal level of Government Contract Work Breakdown Structure (G-CWBS) reporting for the purposes of Earned Value Management was inspected. The G-Score Metric was proposed, which can quantitatively grade a G-CWBS, based on a new method of calculating an Estimate At Completion (EAC) cost for each reported element. A random program generator created in R replicated the characteristics of DOD program artifacts retrieved from the Cost Analysis Data Enterprise (CADE) system. The generated artifacts were validated as a population, however validation at the demographic combination level using an artificial neural network was inconclusive. Comparative WBS forms were created for a …
A Recommendation System For Meta-Modeling: A Meta-Learning Based Approach, Can Cui, Mengqi Hu, Jeffery D. Weir, Teresa Wu
A Recommendation System For Meta-Modeling: A Meta-Learning Based Approach, Can Cui, Mengqi Hu, Jeffery D. Weir, Teresa Wu
Faculty Publications
Various meta-modeling techniques have been developed to replace computationally expensive simulation models. The performance of these meta-modeling techniques on different models is varied which makes existing model selection/recommendation approaches (e.g., trial-and-error, ensemble) problematic. To address these research gaps, we propose a general meta-modeling recommendation system using meta-learning which can automate the meta-modeling recommendation process by intelligently adapting the learning bias to problem characterizations. The proposed intelligent recommendation system includes four modules: (1) problem module, (2) meta-feature module which includes a comprehensive set of meta-features to characterize the geometrical properties of problems, (3) meta-learner module which compares the performance of instance-based …
Simulating Longer Vectors Of Correlated Binary Random Variables Via Multinomial Sampling, Justine Shults
Simulating Longer Vectors Of Correlated Binary Random Variables Via Multinomial Sampling, Justine Shults
UPenn Biostatistics Working Papers
The ability to simulate correlated binary data is important for sample size calculation and comparison of methods for analysis of clustered and longitudinal data with dichotomous outcomes. One available approach for simulating length n vectors of dichotomous random variables is to sample from the multinomial distribution of all possible length n permutations of zeros and ones. However, the multinomial sampling method has only been implemented in general form (without first making restrictive assumptions) for vectors of length 2 and 3, because specifying the multinomial distribution is very challenging for longer vectors. I overcome this difficulty by presenting an algorithm for …