Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Biostatistics (151)
- Medicine and Health Sciences (141)
- Social and Behavioral Sciences (134)
- Applied Statistics (131)
- Public Health (100)
-
- Mathematics (95)
- Life Sciences (94)
- Epidemiology (81)
- Statistical Models (72)
- Computer Sciences (68)
- Applied Mathematics (51)
- Engineering (49)
- Environmental Public Health (43)
- Statistical Methodology (42)
- Public Affairs, Public Policy and Public Administration (41)
- Nutrition (40)
- Health Services Research (37)
- Statistical Theory (36)
- Public Health Education and Promotion (35)
- Occupational Health and Industrial Hygiene (34)
- Health Policy (32)
- Women's Health (32)
- Multivariate Analysis (29)
- Business (28)
- Education (28)
- Probability (28)
- Categorical Data Analysis (27)
- Medical Specialties (26)
- Institution
-
- University of Kentucky (45)
- Wayne State University (33)
- Universitas Indonesia (32)
- Marquette University (29)
- West Virginia University (25)
-
- University of South Carolina (21)
- Southern Methodist University (19)
- Missouri University of Science and Technology (15)
- Himmelfarb Health Sciences Library, The George Washington University (14)
- Old Dominion University (13)
- Utah State University (13)
- COBRA (12)
- City University of New York (CUNY) (12)
- Georgia Southern University (11)
- University of Nebraska - Lincoln (11)
- University of South Florida (11)
- Air Force Institute of Technology (10)
- Prairie View A&M University (10)
- Illinois State University (9)
- Kennesaw State University (9)
- Michigan Technological University (9)
- Montclair State University (9)
- University of Nevada, Las Vegas (9)
- University of Texas at El Paso (9)
- Claremont Colleges (8)
- University of Arkansas, Fayetteville (8)
- The University of Akron (7)
- Louisiana State University (6)
- University of New Mexico (6)
- Western University (6)
- Keyword
-
- Statistics (33)
- Humans (12)
- Female (10)
- Machine learning (9)
- Male (9)
-
- Machine Learning (8)
- Regression (7)
- Characterizations (5)
- Epidemiology (5)
- Logistic regression (5)
- Physical Sciences and Mathematics, Statistics (5)
- Prediction (5)
- Probability (5)
- Simulation (5)
- Adolescent (4)
- Aged (4)
- Bayesian (4)
- Big data (4)
- Classification (4)
- Clinical trials (4)
- Clustering (4)
- Deep learning (4)
- Dietary inflammatory index (4)
- EM algorithm (4)
- Education (4)
- Family (4)
- Inflammation (4)
- Item response theory (4)
- MCMC (4)
- Mathematics (4)
- Publication
-
- Kesmas (32)
- Journal of Modern Applied Statistical Methods (30)
- Theses and Dissertations (27)
- Faculty & Staff Scholarship (24)
- Mathematics, Statistics and Computer Science Faculty Research and Publications (23)
-
- Electronic Theses and Dissertations (20)
- SMU Data Science Review (14)
- Epidemiology Faculty Publications (13)
- Mathematics and Statistics Faculty Research & Creative Works (11)
- Applications and Applied Mathematics: An International Journal (AAM) (10)
- Faculty Publications (10)
- Dissertations, Master's Theses and Master's Reports (9)
- Open Access Theses & Dissertations (9)
- USF Tampa Graduate Theses and Dissertations (9)
- Graduate Theses and Dissertations (8)
- Biostatistics Faculty Publications (7)
- Department of Applied Mathematics and Statistics Faculty Scholarship and Creative Works (7)
- Williams Honors College, Honors Research Projects (7)
- All Graduate Theses and Dissertations, Spring 1920 to Summer 2023 (6)
- Publications and Research (6)
- Theses and Dissertations--Statistics (6)
- Biostatistics, Epidemiology & Environmental Health Sciences: Faculty Publications (5)
- Chulalongkorn University Theses and Dissertations (Chula ETD) (5)
- College of Graduate Studies: Theses & Dissertations (5)
- Department of Statistics: Faculty Publications (5)
- Epidemiology and Biostatistics Publications (5)
- Mathematical and Statistical Science Faculty Research and Publications (5)
- Statistical and Data Sciences: Faculty Publications (5)
- Annual Symposium on Biomathematics and Ecology Education and Research (4)
- Internal Medicine Faculty Publications (4)
- Publication Type
- File Type
Articles 151 - 180 of 596
Full-Text Articles in Statistics and Probability
Dynamics Of Paramagnetic And Ferromagnetic Ellipsoidal Particles In Shear Flow Under A Uniform Magnetic Field, Christopher A. Sobecki, Jie Zhang, Yanzhi Zhang, Cheng Wang
Dynamics Of Paramagnetic And Ferromagnetic Ellipsoidal Particles In Shear Flow Under A Uniform Magnetic Field, Christopher A. Sobecki, Jie Zhang, Yanzhi Zhang, Cheng Wang
Mathematics and Statistics Faculty Research & Creative Works
We investigate the two-dimensional dynamic motion of magnetic particles of ellipsoidal shapes in shear flow under the influence of a uniform magnetic field. In the first part, we present a theoretical analysis of the rotational dynamics of the particles in simple shear flow. By considering paramagnetic and ferromagnetic particles, we study the effects of the direction and strength of the magnetic field on the particle rotation. The critical magnetic-field strength, at which particle rotation is impeded, is determined. In a weak-field regime (i.e., below the critical strength) where the particles execute complete rotations, the symmetry property of the rotational velocity …
Wald Confidence Intervals For A Single Poisson Parameter And Binomial Misclassification Parameter When The Data Is Subject To Misclassification, Nishantha Janith Chandrasena Poddiwala Hewage
Wald Confidence Intervals For A Single Poisson Parameter And Binomial Misclassification Parameter When The Data Is Subject To Misclassification, Nishantha Janith Chandrasena Poddiwala Hewage
Electronic Theses and Dissertations
This thesis is based on a Poisson model that uses both error-free data and error-prone data subject to misclassification in the form of false-negative and false-positive counts. We present maximum likelihood estimators (MLEs), Fisher's Information, and Wald statistics for Poisson rate parameter and the two misclassification parameters. Next, we invert the Wald statistics to get asymptotic confidence intervals for Poisson rate parameter and false-negative rate parameter. The coverage and width properties for various sample size and parameter configurations are studied via a simulation study. Finally, we apply the MLEs and confidence intervals to one real data set and another realistic …
A Comparison Of R, Sas, And Python Implementations Of Random Forests, Breckell Soifua
A Comparison Of R, Sas, And Python Implementations Of Random Forests, Breckell Soifua
All Graduate Plan B and other Reports, Spring 1920 to Spring 2023
The Random Forest method is a useful machine learning tool developed by Leo Breiman. There are many existing implementations across different programming languages; the most popular of which exist in R, SAS, and Python. In this paper, we conduct a comprehensive comparison of these implementations with regards to the accuracy, variable importance measurements, and timing. This comparison was done on a variety of real and simulated data with different classification difficulty levels, number of predictors, and sample sizes. The comparison shows unexpectedly different results between the three implementations.
Coastal Wetland Dynamics Under Sea-Level Rise And Wetland Restoration In The Northern Gulf Of Mexico Using Bayesian Multilevel Models And A Web Tool, Tyler Hardy
Master's Theses
There is currently a lack of modeling framework to predict how relative sea-level rise (SLR), combined with restoration activities, affects landscapes of coastal wetlands with uncertainties accounted for at the entire northern Gulf of Mexico (NGOM). I developed such a modeling framework – Bayesian multi-level models to study the spatial pattern of wetland loss in the NGOM, driven by relative RSLR, vegetation productivity, tidal range, coastal slope, and wave height – all interacting with river-borne sediment availability, indicated by hydrological regimes. These interactions have not been comprehensively investigated before. I further modified this model to assess the efficacy of restoration …
A Programme For Risk Assessment And Minimisation Of Progressive Multifocal Leukoencephalopathy Developed For Vedolizumab Clinical Trials, Asit Parikh, Kristin Stephens, Eugene Major, Irving Fox, Catherine Milch, Serap Sankoh, Michael H. Lev, James M. Provenzale, Jesse Shick, Mark Patti, Megan Mcauliffe, Joseph R. Berger, David B. Clifford
A Programme For Risk Assessment And Minimisation Of Progressive Multifocal Leukoencephalopathy Developed For Vedolizumab Clinical Trials, Asit Parikh, Kristin Stephens, Eugene Major, Irving Fox, Catherine Milch, Serap Sankoh, Michael H. Lev, James M. Provenzale, Jesse Shick, Mark Patti, Megan Mcauliffe, Joseph R. Berger, David B. Clifford
Neurology Faculty Publications
Introduction Over the past decade, the potential for drug-associated progressive multifocal leukoencephalopathy (PML) has become an increasingly important consideration in certain drug development programmes, particularly those of immunomodulatory biologics. Whether the risk of PML with an investigational agent is proven (e.g. extrapolated from relevant experience, such as a class effect) or merely theoretical, the serious consequences of acquiring PML require careful risk minimisation and assessment. No single standard for such risk minimisation exists. Vedolizumab is a recently developed monoclonal antibody to α4β7 integrin. Its clinical development necessitated a dedicated PML risk minimisation assessment as part of a global preapproval regulatory …
Bayesian Analytical Approaches For Metabolomics : A Novel Method For Molecular Structure-Informed Metabolite Interaction Modeling, A Novel Diagnostic Model For Differentiating Myocardial Infarction Type, And Approaches For Compound Identification Given Mass Spectrometry Data., Patrick J. Trainor
Electronic Theses and Dissertations
Metabolomics, the study of small molecules in biological systems, has enjoyed great success in enabling researchers to examine disease-associated metabolic dysregulation and has been utilized for the discovery biomarkers of disease and phenotypic states. In spite of recent technological advances in the analytical platforms utilized in metabolomics and the proliferation of tools for the analysis of metabolomics data, significant challenges in metabolomics data analyses remain. In this dissertation, we present three of these challenges and Bayesian methodological solutions for each. In the first part we develop a new methodology to serve a basis for making higher order inferences in metabolomics, …
Generalized Spatiotemporal Modeling And Causal Inference For Assessing Treatment Effects For Multiple Groups For Ordinal Outcome., Soutik Ghosal
Generalized Spatiotemporal Modeling And Causal Inference For Assessing Treatment Effects For Multiple Groups For Ordinal Outcome., Soutik Ghosal
Electronic Theses and Dissertations
This dissertation consists of three projects and can be categorized in two broad research areas: generalized spatiotemporal modeling and causal inference based on observational data. In the first project, I introduce a Bayesian hierarchical mixed effect hurdle model with a nested random effect structure to model the count for primary care providers and understand their spatial and temporal variation. This study further enables us to identify the health professional shortage areas and the possible impacting factors. In the second project, I have unified popular parametric and nonparametric propensity score-based methods to assess the treatment effect of multiple groups for ordinal …
Empirical Bayesian Approach To Testing Multiple Hypotheses With Separate Priors For Left And Right Alternatives, Naveen K. Bansal, Mehdi Maadooliat, Steven J. Schrodi
Empirical Bayesian Approach To Testing Multiple Hypotheses With Separate Priors For Left And Right Alternatives, Naveen K. Bansal, Mehdi Maadooliat, Steven J. Schrodi
Mathematics, Statistics and Computer Science Faculty Research and Publications
We consider a multiple hypotheses problem with directional alternatives in a decision theoretic framework. We obtain an empirical Bayes rule subject to a constraint on mixed directional false discovery rate (mdFDR≤α) under the semiparametric setting where the distribution of the test statistic is parametric, but the prior distribution is nonparametric. We proposed separate priors for the left tail and right tail alternatives as it may be required for many applications. The proposed Bayes rule is compared through simulation against rules proposed by Benjamini and Yekutieli and Efron. We illustrate the proposed methodology for two sets of …
Reducing Effects Of Dispersal On The Bias Of 2-Sample Mark-Recapture Estimators Of Stream Fish Abundance, James N. Mcnair, Carl R. Ruetz Iii, Ariana Carlson, Jiyeon Suh
Reducing Effects Of Dispersal On The Bias Of 2-Sample Mark-Recapture Estimators Of Stream Fish Abundance, James N. Mcnair, Carl R. Ruetz Iii, Ariana Carlson, Jiyeon Suh
Open Access Publishing Support Funded Articles
The 2-sample mark-recapture method with Chapman’s estimator is often used by inland fishery managers to estimate the reach-scale abundance of stream fish. An important assumption of this method is that no dispersal into or out of the study reach occurs between the two samples. Violations of this assumption are probably common in practice, but their effect on bias (systematic error) of abundance estimates is poorly understood, especially in small populations. Estimation methods permitting dispersal exist but, for logistical reasons, often are infeasible for routine assessments in streams. The purpose of this paper is to extend available results regarding effects of …
Surprise Vs. Probability As A Metric For Proof, Edward K. Cheng, Matthew Ginther
Surprise Vs. Probability As A Metric For Proof, Edward K. Cheng, Matthew Ginther
Vanderbilt Law School Faculty Publications
In this Symposium issue celebrating his career, Professor Michael Risinger in Leveraging Surprise proposes using "the fundamental emotion of surprise" as a way of measuring belief for purposes of legal proof. More specifically, Professor Risinger argues that we should not conceive of the burden of proof in terms of probabilities such as 51%, 95%, or even "beyond a reasonable doubt." Rather, the legal system should reference the threshold using "words of estimative surprise" -asking jurors how surprised they would be if the fact in question were not true. Toward this goal (and being averse to cardinality), he suggests categories such …
Distribution Of A Sum Of Random Variables When The Sample Size Is A Poisson Distribution, Mark Pfister
Distribution Of A Sum Of Random Variables When The Sample Size Is A Poisson Distribution, Mark Pfister
Electronic Theses and Dissertations
A probability distribution is a statistical function that describes the probability of possible outcomes in an experiment or occurrence. There are many different probability distributions that give the probability of an event happening, given some sample size n. An important question in statistics is to determine the distribution of the sum of independent random variables when the sample size n is fixed. For example, it is known that the sum of n independent Bernoulli random variables with success probability p is a Binomial distribution with parameters n and p: However, this is not true when the sample size …
The Expected Number Of Patterns In A Random Generated Permutation On [N] = {1,2,...,N}, Evelyn Fokuoh
The Expected Number Of Patterns In A Random Generated Permutation On [N] = {1,2,...,N}, Evelyn Fokuoh
Electronic Theses and Dissertations
Previous work by Flaxman (2004) and Biers-Ariel et al. (2018) focused on the number of distinct words embedded in a string of words of length n. In this thesis, we will extend this work to permutations, focusing on the maximum number of distinct permutations contained in a permutation on [n] = {1,2,...,n} and on the expected number of distinct permutations contained in a random permutation on [n]. We further considered the problem where repetition of subsequences are as a result of the occurrence of (Type A and/or Type B) replications. Our method of enumerating the Type A replications causes double …
Comparison Of Correlation, Partial Correlation, And Conditional Mutual Information For Interaction Effects Screening In Generalized Linear Models, Ji Li
Graduate Theses and Dissertations
Numerous screening techniques have been developed in recent years for genome-wide association studies (GWASs) (Moore et al., 2010). In this thesis, a novel model-free screening method was developed and validated by an extensive simulation study. Many screening methods were mainly focused on main effects, while very few studies considered the models containing both main effects and interaction effects. In this work, the interaction effects were fully considered and three different methods (Pearson’s Correlation Coefficient, Partial Correlation, and Conditional Mutual Information) were tested and their prediction accuracies were compared.
Pearson’s Correlation Coefficient method, which is a direct interaction screening (DIS) procedure, …
Clustering Mixed Data: An Extension Of The Gower Coefficient With Weighted L2 Distance, Augustine Oppong
Clustering Mixed Data: An Extension Of The Gower Coefficient With Weighted L2 Distance, Augustine Oppong
Electronic Theses and Dissertations
Sorting out data into partitions is increasing becoming complex as the constituents of data is growing outward everyday. Mixed data comprises continuous, categorical, directional functional and other types of variables. Clustering mixed data is based on special dissimilarities of the variables. Some data types may influence the clustering solution. Assigning appropriate weight to the functional data may improve the performance of the clustering algorithm. In this paper we use the extension of the Gower coefficient with judciously chosen weight for the L2 to cluster mixed data.The benefits of weighting are demonstrated both in in applications to the Buoy data set …
Adapting To Sparsity And Heavy Tailed Data, Mohamed Abdelkader Abba
Adapting To Sparsity And Heavy Tailed Data, Mohamed Abdelkader Abba
Graduate Theses and Dissertations
The Lasso and the Horseshoe, gold-standards in the frequentist and Bayesian paradigms, critically depend on learning the error variance. This causes a lack of scale invariance and adaptability to heavy-tailed data. The √ Lasso [Belloni et al., 2011] attempt to correct this by using the `1 norm on both the likelihood and the penalty for the objective function. In contrast, there is essentially no methods for uncertainty quantification or automatic parameter tuning via a formal Bayesian treatment of an unknown error distribution. On the other hand, Bayesian shrinkage priors lacking a local shrinkage term fails to adapt to the large …
Generalized Non-Inferential Approach To Modeling Restricted Discrete Choice For The Case Of The Spatial Random Utility, Elena Labzina
Generalized Non-Inferential Approach To Modeling Restricted Discrete Choice For The Case Of The Spatial Random Utility, Elena Labzina
Arts & Sciences Graduate Student Theses and Dissertations
Multinomial logistic regression model (MNL) is a powerful and easily tractable way for measuring the probabilistic impact of input variables on individual categorical choices. Crucially, the standard MNL assumes that all subjects of the study have the same choice sets. In the meanwhile, especially in political science and economics, this condition is frequently violated. Probably, the most graphical example of varying choice sets (VCS) is partially contested elections. Furthermore, the MNL implicitly implies the Independence of the Irregular Alternatives (IIA) assumption by requiring i.i.d errors that contrasts the MNL and the multinomial probit (MNP) and mixed logit (MXL) models. In …
Implementing The Use Of Personal Activity Data In An Introductory Statistics Course, Lacy Christensen
Implementing The Use Of Personal Activity Data In An Introductory Statistics Course, Lacy Christensen
All Graduate Theses and Dissertations, Spring 1920 to Summer 2023
Integrating real data into a classroom is one of the recommendations in the Guidelines for Assessment and Instruction in Statistics Education (GAISE) college report which lays out guidelines for an introductory statistics course (Committee, GAISE College Report ASA Revision, 2016). In order to assess the effect of using real data in a classroom, the students received physical activity trackers to wear during an undergraduate introductory statistics course taught in the summer. This tracker, a Fitbit, enabled students to monitor and record their steps, calories, and active time throughout the class. Collecting personal activity data (PAD) creates a large database which …
Design, Development And Construct Validation Of The Children’S Dietary Inflammatory Index, Samira Khan, Micheal D. Wirth, Andrew Ortaglia, Christian R. Alvarado, Nitin Shivappa, Thomas Hurley, James R. Hébert
Design, Development And Construct Validation Of The Children’S Dietary Inflammatory Index, Samira Khan, Micheal D. Wirth, Andrew Ortaglia, Christian R. Alvarado, Nitin Shivappa, Thomas Hurley, James R. Hébert
Faculty Publications
Objective: To design and validate a literature-derived, population-based Children’s Dietary Inflammatory Index (C-DII)TM. Design: The C-DII was developed based on a review of literature through 2010. Dietary data obtained from children in 16 different countries were used to create a reference database for computing C-DII scores based on consumption of macronutrients, vitamins, minerals, and whole foods. Construct validation was performed using quantile regression to assess the association between C-reactive protein (CRP) concentrations and C-DII scores. Data Sources: All data used for construct validation were obtained from children between six and 14 years of age (n = 3300) who participated in …
A Math Research Project Inspired By Twin Motherhood, Tiffany N. Kolba
A Math Research Project Inspired By Twin Motherhood, Tiffany N. Kolba
Journal of Humanistic Mathematics
The phenomenon of twins, triplets, quadruplets, and other higher order multiples has fascinated humans for centuries and has even captured the attention of mathematicians who have sought to model the probabilities of multiple births. However, there has not been extensive research into the phenomenon of polyovulation, which is one of the biological mechanisms that produces multiple births. In this paper, I describe how my own experience becoming a mother to twins led me on a quest to better understand the scientific processes going on inside my own body and motivated me to conduct research on polyovulation frequencies. An overview of …
Pretrial Release And Failure-To-Appear In Mclean County, Il, Jonathan Monsma
Pretrial Release And Failure-To-Appear In Mclean County, Il, Jonathan Monsma
Student Research – Stevenson Center
Actuarial risk assessment tools increasingly have been employed in jurisdictions across the U.S. to assist courts in the decision of whether someone charged with a crime should be detained or released prior to their trial. These tools should be continually monitored and researched by independent 3rd parties to ensure that these powerful tools are being administered properly and used in the most proficient way as to provide socially optimal results. McLean County, Illinois began using the Public Safety Assessment-CourtTM (PSA-Court or simply PSA) risk assessment tool beginning in 2016. This study culls data from the McLean County Jail …
A Distance Based Method For Solving Multi-Objective Optimization Problems, Murshid Kamal, Syed Aqib Jalil, Syed Mohd Muneeb, Irfan Ali
A Distance Based Method For Solving Multi-Objective Optimization Problems, Murshid Kamal, Syed Aqib Jalil, Syed Mohd Muneeb, Irfan Ali
Journal of Modern Applied Statistical Methods
A new model for the weighted method of goal programming is proposed based on minimizing the distances between ideal objectives to feasible objective space. It provides the best compromised solution for Multi Objective Linear Programming Problems (MOLPP). The proposed model tackles MOLPP by solving a series of single objective sub-problems, where the objectives are transformed into constraints. The compromise solution so obtained may be improved by defining priorities in terms of the weight. A criterion is also proposed for deciding the best compromise solution. Applications of the algorithm are discussed for transportation and assignment problems involving multiple and conflicting objectives. …
Goalie Analytics: Statistical Evaluation Of Context-Specific Goalie Performance Measures In The National Hockey League, Marc Naples, Logan Gage, Amy Nussbaum
Goalie Analytics: Statistical Evaluation Of Context-Specific Goalie Performance Measures In The National Hockey League, Marc Naples, Logan Gage, Amy Nussbaum
SMU Data Science Review
In this paper, we attempt to improve upon the classic formulation of save percentage in the NHL by controlling the context of the shots and use alternative measures than save percentage. In particular, we find save percentage to be both a weakly repeatable skill and predictor of future performance, and we seek other goalie performance calculations that are more robust. To do so, we use three primary tests to test intra-season consistency, intra-season predictability, and inter-season consistency, and extend the analysis to disentangle team effects on goalie statistics. We find that there are multiple ways to improve upon classic save …
Fuel Flow Reduction Impact Analysis Of Drag Reducing Film Applied To Aircraft Wings, Damon Resnick, Chris Donlan, Nimish Sakalle, Cody Pinkerman
Fuel Flow Reduction Impact Analysis Of Drag Reducing Film Applied To Aircraft Wings, Damon Resnick, Chris Donlan, Nimish Sakalle, Cody Pinkerman
SMU Data Science Review
In this paper, we present an analysis of flight data in order to determine whether the application of the Edge Aerodynamix Conformal Vortex Generator (CVG), applied to the wings of aircraft, reduces fuel flow during cruising conditions of flight. The CVG is a special treatment and film applied to the wings of an aircraft to protect the wings and reduce the non-laminar flow of air around the wings during flight. It is thought that by reducing the non-laminar flow or vortices around and directly behind the wings that an aircraft will move more smoothly through the air and provide a …
Data Center Application Security: Lateral Movement Detection Of Malware Using Behavioral Models, Harinder Pal Singh Bhasin, Elizabeth Ramsdell, Albert Alva, Rajiv Sreedhar, Medha Bhadkamkar
Data Center Application Security: Lateral Movement Detection Of Malware Using Behavioral Models, Harinder Pal Singh Bhasin, Elizabeth Ramsdell, Albert Alva, Rajiv Sreedhar, Medha Bhadkamkar
SMU Data Science Review
Data center security traditionally is implemented at the external network access points, i.e., the perimeter of the data center network, and focuses on preventing malicious software from entering the data center. However, these defenses do not cover all possible entry points for malicious software, and they are not 100% effective at preventing infiltration through the connection points. Therefore, security is required within the data center to detect malicious software activity including its lateral movement within the data center. In this paper, we present a machine learning-based network traffic analysis approach to detect the lateral movement of malicious software within the …
Predictions Generated From A Simulation Engine For Gene Expression Micro-Arrays For Use In Research Laboratories, Gopinath R. Mavankal, John Blevins, Dominique Edwards, Monnie Mcgee, Andrew Hardin
Predictions Generated From A Simulation Engine For Gene Expression Micro-Arrays For Use In Research Laboratories, Gopinath R. Mavankal, John Blevins, Dominique Edwards, Monnie Mcgee, Andrew Hardin
SMU Data Science Review
In this paper we introduce the technical components, the biology and data science involved in the use of microarray technology in biological and clinical research. We discuss how laborious experimental protocols involved in obtaining this data used in laboratories could benefit from using simulations of the data. We discuss the approach used in the simulation engine from [7]. We use this simulation engine to generate a prediction tool in Power BI, a Microsoft, business intelligence tool for analytics and data visualization [22]. This tool could be used in any laboratory using micro-arrays to improve experimental design by comparing how predicted …
Data Scientist’S Analysis Toolbox: Comparison Of Python, R, And Sas Performance, Jim Brittain, Mariana Cendon, Jennifer Nizzi, John Pleis
Data Scientist’S Analysis Toolbox: Comparison Of Python, R, And Sas Performance, Jim Brittain, Mariana Cendon, Jennifer Nizzi, John Pleis
SMU Data Science Review
A quantitative analysis will be performed on experiments utilizing three different tools used for Data Science. The analysis will include replication of analysis along with comparisons of code length, output, and results. Qualitative data will supplement the quantitative findings. The conclusion will provide data support guidance on the correct tool to use for common situations in the field of Data Science.
Predicting Game Day Outcomes In National Football League Games, Josh Klein, Anna Frowein, Chris Irwin
Predicting Game Day Outcomes In National Football League Games, Josh Klein, Anna Frowein, Chris Irwin
SMU Data Science Review
In this paper, we present a model for predicting the game day outcomes of National Football League games. 3 of the most popular sources for game day predictions are analyzed for comparison. Player data and outcomes from previous games are used, but we also incorporate several weather factors into our models. Over 1,700 games were incorporated and 3 separate models are created using simple regression, principal component analysis, and a recursive model. We also discuss the ethicality of using data science techniques by individuals with the knowledge in order to gain an advantage over a population lacking this specialized training.
Estimation Of Finite Population Mean By Using Minimum And Maximum Values In Stratified Random Sampling, Umer Daraz, Javid Shabbir, Hina Khan
Estimation Of Finite Population Mean By Using Minimum And Maximum Values In Stratified Random Sampling, Umer Daraz, Javid Shabbir, Hina Khan
Journal of Modern Applied Statistical Methods
In this paper we have suggested an improved class of ratio type estimators in estimating the finite population mean when information on minimum and maximum values of the auxiliary variable is known. The properties of the suggested class of estimators in terms of bias and mean square error are obtained up to first order of approximation. Two data sets are used for efficiency comparisons.
Cancerin: A Computational Pipeline To Infer Cancer-Associated Cerna Interaction Networks, Duc Do, Serdar Bozdag
Cancerin: A Computational Pipeline To Infer Cancer-Associated Cerna Interaction Networks, Duc Do, Serdar Bozdag
Mathematics, Statistics and Computer Science Faculty Research and Publications
MicroRNAs (miRNAs) inhibit expression of target genes by binding to their RNA transcripts. It has been recently shown that RNA transcripts targeted by the same miRNA could “compete” for the miRNA molecules and thereby indirectly regulate each other. Experimental evidence has suggested that the aberration of such miRNA-mediated interaction between RNAs—called competing endogenous RNA (ceRNA) interaction—can play important roles in tumorigenesis. Given the difficulty of deciphering context-specific miRNA binding, and the existence of various gene regulatory factors such as DNA methylation and copy number alteration, inferring context-specific ceRNA interactions accurately is a computationally challenging task. Here we propose a computational …
A Bayesian Beta-Mixture Model For Nonparametric Irt (Bbm-Irt), Ethan A. Arenson, George Karabatsos
A Bayesian Beta-Mixture Model For Nonparametric Irt (Bbm-Irt), Ethan A. Arenson, George Karabatsos
Journal of Modern Applied Statistical Methods
Item response models typically assume that the item characteristic (step) curves follow a logistic or normal cumulative distribution function, which are strictly monotone functions of person test ability. Such assumptions can be overly-restrictive for real item response data. A simple and more flexible Bayesian nonparametric IRT model for dichotomous items is introduced, which constructs monotone item characteristic (step) curves by a finite mixture of beta distributions, which can support the entire space of monotone curves to any desired degree of accuracy. An adaptive random-walk Metropolis-Hastings algorithm is proposed to estimate the posterior distribution of the model parameters. The Bayesian IRT …