Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

2019

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 301 - 330 of 596

Full-Text Articles in Statistics and Probability

K-Tuple Sampling From Partially Rank-Ordered Sets, Marvin Javier May 2019

K-Tuple Sampling From Partially Rank-Ordered Sets, Marvin Javier

UNLV Theses, Dissertations, Professional Papers, and Capstones

With the introduction of Ranked Set Sampling (RSS), McIntyre (1952) demonstrated that using ranking information to select units for measurement can lead to estimators with reduced variance when compared to their counterparts based on a simple random sample of the same size. This is done by selecting a set of units, and without direct measurement, ranking the units in the set before identifying one unit for measurement. This ranking of the units can be done through judgement ranking (such as visual assessment), or by using a correlated auxiliary variable.

In its original form, RSS does not allow for ties when …


Health Disparities Among Sexual And Gender Minorities, Jennifer Keeley May 2019

Health Disparities Among Sexual And Gender Minorities, Jennifer Keeley

UNLV Theses, Dissertations, Professional Papers, and Capstones

Decades of research has shown that sexual and gender minorities (SGMs) experience adverse health and mental health outcomes to a greater extent than their heterosexual peers. The need to better understand and eliminate health disparities in the SGM population was recognized by the National Institute on Minority Health and Health Disparities (NIMHD) at NIH. The Secretary of Health at the Department of Health and Human Services approved the designation of the SGM population as a health disparities population in 2016 and called for SGM studies to examine the health needs of the SGM population across SGM subgroups via large representative …


The Effects Of Injury On Running Status And Gait Biomechanics, Kristyne Wiegand May 2019

The Effects Of Injury On Running Status And Gait Biomechanics, Kristyne Wiegand

UNLV Theses, Dissertations, Professional Papers, and Capstones

Over the past several decades, endurance running has grown steadily as a popular form of physical activity. Running is easily accessible, does not require expensive equipment, and can be performed without specific skill training. Individuals who run also experience health benefits like increased cardiovascular health and reduced risk of all-cause morbidity. Despite these benefits, running is also associated with high rates of musculoskeletal injury. Although researchers have attempted to identify injury risks and mitigate the incidence of running injury, there is still no consensus as to why runners become injured. Research has also attempted to identify biomechanical movement patterns that …


Dynamic Sampling Versions Of Popular Spc Charts For Big Data Analysis, Samuel Anyaso-Samuel May 2019

Dynamic Sampling Versions Of Popular Spc Charts For Big Data Analysis, Samuel Anyaso-Samuel

Boise State University Theses and Dissertations

The statistical process control (SPC) chart is an effective tool for the analysis, interpretation, and visualization of data from sequential processes. Commonly used SPC charts such as the Shewhart, CUSUM and EWMA charts are widely implemented in detecting distributional shifts in various processes. With recent scientific and technological advancements, massive amounts of data continue to be generated by production, medical, agricultural and many other industrial processes. Conventional SPC charts have significant drawbacks in monitoring such processes, specifically when the velocity of the data flow is greater than the run time of the monitoring procedure. In the literature, dynamic sampling control …


Contributions To Mcmc Methods In Constrained Domains With Applications To Neuroimaging, Sharang Chaudhry May 2019

Contributions To Mcmc Methods In Constrained Domains With Applications To Neuroimaging, Sharang Chaudhry

UNLV Theses, Dissertations, Professional Papers, and Capstones

Markov chain Monte Carlo (MCMC) methods form a rich class of computational techniques that help its user ascertain samples from target distributions when direct sampling is not possible or when their closed forms are intractable. Over the years, MCMC methods have been used in innumerable situations due to their flexibility and generalizability, even in situations involving nonlinear and/or highly parametrized models. In this dissertation, two major works relating to MCMC methods are presented.

The first involves the development of a method to identify the number and directions of nerve fibers using diffusion-weighted MRI measurements. For this, the biological problem is …


Is Animal-Based Biomedical Research Being Used In Its Original Context?, Constança Carvalho, Daniel Alves, Andrew Knight, Luís Vicente Apr 2019

Is Animal-Based Biomedical Research Being Used In Its Original Context?, Constança Carvalho, Daniel Alves, Andrew Knight, Luís Vicente

Validation of Animal Experimentation Collection

No abstract provided.


Critically Evaluating Animal Research, Andrew Knight Apr 2019

Critically Evaluating Animal Research, Andrew Knight

Validation of Animal Experimentation Collection

No abstract provided.


Deep Learning, Medical Physics And Cargo Cult Science., Miguel Romero Phd, Gilmer Valdes Phd, Timothy Solberg Phd, Yannet Interian Phd Apr 2019

Deep Learning, Medical Physics And Cargo Cult Science., Miguel Romero Phd, Gilmer Valdes Phd, Timothy Solberg Phd, Yannet Interian Phd

Creative Activity and Research Day - CARD

Deep learning algorithms have become widely popular, with considerable success in fields where datasets have hundreds of thousands or million points. As deep learning is increasingly applied to the fields of medical physics and radiation oncology, a reasonable question follows: are these techniques the best approach, given the unique conditions in our field? In this study, we investigate the dependence of dataset size on the performance of deep learning algorithms compared with more traditional radiomics-based methods.


Deep Neural Network Architectures For Music Genre Classification, Kai Middlebrook, Shyam Sudhakaran, Kunal Sonar, David Guy Brizan Apr 2019

Deep Neural Network Architectures For Music Genre Classification, Kai Middlebrook, Shyam Sudhakaran, Kunal Sonar, David Guy Brizan

Creative Activity and Research Day - CARD

With the recent advancements in technology, many tasks in fields such as computer vision, natural language processing, and signal processing have been solved using deep learning architectures. In the audio domain, these architectures have been used to learn musical features of songs to predict: moods, genres, and instruments. In the case of genre classification, deep learning models were applied to popular datasets--which are explicitly chosen to represent their genres--and achieved state-of-the-art results. However, these results have not been reproduced on less refined datasets. To this end, we introduce an un-curated dataset which contains genre labels and 30-second audio previews for …


Variance Heterogeneity In Psychological Research: A Monte Carlo Study Of The Consequences For Meta-Analysis, Bruce E. Blaine Apr 2019

Variance Heterogeneity In Psychological Research: A Monte Carlo Study Of The Consequences For Meta-Analysis, Bruce E. Blaine

Statistics Faculty/Staff Publications

Variance heterogeneity is common in psychological research. Surveys of psychological research show that variance ratios (VRs) in two-group studies average around 2.5, with a substantial minority of studies having much higher VRs. Research has established that variance heterogeneity disturbs Type I error rates of parametric tests in primary research. Fixed-effects meta-analysis is a common statistical method in psychology for synthesizing primary research, and plays an important role in cumulative science and evidence-based practice. Little is known about the consequences of variance heterogeneity for meta-analytic estimates. The present research reports a Monte Carlo study in which the results of k = …


A Gene-Based Recessive Diplotype Exome Scan Discovers Fgf6, A Novel Hepcidin-Regulating Iron-Metabolism Gene, Shicheng Guo, Shuai Jiang, Narendranath Epperla, Yanyun Ma, Mehdi Maadooliat, Zhan Ye, Brent Olson, Minghua Wang, Terrie Kitchner, Jeffrey Joyce, Peng An, Fudi Wang, Robert Strenn, Joseph J. Mazza, Jennifer K. Meece, Wenyu Wu, Li Jin, Judith A. Smith, Jiucan Wang, Steven J. Schrodi Apr 2019

A Gene-Based Recessive Diplotype Exome Scan Discovers Fgf6, A Novel Hepcidin-Regulating Iron-Metabolism Gene, Shicheng Guo, Shuai Jiang, Narendranath Epperla, Yanyun Ma, Mehdi Maadooliat, Zhan Ye, Brent Olson, Minghua Wang, Terrie Kitchner, Jeffrey Joyce, Peng An, Fudi Wang, Robert Strenn, Joseph J. Mazza, Jennifer K. Meece, Wenyu Wu, Li Jin, Judith A. Smith, Jiucan Wang, Steven J. Schrodi

Mathematical and Statistical Science Faculty Research and Publications

Standard analyses applied to genome-wide association data are well designed to detect additive effects of moderate strength. However, the power for standard genome-wide association study (GWAS) analyses to identify effects from recessive diplotypes is not typically high. We proposed and conducted a gene-based compound heterozygosity test to reveal additional genes underlying complex diseases. With this approach applied to iron overload, a strong association signal was identified between the fibroblast growth factor–encoding gene, FGF6, and hemochromatosis in the central Wisconsin population. Functional validation showed that fibroblast growth factor 6 protein (FGF-6) regulates iron homeostasis and induces transcriptional regulation of hepcidin. …


The Andersen Likelihood Ratio Test With A Random Split Criterion Lacks Power, Georg Krammer Apr 2019

The Andersen Likelihood Ratio Test With A Random Split Criterion Lacks Power, Georg Krammer

Journal of Modern Applied Statistical Methods

The Andersen LRT uses sample characteristics as split criteria to evaluate Rasch model fit, or theory driven hypothesis testing for a test. The power and Type I error of a random split criterion was evaluated with a simulation study. Results consistently show a random split criterion lacks power.


Deep Embedding Kernel, Linh Le Apr 2019

Deep Embedding Kernel, Linh Le

Doctor of Data Science and Analytics Dissertations

Kernel methods and deep learning are two major branches of machine learning that have achieved numerous successes in both analytics and artificial intelligence. While having their own unique characteristics, both branches work through mapping data to a feature space that is supposedly more favorable towards the given task. This dissertation addresses the strengths and weaknesses of each mapping method through combining them and forming a family of novel deep architectures that center around the Deep Embedding Kernel (DEK). In short, DEK is a realization of a kernel function through a newly deep architecture. The mapping in DEK is both implicit …


Weighted Version Of Generalized Inverse Weibull Distribution, Sofi Mudiasir, S. P. Ahmad Apr 2019

Weighted Version Of Generalized Inverse Weibull Distribution, Sofi Mudiasir, S. P. Ahmad

Journal of Modern Applied Statistical Methods

Weighted distributions are used in many fields, such as medicine, ecology, and reliability. A weighted version of the generalized inverse Weibull distribution, known as weighted generalized inverse Weibull distribution (WGIWD), is proposed. Basic properties including mode, moments, moment generating function, skewness, kurtosis, and Shannon’s entropy are studied. The usefulness of the new model was demonstrated by applying it to a real-life data set. The WGIWD fits better than its submodels, such as length biased generalized inverse Weibull (LGIW), generalized inverse Weibull (GIW), inverse Weibull (IW) and inverse exponential (IE) distributions.


Calibration Of Measurements, Edward Kroc, Bruno D. Zumbo Apr 2019

Calibration Of Measurements, Edward Kroc, Bruno D. Zumbo

Journal of Modern Applied Statistical Methods

Traditional notions of measurement error typically rely on a strong mean-zero assumption on the expectation of the errors conditional on an unobservable “true score” (classical measurement error) or on the data themselves (Berkson measurement error). Weakly calibrated measurements for an unobservable true quantity are defined based on a weaker mean-zero assumption, giving rise to a measurement model of differential error. Applications show it retains many attractive features of estimation and inference when performing a naive data analysis (i.e. when performing an analysis on the error-prone measurements themselves), and other interesting properties not present in the classical or Berkson cases. Applied …


Assessing The Efficacy Of Home-Based Renal Care Using Propensity Scores, Eunice Choi Apr 2019

Assessing The Efficacy Of Home-Based Renal Care Using Propensity Scores, Eunice Choi

Mathematics & Statistics ETDs

This study investigated the efficacy of Home-Based Renal Care (HBRC) in diabetic Zuni Indians with Chronic Kidney Disease (CKD) in New Mexico using propensity scores. Home based intervention as opposed to standard clinical care is a pragmatic treatment approach that incorporates the preference of population in hopes of addressing a cultural barrier to healthcare in this high risk population. This study uses a logistic regression model and a linear regression model to estimate the average effect of HBRC on increasing the likelihood of participants taking a more active role in the management of their chronic condition compared to the control …


Estimation Of Mean With Two-Parameter Ratio-Product-Ratio Estimator In Double Sampling Using Ancillary Information Under Non-Response, Surya K. Pal, Housila P. Singh Apr 2019

Estimation Of Mean With Two-Parameter Ratio-Product-Ratio Estimator In Double Sampling Using Ancillary Information Under Non-Response, Surya K. Pal, Housila P. Singh

Journal of Modern Applied Statistical Methods

Ratio-product-ratio estimators with two parameters in double sampling under non-response are considered along with their properties. Practical conditions are obtained in which the suggested estimators are more proficient than other existing estimators. An example is given.


Capturing Heterogeneity Of Covariate Effects In Hidden Subpopulations In The Presence Of Censoring And Large Number Of Covariates, Farhad Shokoohi, Abbas Khalili, Masoud Asgharian, Shili Lin Apr 2019

Capturing Heterogeneity Of Covariate Effects In Hidden Subpopulations In The Presence Of Censoring And Large Number Of Covariates, Farhad Shokoohi, Abbas Khalili, Masoud Asgharian, Shili Lin

Mathematical Sciences Faculty Research

The advent of modern technology has led to a surge of high-dimensional data in biology and health sciences such as genomics, epigenomics and medicine. The high-grade serous ovarian cancer (HGS-OvCa) data reported by The Cancer Genome Atlas (TCGA) Research Network is one example. The TCGA and other research groups have analyzed several aspects of these data. Here we study the relationship between Disease Free Time (DFT) after surgery among ovarian cancer patients and their DNA methylation profiles of genomic features. Such studies pose additional challenges beyond the typical big data problem due to population substructure and censoring. Despite the availability …


A More Powerful Unconditional Exact Test Of Homogeneity For 2 × C Contingency Table Analysis, Louis Ehwerhemuepha, Heng Sok, Cyril Rakovski Apr 2019

A More Powerful Unconditional Exact Test Of Homogeneity For 2 × C Contingency Table Analysis, Louis Ehwerhemuepha, Heng Sok, Cyril Rakovski

Mathematics, Physics, and Computer Science Faculty Articles and Research

The classical unconditional exact p-value test can be used to compare two multinomial distributions with small samples. This general hypothesis requires parameter estimation under the null which makes the test severely conservative. Similar property has been observed for Fisher's exact test with Barnard and Boschloo providing distinct adjustments that produce more powerful testing approaches. In this study, we develop a novel adjustment for the conservativeness of the unconditional multinomial exact p-value test that produces nominal type I error rate and increased power in comparison to all alternative approaches. We used a large simulation study to empirically estimate the …


Essays On Time Series And Machine Learning Techniques For Risk Management, Michael Kotarinos Apr 2019

Essays On Time Series And Machine Learning Techniques For Risk Management, Michael Kotarinos

USF Tampa Graduate Theses and Dissertations

The Capital Asset Pricing Model combined with the Sharpe ratio is a standard method for choosing assets for selection in a portfolio. However, this method has many structural issues and was designed for a time when high dimensional computing was in its infancy. An alternative to these methods using a mix of Multi-Level Time Series Clustering, the MACBETH algorithm and traditional time series techniques was constructed that minimized data loss and allow for customized portfolio construction for investors with different risk profiles and specialized investment needs. It was shown that these methods are adaptable to cloud computing environments and allow …


Incorporating Pathway Information Into Feature Selection Towards Better Performed Gene Signatures, Suyan Tian, Chi Wang, Bing Wang Apr 2019

Incorporating Pathway Information Into Feature Selection Towards Better Performed Gene Signatures, Suyan Tian, Chi Wang, Bing Wang

Biostatistics Faculty Publications

To analyze gene expression data with sophisticated grouping structures and to extract hidden patterns from such data, feature selection is of critical importance. It is well known that genes do not function in isolation but rather work together within various metabolic, regulatory, and signaling pathways. If the biological knowledge contained within these pathways is taken into account, the resulting method is a pathway-based algorithm. Studies have demonstrated that a pathway-based method usually outperforms its gene-based counterpart in which no biological knowledge is considered. In this article, a pathway-based feature selection is firstly divided into three major categories, namely, pathway-level selection, …


A Grammar For Reproducible And Painless Extract-Transform-Load Operations On Medium Data, Benjamin S. Baumer Apr 2019

A Grammar For Reproducible And Painless Extract-Transform-Load Operations On Medium Data, Benjamin S. Baumer

Statistical and Data Sciences: Faculty Publications

Many interesting datasets available on the Internet are of a medium size—too big to fit into a personal computer’s memory, but not so large that they would not fit comfortably on its hard disk. In the coming years, datasets of this magnitude will inform vital research in a wide array of application domains. However, due to a variety of constraints they are cumbersome to ingest, wrangle, analyze, and share in a reproducible fashion. These obstructions hamper thorough peer-review and thus disrupt the forward progress of science. We propose a predictable and pipeable framework for R (the state-of-the-art statistical computing environment) …


The Housing Bubble, Shelby Brown Apr 2019

The Housing Bubble, Shelby Brown

Mathematics Senior Capstone Papers

The housing market is constantly changing. Fluctuating housing prices and a flooded market have home buyers hesitant to commit and sellers on edge. What if the prospective buyer or seller could take this financial step already knowing the state of the market? The purpose of this project is to attempt to predict the next housing bubble. A multivariable regression analysis is conducted using relevant data including variables such as average property prices, number of foreclosures, etc. in the United States beginning in the year 2009. The trends, patterns, and models created from the regression analysis are compared against data models …


Angry Birds Fly High Again With Data Analytics, Singapore Management University Apr 2019

Angry Birds Fly High Again With Data Analytics, Singapore Management University

Perspectives@SMU

User feedback has transformed Rovio’s culture and game design


Historical Study Of The Relationship Between The Federal Funds Rate And The Inflation Rate, Aaron Wilkins Apr 2019

Historical Study Of The Relationship Between The Federal Funds Rate And The Inflation Rate, Aaron Wilkins

Undergraduate External Publications

It is believed that in order to control high inflation rates, the Federal Reserve Bank (“the Fed”) increases the federal funds rate and when the inflation rate gets low, the Fed takes the opposite approach. (The federal funds rate is a rate of interest that banks charge each other to lend funds and stay above the reserve requirement, set by the government.) This project examines the relationship between the federal funds rate and the inflation rate. Sixty-five years of historical inflation rates and federal funds rates were used as the basis for this exploration. Because a time lag between the …


The Evolution Of Data Science: A New Mode Of Knowledge Production, Jennifer Lewis Priestley, Robert J. Mcgrath Apr 2019

The Evolution Of Data Science: A New Mode Of Knowledge Production, Jennifer Lewis Priestley, Robert J. Mcgrath

Faculty Articles

Is data science a new field of study or simply an extension or specialization of a discipline that already exists, such as statistics, computer science, or mathematics? This article explores the evolution of data science as a potentially new academic discipline, which has evolved as a function of new problem sets that established disciplines have been ill-prepared to address. The authors find that this newly-evolved discipline can be viewed through the lens of a new mode of knowledge production and is characterized by transdisciplinarity collaboration with the private sector and increased accountability. Lessons from this evolution can inform knowledge production …


Measuring Birth Trauma Rates In Maine Using Public Data, Mike Lapika Apr 2019

Measuring Birth Trauma Rates In Maine Using Public Data, Mike Lapika

Thinking Matters Symposium Archive

An increasing number of states are creating databases that collect and organize health insurance claims from public and private health care payers. Since December 2016, at least 18 states have these “all-payer claims databases” (APCDs), including Maine. APCDs are intended to inform cost containment and quality improvement by increasing transparency and informing consumer choice. For this project, we assessed how Maine’s APCD data might be used to produce standardized quality measures across facilities in the state. Specifically, we tested a birth outcome quality measure developed by the Agency for Healthcare Research and Quality (AHRQ), Birth Trauma – Injury to Neonate …


Dice Mythbusters, C. Warren Campbell, William P. Dolan Apr 2019

Dice Mythbusters, C. Warren Campbell, William P. Dolan

Student Research Conference Select Presentations

All dice are unfair because they cannot be manufactured with absolute precision. However, some dice are more unfair than others. Each year hundreds of millions of dice are sold worldwide. Dice commonly used in role playing games are 4-sided (D4), 6-sided (D6), 8-sided (D8), 10-sided (D10), 12-sided (D12), and 20-sided (D20). Most of these are manufactured using plastic mold injection and rock tumbler methods. This method can result in dimensional inaccuracies in the dice and sometimes density inhomogeneities. In 3000-roll tests of eleven D20 dice only three tested fair. In a running chi square test it was shown that for …


Sensitivity Analyses For Tumor Growth Models, Ruchini Dilinika Mendis Apr 2019

Sensitivity Analyses For Tumor Growth Models, Ruchini Dilinika Mendis

Masters Theses & Specialist Projects

This study consists of the sensitivity analysis for two previously developed tumor growth models: Gompertz model and quotient model. The two models are considered in both continuous and discrete time. In continuous time, model parameters are estimated using least-square method, while in discrete time, the partial-sum method is used. Moreover, frequentist and Bayesian methods are used to construct confidence intervals and credible intervals for the model parameters. We apply the Markov Chain Monte Carlo (MCMC) techniques with the Random Walk Metropolis algorithm with Non-informative Prior and the Delayed Rejection Adoptive Metropolis (DRAM) algorithm to construct parameters' posterior distributions and then …


Beach Composition Preferences For Nesting Populations Of Leatherback Sea Turtles (Dermochelys Coriacea), Armila Beach, Guna Yala Comarca, Scott Campbell Apr 2019

Beach Composition Preferences For Nesting Populations Of Leatherback Sea Turtles (Dermochelys Coriacea), Armila Beach, Guna Yala Comarca, Scott Campbell

Independent Study Project (ISP) Collection

Sea turtles play a critical role in marine ecosystems all over the world, including the Caribbean Sea. However, many sea turtle species are under threat due to anthropogenic impacts, such as habitat destruction and fisheries bycatch. This has caused significant declines in sea turtle populations around the world, which in turn has impacted marine ecosystems where sea turtles play critical roles in proper ecosystem functioning. A crucial part of the sea turtle life cycle that has been threatened by anthropogenic factors is nesting. Sea turtles rely on unspoiled beaches with particular physical characteristics for laying their eggs. One of the …