Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 61 - 90 of 113

Full-Text Articles in Numerical Analysis and Scientific Computing

An Evaluation Of Training Size Impact On Validation Accuracy For Optimized Convolutional Neural Networks, Jostein Barry-Straume, Adam Tschannen, Daniel W. Engels, Edward Fine Jan 2019

An Evaluation Of Training Size Impact On Validation Accuracy For Optimized Convolutional Neural Networks, Jostein Barry-Straume, Adam Tschannen, Daniel W. Engels, Edward Fine

SMU Data Science Review

In this paper, we present an evaluation of training size impact on validation accuracy for an optimized Convolutional Neural Network (CNN). CNNs are currently the state-of-the-art architecture for object classification tasks. We used Amazon’s machine learning ecosystem to train and test 648 models to find the optimal hyperparameters with which to apply a CNN towards the Fashion-MNIST (Mixed National Institute of Standards and Technology) dataset. We were able to realize a validation accuracy of 90% by using only 40% of the original data. We found that hidden layers appear to have had zero impact on validation accuracy, whereas the neural …


Improving Vix Futures Forecasts Using Machine Learning Methods, James Hosker, Slobodan Djurdjevic, Hieu Nguyen, Robert Slater Jan 2019

Improving Vix Futures Forecasts Using Machine Learning Methods, James Hosker, Slobodan Djurdjevic, Hieu Nguyen, Robert Slater

SMU Data Science Review

The problem of forecasting market volatility is a difficult task for most fund managers. Volatility forecasts are used for risk management, alpha (risk) trading, and the reduction of trading friction. Improving the forecasts of future market volatility assists fund managers in adding or reducing risk in their portfolios as well as in increasing hedges to protect their portfolios in anticipation of a market sell-off event. Our analysis compares three existing financial models that forecast future market volatility using the Chicago Board Options Exchange Volatility Index (VIX) to six machine/deep learning supervised regression methods. This analysis determines which models provide best …


Modeling Stochastically Intransitive Relationships In Paired Comparison Data, Ryan Patrick Alexander Mcshane Jan 2019

Modeling Stochastically Intransitive Relationships In Paired Comparison Data, Ryan Patrick Alexander Mcshane

Statistical Science Theses and Dissertations

If the Warriors beat the Rockets and the Rockets beat the Spurs, does that mean that the Warriors are better than the Spurs? Sophisticated fans would argue that the Warriors are better by the transitive property, but could Spurs fans make a legitimate argument that their team is better despite this chain of evidence?

We first explore the nature of intransitive (rock-scissors-paper) relationships with a graph theoretic approach to the method of paired comparisons framework popularized by Kendall and Smith (1940). Then, we focus on the setting where all pairs of items, teams, players, or objects have been compared to …


On Cluster Robust Models, José Bayoán Santiago Calderón Jan 2019

On Cluster Robust Models, José Bayoán Santiago Calderón

CGU Theses & Dissertations

Cluster robust models are a kind of statistical models that attempt to estimate parameters considering potential heterogeneity in treatment effects. Absent heterogeneity in treatment effects, the partial and average treatment effect are the same. When heterogeneity in treatment effects occurs, the average treatment effect is a function of the various partial treatment effects and the composition of the population of interest. The first chapter explores the performance of common estimators as a function of the presence of heterogeneity in treatment effects and other characteristics that may influence their performance for estimating average treatment effects. The second chapter examines various approaches …


Microarray Data Analysis And Classification Of Cancers, Grant Gates Jan 2019

Microarray Data Analysis And Classification Of Cancers, Grant Gates

Williams Honors College, Honors Research Projects

When it comes to cancer, there is no standardized approach for identifying new cancer classes nor is there a standardized approach for assigning cancer tumors to existing classes. These two ideas are known as class discovery and class prediction. For a cancer patient to receive proper treatment, it is important that the type of cancer be accurately identified. For my Senior Honors Project, I would like to use this opportunity to research a topic in bioinformatics. Bioinformatics incorporates a few different subjects into one including biology, computer science and statistics. An intricate method for class discovery and class prediction is …


Snap Scholar: The User Experience Of Engaging With Academic Research Through A Tappable Stories Medium, Ieva Burk Jan 2019

Snap Scholar: The User Experience Of Engaging With Academic Research Through A Tappable Stories Medium, Ieva Burk

CMC Senior Theses

With the shift to learn and consume information through our mobile devices, most academic research is still only presented in long-form text. The Stanford Scholar Initiative has explored the segment of content creation and consumption of academic research through video. However, there has been another popular shift in presenting information from various social media platforms and media outlets in the past few years. Snapchat and Instagram have introduced the concept of tappable “Stories” that have gained popularity in the realm of content consumption.

To accelerate the growth of the creation of these research talks, I propose an alternative to video: …


Predicting River Stage Using Recurrent Neural Networks, Eric Rohli Jul 2018

Predicting River Stage Using Recurrent Neural Networks, Eric Rohli

LSU Master's Theses

River stage prediction is an important problem in the water transportation industry. Accurate river stage predictions provide crucial information to barge and tow boat operators, port terminal captains, and lock management officials. Shallow river levels caused by prolonged drought impact the loading capacity of barges and tow boats. High river levels caused by excessive rainfall or snowmelt allow for greater tow capacities but make downstream transportation and lock management risky. Current academic river height prediction systems utilize either time series statistical analysis or machine learning algorithms to forecast future river heights, but systems that combine these two areas often limit …


Physical Applications Of The Geometric Minimum Action Method, George L. Poppe Jr. May 2018

Physical Applications Of The Geometric Minimum Action Method, George L. Poppe Jr.

Dissertations, Theses, and Capstone Projects

This thesis extends the landscape of rare events problems solved on stochastic systems by means of the \textit{geometric minimum action method} (gMAM). These include partial differential equations (PDEs) such as the real Ginzburg-Landau equation (RGLE), the linear Schroedinger equation, along with various forms of the nonlinear Schroedinger equation (NLSE) including an application towards an ultra-short pulse mode-locked laser system (MLL).

Additionally we develop analytical tools that can be used alongside numerics to validate those solutions. This includes the use of instanton methods in deriving state transitions for the linear Schroedinger equation and the cubic diffusive NLSE.

These analytical solutions are …


Understanding Natural Keyboard Typing Using Convolutional Neural Networks On Mobile Sensor Data, Travis Siems Apr 2018

Understanding Natural Keyboard Typing Using Convolutional Neural Networks On Mobile Sensor Data, Travis Siems

Computer Science and Engineering Theses and Dissertations

Mobile phones and other devices with embedded sensors are becoming increasingly ubiquitous. Audio and motion sensor data may be able to detect information that we did not think possible. Some researchers have created models that can predict computer keyboard typing from a nearby mobile device; however, certain limitations to their experiment setup and methods compelled us to be skeptical of the models’ realistic prediction capability. We investigate the possibility of understanding natural keyboard typing from mobile phones by performing a well-designed data collection experiment that encourages natural typing and interactions. This data collection helps capture realistic vulnerabilities of the security …


Predicting The Next Us President By Simulating The Electoral College, Boyan Kostadinov Jan 2018

Predicting The Next Us President By Simulating The Electoral College, Boyan Kostadinov

Publications and Research

We develop a simulation model for predicting the outcome of the US Presidential election based on simulating the distribution of the Electoral College. The simulation model has two parts: (a) estimating the probabilities for a given candidate to win each state and DC, based on state polls, and (b) estimating the probability that a given candidate will win at least 270 electoral votes, and thus win the White House. All simulations are coded using the high-level, open-source programming language R. One of the goals of this paper is to promote computational thinking in any STEM field by illustrating how probabilistic …


The Accuracy, Fairness, And Limits Of Predicting Recidivism, Julie Dressel, Hany Farid Jan 2018

The Accuracy, Fairness, And Limits Of Predicting Recidivism, Julie Dressel, Hany Farid

Dartmouth Scholarship

Algorithms for predicting recidivism are commonly used to assess a criminal defendant’s likelihood of committing a crime. These predictions are used in pretrial, parole, and sentencing decisions. Proponents of these systems argue that big data and advanced machine learning make these analyses more accurate and less biased than humans. We show, however, that the widely used commercial risk assessment software COMPAS is no more accurate or fair than predictions made by people with little or no criminal justice expertise. We further show that a simple linear predictor provided with only two features is nearly equivalent to COMPAS with its 137 …


Non-Linear Machine Learning With Active Sampling For Mox Drift Compensation, Tamara Matthews, Muhammad Iqbal, Horacio Gonzalez-Velez Jan 2018

Non-Linear Machine Learning With Active Sampling For Mox Drift Compensation, Tamara Matthews, Muhammad Iqbal, Horacio Gonzalez-Velez

Conference papers

Abstract—Metal oxide (MOX) gas detectors based on SnO2 provide low-cost solutions for real-time sensing of complex gas mixtures for indoor ambient monitoring. With high sensitivity under ideal conditions, MOX detectors may have poor longterm response accuracy due to environmental factors (humidity and temperature) along with sensor aging, leading to calibration drifts. Finding a simple and efficient solution to correct such calibration drifts has been the subject of numerous studies but remains an open problem. In this work, we present an efficient approach to MOX calibration using active and transfer sampling techniques coupled with non-linear machine learning algorithms, namely neural networks, …


Algorithms For Reconstruction Of Gene Regulatory Networks From High -Throughput Gene Expression Data, Wenping Deng Jan 2018

Algorithms For Reconstruction Of Gene Regulatory Networks From High -Throughput Gene Expression Data, Wenping Deng

Dissertations, Master's Theses and Master's Reports

Understanding gene interactions in complex living systems is one of the central tasks in system biology. With the availability of microarray and RNA-Seq technologies, a multitude of gene expression datasets has been generated towards novel biological knowledge discovery through statistical analysis and reconstruction of gene regulatory networks (GRN). Reconstruction of GRNs can reveal the interrelationships among genes and identify the hierarchies of genes and hubs in networks. The new algorithms I developed in this dissertation are specifically focused on the reconstruction of GRNs with increased accuracy from microarray and RNA-Seq high-throughput gene expression data sets.

The first algorithm (Chapter 2) …


Wildfire Emissions In The Context Of Global Change And The Implications For Mercury Pollution, Aditya Kumar Jan 2018

Wildfire Emissions In The Context Of Global Change And The Implications For Mercury Pollution, Aditya Kumar

Dissertations, Master's Theses and Master's Reports

Wildfires are episodic disturbances that exert a significant influence on the Earth system. They emit substantial amounts of atmospheric pollutants, which can impact atmospheric chemistry/composition and the Earth’s climate at the global and regional scales. This work presents a collection of studies aimed at better estimating wildfire emissions of atmospheric pollutants, quantifying their impacts on remote ecosystems and determining the implications of 2000s-2050s global environmental change (land use/land cover, climate) for wildfire emissions following the Intergovernmental Panel on Climate Change (IPCC) A1B socioeconomic scenario.

A global fire emissions model is developed to compile global wildfire emission inventories for major atmospheric …


Perfect Ratings With Negative Comments: Learning From Contradictory Patient Survey Responses, Andrew S. Gallan, Marina Girju, Roxana Girju Nov 2017

Perfect Ratings With Negative Comments: Learning From Contradictory Patient Survey Responses, Andrew S. Gallan, Marina Girju, Roxana Girju

Patient Experience Journal

This research explores why patients give perfect domain scores yet provide negative comments on surveys. In order to explore this phenomenon, vendor-supplied in-patient survey data from eleven different hospitals of a major U.S. health care system were utilized. The dataset included survey scores and comments from 56,900 patients, collected from January 2015 through October 2016. Of the total number of responses, 30,485 (54%) contained at least one comment. For our analysis, we use a two-step approach: a quantitative analysis on the domain scores augmented by a qualitative text analysis of patients’ comments. To focus the research, we start by building …


Improving The Accuracy For The Long-Term Hydrologic Impact Assessment (L-Thia) Model, Anqi Zhang, Lawrence Theller, Bernard A. Engel Aug 2017

Improving The Accuracy For The Long-Term Hydrologic Impact Assessment (L-Thia) Model, Anqi Zhang, Lawrence Theller, Bernard A. Engel

The Summer Undergraduate Research Fellowship (SURF) Symposium

Urbanization increases runoff by changing land use types from less impervious to impervious covers. Improving the accuracy of a runoff assessment model, the Long-Term Hydrologic Impact Assessment (L-THIA) Model, can help us to better evaluate the potential uses of Low Impact Development (LID) practices aimed at reducing runoff, as well as to identify appropriate runoff and water quality mitigation methods. Several versions of the model have been built over time, and inconsistencies have been introduced between the models. To improve the accuracy and consistency of the model, the equations and parameters (primarily curve numbers in the case of this model) …


The Acquisition And Analysis Of Electroencephalogram Data For The Classification Of Benign Partial Epilepsy Of Childhood With Centrotemporal Spikes, Jessica A. Scarborough May 2017

The Acquisition And Analysis Of Electroencephalogram Data For The Classification Of Benign Partial Epilepsy Of Childhood With Centrotemporal Spikes, Jessica A. Scarborough

Master's Theses

In this thesis, I will expand upon each step in the process of acquiring and analyzing electroencephalogram (EEG) for the classification of benign childhood epilepsy with centrotemporal spikes. Despite huge advancements in the field of health informatics—natural language processing, machine learning, predictive modeling—there are significant barriers to the access of clinical data. These barriers include information blocking, privacy policy concerns, and a lack of stakeholder support. We will see that these roadblocks are all responsible for stunting biomedical research in some way, including my own experiences in acquiring the data for the second chapter of this thesis.

This second chapter …


Network Exploration Of Correlated Multivariate Protein Data For Alzheimer's Disease Association, Matthew J. Lane Apr 2017

Network Exploration Of Correlated Multivariate Protein Data For Alzheimer's Disease Association, Matthew J. Lane

Theses

Alzheimer Disease (AD) is difficult to diagnose by using genetic testing or other traditional methods. Unlike diseases with simple genetic risk components, there exists no single marker determining as to whether someone will develop AD. Furthermore, AD is highly heterogeneous and different subgroups of individuals develop the disease due to differing factors. Traditional diagnostic methods using perceivable cognitive deficiencies are often too little too late due to the brain having suffered damage from decades of disease progression. In order to observe AD at early stages prior to the observation of cognitive deficiencies, biomarkers with greater accuracy are required. By using …


Modeling Volatility Of Financial Time Series Using Arc Length, Benjamin H. Hoerlein Jan 2017

Modeling Volatility Of Financial Time Series Using Arc Length, Benjamin H. Hoerlein

College of Graduate Studies: Theses & Dissertations

This thesis explores how arc length can be modeled and used to measure the risk involved with a financial time series. Having arc length as a measure of volatility can help an investor in sorting which stocks are safer/riskier to invest in. A Gamma autoregressive model of order one(GAR(1)) is proposed to model arc length series. Kernel regression based bias correction is studied when model parameters are estimated using method of moment procedure. As an application, a model-based clustering involving thirty different stocks is presented using k-means++ and hierarchical clustering techniques.


Nondestructive Testing And Structural Health Monitoring Based On Adams And Svm Techniques, Gang Jiang, Yi Ming Deng, Ji Tai Niu Oct 2016

Nondestructive Testing And Structural Health Monitoring Based On Adams And Svm Techniques, Gang Jiang, Yi Ming Deng, Ji Tai Niu

The 8th International Conference on Physical and Numerical Simulation of Materials Processing

No abstract provided.


A Dual-Porosity-Stokes Model And Finite Element Method For Coupling Dual-Porosity Flow And Free Flow, Jiangyong Hou, Meilan Qiu, Xiaoming He, Chaohua Guo, Mingzhen Wei, Baojun Bai Oct 2016

A Dual-Porosity-Stokes Model And Finite Element Method For Coupling Dual-Porosity Flow And Free Flow, Jiangyong Hou, Meilan Qiu, Xiaoming He, Chaohua Guo, Mingzhen Wei, Baojun Bai

Mathematics and Statistics Faculty Research & Creative Works

In this paper, we propose and numerically solve a new model considering confined flow in dual-porosity media coupled with free flow in embedded macrofractures and conduits. Such situation arises, for example, for fluid flows in hydraulic fractured tight/shale oil/gas reservoirs. The flow in dual-porosity media, which consists of both matrix and microfractures, is described by a dual-porosity model. And the flow in the macrofractures and conduits is governed by the Stokes equation. Then the two models are coupled through four physically valid interface conditions on the interface between dual-porosity media and macrofractures/conduits, which play a key role in a physically …


Optimizing The Mix Of Games And Their Locations On The Casino Floor, Jason D. Fiege, Anastasia D. Baran Jun 2016

Optimizing The Mix Of Games And Their Locations On The Casino Floor, Jason D. Fiege, Anastasia D. Baran

International Conference on Gambling & Risk Taking

We present a mathematical framework and computational approach that aims to optimize the mix and locations of slot machine types and denominations, plus other games to maximize the overall performance of the gaming floor. This problem belongs to a larger class of spatial resource optimization problems, concerned with optimizing the allocation and spatial distribution of finite resources, subject to various constraints. We introduce a powerful multi-objective evolutionary optimization and data-modelling platform, developed by the presenter since 2002, and show how this software can be used for casino floor optimization. We begin by extending a linear formulation of the casino floor …


Stationary And Time-Dependent Optimization Of The Casino Floor Slot Machine Mix, Anastasia D. Baran, Jason D. Fiege Jun 2016

Stationary And Time-Dependent Optimization Of The Casino Floor Slot Machine Mix, Anastasia D. Baran, Jason D. Fiege

International Conference on Gambling & Risk Taking

Modeling and optimizing the performance of a mix of slot machines on a gaming floor can be addressed at various levels of coarseness, and may or may not consider time-dependent trends. For example, a model might consider only time-averaged, aggregate data for all machines of a given type; time-dependent aggregate data; time-averaged data for individual machines; or fully time dependent data for individual machines. Fine-grained, time-dependent data for individual machines offers the most potential for detailed analysis and improvements to the casino floor performance, but also suffers the greatest amount of statistical noise. We present a theoretical analysis of single …


Method For Determining Time-Resolved Heat Transfer Coefficient And Adiabatic Effectiveness Waveforms With Unsteady Film Cooling, James L. Rutledge, Jonathan F. Mccall Apr 2016

Method For Determining Time-Resolved Heat Transfer Coefficient And Adiabatic Effectiveness Waveforms With Unsteady Film Cooling, James L. Rutledge, Jonathan F. Mccall

AFIT Patents

A new method for determining heat transfer coefficient (h) and adiabatic effectiveness (η) waveforms h(t) and η(t) from a single test uses a novel inverse heat transfer methodology to use surface temperature histories obtained using prior art approaches to approximate the h(t) and η(t) waveforms. The method best curve fits the data to a pair of truncated Fourier series.


Hpcnmf: A High-Performance Toolbox For Non-Negative Matrix Factorization, Karthik Devarajan, Guoli Wang Feb 2016

Hpcnmf: A High-Performance Toolbox For Non-Negative Matrix Factorization, Karthik Devarajan, Guoli Wang

COBRA Preprint Series

Non-negative matrix factorization (NMF) is a widely used machine learning algorithm for dimension reduction of large-scale data. It has found successful applications in a variety of fields such as computational biology, neuroscience, natural language processing, information retrieval, image processing and speech recognition. In bioinformatics, for example, it has been used to extract patterns and profiles from genomic and text-mining data as well as in protein sequence and structure analysis. While the scientific performance of NMF is very promising in dealing with high dimensional data sets and complex data structures, its computational cost is high and sometimes could be critical for …


A Gene-Based Association Method For Mapping Traits Using Reference Transcriptome Data, Eric R. Gamazon, Heather Wheeler, Kaanan P. Shah, Sahar V. Mozaffari, Keston Aquino-Michaels, Robert J. Carroll, Anne E. Eyler, Joshua C. Denny, Gtex Consortium, Dan L. Nicolae, Nancy J. Cox, Hae Kyung Im Sep 2015

A Gene-Based Association Method For Mapping Traits Using Reference Transcriptome Data, Eric R. Gamazon, Heather Wheeler, Kaanan P. Shah, Sahar V. Mozaffari, Keston Aquino-Michaels, Robert J. Carroll, Anne E. Eyler, Joshua C. Denny, Gtex Consortium, Dan L. Nicolae, Nancy J. Cox, Hae Kyung Im

Bioinformatics Faculty Publications

Genome-wide association studies (GWAS) have identified thousands of variants robustly associated with complex traits. However, the biological mechanisms underlying these associations are, in general, not well understood. We propose a gene-based association method called PrediXcan that directly tests the molecular mechanisms through which genetic variation affects phenotype. The approach estimates the component of gene expression determined by an individual’s genetic profile and correlates ‘imputed’ gene expression with the phenotype under investigation to identify genes involved in the etiology of the phenotype. Genetically regulated gene expression is estimated using whole-genome tissue-dependent prediction models trained with reference transcriptome data sets. PrediXcan enjoys …


Mapping Open Water Bodeis With Optical Remote Sensing, Mary Ellen O'Donnell, Erika Podest Aug 2015

Mapping Open Water Bodeis With Optical Remote Sensing, Mary Ellen O'Donnell, Erika Podest

STAR Program Research Presentations

There is interest in mapping open water bodies using remote sensing data. Coverage and persistence of open water is currently a poorly measured variable due to its spatial and temporal variability across landscapes, especially in remote areas. The presence and persistence of open water is one of the primary indicators of conditions suitable for mosquito breeding habitats. Predicting the risk of mosquito caused disease outbreaks is a required step towards their control and eradication. Satellite observations can provide needed data to support agency decisions for deployment of preventative measures and control resources. This study, which will try to map open …


A Domain Decomposition Method For The Steady-State Navier-Stokes-Darcy Model With Beavers-Joseph Interface Condition, Xiaoming He, Jian Li, Yanping Lin, Ju Ming Jan 2015

A Domain Decomposition Method For The Steady-State Navier-Stokes-Darcy Model With Beavers-Joseph Interface Condition, Xiaoming He, Jian Li, Yanping Lin, Ju Ming

Mathematics and Statistics Faculty Research & Creative Works

This paper proposes and analyzes a Robin-type multiphysics domain decomposition method (DDM) for the steady-state Navier-Stokes-Darcy model with three interface conditions. In addition to the two regular interface conditions for the mass conservation and the force balance, the Beavers-Joseph condition is used as the interface condition in the tangential direction. The major mathematical difficulty in adopting the Beavers-Joseph condition is that it creates an indefinite leading order contribution to the total energy budget of the system [Y. Cao et al., Comm. Math. Sci., 8 (2010), pp. 1-25; Y. Cao et al., SIAM J. Numer. Anal., 47 (2010), pp. 4239-4256]. In …


New Pod Error Expressions, Error Bounds, And Asymptotic Results For Reduced Order Model Of Parabolic Pdes, John R. Singler Apr 2014

New Pod Error Expressions, Error Bounds, And Asymptotic Results For Reduced Order Model Of Parabolic Pdes, John R. Singler

Mathematics and Statistics Faculty Research & Creative Works

The derivations of existing error bounds for reduced order models of time varying partialdi erential equations (PDEs) constructed using proper orthogonal decomposition (POD) haverelied on bounding the error between the POD data and various POD projections of that data.Furthermore, the asymptotic behavior of the model reduction error bounds depends on theasymptotic behavior of the POD data approximation error bounds. We consider time varyingdata taking values in two di erent Hilbert spacesHandV, withVH, and prove exactexpressions for the POD data approximation errors considering four di erent POD projectionsand the two di erent Hilbert space error norms. Furthermore, the exact error expressions …


Computational Pain Quantification And The Effects Of Age, Gender, Culture And Cause, Colin R. Ostberg Apr 2014

Computational Pain Quantification And The Effects Of Age, Gender, Culture And Cause, Colin R. Ostberg

Master's Theses (2009 -)

Chronic pain affects more than 100 million Americans and more than 1.5 billion people worldwide. Pain is a multidimensional construct, expressed through a variety of means. Facial expressions are one such type of pain expression. Automatic facial expression recognition, and in particular pain expression recognition, are fields that have been studied extensively. However, nothing has explored the possibility of an automatic pain quantification algorithm, able to output pain levels based upon a facial image. Developed for a remote monitoring context, a computational pain quantification algorithm has been developed and validated by two distinct sets of data. The second set of …