A Pure-Jump Market-Making Model For High-Frequency Trading,
2015
Purdue University
A Pure-Jump Market-Making Model For High-Frequency Trading, Chi Wai Law
Open Access Dissertations
We propose a new market-making model which incorporates a number of realistic features relevant for high-frequency trading. In particular, we model the dependency structure of prices and order arrivals with novel self- and cross-exciting point processes. Furthermore, instead of assuming the bid and ask prices can be adjusted continuously by the market maker, we formulate the market maker's decisions as an optimal switching problem. Moreover, the risk of overtrading has been taken into consideration by allowing each order to have different size, and the market maker can make use of market orders, which are treated as impulse control, to get …
Overcoming Uncertainty For Within-Network Relational Machine Learning,
2015
Purdue University
Overcoming Uncertainty For Within-Network Relational Machine Learning, Joseph J. Pfeiffer
Open Access Dissertations
People increasingly communicate through email and social networks to maintain friendships and conduct business, as well as share online content such as pictures, videos and products. Relational machine learning (RML) utilizes a set of observed attributes and network structure to predict corresponding labels for items; for example, to predict individuals engaged in securities fraud, we can utilize phone calls and workplace information to make joint predictions over the individuals. However, in large scale and partially observed network domains, missing labels and edges can significantly impact standard relational machine learning methods by introducing bias into the learning and inference processes. In …
Stability Of Machine Learning Algorithms,
2015
Purdue University
Stability Of Machine Learning Algorithms, Wei Sun
Open Access Dissertations
In the literature, the predictive accuracy is often the primary criterion for evaluating a learning algorithm. In this thesis, I will introduce novel concepts of stability into the machine learning community. A learning algorithm is said to be stable if it produces consistent predictions with respect to small perturbation of training samples. Stability is an important aspect of a learning procedure because unstable predictions can potentially reduce users' trust in the system and also harm the reproducibility of scientific conclusions. As a prototypical example, stability of the classification procedure will be discussed extensively. In particular, I will present two new …
Divide And Recombine For Large Complex Data: The Subset Likelihood Modeling Approach To Recombination,
2015
Purdue University
Divide And Recombine For Large Complex Data: The Subset Likelihood Modeling Approach To Recombination, Philip Gautier
Open Access Dissertations
Divide and recombine (D&R) is a statistical framework for the analysis of large complex data. The data are divided into subsets. Numeric and visualization methods, which collectively are analytic methods, are applied to each subset. For each analytic method, the outputs of the application of the method to the subsets are recombined. So each analytic method has associated with it a division method and a recombination method. Here we study D&R methods for likelihood-based model fitting. We introduce a notion of likelihood analysis and modeling. We divide the data and fit a likelihood model on each subset. The fitted model …
In Defense Of Empirical Legal Studies,
2015
University of Georgia
In Defense Of Empirical Legal Studies, Christina L. Boyd
Buffalo Law Review
No abstract provided.
A Bad Proposal,
2015
Ethics and Public Policy Center
Two Worlds, Neither Perfect: A Comment On The Tension Between Legal And Empirical Studies,
2015
University of Iowa
Two Worlds, Neither Perfect: A Comment On The Tension Between Legal And Empirical Studies, Timothy M. Hagle
Buffalo Law Review
No abstract provided.
Numbers, Motivated Reasoning, And Empirical Legal Scholarship,
2015
IIT Chicago-Kent College of Law
Numbers, Motivated Reasoning, And Empirical Legal Scholarship, Carolyn Shapiro
Buffalo Law Review
No abstract provided.
Global Network Inference From Ego Network Samples: Testing A Simulation Approach,
2015
University of Nebraska–Lincoln
Global Network Inference From Ego Network Samples: Testing A Simulation Approach, Jeffrey A. Smith
Department of Sociology: Faculty Publications
Network sampling poses a radical idea: that it is possible to measure global network structure without the full population coverage assumed in most network studies. Network sampling is only useful, however, if a researcher can produce accurate global network estimates. This article explores the practicality of making network inference, focusing on the approach introduced in Smith (2012). The method uses sampled ego network data and simulation techniques to make inference about the global features of the true, unknown network. The validity check here includes more difficult scenarios than previous tests, including those that go beyond the initial scope conditions of …
Relationship Between High School Math Course Selection And Retention Rates At Otterbein University,
2015
Otterbein University
Relationship Between High School Math Course Selection And Retention Rates At Otterbein University, Lauren A. Fisher
Undergraduate Honors Thesis Projects
Binary logistic regression was used to study the relationship between high school math course selection and retention rates at Otterbein University. Graduation rates from postsecondary institutions are low in the United States and, more specifically, at Otterbein. This study is important in helping to determine what can raise retention rates, and ultimately, graduation rates. It directs focus toward high school math course selection and what should be changed before entering a post-secondary institution. Otterbein will have a better idea of what type of students to recruit and which students may be good candidates with some extra help. Recruiting is expensive, …
Zero-Inflated Models To Identify Transcription Factor Binding Sites In Chip-Seq Experiments,
2015
Old Dominion University
Zero-Inflated Models To Identify Transcription Factor Binding Sites In Chip-Seq Experiments, Sameera Dhananjaya Viswakula
Mathematics & Statistics Theses & Dissertations
It is essential to determine the protein-DNA binding sites to understand many biological processes. A transcription factor is a particular type of protein that binds to DNA and controls gene regulation in living organisms. Chromatin immunoprecipitation followed by highthroughput sequencing (ChIP-seq) is considered the gold standard in locating these binding sites and programs use to identify DNA-transcription factor binding sites are known as peak-callers. ChIP-seq data are known to exhibit considerable background noise and other biases. In this study, we propose a negative binomial model (NB), a zero-inflated Poisson model (ZIP) and a zero-inflated negative binomial model (ZINB) for peak-calling. …
Students Learning From Atlanta Public Schools Cheating Scandal,
2015
Dordt College
Students Learning From Atlanta Public Schools Cheating Scandal, Thomas M. Van Soelen
Faculty Work Comprehensive List
Access full-text article on publisher's site:
Dialectical Behavior Therapy For High Suicide Risk In Individuals With Borderline Personality Disorder: A Randomized Clinical Trial And Component Analysis,
2015
University of Washington - Seattle Campus
Dialectical Behavior Therapy For High Suicide Risk In Individuals With Borderline Personality Disorder: A Randomized Clinical Trial And Component Analysis, Marsha M. Linehan, Kathryn E. Korslund, Melanie S. Harned, Robert J. Gallop, Anita Lungu, Andrada D. Neacsiu, Joshua Mcdavid, Katherine Anne Comtois, Angela M. Murray-Gregory
Mathematics Faculty Publications
No abstract provided.
Targeted Estimation And Inference For The Sample Average Treatment Effect,
2015
Division of Biostatistics, University of California, Berkeley - the SEARCH Consortium
Targeted Estimation And Inference For The Sample Average Treatment Effect, Laura B. Balzer, Maya L. Petersen, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
While the population average treatment effect has been the subject of extensive methods and applied research, less consideration has been given to the sample average treatment effect: the mean difference in the counterfactual outcomes for the study units. The sample parameter is easily interpretable and is arguably the most relevant when the study units are not representative of a greater population or when the exposure's impact is heterogeneous. Formally, the sample effect is not identifiable from the observed data distribution. Nonetheless, targeted maximum likelihood estimation (TMLE) can provide an asymptotically unbiased and efficient estimate of both the population and sample …
Directed Energy Planetary Defense,
2015
University of California - Santa Barbara
Directed Energy Planetary Defense, Kelly Kosmo, Philip Lubin, Gary B. Hughes, Janelle Griswold, Qicheng Zhang, Travis Brashears
Statistics
Directed Energy (DE) systems offer the potential for true planetary defense from small to km class threats. Directed energy has evolved dramatically recently and is on an extremely rapid ascent technologically. It is now feasible to consider DE systems for threats from asteroids and comets. DE-STAR (Directed Energy System for Targeting of Asteroids and exploration) is a phased-array laser directed energy system intended for illumination, deflection and compositional analysis of asteroids [1]. It can be configured either as a stand-on or a distant stand-off system. A system of appropriate size would be capable of projecting a laser spot onto the …
Spectral Gene Set Enrichment (Sgse),
2015
Dartmouth College
Spectral Gene Set Enrichment (Sgse), H Robert Frost, Zhigang Li, Jason H. Moore
Dartmouth Scholarship
Gene set testing is typically performed in a supervised context to quantify the association between groups of genes and a clinical phenotype. In many cases, however, a gene set-based interpretation of genomic data is desired in the absence of a phenotype variable. Although methods exist for unsupervised gene set testing, they predominantly compute enrichment relative to clusters of the genomic variables with performance strongly dependent on the clustering algorithm and number of clusters. We propose a novel method, spectral gene set enrichment (SGSE), for unsupervised competitive testing of the association between gene sets and empirical data sources. SGSE first computes …
Predicting Successful Long-Term Weight Loss From Short-Term Weight-Loss Outcomes: New Insights From A Dynamic Energy Balance Model (The Pounds Lost Study),
2015
United States Military Academy
Predicting Successful Long-Term Weight Loss From Short-Term Weight-Loss Outcomes: New Insights From A Dynamic Energy Balance Model (The Pounds Lost Study), Diana Thomas, W Andrada Ivanescu, Corby K. Martin, Steven B. Heymsfield, Kaitlyn Marshall, Victoria E. Bodrato, Donald Williamson, Stephen Anton, Frank M. Sacks, Donna Ryan, George A. Bray
Department of Applied Mathematics and Statistics Faculty Scholarship and Creative Works
Background: Currently, early weight-loss predictions of long-term weight-loss success rely on fixed percent-weight-loss thresholds.
Objective: The objective was to develop thresholds during the first 3 mo of intervention that include the influence of age, sex, baseline weight, percent weight loss, and deviations from expected weight to predict whether a participant is likely to lose 5% or more body weight by year 1.
Design: Data consisting of month 1, 2, 3, and 12 treatment weights were obtained from the 2-y Preventing Obesity Using Novel Dietary Strategies (POUNDS Lost) intervention. Logistic regression models that included covariates of age, height, sex, baseline weight, …
Generalized Least-Squares Regressions V: Multiple Variables,
2015
CUNY Kingsborough Community College
Generalized Least-Squares Regressions V: Multiple Variables, Nataniel Greene
Publications and Research
The multivariate theory of generalized least-squares is formulated here using the notion of generalized means. The multivariate generalized least-squares problem seeks an m dimensional hyperplane which minimizes the average generalized mean of the square deviations between the data and the hyperplane in m + 1 variables. The numerical examples presented suggest that a multivariate generalized least-squares method can be preferable to ordinary least-squares especially in situations where the data are ill- conditioned.
Estimation Of Heterogeneous Panels With Structural Breaks,
2015
Syracuse University
Estimation Of Heterogeneous Panels With Structural Breaks, Badi Baltagi
Center for Policy Research
This paper extends Pesaran's (2006) work on common correlated effects (CCE) estimators for large heterogeneous panels with a general multifactor error structure by allowing for unknown common structural breaks. Structural breaks due to new policy implementation or major technological shocks, are more likely to occur over a longer time span. Consequently, ignoring structural breaks may lead to inconsistent estimation and invalid inference. We propose a general framework that includes heterogeneous panel data models and structural break models as special cases. The least squares method proposed by Bai (1997a, 2010) is applied to estimate the common change points, and the consistency …
Integrating Data Transformation In Principal Components Analysis,
2015
Marquette University
Integrating Data Transformation In Principal Components Analysis, Mehdi Maadooliat, Jianhua Z. Huang, Jianhua Hu
Mathematics, Statistics and Computer Science Faculty Research and Publications
Principal component analysis (PCA) is a popular dimension-reduction method to reduce the complexity and obtain the informative aspects of high-dimensional datasets. When the data distribution is skewed, data transformation is commonly used prior to applying PCA. Such transformation is usually obtained from previous studies, prior knowledge, or trial-and-error. In this work, we develop a model-based method that integrates data transformation in PCA and finds an appropriate data transformation using the maximum profile likelihood. Extensions of the method to handle functional data and missing values are also developed. Several numerical algorithms are provided for efficient computation. The proposed method is illustrated …
