Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Computer Sciences (60)
- Applied Statistics (42)
- Data Science (36)
- Artificial Intelligence and Robotics (29)
- Biostatistics (28)
-
- Statistical Models (28)
- Mathematics (26)
- Engineering (25)
- Life Sciences (21)
- Statistical Methodology (18)
- Applied Mathematics (17)
- Social and Behavioral Sciences (16)
- Business (15)
- Categorical Data Analysis (15)
- Medicine and Health Sciences (15)
- Multivariate Analysis (11)
- Bioinformatics (9)
- Clinical Trials (8)
- Longitudinal Data Analysis and Time Series (8)
- Other Statistics and Probability (8)
- Other Computer Sciences (7)
- Civil and Environmental Engineering (6)
- Economics (6)
- Electrical and Computer Engineering (6)
- Finance and Financial Management (6)
- Public Health (6)
- Analysis (5)
- Databases and Information Systems (5)
- Institution
-
- Southern Methodist University (7)
- University of New Mexico (7)
- New Jersey Institute of Technology (6)
- City University of New York (CUNY) (5)
- Georgia Southern University (5)
-
- Kennesaw State University (5)
- Old Dominion University (5)
- COBRA (4)
- California Polytechnic State University, San Luis Obispo (4)
- Dartmouth College (4)
- University of Nebraska - Lincoln (4)
- University of South Florida (4)
- Utah State University (4)
- Air Force Institute of Technology (3)
- Changsha University of Science and Technology (3)
- Marshall University (3)
- Minnesota State University, Mankato (3)
- Missouri University of Science and Technology (3)
- University of Denver (3)
- University of Louisville (3)
- University of Montana (3)
- University of New Hampshire (3)
- University of South Carolina (3)
- University of Texas at El Paso (3)
- Virginia Commonwealth University (3)
- Western Kentucky University (3)
- Claremont Colleges (2)
- East Tennessee State University (2)
- LSU Health New Orleans (2)
- Louisiana State University (2)
- Publication Year
- Publication
-
- Theses and Dissertations (11)
- Electronic Theses and Dissertations (8)
- SMU Data Science Review (7)
- Dissertations (6)
- College of Graduate Studies: Theses & Dissertations (5)
-
- Mathematics & Statistics ETDs (5)
- Dissertations, Theses, and Capstone Projects (4)
- Master's Theses (4)
- USF Tampa Graduate Theses and Dissertations (4)
- All Graduate Theses, Dissertations, and Other Capstone Projects (3)
- Doctoral Dissertations (3)
- Graduate Student Theses, Dissertations, & Professional Papers (3)
- Journal of China & Foreign Highway (3)
- Open Access Theses & Dissertations (3)
- Theses, Dissertations and Capstones (3)
- U.C. Berkeley Division of Biostatistics Working Paper Series (3)
- All Graduate Theses and Dissertations, Spring 1920 to Summer 2023 (2)
- Doctor of Data Science and Analytics Dissertations (2)
- Economics Faculty Publications (2)
- Faculty Publications (2)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (2)
- Honors Theses and Capstones (2)
- SAML-25 Workshop on Statistical and Machine Learning (2)
- School of Public Health Faculty Publications (2)
- UNLV Theses, Dissertations, Professional Papers, and Capstones (2)
- Williams Honors College, Honors Research Projects (2)
- All Dissertations (1)
- All Faculty Scholarship for the College of the Sciences (1)
- All Graduate Theses and Dissertations, Fall 2023 to Present (1)
- Biostatistics Faculty Publications (1)
- Publication Type
- File Type
Articles 121 - 150 of 153
Full-Text Articles in Statistics and Probability
A Statistical Analysis And Machine Learning Of Genomic Data, Jongyun Jung
A Statistical Analysis And Machine Learning Of Genomic Data, Jongyun Jung
All Graduate Theses, Dissertations, and Other Capstone Projects
Machine learning enables a computer to learn a relationship between two assumingly related types of information. One type of information could thus be used to predict any lack of informaion in the other using the learned relationship. During the last decades, it has become cheaper to collect biological information, which has resulted in increasingly large amounts of data. Biological information such as DNA is currently analyzed by a variety of tools. Although machine learning has already been used in various projects, a flexible tool for analyzing generic biological challenges has not yet been made. The recent advancements in the DNA …
Reliability Analysis For Systems With Outsourced Components, Zhengwei Hu
Reliability Analysis For Systems With Outsourced Components, Zhengwei Hu
Doctoral Dissertations
"The current business model for many industrial firms is to function as system integrators, depending on numerous outsourced components from outside component suppliers. This practice has resulted in tremendous cost savings; it makes system reliability analysis, however, more challenging due to the limited component information available to system designers. The component information is often proprietary to component suppliers. Motivated by the need of system reliability prediction with outsourced components, this work aims to explore feasible ways to accurately predict the system reliability during the system design stage. Four methods are proposed. The first method reconstructs component reliability functions using limited …
Regression Tree Construction For Reinforcement Learning Problems With A General Action Space, Anthony S. Bush Jr
Regression Tree Construction For Reinforcement Learning Problems With A General Action Space, Anthony S. Bush Jr
College of Graduate Studies: Theses & Dissertations
Part of the implementation of Reinforcement Learning is constructing a regression of values against states and actions and using that regression model to optimize over actions for a given state. One such common regression technique is that of a decision tree; or in the case of continuous input, a regression tree. In such a case, we fix the states and optimize over actions; however, standard regression trees do not easily optimize over a subset of the input variables\cite{Card1993}. The technique we propose in this thesis is a hybrid of regression trees and kernel regression. First, a regression tree splits over …
Predicting National Basketball Association Success: A Machine Learning Approach, Adarsh Kannan, Brian Kolovich, Brandon Lawrence, Sohail Rafiqi
Predicting National Basketball Association Success: A Machine Learning Approach, Adarsh Kannan, Brian Kolovich, Brandon Lawrence, Sohail Rafiqi
SMU Data Science Review
In this paper, we present a machine learning based approach to projecting the success of National Basketball Association (NBA) draft prospects. With the proliferation of data, analytics have increasingly be- come a critical component in the assessment of professional and collegiate basketball players. We leverage player biometric data, college statistics, draft selection order, and positional breakdown as modelling features in our prediction algorithms. We found that a player's draft pick and their college statistics are the best predictors of their longevity in the National Basketball Association.
Longitudinal Tracking Of Physiological State With Electromyographic Signals., Robert Warren Stallard
Longitudinal Tracking Of Physiological State With Electromyographic Signals., Robert Warren Stallard
Electronic Theses and Dissertations
Electrophysiological measurements have been used in recent history to classify instantaneous physiological configurations, e.g., hand gestures. This work investigates the feasibility of working with changes in physiological configurations over time (i.e., longitudinally) using a variety of algorithms from the machine learning domain. We demonstrate a high degree of classification accuracy for a binary classification problem derived from electromyography measurements before and after a 35-day bedrest. The problem difficulty is increased with a more dynamic experiment testing for changes in astronaut sensorimotor performance by taking electromyography and force plate measurements before, during, and after a jump from a small platform. A …
Cognitive Virtual Admissions Counselor, Kumar Raja Guvindan Raju, Cory Adams, Raghuram Srinivas
Cognitive Virtual Admissions Counselor, Kumar Raja Guvindan Raju, Cory Adams, Raghuram Srinivas
SMU Data Science Review
Abstract. In this paper, we present a cognitive virtual admissions counselor for the Master of Science in Data Science program at Southern Methodist University. The virtual admissions counselor is a system capable of providing potential students accurate information at the time that they want to know it. After the evaluation of multiple technologies, Amazon’s LEX was selected to serve as the core technology for the virtual counselor chatbot. Student surveys were leveraged to collect and generate training data to deploy the natural language capability. The cognitive virtual admissions counselor platform is currently capable of providing an end-to-end conversational dialog to …
A Convolutional Neural Network Model For Species Classification Of Camera Trap Images, Annie Casey
A Convolutional Neural Network Model For Species Classification Of Camera Trap Images, Annie Casey
Mathematics Undergraduate Theses
The overall purpose of this study was to automate the manual process of tagging species found in camera trap images using machine learning. The basic design of this study was to implement a Convolutional Neural Network model in Python using the Keras and Tensorflow modules that learn to recognize patterns in images in order to classify what species is in a given image and to label it accordingly. Results of the analysis highlight the importance of a large sample size, the degree of accuracy according to various arguments in the model, effectiveness of multiple layers that include Max Pooling, and …
A Comparison Of Machine Learning Techniques For Taxonomic Classification Of Teeth From The Family Bovidae, Gregory J. Matthews, Juliet K. Brophy, Maxwell Luetkemeier, Hongie Gu, George K. Thiruvathukal
A Comparison Of Machine Learning Techniques For Taxonomic Classification Of Teeth From The Family Bovidae, Gregory J. Matthews, Juliet K. Brophy, Maxwell Luetkemeier, Hongie Gu, George K. Thiruvathukal
Mathematics and Statistics: Faculty Publications and Other Works
This study explores the performance of machine learning algorithms on the classification of fossil teeth in the Family Bovidae. Isolated bovid teeth are typically the most common fossils found in southern Africa and they often constitute the basis for paleoenvironmental reconstructions. Taxonomic identification of fossil bovid teeth, however, is often imprecise and subjective. Using modern teeth with known taxons, machine learning algorithms can be trained to classify fossils. Previous work by Brophy et al. [Quantitative morphological analysis of bovid teeth and implications for paleoenvironmental reconstruction of plovers lake, Gauteng Province, South Africa, J. Archaeol. Sci. 41 (2014), pp. …
The Impact Of Data Sovereignty On American Indian Self-Determination: A Framework Proof Of Concept Using Data Science, Joseph Carver Robertson
The Impact Of Data Sovereignty On American Indian Self-Determination: A Framework Proof Of Concept Using Data Science, Joseph Carver Robertson
Electronic Theses and Dissertations
The Data Sovereignty Initiative is a collection of ideas that was designed to create SMART solutions for tribal communities. This concept was to develop a horizontal governance framework to create a strategic act of sovereignty using data science. The core concept of this idea was to present data sovereignty as a way for tribal communities to take ownership of data in order to affect policy and strategic decisions that are data driven in nature. The case studies in this manuscript were developed around statistical theories of spatial statistics, exploratory data analysis, and machine learning. And although these case studies are …
Old English Character Recognition Using Neural Networks, Sattajit Sutradhar
Old English Character Recognition Using Neural Networks, Sattajit Sutradhar
College of Graduate Studies: Theses & Dissertations
Character recognition has been capturing the interest of researchers since the beginning of the twentieth century. While the Optical Character Recognition for printed material is very robust and widespread nowadays, the recognition of handwritten materials lags behind. In our digital era more and more historical, handwritten documents are digitized and made available to the general public. However, these digital copies of handwritten materials lack the automatic content recognition feature of their printed materials counterparts. We are proposing a practical, accurate, and computationally efficient method for Old English character recognition from manuscript images. Our method relies on a modern machine learning …
Comparing Various Machine Learning Statistical Methods Using Variable Differentials To Predict College Basketball, Nicholas Bennett
Comparing Various Machine Learning Statistical Methods Using Variable Differentials To Predict College Basketball, Nicholas Bennett
Williams Honors College, Honors Research Projects
The purpose of this Senior Honors Project is to research, study, and demonstrate newfound knowledge of various machine learning statistical techniques that are not covered in the University of Akron’s statistics major curriculum. This report will be an overview of three machine-learning methods that were used to predict NCAA Basketball results, specifically, the March Madness tournament. The variables used for these methods, models, and tests will include numerous variables kept throughout the season for each team, along with a couple variables that are used by the selection committee when tournament teams are being picked. The end goal is to find …
Estimating The Optimal Cutoff Point For Logistic Regression, Zheng Zhang
Estimating The Optimal Cutoff Point For Logistic Regression, Zheng Zhang
Open Access Theses & Dissertations
Binary classification is one of the main themes of supervised learning. This research is concerned about determining the optimal cutoff point for the continuous-scaled outcomes (e.g., predicted probabilities) resulting from a classifier such as logistic regression. We make note of the fact that the cutoff point obtained from various methods is a statistic, which can be unstable with substantial variation. Nevertheless, due partly to complexity involved in estimating the cutpoint, there has been no formal study on the variance or standard error of the estimated cutoff point.
In this Thesis, a bootstrap aggregation method is put forward to estimate the …
Nonparametric Variable Importance Assessment Using Machine Learning Techniques, Brian D. Williamson, Peter B. Gilbert, Noah Simon, Marco Carone
Nonparametric Variable Importance Assessment Using Machine Learning Techniques, Brian D. Williamson, Peter B. Gilbert, Noah Simon, Marco Carone
UW Biostatistics Working Paper Series
In a regression setting, it is often of interest to quantify the importance of various features in predicting the response. Commonly, the variable importance measure used is determined by the regression technique employed. For this reason, practitioners often only resort to one of a few regression techniques for which a variable importance measure is naturally defined. Unfortunately, these regression techniques are often sub-optimal for predicting response. Additionally, because the variable importance measures native to different regression techniques generally have a different interpretation, comparisons across techniques can be difficult. In this work, we study a novel variable importance measure that can …
Identification Of Prognostic Genes And Gene Sets For Early-Stage Non-Small Cell Lung Cancer Using Bi-Level Selection Methods, Suyan Tian, Chi Wang, Howard H. Chang, Jianguo Sun
Identification Of Prognostic Genes And Gene Sets For Early-Stage Non-Small Cell Lung Cancer Using Bi-Level Selection Methods, Suyan Tian, Chi Wang, Howard H. Chang, Jianguo Sun
Biostatistics Faculty Publications
In contrast to feature selection and gene set analysis, bi-level selection is a process of selecting not only important gene sets but also important genes within those gene sets. Depending on the order of selections, a bi-level selection method can be classified into three categories – forward selection, which first selects relevant gene sets followed by the selection of relevant individual genes; backward selection which takes the reversed order; and simultaneous selection, which performs the two tasks simultaneously usually with the aids of a penalized regression model. To test the existence of subtype-specific prognostic genes for non-small cell lung cancer …
Application Of Response Surface Methods To Determine Conditions For Optimal Genomic Prediction, Reka Howard, Alicia L. Carriquiry, William D. Beavis
Application Of Response Surface Methods To Determine Conditions For Optimal Genomic Prediction, Reka Howard, Alicia L. Carriquiry, William D. Beavis
Department of Statistics: Faculty Publications
An epistatic genetic architecture can have a significant impact on prediction accuracies of genomic prediction (GP) methods. Machine learning methods predict traits comprised of epistatic genetic architectures more accurately than statistical methods based on additive mixed linear models. The differences between these types of GP methods suggest a diagnostic for revealing genetic architectures underlying traits of interest. In addition to genetic architecture, the performance of GP methods may be influenced by the sample size of the training population, the number of QTL, and the proportion of phenotypic variability due to genotypic variability (heritability). Possible values for these factors and the …
Audio-Based Productivity Forecasting Of Construction Cyclic Activities, Chris A. Sabillon
Audio-Based Productivity Forecasting Of Construction Cyclic Activities, Chris A. Sabillon
College of Graduate Studies: Theses & Dissertations
Due to its high cost, project managers must be able to monitor the performance of construction heavy equipment promptly. This cannot be achieved through traditional management techniques, which are based on direct observation or on estimations from historical data. Some manufacturers have started to integrate their proprietary technologies, but construction contractors are unlikely to have a fleet of entirely new and single manufacturer equipment for this to represent a solution. Third party automated approaches include the use of active sensors such as accelerometers and gyroscopes, passive technologies such as computer vision and image processing, and audio signal processing. Hitherto, most …
Towards Deeper Understanding In Neuroimaging, Rex Devon Hjelm
Towards Deeper Understanding In Neuroimaging, Rex Devon Hjelm
Computer Science ETDs
Neuroimaging is a growing domain of research, with advances in machine learning having tremendous potential to expand understanding in neuroscience and improve public health. Deep neural networks have recently and rapidly achieved historic success in numerous domains, and as a consequence have completely redefined the landscape of automated learners, giving promise of significant advances in numerous domains of research. Despite recent advances and advantages over traditional machine learning methods, deep neural networks have yet to have permeated significantly into neuroscience studies, particularly as a tool for discovery. This dissertation presents well-established and novel tools for unsupervised learning which aid in …
Online Cross-Validation-Based Ensemble Learning, David Benkeser, Samuel D. Lendle, Cheng Ju, Mark J. Van Der Laan
Online Cross-Validation-Based Ensemble Learning, David Benkeser, Samuel D. Lendle, Cheng Ju, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Online estimators update a current estimate with a new incoming batch of data without having to revisit past data thereby providing streaming estimates that are scalable to big data. We develop flexible, ensemble-based online estimators of an infinite-dimensional target parameter, such as a regression function, in the setting where data are generated sequentially by a common conditional data distribution given summary measures of the past. This setting encompasses a wide range of time-series models and as special case, models for independent and identically distributed data. Our estimator considers a large library of candidate online estimators and uses online cross-validation to …
Multiple Imputation Of Missing Data In Structural Equation Models With Mediators And Moderators Using Gradient Boosted Machine Learning, Robert J. Milletich Ii
Multiple Imputation Of Missing Data In Structural Equation Models With Mediators And Moderators Using Gradient Boosted Machine Learning, Robert J. Milletich Ii
Psychology Theses & Dissertations
Mediation and moderated mediation models are two commonly used models for indirect effects analysis. In practice, missing data is a pervasive problem in structural equation modeling with psychological data. Multiple imputation (MI) is one method used to estimate model parameters in the presence of missing data, while accounting for uncertainty due to the missing data. Unfortunately, commonly used MI methods are not equipped to handle categorical variables or nonlinear variables such as interactions. In this study, we introduce a general MI framework that uses the Bayesian bootstrap (BB) method to generate posterior inferences for indirect effects and gradient boosted machine …
Learning From Data: Plant Breeding Applications Of Machine Learning, Alencar Xavier
Learning From Data: Plant Breeding Applications Of Machine Learning, Alencar Xavier
Open Access Dissertations
Increasingly, new sources of data are being incorporated into plant breeding pipelines. Enormous amounts of data from field phenomics and genotyping technologies places data mining and analysis into a completely different level that is challenging from practical and theoretical standpoints. Intelligent decision-making relies on our capability of extracting from data useful information that may help us to achieve our goals more efficiently. Many plant breeders, agronomists and geneticists perform analyses without knowing relevant underlying assumptions, strengths or pitfalls of the employed methods. The study endeavors to assess statistical learning properties and plant breeding applications of supervised and unsupervised machine learning …
Data Driven Sample Generator Model With Application To Classification, Alvaro Emilio Ulloa Cerna
Data Driven Sample Generator Model With Application To Classification, Alvaro Emilio Ulloa Cerna
Mathematics & Statistics ETDs
Despite the rapidly growing interest, progress in the study of relations between physiological abnormalities and mental disorders is hampered by complexity of the human brain and high costs of data collection. The complexity can be captured by machine learning approaches, but they still may require significant amounts of data. In this thesis, we seek to mitigate the latter challenge by developing a data driven sample generator model for the generation of synthetic realistic training data. Our method greatly improves generalization in classification of schizophrenia patients and healthy controls from their structural magnetic resonance images. A feed forward neural network trained …
Privacy And Accountability In Black-Box Medicine, Roger Allan Ford, W. Nicholson Price Ii
Privacy And Accountability In Black-Box Medicine, Roger Allan Ford, W. Nicholson Price Ii
Law Faculty Scholarship
Black-box medicine—the use of big data and sophisticated machine learning techniques for health-care applications—could be the future of personalized medicine. Black-box medicine promises to make it easier to diagnose rare diseases and conditions, identify the most promising treatments, and allocate scarce resources among different patients. But to succeed, it must overcome two separate, but related, problems: patient privacy and algorithmic accountability. Privacy is a problem because researchers need access to huge amounts of patient health information to generate useful medical predictions. And accountability is a problem because black-box algorithms must be verified by outsiders to ensure they are accurate and …
A Data Science Course For Undergraduates: Thinking With Data, Benjamin Baumer
A Data Science Course For Undergraduates: Thinking With Data, Benjamin Baumer
Mathematics Sciences: Faculty Publications
Data science is an emerging interdisciplinary field that combines elements of mathematics, statistics, computer science, and knowledge in a particular application domain for the purpose of extracting meaningful information from the increasingly sophisticated array of data available in many settings. These data tend to be nontraditional, in the sense that they are often live, large, complex, and/or messy. A first course in statistics at the undergraduate level typically introduces students to a variety of techniques to analyze small, neat, and clean datasets. However, whether they pursue more formal training in statistics or not, many of these students will end up …
Convergence Of A Reinforcement Learning Algorithm In Continuous Domains, Stephen Carden
Convergence Of A Reinforcement Learning Algorithm In Continuous Domains, Stephen Carden
All Dissertations
In the field of Reinforcement Learning, Markov Decision Processes with a finite number of states and actions have been well studied, and there exist algorithms capable of producing a sequence of policies which converge to an optimal policy with probability one. Convergence guarantees for problems with continuous states also exist. Until recently, no online algorithm for continuous states and continuous actions has been proven to produce optimal policies. This Dissertation contains the results of research into reinforcement learning algorithms for problems in which both the state and action spaces are continuous. The problems to be solved are introduced formally as …
Variable Importance And Prediction Methods For Longitudinal Problems With Missing Variables, Ivan Diaz, Alan E. Hubbard, Anna Decker, Mitchell Cohen
Variable Importance And Prediction Methods For Longitudinal Problems With Missing Variables, Ivan Diaz, Alan E. Hubbard, Anna Decker, Mitchell Cohen
U.C. Berkeley Division of Biostatistics Working Paper Series
In this paper we present prediction and variable importance (VIM) methods for longitudinal data sets containing both continuous and binary exposures subject to missingness. We demonstrate the use of these methods for prognosis of medical outcomes of severe trauma patients, a field in which current medical practice involves rules of thumb and scoring methods that only use a few variables and ignore the dynamic and high-dimensional nature of trauma recovery. Well-principled prediction and VIM methods can thus provide a tool to make care decisions informed by the high-dimensional patient’s physiological and clinical history. Our VIM parameters can be causally interpreted …
Asymptotically Unbiased Estimator Of The Informational Energy With Knn, Angel Caţaron, Răzvan Andonie, Chinmei Y. Chueh
Asymptotically Unbiased Estimator Of The Informational Energy With Knn, Angel Caţaron, Răzvan Andonie, Chinmei Y. Chueh
All Faculty Scholarship for the College of the Sciences
Motivated by machine learning applications (e.g., classification, function approximation, feature extraction), in previous work, we have introduced a non- parametric estimator of Onicescu’s informational energy. Our method was based on the k-th nearest neighbor distances between the n sample points, where k is a fixed positive integer. In the present contribution, we discuss mathematical properties of this estimator. We show that our estimator is asymptotically unbiased and consistent. We provide further experimental results which illustrate the convergence of the estimator for standard distributions.
Computationally Efficient Confidence Intervals For Cross-Validated Area Under The Roc Curve Estimates, Erin Ledell, Maya L. Petersen, Mark J. Van Der Laan
Computationally Efficient Confidence Intervals For Cross-Validated Area Under The Roc Curve Estimates, Erin Ledell, Maya L. Petersen, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
In binary classification problems, the area under the ROC curve (AUC), is an effective means of measuring the performance of your model. Most often, cross-validation is also used, in order to assess how the results will generalize to an independent data set. In order to evaluate the quality of an estimate for cross-validated AUC, we must obtain an estimate for its variance. For massive data sets, the process of generating a single performance estimate can be computationally expensive. Additionally, when using a complex prediction method, calculating the cross-validated AUC on even a relatively small data set can still require a …
Bayesian And Related Methods: Techniques Based On Bayes' Theorem, Mehmet Vurkaç
Bayesian And Related Methods: Techniques Based On Bayes' Theorem, Mehmet Vurkaç
Systems Science Friday Noon Seminar Series
Bayes' theorem is a simple algebraic consequence of conditional probability. Yet, its consequences are critical to philosophy, society, and technology. Starting from its simple derivation, we will show how its interpretation in terms of base rates (priors) and class-conditional likelihoods illuminates everyday problems in medicine and law, and provides signal processing, communications, machine learning, model selection, and other applications of statistics with powerful classification and estimation tools. Next, we will briefly examine some of the ways in which this theorem can be adopted to include multiple attributes, contexts, hypotheses, and levels of risk. Methods derived from or related to Bayes’ …
Empirical Methods For Predicting Student Retention- A Summary From The Literature, Matt Bogard
Empirical Methods For Predicting Student Retention- A Summary From The Literature, Matt Bogard
Economics Faculty Publications
The vast majority of the literature related to the empirical estimation of retention models includes a discussion of the theoretical retention framework established by Bean, Braxton, Tinto, Pascarella, Terenzini and others (see Bean, 1980; Bean, 2000; Braxton, 2000; Braxton et al, 2004; Chapman and Pascarella, 1983; Pascarell and Ternzini, 1978; St. John and Cabrera, 2000; Tinto, 1975) This body of research provides a starting point for the consideration of which explanatory variables to include in any model specification, as well as identifying possible data sources. The literature separates itself into two major camps including research related to the hypothesis testing …
Empirical Methods-A Review: With An Introduction To Data Mining And Machine Learning, Matt Bogard
Empirical Methods-A Review: With An Introduction To Data Mining And Machine Learning, Matt Bogard
Economics Faculty Publications
This presentation was part of a staff workshop focused on empirical methods and applied research. This includes a basic overview of regression with matrix algebra, maximum likelihood, inference, and model assumptions. Distinctions are made between paradigms related to classical statistical methods and algorithmic approaches. The presentation concludes with a brief discussion of generalization error, data partitioning, decision trees, and neural networks.