Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Applied Statistics (16)
- Computer Sciences (15)
- Medicine and Health Sciences (14)
- Multivariate Analysis (13)
- Statistical Models (13)
-
- Data Science (12)
- Statistical Methodology (11)
- Mathematics (10)
- Public Health (9)
- Statistical Theory (9)
- Biostatistics (8)
- Categorical Data Analysis (8)
- Clinical Epidemiology (8)
- Life Sciences (7)
- Social and Behavioral Sciences (7)
- Other Statistics and Probability (6)
- Applied Mathematics (5)
- Microarrays (5)
- Epidemiology (4)
- Genetics and Genomics (4)
- Artificial Intelligence and Robotics (3)
- Arts and Humanities (3)
- Clinical Trials (3)
- Engineering (3)
- Medical Specialties (3)
- Other Computer Sciences (3)
- Probability (3)
- Theory and Algorithms (3)
- Institution
-
- COBRA (14)
- Southern Methodist University (6)
- Utah State University (5)
- University of Nebraska - Lincoln (4)
- University of South Florida (4)
-
- Old Dominion University (3)
- University of South Carolina (3)
- Wayne State University (3)
- Brigham Young University (2)
- Loyola University Chicago (2)
- Virginia Commonwealth University (2)
- City University of New York (CUNY) (1)
- Department of Primary Industries and Regional Development, Western Australia (1)
- Embry-Riddle Aeronautical University (1)
- Florida Institute of Technology (1)
- Georgia Southern University (1)
- Illinois State University (1)
- Kennesaw State University (1)
- Louisiana Tech University (1)
- Marquette University (1)
- Minnesota State University, Mankato (1)
- Missouri University of Science and Technology (1)
- New Jersey Institute of Technology (1)
- Nova Southeastern University (1)
- Purdue University (1)
- Rose-Hulman Institute of Technology (1)
- Technological University Dublin (1)
- The Texas Medical Center Library (1)
- The University of Southern Mississippi (1)
- Universitas Negeri Malang (1)
- Publication Year
- Publication
-
- UW Biostatistics Working Paper Series (11)
- Theses and Dissertations (9)
- SMU Data Science Review (6)
- USF Tampa Graduate Theses and Dissertations (4)
- All Graduate Plan B and other Reports, Spring 1920 to Spring 2023 (3)
-
- Journal of Modern Applied Statistical Methods (3)
- Department of Statistics: Faculty Publications (2)
- Dissertations (2)
- Electronic Theses and Dissertations (2)
- Mathematics & Statistics Faculty Publications (2)
- U.C. Berkeley Division of Biostatistics Working Paper Series (2)
- All Graduate Theses and Dissertations, Fall 2023 to Present (1)
- All Graduate Theses and Dissertations, Spring 1920 to Summer 2023 (1)
- All Graduate Theses, Dissertations, and Other Capstone Projects (1)
- Arts & Sciences Graduate Student Theses and Dissertations (1)
- Beyond: Undergraduate Research Journal (1)
- College of Graduate Studies: Theses & Dissertations (1)
- Computer Science: Faculty Publications and Other Works (1)
- Dissertations and Doctoral Documents, University of Nebraska-Lincoln, 2023– (1)
- Dissertations and Theses (Open Access) (1)
- Dissertations, Theses, and Capstone Projects (1)
- Doctoral Dissertations (1)
- Electrical & Computer Engineering Faculty Publications (1)
- Honors Program: Senior Projects (Public) (1)
- Journal of the Department of Agriculture, Western Australia, Series 4 (1)
- Knowledge Engineering and Data Science (1)
- Mahurin Honors College Capstone Experience/Thesis Projects (1)
- Mathematics and Statistics Faculty Research & Creative Works (1)
- Mathematics and Statistics: Faculty Publications and Other Works (1)
- Mathematics, Statistics and Computer Science Faculty Research and Publications (1)
- Publication Type
Articles 31 - 60 of 76
Full-Text Articles in Statistics and Probability
An Analysis Of The Success Of Farmers Markets In Kentucky Using Logistic Regression And Support Vector Machines, Jeron Russell
An Analysis Of The Success Of Farmers Markets In Kentucky Using Logistic Regression And Support Vector Machines, Jeron Russell
Mahurin Honors College Capstone Experience/Thesis Projects
The purpose of this research is to look at the relationship that market-specific, economic, and demographic variables have with the success of farmers markets in Kentucky. It additionally seeks to build a tool for predicting farmers market success that could be used by policy makers to aid in decision-making processes concerning farmers markets. Logistic regression and Support Vector Machines (SVMs) are used on data acquired from the Kentucky Department of Agriculture and the American Community Survey in order to analyze the data in a traditional statistical approach as well as a machine learning approach. The results included an SVM model …
Fractional Random Weighted Bootstrapping For Classification On Imbalanced Data With Ensemble Decision Tree Methods, Sean Charles Carter
Fractional Random Weighted Bootstrapping For Classification On Imbalanced Data With Ensemble Decision Tree Methods, Sean Charles Carter
USF Tampa Graduate Theses and Dissertations
Ensemble methods are commonly used for building predictive models for classification. Models that are unstable to perturbations in the training set, such as the decision tree, often see considerable reductions in error when grouped, using bootstrapped resamples of the training data to train many models. The non-parametric bootstrap, however, has limited efficacy when used on severely imbalanced data, especially when the number of observations of one or more classes is exceptionally small. We explore the fractional random weighted bootstrap, which randomly assigns fractional weights to observations, as an alternative resampling pro cedure in training machine learning ensembles, particularly decision tree …
Adaptive Feature Engineering Modeling For Ultrasound Image Classification For Decision Support, Hatwib Mugasa
Adaptive Feature Engineering Modeling For Ultrasound Image Classification For Decision Support, Hatwib Mugasa
Doctoral Dissertations
Ultrasonography is considered a relatively safe option for the diagnosis of benign and malignant cancer lesions due to the low-energy sound waves used. However, the visual interpretation of the ultrasound images is time-consuming and usually has high false alerts due to speckle noise. Improved methods of collection image-based data have been proposed to reduce noise in the images; however, this has proved not to solve the problem due to the complex nature of images and the exponential growth of biomedical datasets. Secondly, the target class in real-world biomedical datasets, that is the focus of interest of a biopsy, is usually …
Machine Learning In Support Of Electric Distribution Asset Failure Prediction, Robert D. Flamenbaum, Thomas Pompo, Christopher Havenstein, Jade Thiemsuwan
Machine Learning In Support Of Electric Distribution Asset Failure Prediction, Robert D. Flamenbaum, Thomas Pompo, Christopher Havenstein, Jade Thiemsuwan
SMU Data Science Review
In this paper, we present novel approaches to predicting as- set failure in the electric distribution system. Failures in overhead power lines and their associated equipment in particular, pose significant finan- cial and environmental threats to electric utilities. Electric device failure furthermore poses a burden on customers and can pose serious risk to life and livelihood. Working with asset data acquired from an electric utility in Southern California, and incorporating environmental and geospatial data from around the region, we applied a Random Forest methodology to predict which overhead distribution lines are most vulnerable to fail- ure. Our results provide evidence …
Machine Learning Pipeline For Exoplanet Classification, George Clayton Sturrock, Brychan Manry, Sohail Rafiqi
Machine Learning Pipeline For Exoplanet Classification, George Clayton Sturrock, Brychan Manry, Sohail Rafiqi
SMU Data Science Review
Planet identification has typically been a tasked performed exclusively by teams of astronomers and astrophysicists using methods and tools accessible only to those with years of academic education and training. NASA’s Exoplanet Exploration program has introduced modern satellites capable of capturing a vast array of data regarding celestial objects of interest to assist with researching these objects. The availability of satellite data has opened up the task of planet identification to individuals capable of writing and interpreting machine learning models. In this study, several classification models and datasets are utilized to assign a probability of an observation being an exoplanet. …
Φ-Divergence Loss-Based Artificial Neural Network, R. L. Salamwade, D. M. Sakate, S. K. Mathur
Φ-Divergence Loss-Based Artificial Neural Network, R. L. Salamwade, D. M. Sakate, S. K. Mathur
Journal of Modern Applied Statistical Methods
Artificial Neural Networks (ANNs) can fit non-linear functions and recognize patterns better than several standard techniques. Performance of ANNs is measured by using loss functions. Phi-divergence estimator is generalization of maximum likelihood estimator and it possesses all its properties. A neural network is proposed which is trained using phi-divergence loss.
Bayesian Classification Methods For Bat Call Identification, Zhongmao Liu
Bayesian Classification Methods For Bat Call Identification, Zhongmao Liu
Arts & Sciences Graduate Student Theses and Dissertations
Bat call classification is widely used in bat population monitoring in the field of ecology. Since bat populations are susceptible to changes in their surroundings, it is essential to monitor bat populations for purposes of bat protection and bio-environment protection. The purpose of this thesis is to compare the performance of several classification methods applied to a data set extracted from audio recordings for different species of bats in Mexico. The methods under comparison are (i) a nonparametric Bayesian approach using a multinomial probit model with Gaussian process prior; (ii) support vector machines (SVM); (iii) naive Bayes; and (iv) Bayesian …
A Comparison Of Machine Learning Techniques For Taxonomic Classification Of Teeth From The Family Bovidae, Gregory J. Matthews, Juliet K. Brophy, Maxwell Luetkemeier, Hongie Gu, George K. Thiruvathukal
A Comparison Of Machine Learning Techniques For Taxonomic Classification Of Teeth From The Family Bovidae, Gregory J. Matthews, Juliet K. Brophy, Maxwell Luetkemeier, Hongie Gu, George K. Thiruvathukal
Mathematics and Statistics: Faculty Publications and Other Works
This study explores the performance of machine learning algorithms on the classification of fossil teeth in the Family Bovidae. Isolated bovid teeth are typically the most common fossils found in southern Africa and they often constitute the basis for paleoenvironmental reconstructions. Taxonomic identification of fossil bovid teeth, however, is often imprecise and subjective. Using modern teeth with known taxons, machine learning algorithms can be trained to classify fossils. Previous work by Brophy et al. [Quantitative morphological analysis of bovid teeth and implications for paleoenvironmental reconstruction of plovers lake, Gauteng Province, South Africa, J. Archaeol. Sci. 41 (2014), pp. …
Multiclass Classification Using Support Vector Machines, Duleep Prasanna W. Rathgamage Don
Multiclass Classification Using Support Vector Machines, Duleep Prasanna W. Rathgamage Don
College of Graduate Studies: Theses & Dissertations
In this thesis, we discuss different SVM methods for multiclass classification and introduce the Divide and Conquer Support Vector Machine (DCSVM) algorithm which relies on data sparsity in high dimensional space and performs a smart partitioning of the whole training data set into disjoint subsets that are easily separable. A single prediction performed between two partitions eliminates one or more classes in a single partition, leaving only a reduced number of candidate classes for subsequent steps. The algorithm continues recursively, reducing the number of classes at each step until a final binary decision is made between the last two classes …
Classification Of High-Dimensional Data Based On Multiple Testing Methods, Chong Ma
Classification Of High-Dimensional Data Based On Multiple Testing Methods, Chong Ma
Theses and Dissertations
Supervised and unsupervised classification are common topics in machine learning in both scientific and industrial fields, which usually involve three tasks: prediction, exploration, and explanation. False discovery rate (FDR) theory has a close connection to classical classification theory, which must be employed in a sophisticated way to achieve good performance in various contexts. The study aims to explore novel supervised classifiers and unsupervised classification approaches for functional data and high-dimensional data in genome study by using FDR, respectively. One work develops a novel classifier for functional data by casting the classification problem into a multiple testing task, which involves using …
Real-Time Classification Of Biomedical Signals, Parkinson’S Analytical Model, Abolfazl Saghafi
Real-Time Classification Of Biomedical Signals, Parkinson’S Analytical Model, Abolfazl Saghafi
USF Tampa Graduate Theses and Dissertations
The reach of technological innovation continues to grow, changing all industries as it evolves. In healthcare, technology is increasingly playing a role in almost all processes, from patient registration to data monitoring, from lab tests to self-care tools. The increase in the amount and diversity of generated clinical data requires development of new technologies and procedures capable of integrating and analyzing the BIG generated information as well as providing support in their interpretation.
To that extent, this dissertation focuses on the analysis and processing of biomedical signals, specifically brain and heart signals, using advanced machine learning techniques. That is, the …
Statistical Methods For Assessing Individual Oocyte Viability Through Gene Expression Profiles, Michael O. Bishop
Statistical Methods For Assessing Individual Oocyte Viability Through Gene Expression Profiles, Michael O. Bishop
All Graduate Plan B and other Reports, Spring 1920 to Spring 2023
Abstract
Statistical Methods for Assessing Individual Oocyte Viability Through Gene Expression Profiles
By
Michael O. Bishop
Utah State University, 2017
Major Professor: Dr. John R. Stevens
Department: Mathematics and Statistics
Oocytes are the precursor cells to the female gamete, or egg. While reproduction may vary from species to species, within humans and most domesticated animals, the oocyte maturation process is fairly similar. As an oocyte matures, there are various processes that take place, all of which have an effect on the viability of the individual oocyte. Barring outside damage that may come to the oocyte, one of the primary reasons …
A Framework For The Statistical Analysis Of Mass Spectrometry Imaging Experiments, Kyle Bemis
A Framework For The Statistical Analysis Of Mass Spectrometry Imaging Experiments, Kyle Bemis
Open Access Dissertations
Mass spectrometry (MS) imaging is a powerful investigation technique for a wide range of biological applications such as molecular histology of tissue, whole body sections, and bacterial films , and biomedical applications such as cancer diagnosis. MS imaging visualizes the spatial distribution of molecular ions in a sample by repeatedly collecting mass spectra across its surface, resulting in complex, high-dimensional imaging datasets. Two of the primary goals of statistical analysis of MS imaging experiments are classification (for supervised experiments), i.e. assigning pixels to pre-defined classes based on their spectral profiles, and segmentation (for unsupervised experiments), i.e. assigning pixels to newly …
Regularized Neural Network To Identify Potential Breast Cancer: A Bayesian Approach, Hansapani S. Rodrigo, Chris P. Tsokos, Taysseer Sharaf
Regularized Neural Network To Identify Potential Breast Cancer: A Bayesian Approach, Hansapani S. Rodrigo, Chris P. Tsokos, Taysseer Sharaf
Journal of Modern Applied Statistical Methods
In the current study, we have exemplified the use of Bayesian neural networks for breast cancer classification using the evidence procedure. The optimal Bayesian network has 81% overall accuracy in correctly classifying the true status of breast cancer patients, 59% sensitivity in correctly detecting the malignancy and 83% specificity in correctly detecting the non-malignancy. The area under the receiver operating characteristic curve (0.7940) shows that this is a moderate classification model.
Advanced Data Analysis - Lecture Notes, Erik B. Erhardt, Edward J. Bedrick, Ronald M. Schrader
Advanced Data Analysis - Lecture Notes, Erik B. Erhardt, Edward J. Bedrick, Ronald M. Schrader
Open Textbooks
Lecture notes for Advanced Data Analysis (ADA1 Stat 427/527 and ADA2 Stat 428/528), Department of Mathematics and Statistics, University of New Mexico, Fall 2016-Spring 2017. Additional material including RMarkdown templates for in-class and homework exercises, datasets, R code, and video lectures are available on the course websites: https://statacumen.com/teaching/ada1 and https://statacumen.com/teaching/ada2 .
Contents
I ADA1: Software
- 0 Introduction to R, Rstudio, and ggplot
II ADA1: Summaries and displays, and one-, two-, and many-way tests of means
- 1 Summarizing and Displaying Data
- 2 Estimation in One-Sample Problems
- 3 Two-Sample Inferences
- 4 Checking Assumptions
- 5 One-Way Analysis of Variance
III ADA1: Nonparametric, categorical, …
Variable Selection For Estimating The Optimal Treatment Regimes In The Presence Of A Large Number Of Covariate, Baqun Zhang, Min Zhang
Variable Selection For Estimating The Optimal Treatment Regimes In The Presence Of A Large Number Of Covariate, Baqun Zhang, Min Zhang
The University of Michigan Department of Biostatistics Working Paper Series
Most of existing methods for optimal treatment regimes, with few exceptions, focus on estimation and are not designed for variable selection with the objective of optimizing treatment decisions. In clinical trials and observational studies, often numerous baseline variables are collected and variable selection is essential for deriving reliable optimal treatment regimes. Although many variable selection methods exist, they mostly focus on selecting variables that are important for prediction (predictive variables) instead of variables that have a qualitative interaction with treatment (prescriptive variables) and hence are important for making treatment decisions. We propose a variable selection method within a general classification …
An Analysis Of Accuracy Using Logistic Regression And Time Series, Edwin Baidoo, Jennifer L. Priestley
An Analysis Of Accuracy Using Logistic Regression And Time Series, Edwin Baidoo, Jennifer L. Priestley
Published and Grey Literature from PhD Candidates
This paper analyzes the accuracy rates for logistic regression and time series models. It also examines a relatively new performance index that takes into consideration the business assumptions of credit markets. Although prior research has focused on evaluation metrics, such as AUC and Gini index, this new measure has a more intuitive interpretation for various managers and decision makers and can be applied to both Logistic and Time Series models.
Better Physical Activity Classification Using Smartphone Acceleration Sensor, Muhammad Arif, Mohsin Bilal, Ahmed Kattan, Sheikh Iqbal Ahamed
Better Physical Activity Classification Using Smartphone Acceleration Sensor, Muhammad Arif, Mohsin Bilal, Ahmed Kattan, Sheikh Iqbal Ahamed
Mathematics, Statistics and Computer Science Faculty Research and Publications
Obesity is becoming one of the serious problems for the health of worldwide population. Social interactions on mobile phones and computers via internet through social e-networks are one of the major causes of lack of physical activities. For the health specialist, it is important to track the record of physical activities of the obese or overweight patients to supervise weight loss control. In this study, acceleration sensor present in the smartphone is used to monitor the physical activity of the user. Physical activities including Walking, Jogging, Sitting, Standing, Walking upstairs and Walking downstairs are classified. Time domain features are extracted …
Identification Of Slums In Mumbai, India: Unsupervised Classification Techniques, Frankie St. Amand
Identification Of Slums In Mumbai, India: Unsupervised Classification Techniques, Frankie St. Amand
Thinking Matters Symposium Archive
Slums are contiguous settlements. Inhabitants lack access to safe water, sanitation and sewage infrastructure, secure housing tenure, uncrowded living space, and permanent, durable housing. Addressing these problematic trends begins with identifying contiguous settlements within Mumbai’s urban fabric. Classifications can be performed using satellite images and remote sensing techniques to yield accurate results. Through literature reviews, socio-cultural analysis, and examination of high resolution satellite imagery, this project aims to develop a systematic, accessible, and reproducible method of classifying Mumbai’s slums.
Regularization Methods For Predicting An Ordinal Response Using Longitudinal High-Dimensional Genomic Data, Jiayi Hou
Theses and Dissertations
Ordinal scales are commonly used to measure health status and disease related outcomes in hospital settings as well as in translational medical research. Notable examples include cancer staging, which is a five-category ordinal scale indicating tumor size, node involvement, and likelihood of metastasizing. Glasgow Coma Scale (GCS), which gives a reliable and objective assessment of conscious status of a patient, is an ordinal scaled measure. In addition, repeated measurements are common in clinical practice for tracking and monitoring the progression of complex diseases. Classical ordinal modeling methods based on the likelihood approach have contributed to the analysis of data in …
Enhancement Of Random Forests Using Trees With Oblique Splits, Andrejus Parfionovas
Enhancement Of Random Forests Using Trees With Oblique Splits, Andrejus Parfionovas
All Graduate Theses and Dissertations, Spring 1920 to Summer 2023
Statistical classification is widely used in many areas where there is a need to make a data-driven decision, or to classify complicated cases or objects. For instance: disease diagnostics (is a patient sick or healthy, based on the blood test results?); weather forecasting (will there be a storm tomorrow, based on today's atmospheric pressure, air temperature, and wind velocity?); speech recognition (what was said over the phone, based on the caller's voice level and articulation); spam detection (can the unsolicited commercial e-mails be identified by their content?); and so on.
Classification trees …
Integrative Biomarker Identification And Classification Using High Throughput Assays, Pan Tong
Integrative Biomarker Identification And Classification Using High Throughput Assays, Pan Tong
Dissertations and Theses (Open Access)
It is well accepted that tumorigenesis is a multi-step procedure involving aberrant functioning of genes regulating cell proliferation, differentiation, apoptosis, genome stability, angiogenesis and motility. To obtain a full understanding of tumorigenesis, it is necessary to collect information on all aspects of cell activity. Recent advances in high throughput technologies allow biologists to generate massive amounts of data, more than might have been imagined decades ago. These advances have made it possible to launch comprehensive projects such as (TCGA) and (ICGC) which systematically characterize the molecular fingerprints of cancer cells using gene expression, methylation, copy number, microRNA and SNP microarrays …
The Net Reclassification Index (Nri): A Misleading Measure Of Prediction Improvement With Miscalibrated Or Overfit Models, Margaret Pepe, Jin Fang, Ziding Feng, Thomas Gerds, Jorgen Hilden
The Net Reclassification Index (Nri): A Misleading Measure Of Prediction Improvement With Miscalibrated Or Overfit Models, Margaret Pepe, Jin Fang, Ziding Feng, Thomas Gerds, Jorgen Hilden
UW Biostatistics Working Paper Series
The Net Reclassification Index (NRI) is a very popular measure for evaluating the improvement in prediction performance gained by adding a marker to a set of baseline predictors. However, the statistical properties of this novel measure have not been explored in depth. We demonstrate the alarming result that the NRI statistic calculated on a large test dataset using risk models derived from a training set is likely to be positive even when the new marker has no predictive information. A related theoretical example is provided in which a miscalibrated risk model that includes an uninformative marker is proven to erroneously …
Borrowing Information Across Populations In Estimating Positive And Negative Predictive Values, Ying Huang, Youyi Fong, John Wei, Ziding Feng
Borrowing Information Across Populations In Estimating Positive And Negative Predictive Values, Ying Huang, Youyi Fong, John Wei, Ziding Feng
UW Biostatistics Working Paper Series
A marker's capacity to predict risk of a disease depends on disease prevalence in the target population and its classification accuracy, i.e. its ability to discriminate diseased subjects from non-diseased subjects. The latter is often considered an intrinsic property of the marker; it is independent of disease prevalence and hence more likely to be similar across populations than risk prediction measures. In this paper, we are interested in evaluating the population-specific performance of a risk prediction marker in terms of positive predictive value (PPV) and negative predictive value (NPV) at given thresholds, when samples are available from the target population …
Class Discovery And Prediction Of Tumor With Microarray Data, Bo Liu
Class Discovery And Prediction Of Tumor With Microarray Data, Bo Liu
All Graduate Theses, Dissertations, and Other Capstone Projects
Current microarray technology is able take a single tissue sample to construct an Affymetrix oglionucleotide array containing (estimated) expression levels of thousands of different genes for that tissue. The objective is to develop a more systematic approach to cancer classification based on Affymetrix oglionucleotide microarrays. For this purpose, I studied published colon cancer microarray data. Colon cancer, with 655,000 deaths worldwide per year, has become the fourth most common form of cancer in the United States and the third leading cause of cancer - related death in the Western world. This research has been focuses in two areas: class discovery, …
An Empirical Approach To Evaluating Sufficient Similarity: Utilization Of Euclidean Distance As A Similarity Measure, Scott Marshall
An Empirical Approach To Evaluating Sufficient Similarity: Utilization Of Euclidean Distance As A Similarity Measure, Scott Marshall
Theses and Dissertations
Individuals are exposed to chemical mixtures while carrying out everyday tasks, with unknown risk associated with exposure. Given the number of resulting mixtures it is not economically feasible to identify or characterize all possible mixtures. When complete dose-response data are not available on a (candidate) mixture of concern, EPA guidelines define a similar mixture based on chemical composition, component proportions and expert biological judgment (EPA, 1986, 2000). Current work in this literature is by Feder et al. (2009), evaluating sufficient similarity in exposure to disinfection by-products of water purification using multivariate statistical techniques and traditional hypothesis testing. The work of …
Cluster And Classification Analysis Of Fossil Invertebrates Within The Bird Spring Formation, Arrow Canyon, Nevada: Implications For Relative Rise And Fall Of Sea-Level, Scott L. Morris
Theses and Dissertations
Carbonate strata preserve indicators of local marine environments through time. Such indicators often include microfossils that have relatively unique conditions under which they can survive, including light, nutrients, salinity, and especially water temperature. As such, microfossils are environmental proxies. When these microfossils are preserved in the rock record, they constitute key components of depositional facies. Spence et al. (2004, 2007) has proposed several approaches for determining the facies of a given stratigraphic succession based upon these proxies. Cluster analysis can be used to determine microfossil groups that represent specific environmental conditions. Identifying which microfossil groups exist through time can indicate …
Statistical Learning And Behrens-Fisher Distribution Methods For Heteroscedastic Data In Microarray Analysis, Nabin K. Manandhr-Shrestha
Statistical Learning And Behrens-Fisher Distribution Methods For Heteroscedastic Data In Microarray Analysis, Nabin K. Manandhr-Shrestha
USF Tampa Graduate Theses and Dissertations
The aim of the present study is to identify the di®erentially expressed genes be- tween two di®erent conditions and apply it in predicting the class of new samples using the microarray data. Microarray data analysis poses many challenges to the statis- ticians because of its high dimensionality and small sample size, dubbed as "small n large p problem". Microarray data has been extensively studied by many statisticians and geneticists. Generally, it is said to follow a normal distribution with equal vari- ances in two conditions, but it is not true in general. Since the number of replications is very small, …
Semiparametric And Nonparametric Methods For Evaluating Risk Prediction Markers In Case-Control Studies, Ying Huang, Margaret Pepe
Semiparametric And Nonparametric Methods For Evaluating Risk Prediction Markers In Case-Control Studies, Ying Huang, Margaret Pepe
UW Biostatistics Working Paper Series
The performance of a well calibrated risk model, Risk(Y)=P(D=1|Y), can be characterized by the population distribution of Risk(Y) and displayed with the predictiveness curve. Better performance is characterized by a wider distribution of Risk(Y), since this corresponds to better risk stratification in the sense that more subjects are identified at low and high risk for the outcome D=1. Although methods have been developed to estimate predictiveness curves from cohort studies, most studies to evaluate novel risk prediction markers employ case-control designs. Here we develop semiparametric and nonparametric methods that accommodate case-control data and assume apriori knowledge of P(D=1). Large and …
Data Mining Methods For Malware Detection, Muazzam Siddiqui
Data Mining Methods For Malware Detection, Muazzam Siddiqui
Electronic Theses and Dissertations
This research investigates the use of data mining methods for malware (malicious programs) detection and proposed a framework as an alternative to the traditional signature detection methods. The traditional approaches using signatures to detect malicious programs fails for the new and unknown malwares case, where signatures are not available. We present a data mining framework to detect malicious programs. We collected, analyzed and processed several thousand malicious and clean programs to find out the best features and build models that can classify a given program into a malware or a clean class. Our research is closely related to information retrieval …