Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Genetics (3)
- Genetics and Genomics (3)
- Life Sciences (3)
- Statistical Methodology (3)
- Statistical Theory (3)
-
- Applied Statistics (2)
- Applied Mathematics (1)
- Chemistry (1)
- Climate (1)
- Earth Sciences (1)
- Environmental Chemistry (1)
- Environmental Indicators and Impact Assessment (1)
- Environmental Monitoring (1)
- Environmental Sciences (1)
- Geology (1)
- Hydrology (1)
- Microarrays (1)
- Numerical Analysis and Computation (1)
- Oceanography and Atmospheric Sciences and Meteorology (1)
- Other Earth Sciences (1)
- Water Resource Management (1)
- Publication
- Publication Type
Articles 1 - 5 of 5
Full-Text Articles in Multivariate Analysis
A Geospatial Assessment Of Groundwater Salinization In A Multi-Aquifer System: Durango, Mexico, Juan Lopez-Sierra
A Geospatial Assessment Of Groundwater Salinization In A Multi-Aquifer System: Durango, Mexico, Juan Lopez-Sierra
Graduate Theses/Dissertations
Groundwater salinization poses a critical environmental concern for water resource sustainability in arid and semi-arid regions. This study evaluates spatial and temporal patterns of groundwater salinity across the state of Durango, Mexico, using total dissolved solids (TDS), sodium adsorption ratio (SAR), as salinity indicators and nitrate-nitrogen (NO₃–N) as an anthropogenic indicator. Groundwater quality data were obtained from (CONAGUA), a Mexican water agency. To assess salinity variations with respect to time, while minimizing interannual sampling bias, two multi-year sampling periods were selected: 2012-2013, and 2020-2021. Final datasets consisted of 122 wells for 2012–2013 and 131 wells for 2020–2021. The wells were …
Theory Of Principal Components For Applications In Exploratory Crime Analysis And Clustering, Daniel Silva
Theory Of Principal Components For Applications In Exploratory Crime Analysis And Clustering, Daniel Silva
All Graduate Theses, Dissertations, and Other Capstone Projects
The purpose of this paper is to develop the theory of principal components analysis succinctly from the fundamentals of matrix algebra and multivariate statistics. Principal components analysis is sometimes used as a descriptive technique to explain the variance-covariance or correlation structure of a dataset. However, most often, it is used as a dimensionality reduction technique to visualize a high dimensional dataset in a lower dimensional space. Principal components analysis accomplishes this by using the first few principal components, provided that they account for a substantial proportion of variation in the original dataset. In the same way, the first few principal …
The Clustering Of Regression Models Method With Applications In Gene Expression Data, Li-Xuan Qin, Steven G. Self
The Clustering Of Regression Models Method With Applications In Gene Expression Data, Li-Xuan Qin, Steven G. Self
UW Biostatistics Working Paper Series
Identification of differentially expressed genes and clustering of genes are two important and complementary objectives addressed with gene expression data. For the differential expression question, many "per-gene" analytic methods have been proposed. These methods can generally be characterized as using a regression function to independently model the observations for each gene; various adjustments for multiplicity are then used to interpret the statistical significance of these per-gene regression models over the collection of genes analyzed. Motivated by this common structure of per-gene models, we propose a new model-based clustering method -- the clustering of regression models method, which groups genes that …
Quantification And Visualization Of Ld Patterns And Identification Of Haplotype Blocks, Yan Wang, Sandrine Dudoit
Quantification And Visualization Of Ld Patterns And Identification Of Haplotype Blocks, Yan Wang, Sandrine Dudoit
U.C. Berkeley Division of Biostatistics Working Paper Series
Classical measures of linkage disequilibrium (LD) between two loci, based only on the joint distribution of alleles at these loci, present noisy patterns. In this paper, we propose a new distance-based LD measure, R, which takes into account multilocus haplotypes around the two loci in order to exploit information from neighboring loci. The LD measure R yields a matrix of pairwise distances between markers, based on the correlation between the lengths of shared haplotypes among chromosomes around these markers. Data analysis demonstrates that visualization of LD patterns through the R matrix reveals more deterministic patterns, with much less noise, than …
A New Partitioning Around Medoids Algorithm, Mark J. Van Der Laan, Katherine S. Pollard, Jennifer Bryan
A New Partitioning Around Medoids Algorithm, Mark J. Van Der Laan, Katherine S. Pollard, Jennifer Bryan
U.C. Berkeley Division of Biostatistics Working Paper Series
Kaufman & Rousseeuw (1990) proposed a clustering algorithm Partitioning Around Medoids (PAM) which maps a distance matrix into a specified number of clusters. A particularly nice property is that PAM allows clustering with respect to any specified distance metric. In addition, the medoids are robust representations of the cluster centers, which is particularly important in the common context that many elements do not belong well to any cluster. Based on our experience in clustering gene expression data, we have noticed that PAM does have problems recognizing relatively small clusters in situations where good partitions around medoids clearly exist. In this …