Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

2016

Discipline
Institution
Keyword
Publication

Articles 1 - 15 of 15

Full-Text Articles in Data Science

Pushback: Critical Data Designers And Pollution Politics, Kim Fortun, Lindsay Poirier, Alli Morgan, Brandon Costelloe-Kuehn, Mike Fortun Dec 2016

Pushback: Critical Data Designers And Pollution Politics, Kim Fortun, Lindsay Poirier, Alli Morgan, Brandon Costelloe-Kuehn, Mike Fortun

Statistical and Data Sciences: Faculty Publications

In this paper, we describe how critical data designers have created projects that ‘push back’ against the eclipse of environmental problems by dominant orders: the pioneering pollution database Scorecard, released by the US NGO Environmental Defense Fund in 1997; the US Environmental Protection Agency’s EnviroAtlas that brings together numerous data sets and provides tools for valuing ecosystem services; and the Houston Clean Air Network’s maps of real-time ozone levels in Houston. Drawing on ethnographic observations and interviews, we analyse how critical data designers turn scientific data and findings into claims and visualisations that are meaningful in contemporary political terms. The …


A Bayesian Framework For The Classification Of Microbial Gene Activity States, Craig Disselkoen, Brian Greco, Kaitlyn Cook, Kristin Koch, Reginald Lerebours, Chase Viss, Joshua Cape, Elizabeth Held, Yonatan Ashenafi, Karen Fischer, Allyson Acosta, Mark Cunningham, Aaron A. Best, Matthew Dejongh, Nathan Tintle Aug 2016

A Bayesian Framework For The Classification Of Microbial Gene Activity States, Craig Disselkoen, Brian Greco, Kaitlyn Cook, Kristin Koch, Reginald Lerebours, Chase Viss, Joshua Cape, Elizabeth Held, Yonatan Ashenafi, Karen Fischer, Allyson Acosta, Mark Cunningham, Aaron A. Best, Matthew Dejongh, Nathan Tintle

Statistical and Data Sciences: Faculty Publications

Numerous methods for classifying gene activity states based on gene expression data have been proposed for use in downstream applications, such as incorporating transcriptomics data into metabolic models in order to improve resulting flux predictions. These methods often attempt to classify gene activity for each gene in each experimental condition as belonging to one of two states: active (the gene product is part of an active cellular mechanism) or inactive (the cellular mechanism is not active). These existing methods of classifying gene activity states suffer from multiple limitations, including enforcing unrealistic constraints on the overall proportions of active and inactive …


Tools And Techniques For Computational Reproducibility, Stephen Piccolo, Michael B. Frampton Jul 2016

Tools And Techniques For Computational Reproducibility, Stephen Piccolo, Michael B. Frampton

Faculty Publications

When reporting research findings, scientists document the steps they followed so that others can verify and build upon the research. When those steps have been described in sufficient detail that others can retrace the steps and obtain similar results, the research is said to be reproducible. Computers play a vital role in many research disciplines and present both opportunities and challenges for reproducibility. Computers can be programmed to execute analysis tasks, and those programs can be repeated and shared with others. The deterministic nature of most computer programs means that the same analysis tasks, applied to the same data, will …


The Global Rock-Art Database Project Towards Machine Learning: Building A Collaborative Open Source Platform For Heritage Management From Information Structure To Information Visualization Using Australian Heritage Examples, Robert Haubt Jan 2016

The Global Rock-Art Database Project Towards Machine Learning: Building A Collaborative Open Source Platform For Heritage Management From Information Structure To Information Visualization Using Australian Heritage Examples, Robert Haubt

Staff Scholarship - Australia & Dubai

This guest talk, presented at Lava Lab at the University of Hawaiʻi, explores the intersection of collaboration, data ontology, and information visualization in advancing machine learning within the Global Rock Art Database project. Drawing on insights from the project’s first four years, the talk emphasizes the critical need for cultural heritage preservation by systematically recording and structuring global rock art data in accessible and sustainable ways. This effort not only supports public education on rock art but also facilitates scholarly research.

Key discussions include advancements in data ontology using the CIDOC Conceptual Reference Model (CIDOC CRM) for semantic data management …


Short Distance Structure, L. B. Weinstein, S. E. Kuhn Jan 2016

Short Distance Structure, L. B. Weinstein, S. E. Kuhn

Physics Faculty Publications

Over the last fifteen years of operation, the Jefferson Lab CLAS Collaboration has performed many experiments using nuclear targets. Because the CLAS detector has a very large acceptance and because it used a very open (i.e., nonspecific) trigger, there is a vast amount of data on many different reaction channels yet to be analyzed.

The goal of the Jefferson Lab Nuclear Data Mining grant was to (1) collect the data from nuclear target experiments using the CLAS detector, (2) collect the associated cuts and corrections used to analyze that data, (3) provide non-expert users with a software environment for easy …


Marine Ecoregion And Deepwater Horizon Oil Spill Affect Recruitment And Population Structure Of A Salt Marsh Snail, Steven C. Pennings, Scott Zengel, Jacob Oehrig, Merryl Alber, T. Dale Bishop, Donald R. Deis, Donna Devlin, A. Randall Hughes, John J. Hutchens, Jr., Whitney M. Kiehn, Caroline R. Mcfarlin, Clay L. Montague, Sean P. Powers, C. Edward Proffitt, Nicholle Rutherford, Camille L. Stagg, Keith Walters Jan 2016

Marine Ecoregion And Deepwater Horizon Oil Spill Affect Recruitment And Population Structure Of A Salt Marsh Snail, Steven C. Pennings, Scott Zengel, Jacob Oehrig, Merryl Alber, T. Dale Bishop, Donald R. Deis, Donna Devlin, A. Randall Hughes, John J. Hutchens, Jr., Whitney M. Kiehn, Caroline R. Mcfarlin, Clay L. Montague, Sean P. Powers, C. Edward Proffitt, Nicholle Rutherford, Camille L. Stagg, Keith Walters

University Faculty and Staff Publications

Marine species with planktonic larvae often have high spatial and temporal variation in recruitment that leads to subsequent variation in the ecology of benthic adults. Using a combination of published and unpublished data, we compared the population structure of the salt marsh snail, Littoraria irrorata, between the South Atlantic Bight and the Gulf Coast of the United States to infer geographic differences in recruitment and to test the hypothesis that the Deepwater Horizon oil spill led to widespread recruitment failure of L. irrorata in Louisiana in 2010. Size-frequency distributions in both ecoregions were bimodal, with troughs in the distributions consistent …


Using Data Analytics To Further Understand The Role That Boredom, Loneliness, Social Anxiety, Social Gratification, And Social Relationships (Brag) Play In A Driver’S Decision To Text, Nathan White, Yair Levy, Steven R. Terrell, Steve Bronsburg Jan 2016

Using Data Analytics To Further Understand The Role That Boredom, Loneliness, Social Anxiety, Social Gratification, And Social Relationships (Brag) Play In A Driver’S Decision To Text, Nathan White, Yair Levy, Steven R. Terrell, Steve Bronsburg

All Faculty Scholarship for the College of Education and Professional Studies

Texting while driving is a growing problem that current efforts have failed to curtail. This behavior has serious, and sometimes fatal, consequences, and the factors that cause a driver to text are not well understood. This study investigates the influence that boredom, social relationships, social anxiety, and social gratification (BRAG) have upon the texting driver. A survey instrument was used to collect data from 297 respondents at a mid-sized regional university in the Pacific Northwest of the United States. The data was evaluated with PLS-SEM, which indicated that social gratification plays a very significant role in a driver’s decision to …


A Hierarchical Statistical Engineering Modeling Methodology, Teddy Steven Cotter Jan 2016

A Hierarchical Statistical Engineering Modeling Methodology, Teddy Steven Cotter

Engineering Management & Systems Engineering Faculty Publications

In the ASEM-IAC 2015, Cotter (2015) proposed a systemic joint deterministic-stochastic dynamic causal Bayesian statistical engineering model that addressed the knowledge gap needed to integrate deterministic mathematical engineering models within a stochastic framework. However, Cotter did not specify the modeling methodology through which statistical engineering models could be developed, diagnosed, and applied to predict systemic mission performance. This paper updates research into the development a hierarchical statistical engineering modeling methodology and sets forth the initial theoretical foundation for the methodology.


Two Influential Primate Classifications Logically Aligned, Nico M. Franz, Naomi M. Pier, Deeann M. Reeder, Mingmin Chen, Shizhuo Yu, Parisa Kianmajd, Shaun Bowers, Bertram Ludäscher Jan 2016

Two Influential Primate Classifications Logically Aligned, Nico M. Franz, Naomi M. Pier, Deeann M. Reeder, Mingmin Chen, Shizhuo Yu, Parisa Kianmajd, Shaun Bowers, Bertram Ludäscher

Computer Science Faculty Scholarship

Classifications and phylogenies of perceived natural entities change in the light of new evidence. Taxonomic changes, translated into Code-compliant names, frequently lead to name:meaning dissociations across succeeding treatments. Classification standards such as the Mammal Species of the World (MSW) may experience significant levels of taxonomic change from one edition to the next, with potential costs to long-term, large-scale information integration. This circumstance challenges the biodiversity and phylogenetic data communities to express taxonomic congruence and incongruence inways that both humans and machines can process, that is, to logically represent taxonomic alignments across multiple classifications.We demonstrate that such alignments are feasible for …


Names Are Not Good Enough: Reasoning Over Taxonomic Change In The Andropogon Complex, Nico M. Franz, Mingmin Chen, Parisa Kianmajd, Shizhuo Yu, Shaun Bowers, Alan S. Weakley, Bertram Ludäscher Jan 2016

Names Are Not Good Enough: Reasoning Over Taxonomic Change In The Andropogon Complex, Nico M. Franz, Mingmin Chen, Parisa Kianmajd, Shizhuo Yu, Shaun Bowers, Alan S. Weakley, Bertram Ludäscher

Computer Science Faculty Scholarship

We present a novel, logic-based solution to the challenge of reconciling the meanings of taxonomic names across multiple biological taxonomies. The challenge arises due to limitations inherent in using type-anchored taxonomic names as identifiers of granular semantic similarities and differences being expressed in original and revised taxonomic classifications. We address this challenge through: (1) the use of taxonomic concept labels – thereby individuating name usages according to particular sources and allowing each taxonomy to be recognized separately; (2) sets of user-provided Region Connection Calculus articulations among concepts (RCC-5: congruence, proper inclusion, inverse proper inclusion, overlap, exclusion); and (3) the use …


A Multistep Approach To Single Nucleotide Polymorphism-Set Analysis: An Evaluation Of Power And Type I Error Of Gene-Based Tests Of Association After Pathway-Based Association Tests, Alessandra Valcarcel, Kelsey Grinde, Kaitlyn Cook, Alden Green, Nathan Tintle Jan 2016

A Multistep Approach To Single Nucleotide Polymorphism-Set Analysis: An Evaluation Of Power And Type I Error Of Gene-Based Tests Of Association After Pathway-Based Association Tests, Alessandra Valcarcel, Kelsey Grinde, Kaitlyn Cook, Alden Green, Nathan Tintle

Statistical and Data Sciences: Faculty Publications

The aggregation of functionally associated variants given a priori biological information can aid in the discovery of rare variants associated with complex diseases. Many methods exist that aggregate rare variants into a set and compute a single p value summarizing association between the set of rare variants and a phenotype of interest. These methods are often called gene-based, rare variant tests of association because the variants in the set are often all contained within the same gene. A reasonable extension of these approaches involves aggregating variants across an even larger set of variants (eg, all variants contained in genes within …


A General Method For Combining Different Family-Based Rare-Variant Tests Of Association To Improve Power And Robustness Of A Wide Range Of Genetic Architectures, Alden Green, Kaitlyn Cook, Kelsey Grinde, Alessandra Valcarcel, Nathan Tintle Jan 2016

A General Method For Combining Different Family-Based Rare-Variant Tests Of Association To Improve Power And Robustness Of A Wide Range Of Genetic Architectures, Alden Green, Kaitlyn Cook, Kelsey Grinde, Alessandra Valcarcel, Nathan Tintle

Statistical and Data Sciences: Faculty Publications

Current rare-variant, gene-based tests of association often suffer from a lack of statistical power to detect genotype-phenotype associations as a result of a lack of prior knowledge of genetic disease models combined with limited observations of extremely rare causal variants in population-based samples. The use of pedigree data, in which rare variants are often more highly concentrated than in population-based data, has been proposed as 1 possible method for enhancing power. Methods for combining multiple gene-based tests of association into a single summary p value are a robust approach to different genetic architectures when little a priori knowledge is available …


Mining Human Activity Using Dimensionality Reduction And Pattern Recognition, Ismail El Moudden, Mounir Ouzir, Badreddine Benyacoub, Souad El Bernoussi Jan 2016

Mining Human Activity Using Dimensionality Reduction And Pattern Recognition, Ismail El Moudden, Mounir Ouzir, Badreddine Benyacoub, Souad El Bernoussi

Research and Infrastructure Service Enterprise (RISE) Faculty Publications

Human activity recognition (HAR) is an emerging research topic in pattern recognition, especially in computer vision. The main objective of human activity recognition is to automatically detect and analyze human activities from the information acquired from different sensors. Human activity prediction using big data remains a challengingly open problem. Several approaches have recently been developed in order to find practical ways to solve high dimensionality of data problems. The aim of this study is to attempt, using data mining techniques, to deal with HAR modeling involving a significant number of variables in order to identify relevant parameters from data and …


An Improved Smote Algorithm Based On Genetic Algorithm For Imbalanced Data Collection, Qiong Gu, Xian-Ming Wang, Zhao Wu, Bing Ning, Chun-Sheng Xin Jan 2016

An Improved Smote Algorithm Based On Genetic Algorithm For Imbalanced Data Collection, Qiong Gu, Xian-Ming Wang, Zhao Wu, Bing Ning, Chun-Sheng Xin

Electrical & Computer Engineering Faculty Publications

Classification of imbalanced data has been recognized as a crucial problem in machine learning and data mining. In an imbalanced dataset, minority class instances are likely to be misclassified. When the synthetic minority over-sampling technique (SMOTE) is applied in imbalanced dataset classification, the same sampling rate is set for all samples of the minority class in the process of synthesizing new samples, this scenario involves blindness. To overcome this problem, an improved SMOTE algorithm based on genetic algorithm (GA), namely, GASMOTE was proposed. First, GASMOTE set different sampling rates for different minority class samples. A combination of the sampling rates …


Results And Challenges In Visualizing Analytic Provenance Of Text Analysis Tasks Using Interaction Logs, Rhema Linder, Alyssa M. Peña, Sampath Jayarathna, Eric D. Ragan Jan 2016

Results And Challenges In Visualizing Analytic Provenance Of Text Analysis Tasks Using Interaction Logs, Rhema Linder, Alyssa M. Peña, Sampath Jayarathna, Eric D. Ragan

Computer Science Faculty Publications

After data analysis, recalling and communicating the steps and rationale followed during the analysis can be difficult. This paper explores the use of interaction logs to generate summaries of an analyst's interest based on interactions with specific data items in a text analysis scenario. Our approach uses data-interaction events as a proxy for user interest in and experience of information. Logging can produce verbose logs that detail all available readable content, so the discussed approach uses topic modeling (LDA) over different time segments to summarize the verbose information and generate visualizations of the history of user interest. Our preliminary results …