Open Access. Powered by Scholars. Published by Universities.®
- Keyword
-
- Ethnography (6)
- Alcohol (3)
- College (3)
- Social network (3)
- Statistics education (3)
-
- Data culture (2)
- Data infrastructure (2)
- Data sharing (2)
- Infrastructure (2)
- Interval censoring (2)
- Metadata (2)
- Open government data (2)
- Science and technology studies (2)
- Social networks (2)
- Statistical computing (2)
- 1Education in Computational Statistics (1)
- Accessibility (1)
- Accountability (1)
- Adjustable range (1)
- Applications of Computational Statistics (1)
- Area coverage (1)
- Articulation (1)
- Assessment (1)
- Associate’s programs (1)
- Attention (1)
- Bachelor’s programs (1)
- Bacteria (1)
- Baseball (1)
- Baserunning (1)
- Bayesian model (1)
Articles 31 - 53 of 53
Full-Text Articles in Data Science
The Data Science Corps Wrangle-Analyze- Visualize Program: Building Data Acumen For Undergraduate Students, Nicholas J. Horton, Benjamin Baumer, Andrew Zieffler, Valerie Barr
The Data Science Corps Wrangle-Analyze- Visualize Program: Building Data Acumen For Undergraduate Students, Nicholas J. Horton, Benjamin Baumer, Andrew Zieffler, Valerie Barr
Statistical and Data Sciences: Faculty Publications
We congratulate Kolaczyk, Wright, and Yajima on their innovative statistics practicum that places “practice” at the center of data science education (Kolaczyk et al., 2021, this issue). Their year-long practicum course focuses on the data science life cycle with engagement with external partners and university consulting projects. We agree that training postgraduates in practice needs to be foregrounded in the curriculum in order for students to develop necessary depth in data science practice.
Creating Optimal Conditions For Reproducible Data Analysis In R With ‘Fertile’, Audrey M. Bertin, Benjamin Baumer
Creating Optimal Conditions For Reproducible Data Analysis In R With ‘Fertile’, Audrey M. Bertin, Benjamin Baumer
Statistical and Data Sciences: Faculty Publications
The advancement of scientific knowledge increasingly depends on ensuring that data-driven research is reproducible: that two people with the same data obtain the same results. However, while the necessity of reproducibility is clear, there are significant behavioral and technical challenges that impede its widespread implementation and no clear consensus on standards of what constitutes reproducibility in published research. We present fertile, an R package that focuses on a series of common mistakes programmers make while conducting data science projects in R, primarily through the RStudio integrated development environment. fertile operates in two modes: proactively, to prevent reproducibility mistakes from happening …
Teaching Computational Machine Learning (Without Statistics), Katherine M. Kinnaird
Teaching Computational Machine Learning (Without Statistics), Katherine M. Kinnaird
Statistical and Data Sciences: Faculty Publications
This paper presents an undergraduate machine learning course that emphasizes algorithmic understanding and programming skills while assuming no statistical training. Emphasizing the development of good habits of mind, this course trains students to be independent machine learning practitioners through an iterative, cyclical framework for teaching concepts while adding increasing depth and nuance. Beginning with unsupervised learning, this course is sequenced as a series of machine learning ideas and concepts with specific algorithms acting as concrete examples. This paper also details course organization including evaluation practices and logistics.
The Influence Of Peer And Parental Norms On First-Generation College Students’ Binge Drinking Trajectories, Graham T. Diguiseppi, Jordan P. Davis, Matthew K. Meisel, Melissa A. Clark, Mya L. Roberson, Miles Q. Ott, Nancy P. Barnett
The Influence Of Peer And Parental Norms On First-Generation College Students’ Binge Drinking Trajectories, Graham T. Diguiseppi, Jordan P. Davis, Matthew K. Meisel, Melissa A. Clark, Mya L. Roberson, Miles Q. Ott, Nancy P. Barnett
Statistical and Data Sciences: Faculty Publications
Introduction: First-generation college students are those whose parents have not completed a four-year college degree. The current study addressed the lack of research on first-generation college students’ alcohol use by comparing the binge drinking trajectories of first-generation and continuing-generation students over their first three semesters. The dynamic influence of peer and parental social norms on students’ binge drinking frequencies were also examined. Methods: 1342 college students (n = 225 first-generation) at one private University completed online surveys. Group differences were examined at Time 1, and latent growth-curve models tested the association between first-generation status and social norms (peer descriptive, peer …
Identification And Description Of Potentially Influential Social Network Members Using The Strategic Player Approach, Miles Q. Ott, Sara G. Balestrieri, Graham Diguiseppi, Melissa A. Clark, Michael Bernstein, Sarah Helseth, Nancy P. Barnett
Identification And Description Of Potentially Influential Social Network Members Using The Strategic Player Approach, Miles Q. Ott, Sara G. Balestrieri, Graham Diguiseppi, Melissa A. Clark, Michael Bernstein, Sarah Helseth, Nancy P. Barnett
Statistical and Data Sciences: Faculty Publications
Background: Diffusion of innovations theory posits that ideas and behaviors can be spread through social network ties. In intervention work, intervening upon certain network members may lead to intervention effects “diffusing” into the network to affect the behavior of network members who did not receive the intervention. The strategic players (SP) method, an extension of Borgatti’s Key Players approach, is used to balance the (sometimes) opposing goals of spreading the intervention to as many members of the target group as possible, while preventing the spread of the intervention to others. Objectives: We sought to test whether members of the SP …
Supp & Mapp: Adaptable Structure-Based Representations For Mir Tasks, Claire Savard, Erin H. Bugbee, Melissa R, Mcguirl, Katherine M. Kinnaird
Supp & Mapp: Adaptable Structure-Based Representations For Mir Tasks, Claire Savard, Erin H. Bugbee, Melissa R, Mcguirl, Katherine M. Kinnaird
Statistical and Data Sciences: Faculty Publications
Accurate and flexible representations of music data are paramount to addressing MIR tasks, yet many of the existing approaches are difficult to interpret or rigid in nature. This work introduces two new song representations for structure-based retrieval methods: Surface Pattern Preservation (SuPP), a continuous song representation, and Matrix Pattern Preservation (MaPP), SuPP’s discrete counterpart. These representations come equipped with several user-defined parameters so that they are adaptable for a range of MIR tasks. Experimental results show MaPP as successful in addressing the cover song task on a set of Mazurka scores, with a mean precision of 0.965 and recall of …
Do Misperceptions Of Peer Drinking Influence Personal Drinking Behavior? Results From A Complete Social Network Of First-Year College Students, Melissa J. Cox, Angelo M. Dibello, Matthew K. Meisel, Miles Q. Ott, Shannon R. Kenney, Melissa A. Clark, Nancy P. Barnett
Do Misperceptions Of Peer Drinking Influence Personal Drinking Behavior? Results From A Complete Social Network Of First-Year College Students, Melissa J. Cox, Angelo M. Dibello, Matthew K. Meisel, Miles Q. Ott, Shannon R. Kenney, Melissa A. Clark, Nancy P. Barnett
Statistical and Data Sciences: Faculty Publications
This study considered the influence of misperceptions of typical versus self-identified important peers' heavy drinking on personal heavy drinking intentions and frequency utilizing data from a complete social network of college students. The study sample included data from 1,313 students (44% male, 57% White, 15% Hispanic/Latinx) collected during the fall and spring semesters of their freshman year. Students provided perceived heavy drinking frequency for a typical student peer and up to 10 identified important peers. Personal past-month heavy drinking frequency was assessed for all participants at both time points. By comparing actual with perceived heavy drinking frequencies, measures of misperceptions …
A Grammar For Reproducible And Painless Extract-Transform-Load Operations On Medium Data, Benjamin S. Baumer
A Grammar For Reproducible And Painless Extract-Transform-Load Operations On Medium Data, Benjamin S. Baumer
Statistical and Data Sciences: Faculty Publications
Many interesting datasets available on the Internet are of a medium size—too big to fit into a personal computer’s memory, but not so large that they would not fit comfortably on its hard disk. In the coming years, datasets of this magnitude will inform vital research in a wide array of application domains. However, due to a variety of constraints they are cumbersome to ingest, wrangle, analyze, and share in a reproducible fashion. These obstructions hamper thorough peer-review and thus disrupt the forward progress of science. We propose a predictable and pipeable framework for R (the state-of-the-art statistical computing environment) …
Enrollment And Assessment Of A First-Year College Class Social Network For A Controlled Trial Of The Indirect Effect Of A Brief Motivational Intervention, Nancy P. Barnett, Melissa A. Clark, Shannon R. Kenney, Graham Diguiseppi, Matthew K. Meisel, Sara Balestrieri, Miles Q. Ott, John Light
Enrollment And Assessment Of A First-Year College Class Social Network For A Controlled Trial Of The Indirect Effect Of A Brief Motivational Intervention, Nancy P. Barnett, Melissa A. Clark, Shannon R. Kenney, Graham Diguiseppi, Matthew K. Meisel, Sara Balestrieri, Miles Q. Ott, John Light
Statistical and Data Sciences: Faculty Publications
Heavy drinking and its consequences among college students represent a serious public health problem, and peer social networks are a robust predictor of drinking-related risk behaviors. In a recent trial, we administered a Brief Motivational Intervention (BMI) to a small number of first-year college students to assess the indirect effects of the intervention on peers not receiving the intervention. Objectives: To present the research design, describe the methods used to successfully enroll a high proportion of a first-year college class network, and document participant characteristics. Methods: Prior to study enrollment, we consulted with a student advisory group and campus stakeholders …
Data Sharing At Scale: A Heuristic For Affirming Data Cultures, Lindsay Poirier, Brandon Costelloe-Kuehn
Data Sharing At Scale: A Heuristic For Affirming Data Cultures, Lindsay Poirier, Brandon Costelloe-Kuehn
Statistical and Data Sciences: Faculty Publications
Addressing the most pressing contemporary social, environmental, and technological challenges will require integrating insights and sharing data across disciplines, geographies, and cultures. Strengthening international data sharing networks will not only demand advancing technical, legal, and logistical infrastructure for publishing data in open, accessible formats; it will also require recognizing, respecting, and learning to work across diverse data cultures. This essay introduces a heuristic for pursuing richer characterizations of the “data cultures” at play in international, interdisciplinary data sharing. The heuristic prompts cultural analysts to query the contexts of data sharing for a particular discipline, institution, geography, or project at seven …
Classification As Catachresis: Double Binds Of Representing Difference With Semiotic Infrastructure, Lindsay Poirier
Classification As Catachresis: Double Binds Of Representing Difference With Semiotic Infrastructure, Lindsay Poirier
Statistical and Data Sciences: Faculty Publications
Background; This article explores the results of a three-year ethnographic study of how semiotic infrastructures-or digital standards and frameworks such as taxonomies, schemas, and ontologies that encode the meaning of data-are designed. Analysis: It examines debates over best practices in semiotic infrastructure design, such as how much complexity adopted languages should characterize versus how restrictive they should be. It also discusses political and pragmatic considerations that impact what and how information is represented in an information system. Conclusion and implications: This article suggests that all databased representations are forms of data power, and that examining semiotic infrastructure design provides insight …
Relationships Between Social Network Characteristics, Alcohol Use, And Alcohol-Related Consequences In A Large Network Of First-Year College Students: How Do Peer Drinking Norms Fit In?, Graham T. Diguiseppi, Matthew K. Meisel, Sara G. Balestrieri, Miles Q. Ott, Melissa A. Clark, Nancy P. Barnett
Relationships Between Social Network Characteristics, Alcohol Use, And Alcohol-Related Consequences In A Large Network Of First-Year College Students: How Do Peer Drinking Norms Fit In?, Graham T. Diguiseppi, Matthew K. Meisel, Sara G. Balestrieri, Miles Q. Ott, Melissa A. Clark, Nancy P. Barnett
Statistical and Data Sciences: Faculty Publications
A burgeoning area of research is using social network analysis to investigate college students' substance use behaviors. However, little research has incorporated students' perceived peer drinking norms into these analyses. The present study investigated the association between social network characteristics, alcohol use, and alcohol-related consequences among first-year college students (N 1,342; 81% of the first-year class) at one university. The moderating role of descriptive norms was also examined. Network characteristics and descriptive norms were derived from participants' nominations of up to 10 other students who were important to them; individual network characteristics included popularity (indegree), network expansiveness (outdegree), relationship reciprocity, …
Strategic Players For Identifying Optimal Social Network Intervention Subjects, Miles Q. Ott, John M. Light, Melissa A. Clark, Nancy P. Barnett
Strategic Players For Identifying Optimal Social Network Intervention Subjects, Miles Q. Ott, John M. Light, Melissa A. Clark, Nancy P. Barnett
Statistical and Data Sciences: Faculty Publications
We present a method whereby social network ties are used to identify behavioral leaders who are situated in the network such that these individuals are: 1) able to influence other individuals who are in need of and most receptive to intervention, thereby optimizing the impact of the intervention; and 2) not embedded with ties to individuals that are likely to be behaviorally antagonistic to the intervention or that would compromise the optimal impact of intervention. In this study we developed a method that we call Strategic Players, which is a solution for identifying a set of players who are close …
Devious Design: Digital Infrastructure Challenges For Experimental Ethnography, Lindsay Poirier
Devious Design: Digital Infrastructure Challenges For Experimental Ethnography, Lindsay Poirier
Statistical and Data Sciences: Faculty Publications
No abstract provided.
Pushback: Critical Data Designers And Pollution Politics, Kim Fortun, Lindsay Poirier, Alli Morgan, Brandon Costelloe-Kuehn, Mike Fortun
Pushback: Critical Data Designers And Pollution Politics, Kim Fortun, Lindsay Poirier, Alli Morgan, Brandon Costelloe-Kuehn, Mike Fortun
Statistical and Data Sciences: Faculty Publications
In this paper, we describe how critical data designers have created projects that ‘push back’ against the eclipse of environmental problems by dominant orders: the pioneering pollution database Scorecard, released by the US NGO Environmental Defense Fund in 1997; the US Environmental Protection Agency’s EnviroAtlas that brings together numerous data sets and provides tools for valuing ecosystem services; and the Houston Clean Air Network’s maps of real-time ozone levels in Houston. Drawing on ethnographic observations and interviews, we analyse how critical data designers turn scientific data and findings into claims and visualisations that are meaningful in contemporary political terms. The …
A Bayesian Framework For The Classification Of Microbial Gene Activity States, Craig Disselkoen, Brian Greco, Kaitlyn Cook, Kristin Koch, Reginald Lerebours, Chase Viss, Joshua Cape, Elizabeth Held, Yonatan Ashenafi, Karen Fischer, Allyson Acosta, Mark Cunningham, Aaron A. Best, Matthew Dejongh, Nathan Tintle
A Bayesian Framework For The Classification Of Microbial Gene Activity States, Craig Disselkoen, Brian Greco, Kaitlyn Cook, Kristin Koch, Reginald Lerebours, Chase Viss, Joshua Cape, Elizabeth Held, Yonatan Ashenafi, Karen Fischer, Allyson Acosta, Mark Cunningham, Aaron A. Best, Matthew Dejongh, Nathan Tintle
Statistical and Data Sciences: Faculty Publications
Numerous methods for classifying gene activity states based on gene expression data have been proposed for use in downstream applications, such as incorporating transcriptomics data into metabolic models in order to improve resulting flux predictions. These methods often attempt to classify gene activity for each gene in each experimental condition as belonging to one of two states: active (the gene product is part of an active cellular mechanism) or inactive (the cellular mechanism is not active). These existing methods of classifying gene activity states suffer from multiple limitations, including enforcing unrealistic constraints on the overall proportions of active and inactive …
A Multistep Approach To Single Nucleotide Polymorphism-Set Analysis: An Evaluation Of Power And Type I Error Of Gene-Based Tests Of Association After Pathway-Based Association Tests, Alessandra Valcarcel, Kelsey Grinde, Kaitlyn Cook, Alden Green, Nathan Tintle
A Multistep Approach To Single Nucleotide Polymorphism-Set Analysis: An Evaluation Of Power And Type I Error Of Gene-Based Tests Of Association After Pathway-Based Association Tests, Alessandra Valcarcel, Kelsey Grinde, Kaitlyn Cook, Alden Green, Nathan Tintle
Statistical and Data Sciences: Faculty Publications
The aggregation of functionally associated variants given a priori biological information can aid in the discovery of rare variants associated with complex diseases. Many methods exist that aggregate rare variants into a set and compute a single p value summarizing association between the set of rare variants and a phenotype of interest. These methods are often called gene-based, rare variant tests of association because the variants in the set are often all contained within the same gene. A reasonable extension of these approaches involves aggregating variants across an even larger set of variants (eg, all variants contained in genes within …
A General Method For Combining Different Family-Based Rare-Variant Tests Of Association To Improve Power And Robustness Of A Wide Range Of Genetic Architectures, Alden Green, Kaitlyn Cook, Kelsey Grinde, Alessandra Valcarcel, Nathan Tintle
A General Method For Combining Different Family-Based Rare-Variant Tests Of Association To Improve Power And Robustness Of A Wide Range Of Genetic Architectures, Alden Green, Kaitlyn Cook, Kelsey Grinde, Alessandra Valcarcel, Nathan Tintle
Statistical and Data Sciences: Faculty Publications
Current rare-variant, gene-based tests of association often suffer from a lack of statistical power to detect genotype-phenotype associations as a result of a lack of prior knowledge of genetic disease models combined with limited observations of extremely rare causal variants in population-based samples. The use of pedigree data, in which rare variants are often more highly concentrated than in population-based data, has been proposed as 1 possible method for enhancing power. Methods for combining multiple gene-based tests of association into a single summary p value are a robust approach to different genetic architectures when little a priori knowledge is available …
Data Science In Statistics Curricula: Preparing Students To “Think With Data”, J. Hardin, R. Hoerl, Nicholas J. Horton, D. Nolan, B. Baumer, O. Hall-Holt, P. Murrell, R. Peng, P. Roback, D. Temple Lang, M. D. Ward
Data Science In Statistics Curricula: Preparing Students To “Think With Data”, J. Hardin, R. Hoerl, Nicholas J. Horton, D. Nolan, B. Baumer, O. Hall-Holt, P. Murrell, R. Peng, P. Roback, D. Temple Lang, M. D. Ward
Statistical and Data Sciences: Faculty Publications
A growing number of students are completing undergraduate degrees in statistics and entering the workforce as data analysts. In these positions, they are expected to understand how to use databases and other data warehouses, scrape data from Internet sources, program solutions to complex problems in multiple languages, and think algorithmically as well as statistically. These data science topics have not traditionally been a major component of undergraduate programs in statistics. Consequently, a curricular shift is needed to address additional learning outcomes. The goal of this article is to motivate the importance of data science proficiency and to provide examples and …
Evaluating The Impact Of Genotype Errors On Rare Variant Tests Of Association, Kaitlyn Cook, Alejandra Benitez, Casey Fu, Nathan Tintle
Evaluating The Impact Of Genotype Errors On Rare Variant Tests Of Association, Kaitlyn Cook, Alejandra Benitez, Casey Fu, Nathan Tintle
Statistical and Data Sciences: Faculty Publications
The new class of rare variant tests has usually been evaluated assuming perfect genotype information. In reality, rare variant genotypes may be incorrect, and so rare variant tests should be robust to imperfect data. Errors and uncertainty in SNP genotyping are already known to dramatically impact statistical power for single marker tests on common variants and, in some cases, inflate the type I error rate. Recent results show that uncertainty in genotype calls derived from sequencing reads are dependent on several factors, including read depth, calling algorithm, number of alleles present in the sample, and the frequency at which an …
As Strong As The Weakest Link: Mining Diverse Cliques In Weighted Graphs, Petko Bogdanov, Ben Baumer, Prithwish Basu, Amotz Bar-Noy, Ambuj K. Singh
As Strong As The Weakest Link: Mining Diverse Cliques In Weighted Graphs, Petko Bogdanov, Ben Baumer, Prithwish Basu, Amotz Bar-Noy, Ambuj K. Singh
Statistical and Data Sciences: Faculty Publications
Mining for cliques in networks provides an essential tool for the discovery of strong associations among entities. Applications vary, from extracting core subgroups in team performance data arising in sports, entertainment, research and business; to the discovery of functional complexes in high-throughput gene interaction data. A challenge in all of these scenarios is the large size of real-world networks and the computational complexity associated with clique enumeration. Furthermore, when mining for multiple cliques within the same network, the results need to be diversified in order to extract meaningful information that is both comprehensive and representative of the whole dataset. We …
Maximizing Network Lifetime On The Line With Adjustable Sensing Ranges, Amotz Bar-Noy, Ben Baumer
Maximizing Network Lifetime On The Line With Adjustable Sensing Ranges, Amotz Bar-Noy, Ben Baumer
Statistical and Data Sciences: Faculty Publications
Given n sensors on a line, each of which is equipped with a unit battery charge and an adjustable sensing radius, what schedule will maximize the lifetime of a network that covers the entire line? Trivially, any reasonable algorithm is at least a 1/2-approximation, but we prove tighter bounds for several natural algorithms. We focus on developing a linear time algorithm that maximizes the expected lifetime under a random uniform model of sensor distribution. We demonstrate one such algorithm that achieves an average-case approximation ratio of almost 0.9. Most of the algorithms that we consider come from a family based …
Parsing The Relationship Between Baserunning And Batting Abilities Within Lineups, Ben S. Baumer, James Piette, Brad Null
Parsing The Relationship Between Baserunning And Batting Abilities Within Lineups, Ben S. Baumer, James Piette, Brad Null
Statistical and Data Sciences: Faculty Publications
A baseball team's offensive prowess is a function of two types of abilities: batting and baserunning. While each has been studied extensively in isolation, the effects of their interaction is not well understood. We model offensive output as a scalar function f of an individual player's batting and baserunning profile z. Each of these profiles is in turn estimated from Retrosheet data using heirarchical Bayesian models. We then use the SimulOutCome simulation engine as a method to generate values of f(z) over a fine grid of points. Finally, for each of several methods of taking the extra base, we graphically …