Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

Articles 1 - 30 of 60

Full-Text Articles in Data Science

Tidychangepoint: A Unified Framework For Analyzing Changepoint Detection In Univariate Time Series, Ben Baumer, Biviana Marcela Suárez Sierra Feb 2026

Tidychangepoint: A Unified Framework For Analyzing Changepoint Detection In Univariate Time Series, Ben Baumer, Biviana Marcela Suárez Sierra

Statistical and Data Sciences: Faculty Publications

We present tidychangepoint, a new R package for changepoint detection analysis. Most R packages for segmenting univariate time series focus on providing one or two algorithms for changepoint detection that work with a small set of models and penalized objective functions, and all of them return a custom, nonstandard object type. This makes comparing results across various algorithms, models, and penalized objective functions unnecessarily difficult. tidychangepoint solves this problem by wrapping functions from a variety of existing packages and storing the results in a common S3 class called tidycpt. The package then provides functionality for easily extracting comparable numeric or …


Examining The Intersectional And Structural Issues Of Routine Healthcare Utilization And Access Inequities For Lgb People With Chronic Diseases, Shiya Cao, Mehreen Mirza, Sophia Silovsky, Nicole Tresvalles, Lucia Qin, Sarah Susnea Dec 2025

Examining The Intersectional And Structural Issues Of Routine Healthcare Utilization And Access Inequities For Lgb People With Chronic Diseases, Shiya Cao, Mehreen Mirza, Sophia Silovsky, Nicole Tresvalles, Lucia Qin, Sarah Susnea

Statistical and Data Sciences: Faculty Publications

In the United States, although the gaps in health insurance coverage by sexual orientation have been closing since the implementation of the Affordable Care Act and legalization of same-sex marriage, the LGB group (i.e., lesbian, gay, bisexual) continues to report healthcare utilization and access inequities such as more delayed or unmet care. The extant research has often examined healthcare utilization and access inequities due to affordability (e.g., out-of-pocket costs). However, healthcare utilization and access inequities are only partially explained by cost reasons; there are non-cost reasons that have not been adequately empirically examined. The present study innovatively includes discrimination structural …


A Review Of Research And Practices On Teaching Data Visualizations For Blind And Visually Impaired Students, Shiya Cao Oct 2025

A Review Of Research And Practices On Teaching Data Visualizations For Blind And Visually Impaired Students, Shiya Cao

Statistical and Data Sciences: Faculty Publications

Around 36 million people in the world are blind and an additional 217 million have moderate to severe vision impairment. In higher education, four percent of 54,204 undergraduates who participated in the 2022 American College Health Association survey reported to be blind or have low vision. Those students frequently do not have access to data visualizations we generally teach and use in postsecondary statistics and data science classes. The design of those visualizations is premised on implicit assumptions about the user’s visual ability. Making data visualizations accessible to blind and visually impaired (BVI) people would help improve equity in higher …


Educational Opportunities Of Participatory Gis For Accessibility On A College Campus, Shiya Cao, Heather Rosenfeld, Sarah Susnea Aug 2025

Educational Opportunities Of Participatory Gis For Accessibility On A College Campus, Shiya Cao, Heather Rosenfeld, Sarah Susnea

Statistical and Data Sciences: Faculty Publications

The educational benefits of Participatory GIS (PGIS) in geographic higher education have received limited direct attention, often because of the complexities of integrating PGIS into university curricula. While a few exceptions found important educational benefits of PGIS, extant studies focused primarily on the educational benefits for students who worked in the research teams, instead of participants who contributed their local knowledge and perspectives to mapping. Our research aims to understand the educational benefits of PGIS for participants in a campus accessibility mapping project using the modes of experiential learning, positionality, and service learning. Through this, we also provide strategies for …


Making Plans Findable, Accessible, Interoperable, And Reusable With Data Infrastructure: A Search Engine For Constructing, Analyzing, And Visualizing Planning Documents, Lindsay Poirier, Dexter Antonio, Makenna Dettmann, Tiffany Eng, Jennifer Ganata, Sujoy Ghosh, Mirthala Lopez, Ranesh Karma, Asiya Natekal, Catherine Brinkley Oct 2024

Making Plans Findable, Accessible, Interoperable, And Reusable With Data Infrastructure: A Search Engine For Constructing, Analyzing, And Visualizing Planning Documents, Lindsay Poirier, Dexter Antonio, Makenna Dettmann, Tiffany Eng, Jennifer Ganata, Sujoy Ghosh, Mirthala Lopez, Ranesh Karma, Asiya Natekal, Catherine Brinkley

Statistical and Data Sciences: Faculty Publications

Local land-use plans help guide future development, but it is often difficult to compare content across jurisdictions, making regional coordination and plan evaluation challenging. This research reviews federal, state, and local data infrastructure guidance for land-use plans and compares such guidance to compliance with a California use-case. Findings indicate a number of obstacles to fostering data sharing and comparative analysis of plans: there is currently no central repository of land-use plans; plans are not uniform in format and are often out of date; many plans are not machine-readable thereby inhibiting text extraction, and planning language varies so greatly that there …


Enacting Data Context: Fixing Meaning In Transparency Data Initiatives, Lindsay Poirier Jul 2024

Enacting Data Context: Fixing Meaning In Transparency Data Initiatives, Lindsay Poirier

Statistical and Data Sciences: Faculty Publications

This article documents the “context cultures” underpinning efforts to develop regulations for collecting and reporting data in a United States public database known as Open Payments. Open Payments is a dataset published annually by the US Center for Medicare and Medicaid Services that documents the transfers of value from pharmaceutical and medical device manufacturers to physicians, prescribing non-physicians, and teaching hospitals. In the article, I show context became a manifold concern as differentially-situated actors engaged in modes of public advocacy and social action around not only what data meant, but also what it meant to make data meaningful. I show …


Per- And Polyfluoroalkyl Substance Exposure Risks In Us Carceral Facilities, 2022, Lindsay Poirier, Derrick Salvatore, Phil Brown, Alissa Cordner, Kira Mok, Nicholas Shapiro May 2024

Per- And Polyfluoroalkyl Substance Exposure Risks In Us Carceral Facilities, 2022, Lindsay Poirier, Derrick Salvatore, Phil Brown, Alissa Cordner, Kira Mok, Nicholas Shapiro

Statistical and Data Sciences: Faculty Publications

Objectives. To assess the US incarcerated population’s risk of exposure to per- and polyfluoroalkyl substances (PFASs). Methods. We assessed how many of the 6118 US carceral facilities were located in the same hydrologic unit code watershed boundaries as known or likely locations of PFAS contamination. We conducted geospatial analyses on data aggregated from Environmental Protection Agency databases and a PFAS site tracker in 2022 to model the hydrologically feasible known and presumptive PFAS contamination sites for nearly 2 million incarcerated people. Results. Findings indicate that 5% (∼310) of US carceral facilities have at least 1 known source of PFAS contamination …


Examining Information Systems Use To Facilitate The Workplace Accommodation Process, Shiya Cao Jan 2024

Examining Information Systems Use To Facilitate The Workplace Accommodation Process, Shiya Cao

Statistical and Data Sciences: Faculty Publications

BACKGROUND: The workplace accommodation process is often affected by ineffective and inefficient communications and information exchanges among disabled employees and other stakeholders. Information systems (IS) can play a key role in facilitating a more effective and efficient accommodation process since IS has been shown to facilitate business processes and effect positive organizational changes.

OBJECTIVE: Since there is little to no research that exists on IS use to facilitate the workplace accommodation process, this paper, as a critical first step, examines how IS have been used in the accommodation process.

METHODS: Thirty-six interviews were conducted with disabled employees from various organizations. …


Questions (And Answers) For Incorporating Nontraditional Grading In Your Statistics Courses, Brenna Curley Jan 2024

Questions (And Answers) For Incorporating Nontraditional Grading In Your Statistics Courses, Brenna Curley

Statistical and Data Sciences: Faculty Publications

Nontraditional grading methods have recently become more common, and as with any large pedagogical shift, there are a number of questions to consider when applying a new grading scheme to a course. This article summarizes four types of nontraditional grading and shares experiences from the authors who have applied them to a variety of courses in statistics. This article is structured as a set of questions and answers, seeking to address many of the concerns and considerations that one may face as they transition a course’s grading structure. Supplementary materials for this article are available online.


Population Modeling With Machine Learning Can Enhance Measures Of Mental Health - Open-Data Replication, Ty Easley, Ruiqi Chen, Kayla Hannon, Rosie Dutt, Janine Bijsterbosch Jun 2023

Population Modeling With Machine Learning Can Enhance Measures Of Mental Health - Open-Data Replication, Ty Easley, Ruiqi Chen, Kayla Hannon, Rosie Dutt, Janine Bijsterbosch

Statistical and Data Sciences: Faculty Publications

Efforts to predict trait phenotypes based on functional MRI data from large cohorts have been hampered by low prediction accuracy and/or small effect sizes. Although these findings are highly replicable, the small effect sizes are somewhat surprising given the presumed brain basis of phenotypic traits such as neuroticism and fluid intelligence. We aim to replicate previous work and additionally test multiple data manipulations that may improve prediction accuracy by addressing data pollution challenges. Specifically, we added additional fMRI features, averaged the target phenotype across multiple measurements to obtain more accurate estimates of the underlying trait, balanced the target phenotype's distribution …


Big Ideas In Sports Analytics And Statistical Tools For Their Investigation, Benjamin S. Baumer, Gregory J. Matthews, Quang Nguyen May 2023

Big Ideas In Sports Analytics And Statistical Tools For Their Investigation, Benjamin S. Baumer, Gregory J. Matthews, Quang Nguyen

Statistical and Data Sciences: Faculty Publications

Sports analytics—broadly defined as the pursuit of improvement in athletic performance through the analysis of data—has expanded its footprint both in the professional sports industry and in academia over the past 30 years. In this article, we connect four big ideas that are common across multiple sports: the expected value of a game state, win probability, measures of team strength, and the use of sports betting market data. For each, we explore both the shared similarities and individual idiosyncracies of analytical approaches in each sport. While our focus is on the concepts underlying each type of analysis, any implementation necessarily …


A Diffusion Network Event History Estimator, Jeffrey J. Harden, Bruce A. Desmarais, Mark Brockway, Frederick J. Boehmke, Scott J. Lacombe, Fridolin Linder, Hanna Wallach Apr 2023

A Diffusion Network Event History Estimator, Jeffrey J. Harden, Bruce A. Desmarais, Mark Brockway, Frederick J. Boehmke, Scott J. Lacombe, Fridolin Linder, Hanna Wallach

Government: Faculty Publications

Research on the diffusion of political decisions across jurisdictions typically accounts for units’ influence over each other with (1) observable measures or (2) by inferring latent network ties from past decisions. The former approach assumes that interdependence is static and perfectly captured by the data. The latter mitigates these issues but requires analytical tools that are separate from the main empirical methods for studying diffusion. As a solution, we introduce network event history analysis (NEHA), which incorporates latent network inference into conventional discrete-time event history models. We demonstrate NEHA’s unique methodological and substantive benefits in applications to policy adoption in …


Institutional Design And Policy Responsiveness In Us States, Scott J. Lacombe Mar 2023

Institutional Design And Policy Responsiveness In Us States, Scott J. Lacombe

Government: Faculty Publications

There is significant disagreement on the moderating role of institutions on policy responsive- ness, yet overwhelmingly research in state politics has focused on single institutions. This project leverages a new aggregate scale of state institutions to evaluate if the collective insti- tutional context moderates the influence of public opinion on policy. I use a recently released latent scale of institutional context and find that high levels of accountability pressure strongly strengthen public opinion’s influence on policy for both economic and social policy, while the strength of a state’s checks and balance system is largely unrelated to policy responsiveness. These results …


Data Science Transfer Pathways From Associate's To Bachelor's Programs, Benjamin S. Baumer, Nicholas J. Horton Jan 2023

Data Science Transfer Pathways From Associate's To Bachelor's Programs, Benjamin S. Baumer, Nicholas J. Horton

Statistical and Data Sciences: Faculty Publications

A substantial fraction of students who complete their college education at a public university in the United States begin their journey at one of the 935 public 2-year colleges. While the number of 4-year colleges offering bachelor’s degrees in data science continues to increase, data science instruction at many 2-year colleges lags behind. A major impediment is the relative paucity of introductory data science courses that serve multiple student audiences and can easily transfer. In addition, the lack of predefined transfer pathways (or articulation agreements) for data science creates a growing disconnect that leaves students who want to study data …


Psychometric Properties Of A Combined Go/No-Go And Continuous Performance Task Across Childhood, Caron A.C. Clark, Kaitlyn Cook, Rui Wang, Michael Rueschman, Jerilynn Radcliffe, Susan Redline, H. Gerry Taylor Jan 2023

Psychometric Properties Of A Combined Go/No-Go And Continuous Performance Task Across Childhood, Caron A.C. Clark, Kaitlyn Cook, Rui Wang, Michael Rueschman, Jerilynn Radcliffe, Susan Redline, H. Gerry Taylor

Statistical and Data Sciences: Faculty Publications

Despite the critical importance of attention for children’s self-regulation and mental health, there are few task-based measures of this construct appropriate for use across a wide childhood age range including very young children. Three versions of a combined go/no-go and continuous performance task (GNG/CPT) were created with varying length and timing parameters to maximize their appropriateness for age groups spanning early to middle childhood. As part of the baseline assessment of a clinical trial, 452 children aged 3–12 years (50% male, 50% female; 52% White, non-Hispanic, 27% Black, 16% Hispanic/Latinx; 6% other ethnicity/race) completed the task. Confirmatory factor analysis indicated …


Attending To The Cultures Of Data Science Work, Lindsay Poirier Jan 2023

Attending To The Cultures Of Data Science Work, Lindsay Poirier

Statistical and Data Sciences: Faculty Publications

This essay reflects on the shifting attention to the “social” and the “cultural” in data science communities. While recently the “social” and the “cultural” have been prioritized in data science discourse, social and cultural concerns that get raised in data science are almost always outwardly focused – applying to the communities that data scientists seek to support more so than more computationally-focused data science communities. I argue that data science communities have a responsibility to attend not only to the cultures that orient the work of domain communities, but also to the cultures that orient their own work. I describe …


Evaluation Of Edison's Data Science Competency Framework Through A Comparative Literature Analysis, Karl R. B. Schmitt, Linda Clark, Katherine M. Kinnaird, Ruth E. H. Wertz, Björn Sandstede Jan 2023

Evaluation Of Edison's Data Science Competency Framework Through A Comparative Literature Analysis, Karl R. B. Schmitt, Linda Clark, Katherine M. Kinnaird, Ruth E. H. Wertz, Björn Sandstede

Statistical and Data Sciences: Faculty Publications

During the emergence of Data Science as a distinct discipline, discussions of what exactly constitutes Data Science have been a source of contention, with no clear resolution. These disagreements have been exacerbated by the lack of a clear single disciplinary 'parent.' Many early efforts at defining curricula and courses exist, with the EDISON Project's Data Science Framework (EDISON-DSF) from the European Union being the most complete. The EDISON-DSF includes both a Data Science Body of Knowledge (DS-BoK) and Competency Framework (CF-DS). This paper takes a critical look at how EDISON's CF-DS compares to recent work and other published curricular or …


A Multistate Competing Risks Framework For Preconception Prediction Of Pregnancy Outcomes, Kaitlyn Cook, Neil J. Perkins, Enrique Schisterman, Sebastien Haneuse Dec 2022

A Multistate Competing Risks Framework For Preconception Prediction Of Pregnancy Outcomes, Kaitlyn Cook, Neil J. Perkins, Enrique Schisterman, Sebastien Haneuse

Statistical and Data Sciences: Faculty Publications

Background: Preconception pregnancy risk profiles—characterizing the likelihood that a pregnancy attempt results in a full-term birth, preterm birth, clinical pregnancy loss, or failure to conceive—can provide critical information during the early stages of a pregnancy attempt, when obstetricians are best positioned to intervene to improve the chances of successful conception and full-term live birth. Yet the task of constructing and validating risk assessment tools for this earlier intervention window is complicated by several statistical features: the final outcome of the pregnancy attempt is multinomial in nature, and it summarizes the results of two intermediate stages, conception and gestation, whose outcomes …


Extracellular Dnases Facilitate Antagonism And Coexistence In Bacterial Competitor-Sensing Interference Competition, Aoi Ogawa, Christophe Golé, Maria Bermudez, Odrine Habarugira, Gabrielle Joslin, Taylor Mccain, Autumn Mineo, Jennifer Wise, Julie Xiong, Katherine Yan, Jan A.C. Vriezen Nov 2022

Extracellular Dnases Facilitate Antagonism And Coexistence In Bacterial Competitor-Sensing Interference Competition, Aoi Ogawa, Christophe Golé, Maria Bermudez, Odrine Habarugira, Gabrielle Joslin, Taylor Mccain, Autumn Mineo, Jennifer Wise, Julie Xiong, Katherine Yan, Jan A.C. Vriezen

Biological Sciences: Faculty Publications

Over the last 4 decades, the rate of discovery of novel antibiotics has decreased drastically, ending the era of fortuitous antibiotic discovery. A better understanding of the biology of bacteriogenic toxins potentially helps to prospect for new antibiotics. To initiate this line of research, we quantified antagonists from two different sites at two different depths of soil and found the relative number of antagonists to correlate with the bacterial load and carbon-to-nitrogen (C/N) ratio of the soil. Consecutive studies show the importance of antagonist interactions between soil isolates and the lack of a predicted role for nutrient availability and, therefore, …


Marginal Proportional Hazards Models For Clustered Interval-Censored Data With Time-Dependent Covariates, Kaitlyn Cook, Wenbin Lu, Rui Wang Oct 2022

Marginal Proportional Hazards Models For Clustered Interval-Censored Data With Time-Dependent Covariates, Kaitlyn Cook, Wenbin Lu, Rui Wang

Statistical and Data Sciences: Faculty Publications

The Botswana Combination Prevention Project was a cluster-randomized HIV prevention trial whose follow-up period coincided with Botswana’s national adoption of a universal test-and-treat strategy for HIV management. Of interest is whether, and to what extent, this change in policy (i) modified the observed preventative effects of the study intervention and (ii) was associated with a reduction in the population-level incidence of HIV in Botswana. To address these questions, we propose a stratified proportional hazards model for clustered intervalcensored data with time-dependent covariates and develop a composite expectation maximization algorithm that facilitates estimation of model parameters without placing parametric assumptions on …


The Link Between Democratic Institutions And Population Health In The American States, Julianna Pacheco, Scott Lacombe Oct 2022

The Link Between Democratic Institutions And Population Health In The American States, Julianna Pacheco, Scott Lacombe

Government: Faculty Publications

Context: This project investigates the role of state-level institutions in explaining variation in population health in the American states. Although cross-national research has established the positive effects of democracy on population health, little attention has been given to subnational units. The authors leverage a new data set to understand how political accountability and a system of checks and balances are associated with state population health. Methods: The authors estimate error correction models and two-way fixed effects models to estimate how the strength of state-level democratic institutions is associated with infant mortality rates, life expectancy, and midlife mortality. Findings: The authors …


Implementing Github Actions Continuous Integration To Reduce Error Rates In Ecological Data Collection, Albert Y. Kim, Valentine Herrmann, Ross Barreto, Brianna Calkins, Erika Gonzalez-Akre, Daniel J. Johnson, Jennifer A. Jordan, Lukas Magee, Ian R. Mcgregor, Nicolle Montero, Karl Novak, Teagan Rogers, Jessica Shue, Kristina J. Anderson-Teixeira Sep 2022

Implementing Github Actions Continuous Integration To Reduce Error Rates In Ecological Data Collection, Albert Y. Kim, Valentine Herrmann, Ross Barreto, Brianna Calkins, Erika Gonzalez-Akre, Daniel J. Johnson, Jennifer A. Jordan, Lukas Magee, Ian R. Mcgregor, Nicolle Montero, Karl Novak, Teagan Rogers, Jessica Shue, Kristina J. Anderson-Teixeira

Statistical and Data Sciences: Faculty Publications

Accurate field data are essential to understanding ecological systems and forecasting their responses to global change. Yet, data collection errors are common, and data analysis often lags far enough behind its collection that many errors can no longer be corrected, nor can anomalous observations be revisited. Needed is a system in which data quality assurance and control (QA/QC), along with the production of basic data summaries, can be automated immediately following data collection.

Here, we implement and test a system to satisfy these needs. For two annual tree mortality censuses and a dendrometer band survey at two forest research sites, …


Accountable Data: The Politics And Pragmatics Of Disclosure Datasets, Lindsay Poirier Jun 2022

Accountable Data: The Politics And Pragmatics Of Disclosure Datasets, Lindsay Poirier

Statistical and Data Sciences: Faculty Publications

This paper attends specifically to what I call "disclosure datasets"- tabular datasets produced in accordance with laws requiring various kinds of disclosure. For the purposes of this paper, the most significant defining feature of disclosure datasets is that they aggregate information produced and reported by the same institutions they are meant to hold accountable. Through a series of case studies of disclosure datasets in the United States, I specifically draw attention to two concerns with disclosure datasets: First, for disclosure datasets, there is often political and social mobilization around the definitions that determine reporting thresholds, which in turn implicates what …


An Educator’S Perspective Of The Tidyverse, Mine Çetinkaya-Rundel, Johanna Hardin, Benjamin Baumer, Amelia Mcnamara, Nicholas J. Horton, Colin W. Rundel Apr 2022

An Educator’S Perspective Of The Tidyverse, Mine Çetinkaya-Rundel, Johanna Hardin, Benjamin Baumer, Amelia Mcnamara, Nicholas J. Horton, Colin W. Rundel

Statistical and Data Sciences: Faculty Publications

Computing makes up a large and growing component of data science and statistics courses. Many of those courses, especially when taught by faculty who are statisticians by training, teach R as the programming language. A number of instructors have opted to build much of their teaching around use of the tidyverse. The tidyverse, in the words of its developers, “is a collection of R packages that share a high-level design philosophy and low-level grammar and data structures, so that learning one package makes it easier to learn the next” (Wickham et al. 2019). These shared principles have led to the …


Mental Health In The Uk Biobank: A Roadmap To Self-Report Measures And Neuroimaging Correlates, Rosie K. Dutt, Kayla Hannon, Ty O. Easley, Joseph C. Griffis, Wei Zhang, Janine D. Bijsterbosch Feb 2022

Mental Health In The Uk Biobank: A Roadmap To Self-Report Measures And Neuroimaging Correlates, Rosie K. Dutt, Kayla Hannon, Ty O. Easley, Joseph C. Griffis, Wei Zhang, Janine D. Bijsterbosch

Statistical and Data Sciences: Faculty Publications

The UK Biobank (UKB) is a highly promising dataset for brain biomarker research into population mental health due to its unprecedented sample size and extensive phenotypic, imaging, and biological measurements. In this study, we aimed to provide a shared foundation for UKB neuroimaging research into mental health with a focus on anxiety and depression. We compared UKB self-report measures and revealed important timing effects between scan acquisition and separate online acquisition of some mental health measures. To overcome these timing effects, we introduced and validated the Recent Depressive Symptoms (RDS-4) score which we recommend for state-dependent and longitudinal research in …


Data, Knowledge Practices, And Naturecultural Worlds: Vehicle Emissions In The Anthropocene, Lindsay Poirier Jan 2022

Data, Knowledge Practices, And Naturecultural Worlds: Vehicle Emissions In The Anthropocene, Lindsay Poirier

Statistical and Data Sciences: Faculty Books

This chapter details the various techno-cultural assemblages giving rise to data collected to model and measure anthropogenic worlds, arguing that data-based technologies both represent and co-produce the Anthropocene. It begins with a review of scholarship emerging at the intersection of science and technology studies and information studies that advances understanding of data infrastructure and knowledge practices, and their role within the anthropogenic assemblages that shape history. Drawing on a case study describing how vehicle emissions are measured and regulated in the US, I examine the materialities and mutability of technologies designed to produce data about air quality, along with the …


Comparison Of Caregiver- And Child-Reported Quality Of Life In Children With Sleep-Disordered Breathing, Phoebe Kuo Yu, Kaitlyn Cook, Jiayan Liu, Raouf S. Amin, Craig Derkay, Lisa M. Elden, Susan L. Garetz, Alisha S. George, Sally Ibrahim, Stacey L. Ishman, Erin M. Kirkham, S. Kamal Naqvi, Jerilynn Radcliffe, Kristie R. Ross, Gopi B. Shah, Ignacio E. Tapia, H. Gerry Taylor, David A. Zopf, Susan Redline, Cristina M. Baldassari Jan 2022

Comparison Of Caregiver- And Child-Reported Quality Of Life In Children With Sleep-Disordered Breathing, Phoebe Kuo Yu, Kaitlyn Cook, Jiayan Liu, Raouf S. Amin, Craig Derkay, Lisa M. Elden, Susan L. Garetz, Alisha S. George, Sally Ibrahim, Stacey L. Ishman, Erin M. Kirkham, S. Kamal Naqvi, Jerilynn Radcliffe, Kristie R. Ross, Gopi B. Shah, Ignacio E. Tapia, H. Gerry Taylor, David A. Zopf, Susan Redline, Cristina M. Baldassari

Statistical and Data Sciences: Faculty Publications

Objective. Caregivers frequently report poor quality of life(QOL) in children with sleep-disordered breathing (SDB).Our objective is to assess the correlation between care-giver- and child-reported QOL in children with mild SDBand identify factors associated with differences between caregiver and child report.

Study Design. Analysis of baseline data from a multi-institutional randomized trialSetting. Pediatric Adenotonsillectomy Trial for Snoring, where children with mild SDB (obstructive apnea-hypopnea index\3) were randomized to observation or adenotonsillectomy.

Methods. The Pediatric Quality of Life Inventory (Peds QL)assessed baseline global QOL in participating children 5 to12 years old and their caregivers. Caregiver and child scores were compared. Multivariable regression …


The Forestecology R Package For Fitting And Assessing Neighborhood Models Of The Effect Of Interspecific Competition On The Growth Of Trees, Albert Y. Kim, David N. Allen, Simon P. Couch Nov 2021

The Forestecology R Package For Fitting And Assessing Neighborhood Models Of The Effect Of Interspecific Competition On The Growth Of Trees, Albert Y. Kim, David N. Allen, Simon P. Couch

Statistical and Data Sciences: Faculty Publications

Neighborhood competition models are powerful tools to measure the effect of interspecific competition. Statistical methods to ease the application of these models are currently lacking. We present the forestecology package providing methods to (a) specify neighborhood competition models, (b) evaluate the effect of competitor species identity using permutation tests, and (cs) measure model performance using spatial cross-validation. Following Allen and Kim (PLoS One, 15, 2020, e0229930), we implement a Bayesian linear regression neighborhood competition model. We demonstrate the package's functionality using data from the Smithsonian Conservation Biology Institute's large forest dynamics plot, part of the ForestGEO global network of research …


Facilitating Team-Based Data Science: Lessons Learned From The Dsc-Wav Project, Chelsey Legacy, Andrew Zieffler, Benjamin S. Baumer, Valerie Barr, Nicholas J. Horton Oct 2021

Facilitating Team-Based Data Science: Lessons Learned From The Dsc-Wav Project, Chelsey Legacy, Andrew Zieffler, Benjamin S. Baumer, Valerie Barr, Nicholas J. Horton

Statistical and Data Sciences: Faculty Publications

While coursework provides undergraduate data science students with some relevant analytic skills, many are not given the rich experiences with data and computing they need to be successful in the workplace. Additionally, students often have limited exposure to team-based data science and the principles and tools of collaboration that are encountered outside of school. In this paper, we describe the DSC-WAV program, an NSF-funded data science workforce development project in which teams of undergraduate sophomores and juniors work with a local non-profit organization on a data-focused problem. To help students develop a sense of agency and improve confidence in their …


Infer: An R Package For Tidyverse-Friendly Statistical Inference, Simon P. Couch, Andrew P. Bray, Chester Ismay, Evgeni Chasnovski, B. Baumer, Mine Cetinkaya-Rundel Sep 2021

Infer: An R Package For Tidyverse-Friendly Statistical Inference, Simon P. Couch, Andrew P. Bray, Chester Ismay, Evgeni Chasnovski, B. Baumer, Mine Cetinkaya-Rundel

Statistical and Data Sciences: Faculty Publications

infer implements an expressive grammar to perform statistical inference that adheres to the tidyverse design framework (Wickham et al., 2019). Rather than providing methods for specific statistical tests, this package consolidates the principles that are shared among common hypothesis tests and confidence intervals into a set of four main verbs (functions), supplemented with many utilities to visualize and extract value from their outputs.