Open Access. Powered by Scholars. Published by Universities.®
- Keyword
-
- Ethnography (6)
- Alcohol (3)
- College (3)
- Social network (3)
- Statistics education (3)
-
- Data culture (2)
- Data infrastructure (2)
- Data sharing (2)
- Infrastructure (2)
- Interval censoring (2)
- Metadata (2)
- Open government data (2)
- Science and technology studies (2)
- Social networks (2)
- Statistical computing (2)
- 1Education in Computational Statistics (1)
- Accessibility (1)
- Accountability (1)
- Adjustable range (1)
- Applications of Computational Statistics (1)
- Area coverage (1)
- Articulation (1)
- Assessment (1)
- Associate’s programs (1)
- Attention (1)
- Bachelor’s programs (1)
- Bacteria (1)
- Baseball (1)
- Baserunning (1)
- Bayesian model (1)
Articles 1 - 30 of 53
Full-Text Articles in Data Science
Tidychangepoint: A Unified Framework For Analyzing Changepoint Detection In Univariate Time Series, Ben Baumer, Biviana Marcela Suárez Sierra
Tidychangepoint: A Unified Framework For Analyzing Changepoint Detection In Univariate Time Series, Ben Baumer, Biviana Marcela Suárez Sierra
Statistical and Data Sciences: Faculty Publications
We present tidychangepoint, a new R package for changepoint detection analysis. Most R packages for segmenting univariate time series focus on providing one or two algorithms for changepoint detection that work with a small set of models and penalized objective functions, and all of them return a custom, nonstandard object type. This makes comparing results across various algorithms, models, and penalized objective functions unnecessarily difficult. tidychangepoint solves this problem by wrapping functions from a variety of existing packages and storing the results in a common S3 class called tidycpt. The package then provides functionality for easily extracting comparable numeric or …
Examining The Intersectional And Structural Issues Of Routine Healthcare Utilization And Access Inequities For Lgb People With Chronic Diseases, Shiya Cao, Mehreen Mirza, Sophia Silovsky, Nicole Tresvalles, Lucia Qin, Sarah Susnea
Examining The Intersectional And Structural Issues Of Routine Healthcare Utilization And Access Inequities For Lgb People With Chronic Diseases, Shiya Cao, Mehreen Mirza, Sophia Silovsky, Nicole Tresvalles, Lucia Qin, Sarah Susnea
Statistical and Data Sciences: Faculty Publications
In the United States, although the gaps in health insurance coverage by sexual orientation have been closing since the implementation of the Affordable Care Act and legalization of same-sex marriage, the LGB group (i.e., lesbian, gay, bisexual) continues to report healthcare utilization and access inequities such as more delayed or unmet care. The extant research has often examined healthcare utilization and access inequities due to affordability (e.g., out-of-pocket costs). However, healthcare utilization and access inequities are only partially explained by cost reasons; there are non-cost reasons that have not been adequately empirically examined. The present study innovatively includes discrimination structural …
A Review Of Research And Practices On Teaching Data Visualizations For Blind And Visually Impaired Students, Shiya Cao
Statistical and Data Sciences: Faculty Publications
Around 36 million people in the world are blind and an additional 217 million have moderate to severe vision impairment. In higher education, four percent of 54,204 undergraduates who participated in the 2022 American College Health Association survey reported to be blind or have low vision. Those students frequently do not have access to data visualizations we generally teach and use in postsecondary statistics and data science classes. The design of those visualizations is premised on implicit assumptions about the user’s visual ability. Making data visualizations accessible to blind and visually impaired (BVI) people would help improve equity in higher …
Educational Opportunities Of Participatory Gis For Accessibility On A College Campus, Shiya Cao, Heather Rosenfeld, Sarah Susnea
Educational Opportunities Of Participatory Gis For Accessibility On A College Campus, Shiya Cao, Heather Rosenfeld, Sarah Susnea
Statistical and Data Sciences: Faculty Publications
The educational benefits of Participatory GIS (PGIS) in geographic higher education have received limited direct attention, often because of the complexities of integrating PGIS into university curricula. While a few exceptions found important educational benefits of PGIS, extant studies focused primarily on the educational benefits for students who worked in the research teams, instead of participants who contributed their local knowledge and perspectives to mapping. Our research aims to understand the educational benefits of PGIS for participants in a campus accessibility mapping project using the modes of experiential learning, positionality, and service learning. Through this, we also provide strategies for …
Making Plans Findable, Accessible, Interoperable, And Reusable With Data Infrastructure: A Search Engine For Constructing, Analyzing, And Visualizing Planning Documents, Lindsay Poirier, Dexter Antonio, Makenna Dettmann, Tiffany Eng, Jennifer Ganata, Sujoy Ghosh, Mirthala Lopez, Ranesh Karma, Asiya Natekal, Catherine Brinkley
Making Plans Findable, Accessible, Interoperable, And Reusable With Data Infrastructure: A Search Engine For Constructing, Analyzing, And Visualizing Planning Documents, Lindsay Poirier, Dexter Antonio, Makenna Dettmann, Tiffany Eng, Jennifer Ganata, Sujoy Ghosh, Mirthala Lopez, Ranesh Karma, Asiya Natekal, Catherine Brinkley
Statistical and Data Sciences: Faculty Publications
Local land-use plans help guide future development, but it is often difficult to compare content across jurisdictions, making regional coordination and plan evaluation challenging. This research reviews federal, state, and local data infrastructure guidance for land-use plans and compares such guidance to compliance with a California use-case. Findings indicate a number of obstacles to fostering data sharing and comparative analysis of plans: there is currently no central repository of land-use plans; plans are not uniform in format and are often out of date; many plans are not machine-readable thereby inhibiting text extraction, and planning language varies so greatly that there …
Enacting Data Context: Fixing Meaning In Transparency Data Initiatives, Lindsay Poirier
Enacting Data Context: Fixing Meaning In Transparency Data Initiatives, Lindsay Poirier
Statistical and Data Sciences: Faculty Publications
This article documents the “context cultures” underpinning efforts to develop regulations for collecting and reporting data in a United States public database known as Open Payments. Open Payments is a dataset published annually by the US Center for Medicare and Medicaid Services that documents the transfers of value from pharmaceutical and medical device manufacturers to physicians, prescribing non-physicians, and teaching hospitals. In the article, I show context became a manifold concern as differentially-situated actors engaged in modes of public advocacy and social action around not only what data meant, but also what it meant to make data meaningful. I show …
Per- And Polyfluoroalkyl Substance Exposure Risks In Us Carceral Facilities, 2022, Lindsay Poirier, Derrick Salvatore, Phil Brown, Alissa Cordner, Kira Mok, Nicholas Shapiro
Per- And Polyfluoroalkyl Substance Exposure Risks In Us Carceral Facilities, 2022, Lindsay Poirier, Derrick Salvatore, Phil Brown, Alissa Cordner, Kira Mok, Nicholas Shapiro
Statistical and Data Sciences: Faculty Publications
Objectives. To assess the US incarcerated population’s risk of exposure to per- and polyfluoroalkyl substances (PFASs). Methods. We assessed how many of the 6118 US carceral facilities were located in the same hydrologic unit code watershed boundaries as known or likely locations of PFAS contamination. We conducted geospatial analyses on data aggregated from Environmental Protection Agency databases and a PFAS site tracker in 2022 to model the hydrologically feasible known and presumptive PFAS contamination sites for nearly 2 million incarcerated people. Results. Findings indicate that 5% (∼310) of US carceral facilities have at least 1 known source of PFAS contamination …
Examining Information Systems Use To Facilitate The Workplace Accommodation Process, Shiya Cao
Examining Information Systems Use To Facilitate The Workplace Accommodation Process, Shiya Cao
Statistical and Data Sciences: Faculty Publications
BACKGROUND: The workplace accommodation process is often affected by ineffective and inefficient communications and information exchanges among disabled employees and other stakeholders. Information systems (IS) can play a key role in facilitating a more effective and efficient accommodation process since IS has been shown to facilitate business processes and effect positive organizational changes.
OBJECTIVE: Since there is little to no research that exists on IS use to facilitate the workplace accommodation process, this paper, as a critical first step, examines how IS have been used in the accommodation process.
METHODS: Thirty-six interviews were conducted with disabled employees from various organizations. …
Questions (And Answers) For Incorporating Nontraditional Grading In Your Statistics Courses, Brenna Curley
Questions (And Answers) For Incorporating Nontraditional Grading In Your Statistics Courses, Brenna Curley
Statistical and Data Sciences: Faculty Publications
Nontraditional grading methods have recently become more common, and as with any large pedagogical shift, there are a number of questions to consider when applying a new grading scheme to a course. This article summarizes four types of nontraditional grading and shares experiences from the authors who have applied them to a variety of courses in statistics. This article is structured as a set of questions and answers, seeking to address many of the concerns and considerations that one may face as they transition a course’s grading structure. Supplementary materials for this article are available online.
Population Modeling With Machine Learning Can Enhance Measures Of Mental Health - Open-Data Replication, Ty Easley, Ruiqi Chen, Kayla Hannon, Rosie Dutt, Janine Bijsterbosch
Population Modeling With Machine Learning Can Enhance Measures Of Mental Health - Open-Data Replication, Ty Easley, Ruiqi Chen, Kayla Hannon, Rosie Dutt, Janine Bijsterbosch
Statistical and Data Sciences: Faculty Publications
Efforts to predict trait phenotypes based on functional MRI data from large cohorts have been hampered by low prediction accuracy and/or small effect sizes. Although these findings are highly replicable, the small effect sizes are somewhat surprising given the presumed brain basis of phenotypic traits such as neuroticism and fluid intelligence. We aim to replicate previous work and additionally test multiple data manipulations that may improve prediction accuracy by addressing data pollution challenges. Specifically, we added additional fMRI features, averaged the target phenotype across multiple measurements to obtain more accurate estimates of the underlying trait, balanced the target phenotype's distribution …
Big Ideas In Sports Analytics And Statistical Tools For Their Investigation, Benjamin S. Baumer, Gregory J. Matthews, Quang Nguyen
Big Ideas In Sports Analytics And Statistical Tools For Their Investigation, Benjamin S. Baumer, Gregory J. Matthews, Quang Nguyen
Statistical and Data Sciences: Faculty Publications
Sports analytics—broadly defined as the pursuit of improvement in athletic performance through the analysis of data—has expanded its footprint both in the professional sports industry and in academia over the past 30 years. In this article, we connect four big ideas that are common across multiple sports: the expected value of a game state, win probability, measures of team strength, and the use of sports betting market data. For each, we explore both the shared similarities and individual idiosyncracies of analytical approaches in each sport. While our focus is on the concepts underlying each type of analysis, any implementation necessarily …
Data Science Transfer Pathways From Associate's To Bachelor's Programs, Benjamin S. Baumer, Nicholas J. Horton
Data Science Transfer Pathways From Associate's To Bachelor's Programs, Benjamin S. Baumer, Nicholas J. Horton
Statistical and Data Sciences: Faculty Publications
A substantial fraction of students who complete their college education at a public university in the United States begin their journey at one of the 935 public 2-year colleges. While the number of 4-year colleges offering bachelor’s degrees in data science continues to increase, data science instruction at many 2-year colleges lags behind. A major impediment is the relative paucity of introductory data science courses that serve multiple student audiences and can easily transfer. In addition, the lack of predefined transfer pathways (or articulation agreements) for data science creates a growing disconnect that leaves students who want to study data …
Psychometric Properties Of A Combined Go/No-Go And Continuous Performance Task Across Childhood, Caron A.C. Clark, Kaitlyn Cook, Rui Wang, Michael Rueschman, Jerilynn Radcliffe, Susan Redline, H. Gerry Taylor
Psychometric Properties Of A Combined Go/No-Go And Continuous Performance Task Across Childhood, Caron A.C. Clark, Kaitlyn Cook, Rui Wang, Michael Rueschman, Jerilynn Radcliffe, Susan Redline, H. Gerry Taylor
Statistical and Data Sciences: Faculty Publications
Despite the critical importance of attention for children’s self-regulation and mental health, there are few task-based measures of this construct appropriate for use across a wide childhood age range including very young children. Three versions of a combined go/no-go and continuous performance task (GNG/CPT) were created with varying length and timing parameters to maximize their appropriateness for age groups spanning early to middle childhood. As part of the baseline assessment of a clinical trial, 452 children aged 3–12 years (50% male, 50% female; 52% White, non-Hispanic, 27% Black, 16% Hispanic/Latinx; 6% other ethnicity/race) completed the task. Confirmatory factor analysis indicated …
Attending To The Cultures Of Data Science Work, Lindsay Poirier
Attending To The Cultures Of Data Science Work, Lindsay Poirier
Statistical and Data Sciences: Faculty Publications
This essay reflects on the shifting attention to the “social” and the “cultural” in data science communities. While recently the “social” and the “cultural” have been prioritized in data science discourse, social and cultural concerns that get raised in data science are almost always outwardly focused – applying to the communities that data scientists seek to support more so than more computationally-focused data science communities. I argue that data science communities have a responsibility to attend not only to the cultures that orient the work of domain communities, but also to the cultures that orient their own work. I describe …
Evaluation Of Edison's Data Science Competency Framework Through A Comparative Literature Analysis, Karl R. B. Schmitt, Linda Clark, Katherine M. Kinnaird, Ruth E. H. Wertz, Björn Sandstede
Evaluation Of Edison's Data Science Competency Framework Through A Comparative Literature Analysis, Karl R. B. Schmitt, Linda Clark, Katherine M. Kinnaird, Ruth E. H. Wertz, Björn Sandstede
Statistical and Data Sciences: Faculty Publications
During the emergence of Data Science as a distinct discipline, discussions of what exactly constitutes Data Science have been a source of contention, with no clear resolution. These disagreements have been exacerbated by the lack of a clear single disciplinary 'parent.' Many early efforts at defining curricula and courses exist, with the EDISON Project's Data Science Framework (EDISON-DSF) from the European Union being the most complete. The EDISON-DSF includes both a Data Science Body of Knowledge (DS-BoK) and Competency Framework (CF-DS). This paper takes a critical look at how EDISON's CF-DS compares to recent work and other published curricular or …
A Multistate Competing Risks Framework For Preconception Prediction Of Pregnancy Outcomes, Kaitlyn Cook, Neil J. Perkins, Enrique Schisterman, Sebastien Haneuse
A Multistate Competing Risks Framework For Preconception Prediction Of Pregnancy Outcomes, Kaitlyn Cook, Neil J. Perkins, Enrique Schisterman, Sebastien Haneuse
Statistical and Data Sciences: Faculty Publications
Background: Preconception pregnancy risk profiles—characterizing the likelihood that a pregnancy attempt results in a full-term birth, preterm birth, clinical pregnancy loss, or failure to conceive—can provide critical information during the early stages of a pregnancy attempt, when obstetricians are best positioned to intervene to improve the chances of successful conception and full-term live birth. Yet the task of constructing and validating risk assessment tools for this earlier intervention window is complicated by several statistical features: the final outcome of the pregnancy attempt is multinomial in nature, and it summarizes the results of two intermediate stages, conception and gestation, whose outcomes …
Marginal Proportional Hazards Models For Clustered Interval-Censored Data With Time-Dependent Covariates, Kaitlyn Cook, Wenbin Lu, Rui Wang
Marginal Proportional Hazards Models For Clustered Interval-Censored Data With Time-Dependent Covariates, Kaitlyn Cook, Wenbin Lu, Rui Wang
Statistical and Data Sciences: Faculty Publications
The Botswana Combination Prevention Project was a cluster-randomized HIV prevention trial whose follow-up period coincided with Botswana’s national adoption of a universal test-and-treat strategy for HIV management. Of interest is whether, and to what extent, this change in policy (i) modified the observed preventative effects of the study intervention and (ii) was associated with a reduction in the population-level incidence of HIV in Botswana. To address these questions, we propose a stratified proportional hazards model for clustered intervalcensored data with time-dependent covariates and develop a composite expectation maximization algorithm that facilitates estimation of model parameters without placing parametric assumptions on …
Implementing Github Actions Continuous Integration To Reduce Error Rates In Ecological Data Collection, Albert Y. Kim, Valentine Herrmann, Ross Barreto, Brianna Calkins, Erika Gonzalez-Akre, Daniel J. Johnson, Jennifer A. Jordan, Lukas Magee, Ian R. Mcgregor, Nicolle Montero, Karl Novak, Teagan Rogers, Jessica Shue, Kristina J. Anderson-Teixeira
Implementing Github Actions Continuous Integration To Reduce Error Rates In Ecological Data Collection, Albert Y. Kim, Valentine Herrmann, Ross Barreto, Brianna Calkins, Erika Gonzalez-Akre, Daniel J. Johnson, Jennifer A. Jordan, Lukas Magee, Ian R. Mcgregor, Nicolle Montero, Karl Novak, Teagan Rogers, Jessica Shue, Kristina J. Anderson-Teixeira
Statistical and Data Sciences: Faculty Publications
Accurate field data are essential to understanding ecological systems and forecasting their responses to global change. Yet, data collection errors are common, and data analysis often lags far enough behind its collection that many errors can no longer be corrected, nor can anomalous observations be revisited. Needed is a system in which data quality assurance and control (QA/QC), along with the production of basic data summaries, can be automated immediately following data collection.
Here, we implement and test a system to satisfy these needs. For two annual tree mortality censuses and a dendrometer band survey at two forest research sites, …
Accountable Data: The Politics And Pragmatics Of Disclosure Datasets, Lindsay Poirier
Accountable Data: The Politics And Pragmatics Of Disclosure Datasets, Lindsay Poirier
Statistical and Data Sciences: Faculty Publications
This paper attends specifically to what I call "disclosure datasets"- tabular datasets produced in accordance with laws requiring various kinds of disclosure. For the purposes of this paper, the most significant defining feature of disclosure datasets is that they aggregate information produced and reported by the same institutions they are meant to hold accountable. Through a series of case studies of disclosure datasets in the United States, I specifically draw attention to two concerns with disclosure datasets: First, for disclosure datasets, there is often political and social mobilization around the definitions that determine reporting thresholds, which in turn implicates what …
An Educator’S Perspective Of The Tidyverse, Mine Çetinkaya-Rundel, Johanna Hardin, Benjamin Baumer, Amelia Mcnamara, Nicholas J. Horton, Colin W. Rundel
An Educator’S Perspective Of The Tidyverse, Mine Çetinkaya-Rundel, Johanna Hardin, Benjamin Baumer, Amelia Mcnamara, Nicholas J. Horton, Colin W. Rundel
Statistical and Data Sciences: Faculty Publications
Computing makes up a large and growing component of data science and statistics courses. Many of those courses, especially when taught by faculty who are statisticians by training, teach R as the programming language. A number of instructors have opted to build much of their teaching around use of the tidyverse. The tidyverse, in the words of its developers, “is a collection of R packages that share a high-level design philosophy and low-level grammar and data structures, so that learning one package makes it easier to learn the next” (Wickham et al. 2019). These shared principles have led to the …
Mental Health In The Uk Biobank: A Roadmap To Self-Report Measures And Neuroimaging Correlates, Rosie K. Dutt, Kayla Hannon, Ty O. Easley, Joseph C. Griffis, Wei Zhang, Janine D. Bijsterbosch
Mental Health In The Uk Biobank: A Roadmap To Self-Report Measures And Neuroimaging Correlates, Rosie K. Dutt, Kayla Hannon, Ty O. Easley, Joseph C. Griffis, Wei Zhang, Janine D. Bijsterbosch
Statistical and Data Sciences: Faculty Publications
The UK Biobank (UKB) is a highly promising dataset for brain biomarker research into population mental health due to its unprecedented sample size and extensive phenotypic, imaging, and biological measurements. In this study, we aimed to provide a shared foundation for UKB neuroimaging research into mental health with a focus on anxiety and depression. We compared UKB self-report measures and revealed important timing effects between scan acquisition and separate online acquisition of some mental health measures. To overcome these timing effects, we introduced and validated the Recent Depressive Symptoms (RDS-4) score which we recommend for state-dependent and longitudinal research in …
Comparison Of Caregiver- And Child-Reported Quality Of Life In Children With Sleep-Disordered Breathing, Phoebe Kuo Yu, Kaitlyn Cook, Jiayan Liu, Raouf S. Amin, Craig Derkay, Lisa M. Elden, Susan L. Garetz, Alisha S. George, Sally Ibrahim, Stacey L. Ishman, Erin M. Kirkham, S. Kamal Naqvi, Jerilynn Radcliffe, Kristie R. Ross, Gopi B. Shah, Ignacio E. Tapia, H. Gerry Taylor, David A. Zopf, Susan Redline, Cristina M. Baldassari
Comparison Of Caregiver- And Child-Reported Quality Of Life In Children With Sleep-Disordered Breathing, Phoebe Kuo Yu, Kaitlyn Cook, Jiayan Liu, Raouf S. Amin, Craig Derkay, Lisa M. Elden, Susan L. Garetz, Alisha S. George, Sally Ibrahim, Stacey L. Ishman, Erin M. Kirkham, S. Kamal Naqvi, Jerilynn Radcliffe, Kristie R. Ross, Gopi B. Shah, Ignacio E. Tapia, H. Gerry Taylor, David A. Zopf, Susan Redline, Cristina M. Baldassari
Statistical and Data Sciences: Faculty Publications
Objective. Caregivers frequently report poor quality of life(QOL) in children with sleep-disordered breathing (SDB).Our objective is to assess the correlation between care-giver- and child-reported QOL in children with mild SDBand identify factors associated with differences between caregiver and child report.
Study Design. Analysis of baseline data from a multi-institutional randomized trialSetting. Pediatric Adenotonsillectomy Trial for Snoring, where children with mild SDB (obstructive apnea-hypopnea index\3) were randomized to observation or adenotonsillectomy.
Methods. The Pediatric Quality of Life Inventory (Peds QL)assessed baseline global QOL in participating children 5 to12 years old and their caregivers. Caregiver and child scores were compared. Multivariable regression …
The Forestecology R Package For Fitting And Assessing Neighborhood Models Of The Effect Of Interspecific Competition On The Growth Of Trees, Albert Y. Kim, David N. Allen, Simon P. Couch
The Forestecology R Package For Fitting And Assessing Neighborhood Models Of The Effect Of Interspecific Competition On The Growth Of Trees, Albert Y. Kim, David N. Allen, Simon P. Couch
Statistical and Data Sciences: Faculty Publications
Neighborhood competition models are powerful tools to measure the effect of interspecific competition. Statistical methods to ease the application of these models are currently lacking. We present the forestecology package providing methods to (a) specify neighborhood competition models, (b) evaluate the effect of competitor species identity using permutation tests, and (cs) measure model performance using spatial cross-validation. Following Allen and Kim (PLoS One, 15, 2020, e0229930), we implement a Bayesian linear regression neighborhood competition model. We demonstrate the package's functionality using data from the Smithsonian Conservation Biology Institute's large forest dynamics plot, part of the ForestGEO global network of research …
Facilitating Team-Based Data Science: Lessons Learned From The Dsc-Wav Project, Chelsey Legacy, Andrew Zieffler, Benjamin S. Baumer, Valerie Barr, Nicholas J. Horton
Facilitating Team-Based Data Science: Lessons Learned From The Dsc-Wav Project, Chelsey Legacy, Andrew Zieffler, Benjamin S. Baumer, Valerie Barr, Nicholas J. Horton
Statistical and Data Sciences: Faculty Publications
While coursework provides undergraduate data science students with some relevant analytic skills, many are not given the rich experiences with data and computing they need to be successful in the workplace. Additionally, students often have limited exposure to team-based data science and the principles and tools of collaboration that are encountered outside of school. In this paper, we describe the DSC-WAV program, an NSF-funded data science workforce development project in which teams of undergraduate sophomores and juniors work with a local non-profit organization on a data-focused problem. To help students develop a sense of agency and improve confidence in their …
Infer: An R Package For Tidyverse-Friendly Statistical Inference, Simon P. Couch, Andrew P. Bray, Chester Ismay, Evgeni Chasnovski, B. Baumer, Mine Cetinkaya-Rundel
Infer: An R Package For Tidyverse-Friendly Statistical Inference, Simon P. Couch, Andrew P. Bray, Chester Ismay, Evgeni Chasnovski, B. Baumer, Mine Cetinkaya-Rundel
Statistical and Data Sciences: Faculty Publications
infer implements an expressive grammar to perform statistical inference that adheres to the tidyverse design framework (Wickham et al., 2019). Rather than providing methods for specific statistical tests, this package consolidates the principles that are shared among common hypothesis tests and confidence intervals into a set of four main verbs (functions), supplemented with many utilities to visualize and extract value from their outputs.
Estimation Of Conditional Power For Cluster-Randomized Trials With Interval-Censored Endpoints, Kaitlyn Cook, Rui Wang
Estimation Of Conditional Power For Cluster-Randomized Trials With Interval-Censored Endpoints, Kaitlyn Cook, Rui Wang
Statistical and Data Sciences: Faculty Publications
Cluster-randomized trials (CRTs) of infectious disease preventions often yield correlated, interval-censored data: dependencies may exist between observations from the same cluster, and event occurrence may be assessed only at intermittent study visits. This data structure must be accounted for when conducting interim monitoring and futility assessment for CRTs. In this article, we propose a flexible framework for conditional power estimation when outcomes are correlated and interval-censored. Under the assumption that the survival times follow a shared frailty model, we first characterize the correspondence between the marginal and cluster-conditional survival functions, and then use this relationship to semiparametrically estimate the cluster-specific …
An Integrated Magneto-Electrochemical Device For The Rapid Profiling Of Tumour Extracellular Vesicles From Blood Plasma, Jongmin Park, Jun Seok Park, Chen Han Huang, Ala Jo, Kaitlyn Cook, Rui Wang, Hsing Ying Lin, Jan Van Deun, Huiyan Li, Jouha Min, Lan Wang, Ghilsuk Yoon, Bob S. Carter, Leonora Balaj, Gyu Seog Choi, Cesar M. Castro, Ralph Weissleder, Hakho Lee
An Integrated Magneto-Electrochemical Device For The Rapid Profiling Of Tumour Extracellular Vesicles From Blood Plasma, Jongmin Park, Jun Seok Park, Chen Han Huang, Ala Jo, Kaitlyn Cook, Rui Wang, Hsing Ying Lin, Jan Van Deun, Huiyan Li, Jouha Min, Lan Wang, Ghilsuk Yoon, Bob S. Carter, Leonora Balaj, Gyu Seog Choi, Cesar M. Castro, Ralph Weissleder, Hakho Lee
Statistical and Data Sciences: Faculty Publications
Assays for cancer diagnosis via the analysis of biomarkers on circulating extracellular vesicles (EVs) typically have lengthy sample workups, limited throughput or insufficient sensitivity, or do not use clinically validated biomarkers. Here we report the development and performance of a 96-well assay that integrates the enrichment of EVs by antibody-coated magnetic beads and the electrochemical detection, in less than one hour of total assay time, of EV-bound proteins after enzymatic amplification. By using the assay with a combination of antibodies for clinically relevant tumour biomarkers (EGFR, EpCAM, CD24 and GPA33) of colorectal cancer (CRC), we classified plasma samples from 102 …
Reading Datasets: Strategies For Interpreting The Politics Of Data Signification, Lindsay Poirier
Reading Datasets: Strategies For Interpreting The Politics Of Data Signification, Lindsay Poirier
Statistical and Data Sciences: Faculty Publications
All datasets emerge from and are enmeshed in power-laden semiotic systems. While emerging data ethics curriculum is supporting data science students in identifying data biases and their consequences, critical attention to the cultural histories and vested interests animating data semantics is needed to elucidate the assumptions and political commitments on which data rest, along with the externalities they produce. In this article, I introduce three modes of reading that can be engaged when studying datasets—a denotative reading (extrapolating the literal meaning of values in a dataset), a connotative reading (tracing the socio-political provenance of data semantics), and a deconstructive reading …
Automatic Hierarchy Expansion For Improved Structure And Chord Evaluation, Katherine M. Kinnaird, Brian Mcfee
Automatic Hierarchy Expansion For Improved Structure And Chord Evaluation, Katherine M. Kinnaird, Brian Mcfee
Statistical and Data Sciences: Faculty Publications
No abstract provided.
Moving Ethnography: Infrastructuring Doubletakes And Switchbacks In Experimental Collaborative Methods, Aalok Khandekar, Brandon Costelloe-Kuehn, Lindsay Poirier, Alli Morgan, Alison Kenner, Kim Fortun, Mike Fortun
Moving Ethnography: Infrastructuring Doubletakes And Switchbacks In Experimental Collaborative Methods, Aalok Khandekar, Brandon Costelloe-Kuehn, Lindsay Poirier, Alli Morgan, Alison Kenner, Kim Fortun, Mike Fortun
Statistical and Data Sciences: Faculty Publications
In this article, we describe how our work at a particular nexus of STS, ethnography, and critical theory—informed by experimental sensibilities in both the arts and sciences—transformed as we built and learned to use collaborative workflows and supporting digital infrastructure. Responding to the call of this special issue to be “ethnographic about ethnography,” we describe what we have learned about our own methods and collaborative practices through building digital infrastructure to support them. Supporting and accounting for how experimental ethnographic projects move—through different points in a research workflow, with many switchbacks, with project designs constantly changing as the research develops—was …