Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Computer Sciences (1157)
- Medicine and Health Sciences (780)
- Life Sciences (765)
- Bioinformatics (568)
- Statistics and Probability (550)
-
- Biomedical Informatics (530)
- Engineering (528)
- Artificial Intelligence and Robotics (526)
- Social and Behavioral Sciences (520)
- Databases and Information Systems (212)
- Computer Engineering (208)
- Electrical and Computer Engineering (204)
- Applied Statistics (194)
- Medical Sciences (190)
- Business (189)
- Statistical Models (181)
- Applied Mathematics (175)
- Medical Specialties (173)
- Environmental Sciences (149)
- Theory and Algorithms (149)
- Mathematics (144)
- Other Computer Sciences (127)
- Data Storage Systems (123)
- Systems and Communications (120)
- Numerical Analysis and Scientific Computing (116)
- Public Health (116)
- Public Affairs, Public Policy and Public Administration (109)
- Statistical Methodology (109)
- Institution
-
- The Texas Medical Center Library (523)
- Old Dominion University (173)
- Southern Methodist University (144)
- Universitas Negeri Malang (113)
- City University of New York (CUNY) (101)
-
- CCT College Dublin (91)
- Chapman University (67)
- Kennesaw State University (63)
- University of Central Florida (62)
- Smith College (60)
- Air Force Institute of Technology (57)
- Embry-Riddle Aeronautical University (52)
- Singapore Management University (45)
- University of Arkansas, Fayetteville (45)
- Chinese Academy of Sciences (44)
- Purdue University (44)
- California Polytechnic State University, San Luis Obispo (39)
- Technological University Dublin (39)
- Illinois State University (38)
- University of Kentucky (38)
- University of Nebraska - Lincoln (38)
- New Jersey Institute of Technology (37)
- West Virginia University (37)
- Claremont Colleges (36)
- Virginia Commonwealth University (35)
- Clemson University (32)
- Dartmouth College (31)
- University of Texas at Arlington (27)
- East Tennessee State University (26)
- Minnesota State University, Mankato (26)
- Keyword
-
- Humans (278)
- Machine learning (241)
- Machine Learning (218)
- Deep learning (115)
- Computer Science (107)
-
- Deep Learning (94)
- Artificial Intelligence (65)
- Data science (58)
- Data Science (57)
- Natural Language Processing (56)
- COVID-19 (55)
- Artificial intelligence (53)
- Female (52)
- Male (50)
- Classification (49)
- Natural language processing (47)
- Animals (41)
- Data (41)
- Electronic Health Records (41)
- Neural Networks (40)
- Algorithms (38)
- Big data (37)
- Data mining (37)
- Statistics (36)
- Clustering (32)
- Computer science (31)
- Adult (30)
- NLP (30)
- Neural networks (30)
- Random Forest (30)
- Publication Year
- Publication
-
- Faculty, Staff and Student Publications (508)
- SMU Data Science Review (124)
- Knowledge Engineering and Data Science (113)
- Theses and Dissertations (111)
- ICT (91)
-
- Data Science and Data Mining (53)
- Dissertations (53)
- Statistical and Data Sciences: Faculty Publications (53)
- Electronic Theses and Dissertations (49)
- Dissertations, Theses, and Capstone Projects (45)
- Bulletin of Chinese Academy of Sciences (Chinese Version) (44)
- Research Collection School Of Computing and Information Systems (37)
- Master's Theses (35)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (34)
- Data Science Undergraduate Honors Theses (31)
- Annual Symposium on Biomathematics and Ecology Education and Research (30)
- Computer Science Faculty Publications (30)
- Publications and Research (30)
- Computational and Data Sciences (PhD) Dissertations (25)
- All Graduate Theses, Dissertations, and Other Capstone Projects (24)
- Symposium of Student Scholars (24)
- All Dissertations (23)
- Articles (23)
- Electrical & Computer Engineering Faculty Publications (22)
- CBN Journal of Applied Statistics (JAS) (21)
- College of Graduate Studies: Theses & Dissertations (20)
- CMC Senior Theses (19)
- Electronic Theses, Projects, and Dissertations (19)
- Theses (19)
- Faculty Publications (18)
- Publication Type
- File Type
Articles 2761 - 2790 of 3244
Full-Text Articles in Data Science
A Data Exploration Of Jeopardy! From 1984 To The Present, Brian S. Hamilton
A Data Exploration Of Jeopardy! From 1984 To The Present, Brian S. Hamilton
Dissertations, Theses, and Capstone Projects
The gameshow Jeopardy! has been around in its current iteration—hosted by Alex Trebek—since 1984. During this time, it has accumulated data on clues, contestants, and possible strategies on how to win. Using a crowd-sourced archive called J! Archive, this project seeks to find trends in the topics that the game covers and take a deeper look into the performance of its contestants. It employs topic modeling, a text-analysis method, to organize the hundreds of thousands of archived clues and statistical analysis to rate the performance of contestants by gender. Using web-based visualization tools, the data is shown in an …
Sensory Stressors Impact Species Responses Across Local And Continental Scales, Ashley A. Wilson
Sensory Stressors Impact Species Responses Across Local And Continental Scales, Ashley A. Wilson
Master's Theses
Pervasive growth in industrialization and advances in technology now exposes much of the world to anthropogenic night light and noise (ANLN), which pose a global environmental challenge in terrestrial environments. An estimated one-tenth of the planet’s land area experiences artificial light at night — and that rises to 23% if skyglow is included. Moreover, anthropogenic noise is associated with urban development and transportation networks, as the ecological impact of roads alone is estimated to affect one-fifth of the total land cover of the United States and is increasing in space and intensity. Existing research involving impacts of light or noise …
Team Formation Using Recommendation Systems, Shreyas Patil
Team Formation Using Recommendation Systems, Shreyas Patil
Theses
The importance of team formation has been realized since ages, but finding the most effective team out of the available human resources is a problem that persists to the date. Having members with complementary skills, along with a few must-have behavioral traits, such as trust and collaborativeness among the team members are the key ingredients behind team synergy and performance. This thesis designs and implements two different algorithms for the team formation problem using ideas adapted from the recommender systems literature. One of the proposed solutions uses the Glicko-2 rating system to rate the employees’ skills which can easily separate …
The Transcript Profile Changes With Developmental Maturation Of Fetal Lung Type 2 Cells: An Analysis Of Rnaseq Data, Heber C. Nielsen, Volodymyr Orlov, Rebecca Holsapple, Monnie Mcgee
The Transcript Profile Changes With Developmental Maturation Of Fetal Lung Type 2 Cells: An Analysis Of Rnaseq Data, Heber C. Nielsen, Volodymyr Orlov, Rebecca Holsapple, Monnie Mcgee
SMU Data Science Review
In this paper, we utilize next-generation sequencing (NGS) data from the LungMap project to identify and characterize the developmental RNA transcriptome in alveolar epithelial type II cells of embryonic mouse lungs of gestational ages embryonic days 16 (E16) and 18 (E18). Late gestation lung cellular maturation is necessary for survival at birth. Using R and the BioConductor packages for RNAseq analysis, we analyze changes in the mouse lung RNA transcriptome as this maturation process takes place. We particularly identify the cluster of genes whose expression changes markedly between immature (E16) and mature (E18) lungs which can be used to define …
Forecasting Power Consumption In Pennsylvania During The Covid-19 Pandemic: A Sarimax Model With External Covid-19 And Unemployment Variables, Jackson Au, Javier Saldaña Jr., Ben Spanswick, John Santerre
Forecasting Power Consumption In Pennsylvania During The Covid-19 Pandemic: A Sarimax Model With External Covid-19 And Unemployment Variables, Jackson Au, Javier Saldaña Jr., Ben Spanswick, John Santerre
SMU Data Science Review
In this paper, we present how electrical consumption can reveal insight into the novel COVID-19 pandemic spread. We analyze electrical power consumption provided by PPL Electric Utilities, Department of Labor’s unemployment claims, and the COVID-19 cases/deaths for the State of Pennsylvania to study the impact of the pandemic on the infrastructure. Using a SARIMA model as our benchmark and we analyzed the use of a SARIMAX model to forecast the power consumption in Pennsylvania 14 days ahead. Our work quantifies and illuminates the effect that the strict legislation passed to minimize the spread of COVID19 had a on power consumption. …
Compressed Dna Representation For Efficient Amr Classification, John Partee, Robert Hazell, Anjli Solsi, John Santerre
Compressed Dna Representation For Efficient Amr Classification, John Partee, Robert Hazell, Anjli Solsi, John Santerre
SMU Data Science Review
In this paper, we explore a representation methodology for the compression of DNA isolates. Using lossless string compression via tokenization of frequently repeated segments of DNA, we reduce the length of the isolates to be counted as k-mers for classification. With this new representation, we apply a previously established feature sampling method to dramatically reduce the feature space. In understanding the genetic diversity, we also look at conserving biological function across these spaces. Using a random forest model we were able to predict the resistance or susceptibility of bacteria with 85-90\% accuracy, with a 30-50\% reduction in overall isolate length, …
Spoken Language Recognition On Open-Source Datasets, Brady Arendale, Samira Zarandioon, Ryan Goodwin, Douglas Reynolds
Spoken Language Recognition On Open-Source Datasets, Brady Arendale, Samira Zarandioon, Ryan Goodwin, Douglas Reynolds
SMU Data Science Review
The field of speaker and language recognition is constantly being researched and developed, but much of this research is done on private or expensive datasets, making the field more inaccessible than many other areas of machine learning. In addition, many papers make performance claims without comparing their models to other recent research. With the recent development of public multilingual speech corpora such as Mozilla's Common Voice as well as several single-language corpora, we now have the resources to attempt to address both of these problems. We construct an eight-language dataset from Common Voice and a Google Bengali corpus as well …
Predicting Attrition - A Driver For Creating Value, Realizing Strategy, And Refining Key Hr Processes, Kevin Mendonsa, Maureen Stolberg, Vivek Viswanathan, Scott Crum
Predicting Attrition - A Driver For Creating Value, Realizing Strategy, And Refining Key Hr Processes, Kevin Mendonsa, Maureen Stolberg, Vivek Viswanathan, Scott Crum
SMU Data Science Review
Talent is the most important asset for every organization's success. While attrition (or churn) and turnover can refer to both employees and customers, this paper will focus on employee attrition only. Many organizations accept attrition as an inevitable cost of doing business and do nothing to adopt or implement mitigating strategies to combat it. World class companies on the other hand take deliberate measures to understand, control and mitigate attrition (turnover) at every stage. Unmitigated attrition can have a devastating effect on an organization's bottom line and market value. In addition, the “invisible" costs of low employee morale, reduced employee …
An Effective Method For Attribute Subset Selection, Considering The Resource In Pattern Recognition, Bakhtiyorjon Bakirovich Akbaraliev
An Effective Method For Attribute Subset Selection, Considering The Resource In Pattern Recognition, Bakhtiyorjon Bakirovich Akbaraliev
Chemical Technology, Control and Management
An analytical method for determining informative sets of features (INP) is developed, taking into account the resource for criteria based on the use of a measure of dispersion of classified objects. The areas of existence of the solution are defined. The statements and properties for the Fischer-type information criterion are proved, using which the proposed analytical method for determining the INP guarantees optimal results in the sense of maximizing the selected functional. The appropriateness of choosing this type of informative criterion is justified. A method for transforming attributes is proposed. The universality of the method in relation to the type …
The Most-Cited Articles In Data In Brief Journal: A Bibliometric Analysis Using Scopus Data, Lusiana Wulansari, Ansari Saleh Ahmar, Agus Rochmat, Nurmawati, Akbar Iskandar
The Most-Cited Articles In Data In Brief Journal: A Bibliometric Analysis Using Scopus Data, Lusiana Wulansari, Ansari Saleh Ahmar, Agus Rochmat, Nurmawati, Akbar Iskandar
Library Philosophy and Practice (e-journal)
Bibliometric analysis is one of the research approaches that utilizes quantitative and mathematical data to address problems posed in the context of visualization to see patterns in the field of science. In fact, bibliometric analysis may also include a wider overview of the names of the most influential writers in the area of science. This data analysis would discuss the most-cited articles in Data in Brief Journal including the countries, authors. The data was collected on 31st May 2020 of Scopus database. The literature review was conducted using the keyword: ISSN (2352-3409). The bibliometric analysis is visualized utilizing the VosViewer …
Blockchain Technology And Freight Forwarder Exploration Of Implications Focused On Practitioners In Shanghai, Johannes Van Bohemen
Blockchain Technology And Freight Forwarder Exploration Of Implications Focused On Practitioners In Shanghai, Johannes Van Bohemen
World Maritime University Dissertations
No abstract provided.
How Port Logistics Competitiveness Evolves Among Major Ports In China And Europe (1998-2018), Jiawei Wang
How Port Logistics Competitiveness Evolves Among Major Ports In China And Europe (1998-2018), Jiawei Wang
World Maritime University Dissertations
No abstract provided.
Data Visualization And Infographics Design Art404g/Dsp Xxx, Harrison Dekker
Data Visualization And Infographics Design Art404g/Dsp Xxx, Harrison Dekker
Library Impact Statements
No abstract provided.
Multi‑View Clustering For Multi‑Omics Data Using Unifed Embedding, Mohammed Hasanuzzaman, Sayantan Mitra, Sriparna Saha
Multi‑View Clustering For Multi‑Omics Data Using Unifed Embedding, Mohammed Hasanuzzaman, Sayantan Mitra, Sriparna Saha
Articles
In real world applications, data sets are often comprised of multiple views, which provide consensus and complementary information to each other. Embedding learning is an effective strategy for nearest neighbour search and dimensionality reduction in large data sets. This paper attempts to learn a unified probability distribution of the points across different views and generates a unified embedding in a low-dimensional space to optimally preserve neighbourhood identity. Probability distributions generated for each point for each view are combined by conflation method to create a single unified distribution. The goal is to approximate this unified distribution as much as possible when …
Multi‑View Clustering For Multi‑Omics Data Using Unifed Embedding, Mohammed Hasanuzzaman, Sayantan Mitra, Sriparna Saha
Multi‑View Clustering For Multi‑Omics Data Using Unifed Embedding, Mohammed Hasanuzzaman, Sayantan Mitra, Sriparna Saha
Department of Computer Science Publications
In real world applications, data sets are often comprised of multiple views, which provide consensus and complementary information to each other. Embedding learning is an effective strategy for nearest neighbour search and dimensionality reduction in large data sets. This paper attempts to learn a unified probability distribution of the points across different views and generates a unified embedding in a low-dimensional space to optimally preserve neighbourhood identity. Probability distributions generated for each point for each view are combined by conflation method to create a single unified distribution. The goal is to approximate this unified distribution as much as possible when …
Big Data Analytics Applied To Healthcare, Xuejuan Zhang, Boris Vishnevsky
Big Data Analytics Applied To Healthcare, Xuejuan Zhang, Boris Vishnevsky
School of Continuing and Professional Studies Student Papers
In this paper, we review the recent literature related to Big Data Analytics (BDA). We also discuss ways of applying BDA in Healthcare. In Section 1, we discuss the definition of Big Data Analytics and its characteristics. In Section 2, we discuss the healthcare ecosystem's main stakeholders and the data of each main stakeholder. Section 3 discusses the challenges and opportunities of leveraging Big Data Analytics by healthcare stakeholders.
Cell Assembly Detection In Low Firing-Rate Spike Train Data, Phan Minh Duc Truong
Cell Assembly Detection In Low Firing-Rate Spike Train Data, Phan Minh Duc Truong
Mathematics Theses and Dissertations
Cell assemblies, defined as groups of neurons forming temporal spike coordination, are thought to be fundamental units supporting major cognitive functions. However, detecting cell assemblies is challenging since they can occur at a range of time scales and with a range of precisions, from synchronous spikes to co-variations in firing rate. In this dissertation, we use a recently published cell assembly detection (CAD) algorithm that is capable of detecting assemblies at a range of time scales and precisions. We first showed that the CAD method can be applied to sparser spike train data than what have previously been reported. This …
Iba Newsletter [August 2020], Communications Department, Office Of The Registrar
Iba Newsletter [August 2020], Communications Department, Office Of The Registrar
IBA News
No abstract provided.
A Study Of Information Bots And Knowledge Bots, Amartya Hatua
A Study Of Information Bots And Knowledge Bots, Amartya Hatua
Dissertations
In this dissertation, a study of different aspects of information bots and knowledge bots is done. The research contributes to a better understanding of the various characteristics of information bots as well as the different patterns and factors responsible for the information diffusion in a social network. This research also shows how these factors can be used to predict information diffusion for a particular topic in a social network. The second part of the research is focused on strategies for improving the knowledge base of knowledge bots, where two different approaches are studied. In the first approach, knowledge is transferred …
Variable Compact Multi-Point Upscaling Schemes For Anisotropic Diffusion Problems In Three-Dimensions, James Quinlan
Variable Compact Multi-Point Upscaling Schemes For Anisotropic Diffusion Problems In Three-Dimensions, James Quinlan
Dissertations
Simulation is a useful tool to mitigate risk and uncertainty in subsurface flow models that contain geometrically complex features and in which the permeability field is highly heterogeneous. However, due to the level of detail in the underlying geocellular description, an upscaling procedure is needed to generate a coarsened model that is computationally feasible to perform simulations. These procedures require additional attention when coefficients in the system exhibit full-tensor anisotropy due to heterogeneity or not aligned with the computational grid. In this thesis, we generalize a multi-point finite volume scheme in several ways and benchmark it against the industry-standard routines. …
Empirical Studies Of Deep Learning On Information Diffusion On Social Networks And Collective Task Learning For Swarm Robotics, Trung T. Nguyen
Empirical Studies Of Deep Learning On Information Diffusion On Social Networks And Collective Task Learning For Swarm Robotics, Trung T. Nguyen
Dissertations
Researchers in multiple disciplines have recently adopted deep learning because of its ability of high accuracy representation learning from big and complex data. My research goal in this thesis is developing deep learning models for information diffusion analysis on social networks and collective tasks learning in swarm robotics. Firstly, the information diffusion on social networks is modeled as a multivariate time series in three dimensions with ten features. Then, we applied time-series clustering algorithms with Dynamic Time Warping to discover different patterns of our models. Then, we build a prediction model based on LSTM, which outperforms traditional time-series prediction methods. …
Machine Learning Approaches For Improving Prediction Performance Of Structure-Activity Relationship Models, Gabriel Idakwo
Machine Learning Approaches For Improving Prediction Performance Of Structure-Activity Relationship Models, Gabriel Idakwo
Dissertations
In silico bioactivity prediction studies are designed to complement in vivo and in vitro efforts to assess the activity and properties of small molecules. In silico methods such as Quantitative Structure-Activity/Property Relationship (QSAR) are used to correlate the structure of a molecule to its biological property in drug design and toxicological studies. In this body of work, I started with two in-depth reviews into the application of machine learning based approaches and feature reduction methods to QSAR, and then investigated solutions to three common challenges faced in machine learning based QSAR studies.
First, to improve the prediction accuracy of learning …
A Novel Correction For The Adjusted Box-Pierce Test — New Risk Factors For Emergency Department Return Visits Within 72 Hours For Children With Respiratory Conditions — General Pediatric Model For Understanding And Predicting Prolonged Length Of Stay, Sidy Danioko
Computational and Data Sciences (PhD) Dissertations
This thesis represents the results of three research projects that underline the breadth and depth of my interests.
Firstly, I devoted some efforts to the well-known Box-Pierce goodness-of-fit tests for time series models which has been an important research topic over the last few decades. All previously proposed tests are focused on changes of the test statistics. Instead, I adopted a different approach that takes the best performing test and modifying the rejection region. Thus, I developed a semiparametric correction of the Adjusted Box-Pierce test that attains the best I error rates for all sample sizes and lags and outperforms …
Shotgun Metagenomic Sequencing Data Of Sunflower Rhizosphere Microbial Community In South Africa, Olubukola Oluranti Babalola, Temitayo Tosin Alawiye, Carlos M. Rodriguez Lopez, Ayansina Segun Ayangbenro
Shotgun Metagenomic Sequencing Data Of Sunflower Rhizosphere Microbial Community In South Africa, Olubukola Oluranti Babalola, Temitayo Tosin Alawiye, Carlos M. Rodriguez Lopez, Ayansina Segun Ayangbenro
Horticulture Faculty Publications
This dataset presents shotgun metagenomic sequencing of sunflower rhizosphere microbiome in Bloemhof, South Africa. Data were collected to decipher the structure and function in the sunflower microbial community. Illumina HiSeq platform using next generation sequencing of the DNA was carried out. The metagenome comprised 8,991,566 sequences totaling 1,607,022,279 bp size and 66% GC content. The metagenome was deposited into the NCBI database and can be accessed with the SRA accession number SRR10418054. An online metagenome server (MG RAST) using the subsystem database revealed bacteria had the highest taxonomical representation with 98.47%, eukaryote at 1.23%, and archaea at 0.20%. The most …
Human Trafficking In Nepal: Can Big Data Help?, Shushant Khanal
Human Trafficking In Nepal: Can Big Data Help?, Shushant Khanal
Undergraduate Research Journal
This paper provides an overview of human trafficking in Nepal, identifies strategies implemented by the government of the country to handle the problem and possibilities of using big data as a solution to the problem of human trafficking in Nepal. Big data, may be defined as the collection of a large volume of data from the past that is processed using machine learning and artificial intelligence to find a common pattern. The use of big data in tackling the problem of human trafficking is not new in developed countries like the United States but it is still a foreign idea …
Applications Of Artificial Intelligence And Graphy Theory To Cyberbullying, Jesse D. Simpson
Applications Of Artificial Intelligence And Graphy Theory To Cyberbullying, Jesse D. Simpson
Graduate Theses/Dissertations
Cyberbullying is an ongoing and devastating issue in today's online social media. Abusive users engage in cyber-harassment by utilizing social media to send posts, private messages, tweets, or pictures to innocent social media users. Detecting and preventing cases of cyberbullying is crucial. In this work, I analyze multiple machine learning, deep learning, and graph analysis algorithms and explore their applicability and performance in pursuit of a robust system for detecting cyberbullying. First, I evaluate the performance of the machine learning algorithms Support Vector Machine, Naïve Bayes, Random Forest, Decision Tree, and Logistic Regression. This yielded positive results and obtained upwards …
Colleague To Banner Migration: Data Conversion Guide For Institutional Research, Laura Osborn
Colleague To Banner Migration: Data Conversion Guide For Institutional Research, Laura Osborn
Masters Theses & Doctoral Dissertations
When the SDBOR decided to migrate their current student information system into a shared system with HR and Finance, adjustments needed to be made to accommodate for current Banner settings and work around tables that were already populated with HRFIS data. The change in data type of the student identifier from that of a 7-digit numeric field to a 9- digit alpha-numeric field poses problems for running aggregate data calculations. Additional complications include having some information such as first-generation status that was not migrated between the systems, and cases such as college coding where tables that were designed for student …
Understanding Spatial Language In Radiology: Representation Framework, Annotation, And Spatial Relation Extraction From Chest X-Ray Reports Using Deep Learning, Surabhi Datta, Yuqi Si, Laritza Rodriguez, Sonya E Shooshan, Dina Demner-Fushman, Kirk Roberts
Understanding Spatial Language In Radiology: Representation Framework, Annotation, And Spatial Relation Extraction From Chest X-Ray Reports Using Deep Learning, Surabhi Datta, Yuqi Si, Laritza Rodriguez, Sonya E Shooshan, Dina Demner-Fushman, Kirk Roberts
Faculty, Staff and Student Publications
Radiology reports contain a radiologist's interpretations of images, and these images frequently describe spatial relations. Important radiographic findings are mostly described in reference to an anatomical location through spatial prepositions. Such spatial relationships are also linked to various differential diagnoses and often described through uncertainty phrases. Structured representation of this clinically significant spatial information has the potential to be used in a variety of downstream clinical informatics applications. Our focus is to extract these spatial representations from the reports. For this, we first define a representation framework based on the Spatial Role Labeling (SpRL) scheme, which we refer to as …
Data Mining For Structural Damage Identification Using Hybrid Artificial Neural Network Based Algorithm For Beam And Slab Girder, Gordan Meisam
Data Mining For Structural Damage Identification Using Hybrid Artificial Neural Network Based Algorithm For Beam And Slab Girder, Gordan Meisam
Student Works (2020-2029)
One of the approaches for structural health monitoring (SHM) consists of two major components, i.e. a network of sensors to collect the response data and an extraction method to obtain information on the structural health condition. Data mining (DM) is a novel data extraction technology which can employ for development of inverse analysis. Implementation of DM techniques in different areas of civil engineering has recently given very good results. However, application of DM in SHM is not used as much as expected, thus, many challenges are still ahead. Therefore, it is necessary to develop the applicability of DM in SHM. …
Statistical Methods For Resolving Intratumor Heterogeneity With Single-Cell Dna Sequencing, Alexander Davis
Statistical Methods For Resolving Intratumor Heterogeneity With Single-Cell Dna Sequencing, Alexander Davis
Dissertations and Theses (Open Access)
Tumor cells have heterogeneous genotypes, which drives progression and treatment resistance. Such genetic intratumor heterogeneity plays a role in the process of clonal evolution that underlies tumor progression and treatment resistance. Single-cell DNA sequencing is a promising experimental method for studying intratumor heterogeneity, but brings unique statistical challenges in interpreting the resulting data. Researchers lack methods to determine whether sufficiently many cells have been sampled from a tumor. In addition, there are no proven computational methods for determining the ploidy of a cell, a necessary step in the determination of copy number. In this work, software for calculating probabilities from …