Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Databases and Information Systems (143)
- Business (133)
- Management Information Systems (128)
- Engineering (63)
- Artificial Intelligence and Robotics (53)
-
- Life Sciences (38)
- Bioinformatics (31)
- Computer Engineering (31)
- Statistics and Probability (27)
- Electrical and Computer Engineering (23)
- Biostatistics (21)
- Data Science (19)
- Electrical and Electronics (14)
- Social and Behavioral Sciences (14)
- Mathematics (11)
- Applied Mathematics (10)
- Theory and Algorithms (8)
- Biomedical Engineering and Bioengineering (7)
- Numerical Analysis and Scientific Computing (7)
- Software Engineering (7)
- Information Security (6)
- Physics (6)
- Mechanical Engineering (5)
- Medicine and Health Sciences (5)
- Robotics (5)
- Communication (4)
- Education (4)
- Astrophysics and Astronomy (3)
- Keyword
-
- Machine learning (50)
- Deep learning (33)
- Data mining (16)
- Bioinformatics (11)
- Image processing (11)
-
- Computer vision (8)
- Data compression (Computer science) (8)
- Parallel processing (Electronic computers) (8)
- Artificial intelligence (7)
- Big data (7)
- Clustering (7)
- Natural language processing (7)
- Object-oriented databases. (7)
- Pattern recognition systems. (7)
- Cloud computing (6)
- Object-oriented programming (Computer science). (6)
- Office information systems. (6)
- Pattern recognition (6)
- Software engineering (6)
- Algorithms (5)
- Auditing (5)
- Collaboration (5)
- Computer network architectures (5)
- Computer networks (5)
- Information storage and retrieval systems. (5)
- Information systems (5)
- Object-oriented databases (5)
- Optimization (5)
- Parallel programming (Computer science) (5)
- Reinforcement learning (5)
- Publication Year
- Publication
Articles 211 - 240 of 571
Full-Text Articles in Computer Sciences
Cancer Risk Prediction With Next Generation Sequencing Data Using Machine Learning, Nihir Patel
Cancer Risk Prediction With Next Generation Sequencing Data Using Machine Learning, Nihir Patel
Theses
The use of computational biology for next generation sequencing (NGS) analysis is rapidly increasing in genomics research. However, the effectiveness of NGS data to predict disease abundance is yet unclear. This research investigates the problem in the whole exome NGS data of the chronic lymphocytic leukemia (CLL) available at dbGaP. Initially, raw reads from samples are aligned to the human reference genome using burrows wheeler aligner. From the samples, structural variants, namely, Single Nucleotide Polymorphism (SNP) and Insertion Deletion (INDEL) are identified and are filtered using SAMtools as well as with Genome Analyzer Tool Kit (GATK). Subsequently, the variants are …
New Directions For Remote Data Integrity Checking Of Cloud Storage, Bo Chen
New Directions For Remote Data Integrity Checking Of Cloud Storage, Bo Chen
Dissertations
Cloud storage services allow data owners to outsource their data, and thus reduce their workload and cost in data storage and management. However, most data owners today are still reluctant to outsource their data to the cloud storage providers (CSP), simply because they do not trust the CSPs, and have no confidence that the CSPs will secure their valuable data. This dissertation focuses on Remote Data Checking (RDC), a collection of protocols which can allow a client (data owner) to check the integrity of data outsourced at an untrusted server, and thus to audit whether the server fulfills its contractual …
Intent-Based User Segmentation With Query Enhancement, Wei Xiong
Intent-Based User Segmentation With Query Enhancement, Wei Xiong
Dissertations
With the rapid advancement of the internet, accurate prediction of user's online intent underlying their search queries has received increasing attention from the online advertising community. As a rich source of information on web user's behavior, query logs have been leveraged by advertising companies to deliver personalized advertisements. However, a typical query usually contains very few terms, which only carry a small amount of information about a user's interest. The tendency of users to use short and ambiguous queries makes it difficult to fully describe and distinguish a user's intent. In addition, the query feature space is sparse, as only …
Computational Methods For The Analysis Of Next Generation Sequencing Data, Wei Wang
Computational Methods For The Analysis Of Next Generation Sequencing Data, Wei Wang
Dissertations
Recently, next generation sequencing (NGS) technology has emerged as a powerful approach and dramatically transformed biomedical research in an unprecedented scale. NGS is expected to replace the traditional hybridization-based microarray technology because of its affordable cost and high digital resolution. Although NGS has significantly extended the ability to study the human genome and to better understand the biology of genomes, the new technology has required profound changes to the data analysis. There is a substantial need for computational methods that allow a convenient analysis of these overwhelmingly high-throughput data sets and address an increasing number of compelling biological questions which …
Investigation Of New Feature Descriptors For Image Search And Classification, Atreyee Sinha
Investigation Of New Feature Descriptors For Image Search And Classification, Atreyee Sinha
Dissertations
Content-based image search, classification and retrieval is an active and important research area due to its broad applications as well as the complexity of the problem. Understanding the semantics and contents of images for recognition remains one of the most difficult and prevailing problems in the machine intelligence and computer vision community. With large variations in size, pose, illumination and occlusions, image classification is a very challenging task. A good classification framework should address the key issues of discriminatory feature extraction as well as efficient and accurate classification. Towards that end, this dissertation focuses on exploring new image descriptors by …
Risk Prediction With Genomic Data, Bharati Jadhav
Risk Prediction With Genomic Data, Bharati Jadhav
Theses
Genome wide association study (GWAS) is widely used with various machine learning algorithms to predict disease risk. This thesis investigates this widely used approach of GWAS using Single Nucleotide Polymorphism (SNP) genotype data and a novel approach of disease risk prediction with whole exome sequencing data, namely Whole Exome Wide Association Study (WEWAS). It further applies a discriminating machine learning algorithm, namely a Support Vector Machine (SVM) with different Kernel functions. For this study, only SNPs generated using genotyping technology, which focuses more on common variants, are used initially for disease prediction. Later, the whole exome data generated using Next …
Increasing Adolescent Interest In Computing Through The Use Of Social Cognitive Career Theory, Osama Eljabiri
Increasing Adolescent Interest In Computing Through The Use Of Social Cognitive Career Theory, Osama Eljabiri
Dissertations
While empirical research efforts are sufficient to provide evidence of the role of most constructs in the Social Cognitive Career Theory (SCCT), this dissertation shifts the research focus and finds serious shortcomings in defining the construct of computer technology learning experiences design.
The purpose of this dissertation is to investigate whether, and to what extent, the proposed SCCT-enhanced framework can increase self-efficacy and interest of pre-college and college students in computer-based technology through the newly proposed “Learning Experiences” construct; in particular, whether it can reduce the gender gaps.
As a result of a comprehensive literature review, the dissertation connects learning, …
Innovative Local Texture Descriptors With Application To Eye Detection, Jiayu Gu
Innovative Local Texture Descriptors With Application To Eye Detection, Jiayu Gu
Dissertations
Local Binary Patterns (LBP), which is one of the well-known texture descriptors, has broad applications in pattern recognition and computer vision. The attractive properties of LBP are its tolerance to illumination variations and its computational simplicity. However, LBP only compares a pixel with those in its own neighborhood and encodes little information about the relationship of the local texture with the features. This dissertation introduces a new Feature Local Binary Patterns (FLBP) texture descriptor that can compare a pixel with those in its own neighborhood as well as in other neighborhoods and encodes the information of both local texture and …
Using Structural And Semantic Methodologies To Enhance Biomedical Terminologies, Zhe He
Using Structural And Semantic Methodologies To Enhance Biomedical Terminologies, Zhe He
Dissertations
Biomedical terminologies and ontologies underlie various Health Information Systems (HISs), Electronic Health Record (EHR) Systems, Health Information Exchanges (HIEs) and health administrative systems. Moreover, the proliferation of interdisciplinary research efforts in the biomedical field is fueling the need to overcome terminological barriers when integrating knowledge from different fields into a unified research project. Therefore well-developed and well-maintained terminologies are in high demand. Most of the biomedical terminologies are large and complex, which makes it impossible for human experts to manually detect and correct all errors and inconsistencies. Automated and semi-automated Quality Assurance methodologies that focus on areas that are more …
Vehicle Re-Routing Strategies For Congestion Avoidance, Juan Pan
Vehicle Re-Routing Strategies For Congestion Avoidance, Juan Pan
Dissertations
Traffic congestion causes driver frustration and costs billions of dollars annually in lost time and fuel consumption. This dissertation introduces a cost-effective and easily deployable vehicular re-routing system that reduces the effects of traffic congestion. The system collects real-time traffic data from vehicles and road-side sensors, and computes proactive, individually tailored re-routing guidance, which is pushed to vehicles when signs of congestion are observed on their routes. Subsequently, this dissertation proposes and evaluates two classes of re-routing strategies designed to be incorporated into this system, namely, Single Shortest Path strategies and Multiple Shortest Paths Strategies.
These strategies are firstly implemented …
Svmaud: Using Textual Information To Predict The Audience Level Of Written Works Using Support Vector Machines, Todd Will
Dissertations
Information retrieval systems should seek to match resources with the reading ability of the individual user; similarly, an author must choose vocabulary and sentence structures appropriate for his or her audience. Traditional readability formulas, including the popular Flesch-Kincaid Reading Age and the Dale-Chall Reading Ease Score, rely on numerical representations of text characteristics, including syllable counts and sentence lengths, to suggest audience level of resources. However, the author’s chosen vocabulary, sentence structure, and even the page formatting can alter the predicted audience level by several levels, especially in the case of digital library resources. For these reasons, the performance of …
Comparison Of Different Differential Expression Analysis Tools For Rna-Seq Data, Junfei Zhu
Comparison Of Different Differential Expression Analysis Tools For Rna-Seq Data, Junfei Zhu
Theses
In molecular biology research, RNA-seq is a relatively new method for transcriptome profiling. It utilizes the next generation sequencing technology to provide huge amount information about the variety and abundance of RNA present in an organism of interest at a specific state and a given time. One of the most important tasks of RNA-seq analysis is finding genes that are expressed differently in different subject groups. A lot of differential expression analysis tools for RNA-seq have been developed, but there is no golden standard in this field. In this research, four commonly used tools (DESeq, edgeR, limma, and cuffdiff) are …
Structural Indicators For Effective Quality Assurance Of Snomed Ct, Ankur Agrawal
Structural Indicators For Effective Quality Assurance Of Snomed Ct, Ankur Agrawal
Dissertations
The Standardized Nomenclature of Medicine -- Clinical Terms (SNOMED CT -- further abbreviated as SCT) has been endorsed as a premier clinical terminology by many national and international organizations. The US Government has chosen SCT to play a significant role in its initiative to promote Electronic Health Record (EH R) country-wide. However, there is evidence suggesting that, at the moment, SCT is not optimally modeled for its intended use by healthcare practitioners. There is a need to perform quality assurance (QA) of SCT to help expedite its use as a reference terminology for clinical purposes as planned for EH R …
Concept Graphs: Applications To Biomedical Text Categorization And Concept Extraction, Said Bleik
Concept Graphs: Applications To Biomedical Text Categorization And Concept Extraction, Said Bleik
Dissertations
As science advances, the underlying literature grows rapidly providing valuable knowledge mines for researchers and practitioners. The text content that makes up these knowledge collections is often unstructured and, thus, extracting relevant or novel information could be nontrivial and costly. In addition, human knowledge and expertise are being transformed into structured digital information in the form of vocabulary databases and ontologies. These knowledge bases hold substantial hierarchical and semantic relationships of common domain concepts. Consequently, automating learning tasks could be reinforced with those knowledge bases through constructing human-like representations of knowledge. This allows developing algorithms that simulate the human reasoning …
Novel Color And Local Image Descriptors For Content-Based Image Search, Sugata Banerji
Novel Color And Local Image Descriptors For Content-Based Image Search, Sugata Banerji
Dissertations
Content-based image classification, search and retrieval is a rapidly-expanding research area. With the advent of inexpensive digital cameras, cheap data storage, fast computing speeds and ever-increasing data transfer rates, millions of images are stored and shared over the Internet every day. This necessitates the development of systems that can classify these images into various categories without human intervention and on being presented a query image, can identify its contents in order to retrieve similar images.
Towards that end, this dissertation focuses on investigating novel image descriptors based on texture, shape, color, and local information for advancing content-based image search. Specifically, …
Genome Wide Search For Pseudo Knotted Non-Coding Rnas, Meghana S. Vasavada
Genome Wide Search For Pseudo Knotted Non-Coding Rnas, Meghana S. Vasavada
Theses
Non-coding RNAs (ncRNAs) are the functional RNA molecules that are involved in many biological processes including gene regulation, chromosome replication and RNA modification. Searching genomes using computational methods has become an important asset for prediction and annotation of ncRNAs. To annotate an individual genome for a specific family of ncRNAs, a computational tool is interpreted to scan through the genome and align its sequence segments to some structure model for the ncRNA family. With the recent advances in detecting an ncRNA in the genome, heuristic techniques are designed to perform an accurate search and sequence-structure alignment. This study uses a …
A Gpu Program To Compute Snp-Snp Interactions In Genome-Wide Association Studies, Srividya Ramakrishnan
A Gpu Program To Compute Snp-Snp Interactions In Genome-Wide Association Studies, Srividya Ramakrishnan
Theses
With the recent advances in the next generation sequencing technologies, short read sequences of human genome are made more accessible. Paired end sequencing of short reads is currently the most sensitive method for detecting somatic mutations that arise during tumor development. In this study, a novel approach to optimize the detection of structural variants using a new short read alignment program is presented.
Pairwise interaction effects of the Single Nucleotide Polymorphisms (SNPs) have proven to uncover the underlying complex disease traits. Computing the disease risk based on the interaction effects of SNPs on a case - control study is a …
Rna-Sequence Analysis Of Human Melanoma Cells, Jharna Miya
Rna-Sequence Analysis Of Human Melanoma Cells, Jharna Miya
Theses
RNA-sequencing refers to the use of high throughput sequencing technologies that are used to sequence cDNA in order to get the complete information of a sample’s RNA content. The objective of this study is to analyze this data in different aspects and to characterize gene expression. Besides this characterization, the data was also used to investigate the effect of sequencing depth on gene expression measurements.
This research focuses on quantitative measurement of expression levels of genes and their transcripts. In this study, complementary DNA fragments of cultured human melanoma cells are sequenced and a total of 139,501,106 million 200-bp reads …
Performance Comparison Of Five Rna-Seq Alignment Tools, Yuanpeng Lu
Performance Comparison Of Five Rna-Seq Alignment Tools, Yuanpeng Lu
Theses
Aligning millions of short reads to a reference genome is a critical task in high throughput sequencing. In recent years, a large number of mapping algorithms have been developed, all of which have in common that they align a vast number of reads to genomic or transcriptomic sequences. RNA-Seq data is discrete in nature, therefore with reasonable gene models and comparative metrics RNA-Seq data can be simulated to sufficient accuracy to enable meaningful benchmarking of alignment algorithms. To provide guidance in the choice of alignment algorithms, five different alignment tools for RNA-Seq data are evaluated. In order to compare the …
Polyaseeker: A Computational Framework For Identifying Polyadenylation Cleavage Site From Rna-Seq, Xiao Ling
Polyaseeker: A Computational Framework For Identifying Polyadenylation Cleavage Site From Rna-Seq, Xiao Ling
Theses
Alternative polyadenylation (APA) of mRNA plays a crucial role for post-transcriptional gene regulation. Recently, advances in next generation sequencing technology have made it possible to efficiently characterize the transcriptome and identify the 3’end of polyadenylated RNAs. However, no comprehensive bioi nformatic pipelines have fulfilled this goal. The PolyASeeker, a computational framework for identifying polyadenylation cleavage sites from RNA-Seq data is proposed in this thesis. By using the simulated RNA-seq dataset, a novel method is developed to evaluate the performance of the proposed framework versus the traditional A-stretch approach, and compute accurate Precisions and Recalls that previous estimation could not get. …
Eye Detection Using Discriminatory Features And An Efficient Support Vector Machine, Shuo Chen
Eye Detection Using Discriminatory Features And An Efficient Support Vector Machine, Shuo Chen
Dissertations
Accurate and efficient eye detection has broad applications in computer vision, machine learning, and pattern recognition. This dissertation presents a number of accurate and efficient eye detection methods using various discriminatory features and a new efficient Support Vector Machine (eSVM).
This dissertation first introduces five popular image representation methods - the gray-scale image representation, the color image representation, the 2D Haar wavelet image representation, the Histograms of Oriented Gradients (HOG) image representation, and the Local Binary Patterns (LBP) image representation - and then applies these methods to derive five types of discriminatory features. Comparative assessments are then presented to evaluate …
Structural Analysis And Auditing Of Snomed Hierarchies Using Abstraction Networks, Yue Wang
Structural Analysis And Auditing Of Snomed Hierarchies Using Abstraction Networks, Yue Wang
Dissertations
SNOMED is one of the leading healthcare terminologies being used worldwide. Due to its sheer volume and continuing expansion, it is inevitable that errors will make their way into SNOMED. Thus, quality assurance is an important part of its maintenance cycle.
A structural approach is presented in this dissertation, aiming at developing automated techniques that can aid auditors in the discovery of terminology errors more effectively and efficiently. Large SNOMED hierarchies are partitioned, based primarily on their relationships patterns, into concept groups of more manageable sizes. Three related abstraction networks with respect to a SNOMED hierarchy, namely the area taxonomy …
Effects Of Information Importance And Distribution On Information Exchange In Team Decision Making, Babajide James Osatuyi
Effects Of Information Importance And Distribution On Information Exchange In Team Decision Making, Babajide James Osatuyi
Dissertations
Teams in organizations are strategically built with members from domains and experiences so that a wider range of information and options can be pooled. This strategic team structure is based on the assumption that when team members share the information they have, the team as a whole can access a larger pool of information than any one member acting alone, potentially enabling them to make better decisions. However, studies have shown that teams, unlike individuals, sometimes do not effectively share and use the unique information available to them, leading to poorer decisions. Research on information sharing in team decision making …
Registration And Categorization Of Camera Captured Documents, Venkata Gopal Edupuganti
Registration And Categorization Of Camera Captured Documents, Venkata Gopal Edupuganti
Dissertations
Camera captured document image analysis concerns with processing of documents captured with hand-held sensors, smart phones, or other capturing devices using advanced image processing, computer vision, pattern recognition, and machine learning techniques. As there is no constrained capturing in the real world, the captured documents suffer from illumination variation, viewpoint variation, highly variable scale/resolution, background clutter, occlusion, and non-rigid deformations e.g., folds and crumples. Document registration is a problem where the image of a template document whose layout is known is registered with a test document image. Literature in camera captured document mosaicing addressed the registration of captured documents with …
Example Based Texture Synthesis And Quantification Of Texture Quality, Chandralekha De
Example Based Texture Synthesis And Quantification Of Texture Quality, Chandralekha De
Dissertations
Textures have been used effectively to create realistic environments for virtual worlds by reproducing the surface appearances. One of the widely-used methods for creating textures is the example based texture synthesis method. In this method of generating a texture of arbitrary size, an input image from the real world is provided. This input image is used for the basis of generating large textures. Various methods based on the underlying pattern of the image have been used to create these textures; however, the problem of finding an algorithm which provides a good output is still an open research issue. Moreover, the …
Reducing The Risk Of Software Cost Estimation, Shixian Yang
Reducing The Risk Of Software Cost Estimation, Shixian Yang
Theses
Inaccurate cost estimation is a well-known problem in software development. The common cost estimation models are point estimates that are unable to quantify uncertainties. Furthermore, it is difficult to calibrate the uncertainties in cost estimation due to the lack of information. The purpose of this thesis is to prove that probability techniques could be synthesized into COCOMO (Constructive Cost Model) to quantify uncertainties. Another aim is to find out how to get more insight on reducing the risk of cost estimation. In this thesis, some historical data is presented to show the variance in factors of COCOMO. Monte Carlo simulation …
Phenotype Prediction And Feature Selection In Genome-Wide Association Studies, Andrew Roberts
Phenotype Prediction And Feature Selection In Genome-Wide Association Studies, Andrew Roberts
Theses
Genome wide association studies (GWAS) search for correlations between single nucleotide polymorphisms (SNPs) in a subject genome and an observed phenotype. GWAS can be used to generate models for predicting phenotype based on genotype, as well as aiding in identification of specific genes affecting the biological mechanism underlying the phenotype.
In this investigation, phenotype prediction models are constructed from GWAS training data and are evaluated for performance on test data. Three methods are used to rank SNPs by their correlation with the phenotype: the univariate Wald test, a multivariate, support vector machine (SVM) based technique, and a hybrid method where …
Heterogeneity-Aware And Energy-Aware Scheduling And Routing In Wireless Sensor Networks, Mahesh Kumar Vasanthu Somashekar
Heterogeneity-Aware And Energy-Aware Scheduling And Routing In Wireless Sensor Networks, Mahesh Kumar Vasanthu Somashekar
Theses
A Wireless Sensor Network (WSN) is a group of specialized transducers, called sensor nodes, with a communication infrastructure intended to monitor and record conditions at diverse locations. Since WSN applications are usually deployed in an open environment, the network is exposed to rough weather conditions, such as rain and snow. Another problem that WSN applications need to deal with is the energy constraints of sensor nodes. Both problems adversely affect the lifetime of WSN applications. A lot of research has been conducted to prolong the lifetime of WSN applications considering energy constraints of sensor nodes, but not much research has …
A Comparative Analysis Of Machine Learning Algorithms For Genome Wide Association Studies, Neha Singh
A Comparative Analysis Of Machine Learning Algorithms For Genome Wide Association Studies, Neha Singh
Theses
Variations present in human genome play a vital role in the emergence of genetic disorders and abnormal traits. Single Nucleotide Polymorphism (SNP) is considered as the most common source of genetic variations. Genome Wide Association Studies (GWAS) probe these variations present in human population and find their association with complex genetic disorders. Now these days, recent advances in technology and drastic reduction in costs of Genome Wide Association Studies provide the opportunity to have a plethora of genomic data that delivers huge information of these variations to analyze. In fact, there is significant difference in pace of data generation and …
Data Mining Of Tetraloop-Tetraloop Receptors In Rna Xml Files, Sinan Ramazanoglu
Data Mining Of Tetraloop-Tetraloop Receptors In Rna Xml Files, Sinan Ramazanoglu
Theses
RNA (Ribonucleic acid) Motifs are tertiary structures that play an important role in the folding mechanism of the RNA molecule. The overall function of a RNA Motif depends on its specific bp (base pairs) sequence that constitutes the secondary structure. Data mining is a novel method in both discovering potential tertiary structures within DNA (Deoxyribonucleic acid), RNA, and protein molecules and storing the information in databases. The RNA Motif of interest is the tetraloop-tetraloop receptor, which is composed of a highly conserved 11 nt (nucleotide) sequence and a tetraloop with the generic form of GNRA (where N = any base …