Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Artificial Intelligence and Robotics (685)
- Engineering (392)
- Computer Engineering (189)
- Medicine and Health Sciences (163)
- Electrical and Computer Engineering (152)
-
- Social and Behavioral Sciences (140)
- Data Science (130)
- Databases and Information Systems (122)
- Theory and Algorithms (99)
- Information Security (82)
- Life Sciences (79)
- Software Engineering (76)
- Numerical Analysis and Scientific Computing (75)
- Other Computer Sciences (70)
- Business (68)
- Statistics and Probability (60)
- Medical Specialties (58)
- Physics (41)
- Analytical, Diagnostic and Therapeutic Techniques and Equipment (40)
- Arts and Humanities (38)
- Mathematics (33)
- Operations Research, Systems Engineering and Industrial Engineering (33)
- Education (32)
- Diseases (29)
- Graphics and Human Computer Interfaces (29)
- Law (29)
- Applied Mathematics (28)
- Bioinformatics (28)
- Institution
-
- Old Dominion University (184)
- Singapore Management University (145)
- Air Force Institute of Technology (70)
- Zayed University (70)
- TÜBİTAK (65)
-
- New Jersey Institute of Technology (50)
- Brigham Young University (45)
- University of Texas at Arlington (45)
- Edith Cowan University (38)
- Technological University Dublin (38)
- Portland State University (36)
- Chapman University (30)
- San Jose State University (29)
- University of Nebraska - Lincoln (29)
- Wright State University (23)
- City University of New York (CUNY) (22)
- Utah State University (22)
- California Polytechnic State University, San Luis Obispo (19)
- University of Denver (19)
- University at Albany, State University of New York (18)
- University of Kentucky (18)
- Boise State University (17)
- Dartmouth College (17)
- Thomas Jefferson University (17)
- Purdue University (15)
- University of Arkansas, Fayetteville (15)
- University of South Florida (15)
- University of Texas Rio Grande Valley (15)
- Loyola University Chicago (14)
- University of Louisville (14)
- Publication Year
- Publication
-
- Theses and Dissertations (128)
- Research Collection School Of Computing and Information Systems (121)
- All Works (70)
- Turkish Journal of Electrical Engineering and Computer Sciences (65)
- Dissertations (56)
-
- Electrical & Computer Engineering Faculty Publications (45)
- Electronic Theses and Dissertations (41)
- Faculty Publications (41)
- Computer Science Faculty Publications (31)
- Computer Science and Engineering Dissertations - Archive (24)
- Research outputs 2022 to 2026 (23)
- Master's Projects (22)
- Master's Theses (22)
- Browse all Theses and Dissertations (21)
- Dissertations and Theses (21)
- Computer Science and Engineering Theses - Archive (19)
- Conference papers (19)
- Legacy Theses & Dissertations (2009 - 2024) (18)
- Boise State University Theses and Dissertations (16)
- Articles (14)
- Theses (13)
- USF Tampa Graduate Theses and Dissertations (13)
- CCAC Theses and Dissertations (12)
- Computer Science Theses & Dissertations (12)
- Journal of System Simulation (12)
- Computer Science: Faculty Publications and Other Works (11)
- Electrical & Computer Engineering Theses & Dissertations (11)
- Engineering Management & Systems Engineering Faculty Publications (11)
- Graduate Theses and Dissertations (11)
- Mathematics, Physics, and Computer Science Faculty Articles and Research (11)
- Publication Type
- File Type
Articles 1111 - 1140 of 1665
Full-Text Articles in Computer Sciences
Machine Learning Prediction Of Glioblastoma Patient One-Year Survival, Andrew Du '20, Warren Mcgee, Jane Y. Wu
Machine Learning Prediction Of Glioblastoma Patient One-Year Survival, Andrew Du '20, Warren Mcgee, Jane Y. Wu
Student Publications & Research
Glioblastoma (GBM) is a grade IV astrocytoma formed primarily from cancerous astrocytes and sustained by intense angiogenesis. GBM often causes non-specific symptoms, creating difficulty for diagnosis. This study aimed to utilize machine learning techniques to provide an accurate one-year survival prognosis for GBM patients using clinical and genomic data from the Chinese Glioma Genome Atlas. Logistic regression (LR), support vector machines (SVM), random forest (RF), and ensemble models were used to identify and select predictors for GBM survival and to classify patients into those with an overall survival (OS) of less than one year and one year or greater. With …
Development Of Machine Learning Tutorials For R, John Pintar
Development Of Machine Learning Tutorials For R, John Pintar
All Undergraduate Theses and Capstone Projects
Machine learning (ML) techniques developed in computer science have revolutionized nearly every sector of industry. Despite the prevalence and usefulness of ML, students outside of computer science rarely receive training in ML. Students frequently receive training in statistical analysis, often using the software package R, which is free, open source, and has additional downloadable modules. A popular module is the ML package caret, which contains 238 different ML algorithms, each with 0-9 hyperparameters. caret is powerful, flexible, and provides consistent syntax across algorithms. In the hands of an experienced practitioner, this tunability is welcomed and can increase accuracy. However, when …
Multi-Class Twitter Data Categorization And Geocoding With A Novel Computing Framework, Sakib Mahmud Khan, Mashrur Chowdhury, Linh B. Ngo, Amy Apon
Multi-Class Twitter Data Categorization And Geocoding With A Novel Computing Framework, Sakib Mahmud Khan, Mashrur Chowdhury, Linh B. Ngo, Amy Apon
Computer Science Faculty Publications
This study details the progress in transportation data analysis with a novel computing framework in keeping with the continuous evolution of the computing technology. The computing framework combines the Labeled Latent Dirichlet Allocation (L-LDA)-incorporated Support Vector Machine (SVM) classifier with the supporting computing strategy on publicly available Twitter data in determining transportation-related events to provide reliable information to travelers. The analytical approach includes analyzing tweets using text classification and geocoding locations based on string similarity. A case study conducted for the New York City and its surrounding areas demonstrates the feasibility of the analytical approach. Approximately 700,010 tweets are analyzed …
Lung Cancer Subtype Differentiation From Positron Emission Tomography Images, Oğuzhan Ayyildiz, Zafer Aydin, Bülent Yilmaz, Seyhan Karaçavuş, Kübra Şenkaya, Semra İçer, Erdem Arzu Taşdemi̇r, Eser Kaya
Lung Cancer Subtype Differentiation From Positron Emission Tomography Images, Oğuzhan Ayyildiz, Zafer Aydin, Bülent Yilmaz, Seyhan Karaçavuş, Kübra Şenkaya, Semra İçer, Erdem Arzu Taşdemi̇r, Eser Kaya
Turkish Journal of Electrical Engineering and Computer Sciences
Lung cancer is one of the deadly cancer types, and almost 85 % of lung cancers are nonsmall cell lung cancer (NSCLC). In the present study we investigated classification and feature selection methods for the differentiation of two subtypes of NSCLC, namely adenocarcinoma (ADC) and squamous cell carcinoma (SqCC). The major advances in understanding the effects of therapy agents suggest that future targeted therapies will be increasingly subtype specific. We obtained positron emission tomography (PET) images of 93 patients with NSCLC, 39 of which had ADC while the rest had SqCC. Random walk segmentation was applied to delineate three-dimensional tumor …
Revised Polyhedral Conic Functions Algorithm For Supervised Classification, Gürhan Ceylan, Gürkan Öztürk
Revised Polyhedral Conic Functions Algorithm For Supervised Classification, Gürhan Ceylan, Gürkan Öztürk
Turkish Journal of Electrical Engineering and Computer Sciences
In supervised classification, obtaining nonlinear separating functions from an algorithm is crucial for prediction accuracy. This paper analyzes the polyhedral conic functions (PCF) algorithm that generates nonlinear separating functions by only solving simple subproblems. Then, a revised version of the algorithm is developed that achieves better generalization and fast training while maintaining the simplicity and high prediction accuracy of the original PCF algorithm. This is accomplished by making the following modifications to the subproblem: extension of the objective function with a regularization term, relaxation of a hard constraint set and introduction of a new error term. Experimental results show that …
A Machine Learning System For Glaucoma Detection Using Inexpensive Machine Learning, Jon Kilgannon
A Machine Learning System For Glaucoma Detection Using Inexpensive Machine Learning, Jon Kilgannon
West Chester University Master’s Theses
This thesis presents a neural network system which segments images of the retina to calculate the cup-to-disc ratio, one of the diagnostic indicators of the presence or continuing development of glaucoma, a disease of the eye which causes blindness. The neural network is designed to run on commodity hardware and to be run with minimal skill required from the user by packaging the software required to run the network into a Singularity image. The RIGA dataset used to train the network provides images of the retina which have been annotated with the location of the optic cup and disc by …
A Deep Learning Approach To Mapping Irrigation: U-Net Irrmapper, Thomas Henry Colligan Iv
A Deep Learning Approach To Mapping Irrigation: U-Net Irrmapper, Thomas Henry Colligan Iv
Graduate Student Theses, Dissertations, & Professional Papers
Accurate maps of irrigation are essential for understanding and managing water resources in light of a warming climate. We present a new method for mapping irrigation and apply it to the state of Montana over the years 2000-2019. The method is based on an ensemble of convolutional neural networks that only rely on raw Landsat surface reflectance data. The ensemble of networks method learns to mask clouds and ignore Landsat 7 scan-line failures without supervision, reducing the need for preprocessing data or feature engineering. Unlike other approaches to mapping irrigation, the method doesn't use other mapping products like the Cropland …
An Univariable Approach For Forecasting Workload In The Maintenance Industry, Paulo Silva, Fernando Pérez Téllez, John Cardiff
An Univariable Approach For Forecasting Workload In The Maintenance Industry, Paulo Silva, Fernando Pérez Téllez, John Cardiff
Articles
The forecasting of the workload in the maintenance industry is of great value to improve human resources allocation and reduce overwork. In this paper, we discuss the problem and the challenges it pertains. We analyze data from a company operating in the industry and present the results of several forecasting models.
Satire Identification In Turkish News Articles Based On Ensemble Of Classifiers, Aytuğ Onan, Mansur Alp Toçoğlu
Satire Identification In Turkish News Articles Based On Ensemble Of Classifiers, Aytuğ Onan, Mansur Alp Toçoğlu
Turkish Journal of Electrical Engineering and Computer Sciences
Social media and microblogging platforms generally contain elements of figurative and nonliteral language, including satire. The identification of figurative language is a fundamental task for sentiment analysis. It will not be possible to obtain sentiment analysis methods with high classification accuracy if elements of figurative language have not been properly identified. Satirical text is a kind of figurative language, in which irony and humor have been utilized to ridicule or criticize an event or entity. Satirical news is a pervasive issue on social media platforms, which can be deceptive and harmful. This paper presents an ensemble scheme for satirical news …
A Hierarchical Temporal Memory Sequence Classifier For Streaming Data, Jeffrey Barnett
A Hierarchical Temporal Memory Sequence Classifier For Streaming Data, Jeffrey Barnett
CCAC Theses and Dissertations
Real-world data streams often contain concept drift and noise. Additionally, it is often the case that due to their very nature, these real-world data streams also include temporal dependencies between data. Classifying data streams with one or more of these characteristics is exceptionally challenging. Classification of data within data streams is currently the primary focus of research efforts in many fields (i.e., intrusion detection, data mining, machine learning). Hierarchical Temporal Memory (HTM) is a type of sequence memory that exhibits some of the predictive and anomaly detection properties of the neocortex. HTM algorithms conduct training through exposure to a stream …
An Approach To Twitter Event Detection Using The Newsworthiness Metric, Jonathan Adkins
An Approach To Twitter Event Detection Using The Newsworthiness Metric, Jonathan Adkins
CCAC Theses and Dissertations
No abstract provided.
Detecting Rogue Manipulation Of Smart Home Device Settings, David Zeichick
Detecting Rogue Manipulation Of Smart Home Device Settings, David Zeichick
CCAC Theses and Dissertations
Smart home devices control a home’s environmental and security settings. This includes devices that control home thermostats, sprinkler systems, light bulbs, and home appliances. Malicious manipulation of the settings of these devices by an outside adversary has caused emotional distress and could even cause physical harm. For example, researchers have reported that there is a rise in domestic abuse perpetrated via smart home devices; victims have reported their thermostat settings being unwittingly manipulated and being locked out of their house due to their smart lock code being changed. Rapid adoption of smart home devices by consumers has led to an …
Heterogeneous Multi-Layered Network Model For Omics Data Integration And Analysis, Bohyun Lee, Shuo Zhang, Aleksandar Poleksic, Lei Xie
Heterogeneous Multi-Layered Network Model For Omics Data Integration And Analysis, Bohyun Lee, Shuo Zhang, Aleksandar Poleksic, Lei Xie
Faculty Work
Advances in next-generation sequencing and high-throughput techniques have enabled the generation of vast amounts of diverse omics data. These big data provide an unprecedented opportunity in biology, but impose great challenges in data integration, data mining, and knowledge discovery due to the complexity, heterogeneity, dynamics, uncertainty, and high-dimensionality inherited in the omics data. Network has been widely used to represent relations between entities in biological system, such as protein-protein interaction, gene regulation, and brain connectivity (i.e. network construction) as well as to infer novel relations given a reconstructed network (aka link prediction). Particularly, heterogeneous multi-layered network (HMLN) has proven successful …
Uncertainty Learning In Subjective Logic And Pattern Discovery In Network Data, Adilijiang Alimu
Uncertainty Learning In Subjective Logic And Pattern Discovery In Network Data, Adilijiang Alimu
Legacy Theses & Dissertations (2009 - 2024)
Uncertainty caused by unreliable or insufficient data and vulnerable machine learning models
Improving Pain Management In Patients With Sickle Cell Disease Using Machine Learning Techniques, Fan Yang
Improving Pain Management In Patients With Sickle Cell Disease Using Machine Learning Techniques, Fan Yang
Browse all Theses and Dissertations
Sickle cell disease (SCD) is an inherited red blood cell disorder that can cause a multitude of complications throughout a patient's life. Pain is the most common complication and a significant cause of morbidity. Since pain is a highly subjective experience, both medical providers and patients express difficulty in determining ideal treatment and management strategies for pain. Therefore, the development of objective pain assessment and pain forecasting methods is critical to pain management in SCD. On the other hand, the rapidly increasing use of mobile health (mHealth) technology and wearable devices gives the ability to build a remote health intervention …
Disaster Damage Categorization Applying Satellite Images And Machine Learning Algorithm, Farinaz Sabz Ali Pour, Adrian Gheorghe
Disaster Damage Categorization Applying Satellite Images And Machine Learning Algorithm, Farinaz Sabz Ali Pour, Adrian Gheorghe
Engineering Management & Systems Engineering Faculty Publications
Special information has a significant role in disaster management. Land cover mapping can detect short- and long-term changes and monitor the vulnerable habitats. It is an effective evaluation to be included in the disaster management system to protect the conservation areas. The critical visual and statistical information presented to the decision-makers can help in mitigation or adaption before crossing a threshold. This paper aims to contribute in the academic and the practice aspects by offering a potential solution to enhance the disaster data source effectiveness. The key research question that the authors try to answer in this paper is how …
Text Mining Methods For Analyzing Online Health Information And Communication, Sifei Han
Text Mining Methods For Analyzing Online Health Information And Communication, Sifei Han
Theses and Dissertations--Computer Science
The Internet provides an alternative way to share health information. Specifically, social network systems such as Twitter, Facebook, Reddit, and disease specific online support forums are increasingly being used to share information on health related topics. This could be in the form of personal health information disclosure to seek suggestions or answering other patients' questions based on their history. This social media uptake gives a new angle to improve the current health communication landscape with consumer generated content from social platforms. With these online modes of communication, health providers can offer more immediate support to the people seeking advice. Non-profit …
Analyze Informant-Based Questionnaire For The Early Diagnosis Of Senile Dementia Using Deep Learning, Fubao Zhu, Xiaonan Li, Daniel Mcgonigle, Haipeng Tang, Zhuo He, Chaoyang Zhang, Guang-Uei Hung, Pai-Yi Chiu, Weihua Zhou
Analyze Informant-Based Questionnaire For The Early Diagnosis Of Senile Dementia Using Deep Learning, Fubao Zhu, Xiaonan Li, Daniel Mcgonigle, Haipeng Tang, Zhuo He, Chaoyang Zhang, Guang-Uei Hung, Pai-Yi Chiu, Weihua Zhou
Michigan Tech Publications, Part 1
OBJECTIVE: This paper proposes a multiclass deep learning method for the classification of dementia using an informant-based questionnaire.
METHODS: A deep neural network classification model based on Keras framework is proposed in this paper. To evaluate the advantages of our proposed method, we compared the performance of our model with industry-standard machine learning approaches. We enrolled 6,701 individuals, which were randomly divided into training data sets (6030 participants) and test data sets (671 participants). We evaluated each diagnostic model in the test set using accuracy, precision, recall, and F1-Score.
RESULTS: Compared with the seven conventional machine learning algorithms, the DNN …
Pretraining Deep Learning Models For Natural Language Understanding, Han Shao
Pretraining Deep Learning Models For Natural Language Understanding, Han Shao
Honors Papers
Since the first bidirectional deep learn- ing model for natural language understanding, BERT, emerged in 2018, researchers have started to study and use pretrained bidirectional autoencoding or autoregressive models to solve language problems. In this project, I conducted research to fully understand BERT and XLNet and applied their pretrained models to two language tasks: reading comprehension (RACE) and part-of-speech tagging (The Penn Treebank). After experimenting with those released models, I implemented my own version of ELECTRA, a pretrained text encoder as a discriminator instead of a generator to improve compute-efficiency, with BERT as its underlying architecture. To reduce the number …
Machine Learning? In My Election? It's More Likely Than You Think: Voting Rules Via Neural Networks, Daniel Firebanks-Quevedo
Machine Learning? In My Election? It's More Likely Than You Think: Voting Rules Via Neural Networks, Daniel Firebanks-Quevedo
Honors Papers
Impossibility theorems in social choice have represented a barrier in the creation of universal, non-dictatorial, and non-manipulable voting rules, highlighting a key trade-off between social welfare and strategy-proofness. However, a social planner may be concerned with only a particular preference distribution and wonder whether it is possible to better optimize this trade-off. To address this problem, we propose an end-to-end, machine learning-based framework that creates voting rules according to a social planner's constraints, for any type of preference distribution. After experimenting with rank-based social choice rules, we find that automatically-designed rules are less susceptible to manipulation than most existing rules, …
Toward Efficient Automation Of Interpretable Machine Learning Boosting, Nathan Neuhaus
Toward Efficient Automation Of Interpretable Machine Learning Boosting, Nathan Neuhaus
All Master's Theses
Developing efficient automated methods for Interpretable Machine Learning (IML) is an important and long-term goal in the field of Artificial Intelligence. Currently the Machine Learning landscape is dominated by Neural Networks (NNs) and Support Vector Machines (SVMs), models which are often highly accurate. Despite high accuracy, such models are essentially “black boxes” and therefore are too risky for situations like healthcare where real lives are at stake. In such situations, so called “glass-box” models, such as Decision Trees (DTs), Bayesian Networks (BNs), and Logic Relational (LR) models are often preferred, however can succumb to accuracy limitations. Unfortunately, having to choose …
Modulation Of Medical Condition Likelihood By Patient History Similarity, Jonathan Turner, Dympna O'Sullivan, Jon Bird
Modulation Of Medical Condition Likelihood By Patient History Similarity, Jonathan Turner, Dympna O'Sullivan, Jon Bird
Articles
Introduction: We describe an analysis that modulates the simple population prevalence derived likelihood of a particular condition occurring in an individual by matching the individual with other individuals with similar clinical histories and determining the prevalence of the condition within the matched group.
Methods: We have taken clinical event codes and dates from anonymised longitudinal primary care records for 25,979 patients with 749,053 recorded clinical events. Using a nearest neighbour approach, for each patient, the likelihood of a condition occurring was adjusted from the population prevalence to the prevalence of the condition within those patients with the closest …
Language Model Co-Occurrence Linking For Interleaved Activity Discovery, Eoin Rogers, Robert J. Ross, John D. Kelleher
Language Model Co-Occurrence Linking For Interleaved Activity Discovery, Eoin Rogers, Robert J. Ross, John D. Kelleher
Conference papers
As ubiquitous computer and sensor systems become abundant, the potential for automatic identification and tracking of human behaviours becomes all the more evident. Annotating complex human behaviour datasets to achieve ground truth for supervised training can however be extremely labour-intensive, and error prone. One possible solution to this problem is activity discovery: the identification of activities in an unlabelled dataset by means of an unsupervised algorithm. This paper presents a novel approach to activity discovery that utilises deep learning based language production models to construct a hierarchical, tree-like structure over a sequential vector of sensor events. Our approach differs from …
Synthesising Tabular Datasets Using Wasserstein Conditional Gans With Gradient Penalty (Wcgan-Gp), Manhar Singh Walia, Brendan Tierney, Susan Mckeever
Synthesising Tabular Datasets Using Wasserstein Conditional Gans With Gradient Penalty (Wcgan-Gp), Manhar Singh Walia, Brendan Tierney, Susan Mckeever
Conference papers
Deep learning based methods based on Generative Adversarial Networks (GANs) have seen remarkable success in data synthesis of images and text. This study investigates the use of GANs for the generation of tabular mixed dataset. We apply Wasserstein Conditional Generative Adversarial Network (WCGAN-GP) to the task of generating tabular synthetic data that is indistinguishable from the real data, without incurring information leakage. The performance of WCGAN-GP is compared against both the ground truth datasets and SMOTE using three labelled real-world datasets from different domains. Our results for WCGAN-GP show that the synthetic data preserves distributions and relationships of the real …
Design Of A Novel Wearable Ultrasound Vest For Autonomous Monitoring Of The Heart Using Machine Learning, Garrett G. Goodman
Design Of A Novel Wearable Ultrasound Vest For Autonomous Monitoring Of The Heart Using Machine Learning, Garrett G. Goodman
Browse all Theses and Dissertations
As the population of older individuals increases worldwide, the number of people with cardiovascular issues and diseases is also increasing. The rate at which individuals in the United States of America and worldwide that succumb to Cardiovascular Disease (CVD) is rising as well. Approximately 2,303 Americans die to some form of CVD per day according to the American Heart Association. Furthermore, the Center for Disease Control and Prevention states that 647,000 Americans die yearly due to some form of CVD, which equates to one person every 37 seconds. Finally, the World Health Organization reports that the number one cause of …
Comparing Predictive Performance Of Statistical Learning Models On Medical Data, Francis Biney
Comparing Predictive Performance Of Statistical Learning Models On Medical Data, Francis Biney
Open Access Theses & Dissertations
This work investigates the predictive performance of 10 Machine learning models on three medical data including Breast cancer, Heart disease and Prostate cancer. Furthermore, we use the models to identify risk factors that contribute significantly to these diseases.
The models considered include; Logistic regression with L1 and L_2 penalties, Principal component logistic regression(PCR-LR), Partial least squares logistic regression(PLS-LR), Multivariate adaptive regression splines(MARS), Support vector machine with Radial Basis Kernel (SVM-RBK), Random Forest(RF), Gradient Boosting Machines(GBM), Elastic Net (Enet) and Feedforward Neural Network(FFNN). The models were grouped according to their similarities and learning style; i) Linear regularized models: LR-Lasso, LR-Ridge and …
Towards Personalized Medicine: Computational Approaches For Drug Repurposing And Cell Type Identification, Azam Peyvandipour
Towards Personalized Medicine: Computational Approaches For Drug Repurposing And Cell Type Identification, Azam Peyvandipour
Wayne State University Dissertations
The traditional drug discovery process is extremely slow and costly. More than 90% of drugs fail to pass beyond the early stage of development and toxicity tests, and many of the drugs that go through early phases of the clinical trials fail because of adverse reactions, side effects, or lack of efficiency. In spite of unprecedented investments in research and development (R&D), the number of new FDA-approved drugs remains low, reflecting the limitations of the current R&D model.
In this context, finding new disease indications for existing drugs sidesteps these issues and can therefore increase the available therapeutic choices at …
Modelling Interleaved Activities Using Language Models, Eoin Rogers, Robert J. Ross, John D. Kelleher
Modelling Interleaved Activities Using Language Models, Eoin Rogers, Robert J. Ross, John D. Kelleher
Conference papers
We propose a new approach to activity discovery, based on the neural language modelling of streaming sensor events. Our approach proceeds in multiple stages: we build binary links between activities using probability distributions generated by a neural language model trained on the dataset, and combine the binary links to produce complex activities. We then use the activities as sensor events, allowing us to build complex hierarchies of activities. We put an emphasis on dealing with interleaving, which represents a major challenge for many existing activity discovery systems. The system is tested on a realistic dataset, demonstrating it as a promising …
A Probabilistic Machine Learning Framework For Cloud Resource Selection On The Cloud, Syeduzzaman Khan
A Probabilistic Machine Learning Framework For Cloud Resource Selection On The Cloud, Syeduzzaman Khan
University of the Pacific Theses and Dissertations
The execution of the scientific applications on the Cloud comes with great flexibility, scalability, cost-effectiveness, and substantial computing power. Market-leading Cloud service providers such as Amazon Web service (AWS), Azure, Google Cloud Platform (GCP) offer various general purposes, memory-intensive, and compute-intensive Cloud instances for the execution of scientific applications. The scientific community, especially small research institutions and undergraduate universities, face many hurdles while conducting high-performance computing research in the absence of large dedicated clusters. The Cloud provides a lucrative alternative to dedicated clusters, however a wide range of Cloud computing choices makes the instance selection for the end-users. This thesis …
Searching For Needles In The Cosmic Haystack, Thomas Ryan Devine
Searching For Needles In The Cosmic Haystack, Thomas Ryan Devine
Graduate Theses, Dissertations, and Problem Reports (ETD)
Searching for pulsar signals in radio astronomy data sets is a difficult task. The data sets are extremely large, approaching the petabyte scale, and are growing larger as instruments become more advanced. Big Data brings with it big challenges. Processing the data to identify candidate pulsar signals is computationally expensive and must utilize parallelism to be scalable. Labeling benchmarks for supervised classification is costly. To compound the problem, pulsar signals are very rare, e.g., only 0.05% of the instances in one data set represent pulsars. Furthermore, there are many different approaches to candidate classification with no consensus on a best …