Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Computer Sciences (26)
- Artificial Intelligence and Robotics (13)
- Social and Behavioral Sciences (6)
- Theory and Algorithms (5)
- Medicine and Health Sciences (4)
-
- Cybersecurity (3)
- Engineering (3)
- Genetics and Genomics (3)
- Graphics and Human Computer Interfaces (3)
- Information Security (3)
- Library and Information Science (3)
- Life Sciences (3)
- Computational Biology (2)
- Databases and Information Systems (2)
- Electrical and Computer Engineering (2)
- Medical Specialties (2)
- Oncology (2)
- Physics (2)
- Applied Mathematics (1)
- Archival Science (1)
- Biomedical Engineering and Bioengineering (1)
- Business (1)
- Cognition and Perception (1)
- Communication (1)
- Disability and Equity in Education (1)
- Education (1)
- Engineering Physics (1)
- Epidemiology (1)
- Keyword
-
- Machine learning (4)
- Learning systems (3)
- Distributed computing (2)
- Graph neural networks (2)
- Graph perception (2)
-
- Graph theory (2)
- Graph usability (2)
- Low vision (2)
- Machine learning—Mathematical models (2)
- Pattern recognition systems (2)
- Privacy (2)
- Quantum algorithms (2)
- Screen magnifier (2)
- Adaptation models (1)
- Adversarial machine learning (1)
- Algorithms—Optimization (1)
- Analysis and statistical methods (1)
- ArXiv (1)
- Artificial intelligence (1)
- Artificial intelligence—Reliability (1)
- Artificial intelligence—Research (1)
- Aviation industry (1)
- Bayesian analysis (1)
- Biochemical markers (1)
- Bioinformatics (1)
- Biomarkers (1)
- Breast Cancer (1)
- Cancer subtyping (1)
- Catastrophic forgetting (1)
- Class incremental learning (1)
Articles 1 - 30 of 30
Full-Text Articles in Data Science
An Explainable Transformer Framework For Sentiment Analysis In Aviation Workforce Data, Sovon Chakraborty, Protiva Das, Fahmid Al Farid, Fuyad Hasan Bhoyan, Farig Yousuf Sadeque, Jia Uddin, Hezerul Abdul Karim
An Explainable Transformer Framework For Sentiment Analysis In Aviation Workforce Data, Sovon Chakraborty, Protiva Das, Fahmid Al Farid, Fuyad Hasan Bhoyan, Farig Yousuf Sadeque, Jia Uddin, Hezerul Abdul Karim
Computer Science Faculty Publications
Aviation is one of the predominant sectors that contribute significantly to the global economy. With the advent of technology, this industry is witnessing a paradigm shift towards data-driven approaches. The morale of the airline employees is barely noticed, which causes fatigue and depression. Furthermore, these mental health issues can be active reasons for destructive accidents. In this research, the authors are focused on collecting insightful information on aviation employees from Glassdoor.com. Moreover, the authors focus on analyzing the sentiments of the employees of renowned aviation companies. Primarily, the authors scraped necessary data from Glassdoor.com and created a dataset named JetJobJoy …
Quantum Machine Learning Models: Principles, Frameworks, And Computational Challenges, K. A. Jayabalaji, S. Venkata Anand, Dineshkumar Rajendran, Prasanta Chatterjee Biswas, Sardor Omonov, Rubaid Ashfaq
Quantum Machine Learning Models: Principles, Frameworks, And Computational Challenges, K. A. Jayabalaji, S. Venkata Anand, Dineshkumar Rajendran, Prasanta Chatterjee Biswas, Sardor Omonov, Rubaid Ashfaq
Computer Science Faculty Publications
Quantum machine learning (QML) has become an optimistic avenue of harnessing quantum computation in data-driven modeling, especially of issues with high dimensionality and complicated correlations. Current methods are generally based on fixed or over-parameterized quantum circuits, and hence restricted to scalability as well as unproductive optimization in real-world hardware. This chapter introduces a hybrid quantum-classical learning system that is adaptive and provides principled quantum data encoding, architecture-conscious variational circuit design and resource-optimal optimization. The technique is based on the concepts of quantum architecture search and subspace-preserving transformations to trade expressiveness with trainability, and discretize the quantum model into a classical …
Design And Analysis Of Modern Quantum Neural Network Architectures For Intelligent Systems, Lakshmi Chandrakanth Kasireddy, Prabhakara Rao Kapula, Dineshkumar Rajendran, Neha Bharani, Srikanth Pulipeti, Islombek Khushvaktov
Design And Analysis Of Modern Quantum Neural Network Architectures For Intelligent Systems, Lakshmi Chandrakanth Kasireddy, Prabhakara Rao Kapula, Dineshkumar Rajendran, Neha Bharani, Srikanth Pulipeti, Islombek Khushvaktov
Computer Science Faculty Publications
Quantum neural networks (QNNs) offer a principled pathway for integrating quantum computation with machine learning through superposition- and entanglement-based representations. This chapter proposes an architecture-aware design and evaluation framework for modern QNNs, emphasizing robustness and system feasibility alongside predictive performance. Multiple architectures variational QNNs, quantum convolutional neural networks, tensor-network hybrids, and fully quantum models—are assessed under a unified protocol. Experimental analysis shows that the proposed architecture-search–guided QNN achieves 91.8% classification accuracy and an F1-score of 0.914, outperforming fixed-template variational QNNs by approximately 5.6 percentage points. Under depolarizing noise with probability p = 0.10, the proposed model retains 85.3% accuracy, whereas …
Cognitive Prosthetic: An Ai-Enabled Multimodal System For Episodic Recall In Knowledge Work, Lawrence Obiuwevwi, Krzystof J. Rechowicz, Vikas Ashok, Sachin Shetty, Sampath Jayarathna
Cognitive Prosthetic: An Ai-Enabled Multimodal System For Episodic Recall In Knowledge Work, Lawrence Obiuwevwi, Krzystof J. Rechowicz, Vikas Ashok, Sachin Shetty, Sampath Jayarathna
Computer Science Faculty Publications
Modern knowledge workplaces increasingly strain human episodic memory as individuals navigate fragmented attention, overlapping meetings, and multimodal information streams. Existing workplace tools provide partial support through note-taking or analytics but rarely integrate cognitive, physiological, and attentional context into retrievable memory representations. This paper presents the Cognitive Prosthetic Multimodal System (CPMS)—an AI-enabled proof-of-concept designed to support episodic recall in knowledge work through structured episodic capture and natural language retrieval. CPMS synchronizes speech transcripts, physiological signals, and gaze behavior into temporally aligned, JSON-based episodic records processed locally for privacy. Beyond data logging, the system includes a web-based retrieval interface that allows users …
Privshap: A Finer-Granularity Network Linearization Method For Private Inference, Xiangrui Xu, Zhenzhen Wang, Rui Ning, Chunsheng Xiu, Hongyi Wu
Privshap: A Finer-Granularity Network Linearization Method For Private Inference, Xiangrui Xu, Zhenzhen Wang, Rui Ning, Chunsheng Xiu, Hongyi Wu
Computer Science Faculty Publications
Private inference applies cryptographic techniques like homomorphic encryption, garble circuit and secret sharing to keep both sides privacy in a client-server setting during inference. It is often hindered by the high communication overheads, especially at non-linear activation layers such as ReLU. Hence ReLU pruning has been widely recognized as an efficient way to accelerate private inference. Existing approaches to ReLU pruning typically rely on coarse hypothesis, which assume an inverse correlation between the importance of ReLU and linear layers or shallow activation layers have less importance for universal models, to assign the budgets according to the layer while preserving the …
Heterogeneous Clustering Of Multiomics Data For Breast Cancer Subgroup Classification And Detection, Joseph Pateras, Musaddiq Lodi, Pratip Rana, Preetam Ghosh
Heterogeneous Clustering Of Multiomics Data For Breast Cancer Subgroup Classification And Detection, Joseph Pateras, Musaddiq Lodi, Pratip Rana, Preetam Ghosh
Computer Science Faculty Publications
The rapid growth of diverse -omics datasets has made multiomics data integration crucial in cancer research. This study adapts the expectation–maximization routine for the joint latent variable modeling of multiomics patient profiles. By combining this approach with traditional biological feature selection methods, this study optimizes latent distribution, enabling efficient patient clustering from well-studied cancer types with reduced computational expense. The proposed optimization subroutines enhance survival analysis and improve runtime performance. This article presents a framework for distinguishing cancer subtypes and identifying potential biomarkers for breast cancer. Key insights into individual subtype expression and function were obtained through differentially expressed gene …
A Data-Driven Sliding-Window Pairwise Comparative Approach For The Estimation Of Transmission Fitness Of Sars-Cov-2 Variants And The Construction Of The Evolution Fitness Landscape, Md Jubair Pantho, Richard Annan, Landen Alexander Bauder, Sophia Huang, Letu Qingge, Hong Qin
A Data-Driven Sliding-Window Pairwise Comparative Approach For The Estimation Of Transmission Fitness Of Sars-Cov-2 Variants And The Construction Of The Evolution Fitness Landscape, Md Jubair Pantho, Richard Annan, Landen Alexander Bauder, Sophia Huang, Letu Qingge, Hong Qin
Computer Science Faculty Publications
Estimating the transmission fitness of SARS-CoV-2 variants and understanding their evolutionary fitness trends are important for epidemiological forecasting. Existing methods are often constrained by their parametric natures and do not satisfactorily align with the observations during COVID-19. Here, we introduce a sliding-window data-driven pairwise comparison method, the differential population growth rate (DPGR) that uses viral strains as internal controls to mitigate sampling biases. DPGR is applicable in time windows in which the logarithmic ratio of two variant subpopulations is approximately linear. We apply DPGR to genomic surveillance data and focus on variants of concern (VOCs) in multiple countries and regions. …
Benchmarking Batch-Effect Correction Methods Towards The Construction Of A Triple-Negative Breast Cancer Cell Atlas, Peter Scheible, Amy H. Tang, Jing He, Jiangwen Sun
Benchmarking Batch-Effect Correction Methods Towards The Construction Of A Triple-Negative Breast Cancer Cell Atlas, Peter Scheible, Amy H. Tang, Jing He, Jiangwen Sun
Computer Science Faculty Publications
Triple-negative breast cancer (TNBC) requires detailed cellular mapping given its aggressive nature, immense tumor heterogeneity and genetic diversity. We integrated 156,794 cells from six scRNA-seq datasets—including tumors, metastases, and cell lines—to build a TNBC scRNA cell atlas, focusing on batch effect mitigation while maintaining biological and molecular details. Preprocessing f ilters noise, normalizes data, and leverages PCA for integration readiness. We utilized scANVI, a semi-supervised tool, to align datasets, preserving TNBC’s complex tumor heterogeneity via marker annotations [1]. UMAPs demonstrate biological clustering in integrated data, contrasted with datasetdriven unintegrated patterns. Assessments verifying effective batch correction. This method aligns with NASA’s …
Ai For Nuclear Physics: The Exclaim Project, S. Liuti, D. Adams, M. Boër, G. W. Chern, M. Cuic, M. Engelhardt, G. R. Goldstein, B. Kriesten, Y. Li, H. W. Lin, M. Sievert, D. Sivers
Ai For Nuclear Physics: The Exclaim Project, S. Liuti, D. Adams, M. Boër, G. W. Chern, M. Cuic, M. Engelhardt, G. R. Goldstein, B. Kriesten, Y. Li, H. W. Lin, M. Sievert, D. Sivers
Computer Science Faculty Publications
An overview of the recent activity of the newly funded EXCLusives with AI and Machine learning (EXCLAIM) collaboration is presented. The main goal of the collaboration is to develop a framework to implement AI and machine learning techniques in problems emerging from the phenomenology of high energy exclusive scattering processes from nucleons and nuclei, maximizing the information that can be extracted from various sets of experimental data, while implementing theoretical constraints from lattice QCD. A specific perspective embraced by EXCLAIM is to use the methods of theoretical physics to understand the working of ML, beyond its standardized applications to physics …
Neural Topic Modeling Via Contextual And Graph Information Fusion, Jiyuan Liu, Jiaxing Yan, Chunjiang Zhu, Xingyu Liu, Qing Li, Yanghui Rao
Neural Topic Modeling Via Contextual And Graph Information Fusion, Jiyuan Liu, Jiaxing Yan, Chunjiang Zhu, Xingyu Liu, Qing Li, Yanghui Rao
Computer Science Faculty Publications
Topic modeling is a powerful unsupervised tool for knowledge discovery. However, existing work struggles with generating limited-quality topics that are uninformative and incoherent, which hindering interpretable insights from managing textual data. In this paper, we improve the original variational autoencoder framework by incorporating contextual and graph information to address the above issues. First, the encoder utilizes topic fusion techniques to combine contextual and bag-of-words information well, and meanwhile exploits the constraints of topic alignment and topic sharpening to generate informative topics. Second, we develop a simple word co-occurrence graph information fusion strategy that efficiently increases topic coherence. On three benchmark …
Decode The Workload: Training Deep Learning Models For Efficient Compute Cluster Representation, Ahmed Hossam Mohammed, Mark Jones, Diana Mcspadden, Malachi Schram, Bryan Hess, Kishansingh Rajput
Decode The Workload: Training Deep Learning Models For Efficient Compute Cluster Representation, Ahmed Hossam Mohammed, Mark Jones, Diana Mcspadden, Malachi Schram, Bryan Hess, Kishansingh Rajput
Computer Science Faculty Publications
In this study, we address the mounting challenge of monitoring high throughput computing clusters running computationally intensive jobs, which increasingly strains system administrators. We develop autoencoders that analyze traces of Linux kernel CPU metrics to capture salient system features by producing robust compressed embeddings for various downstream tasks. In addition, we employ graph neural networks to incorporate contextual information from surrounding CPUs and assess their performance. We also demonstrate the enhanced job differentiation achieved by increasing the sampling rate of these traces. Our models are evaluated based on their ability to generate meaningful latent representations, detect anomalies, and distinguish between …
S²Il: Structurally Stable Incremental Learning, S. Balasubramanian, P. Yedu Krishna, Talasu Sai Sriram, M. Sai Subramaniam, Manepalli Pranav Phanindra Sai, Ravi Mukkamala
S²Il: Structurally Stable Incremental Learning, S. Balasubramanian, P. Yedu Krishna, Talasu Sai Sriram, M. Sai Subramaniam, Manepalli Pranav Phanindra Sai, Ravi Mukkamala
Computer Science Faculty Publications
Feature Distillation (FD) strategies are proven to be effective in mitigating Catastrophic Forgetting (CF) seen in Class Incremental Learning (CIL). However, current FD approaches enforce strict alignment of feature magnitudes and directions across incremental steps, limiting the model’s ability to adapt to new knowledge. In this paper, we propose Structurally Stable Incremental Learning (S²IL), a FD method for CIL that mitigates forgetting by focusing on preserving the overall spatial patterns of features which promote flexible (plasticity) yet stable representations that preserve old knowledge (stability). We also demonstrate that our proposed method S²IL achieves strong incremental accuracy and outperforms other FD …
Effective Pii Extraction From Llms Through Augmented Few-Shot Learning, Shuai Cheng, Shu Meng, Haitao Xu, Haoran Zhang, Shuai Hao, Chuan Yue, Wenrui Ma, Meng Han, Fang Zhang, Zhao Li
Effective Pii Extraction From Llms Through Augmented Few-Shot Learning, Shuai Cheng, Shu Meng, Haitao Xu, Haoran Zhang, Shuai Hao, Chuan Yue, Wenrui Ma, Meng Han, Fang Zhang, Zhao Li
Computer Science Faculty Publications
Large Language Models (LLMs) exhibit strong natural language processing capabilities but also pose significant privacy risks, particularly regarding the leakage of Personally Identifiable Information (PII) embedded in their training data. Existing PII extraction methods suffer from the limitations of low success rates or impracticality for large-scale PII extraction. In this study, we propose a novel PII extraction approach based on enhanced few-shot learning techniques, which achieves efficient and cost-effective PII retrieval without relying on fine-tuning or jailbreaking. We evaluated our approach on both open-source and closed-source LLMs. The experimental results demonstrate that, for non-targeted PII extraction, the attack success rate …
Energy-Based Deep Incomplete Multi-View Clustering, Ziyu Wang, Yiming Du, Rui Ning, Lusi Li
Energy-Based Deep Incomplete Multi-View Clustering, Ziyu Wang, Yiming Du, Rui Ning, Lusi Li
Computer Science Faculty Publications
Incomplete multi-view clustering (IMVC) deals with real-world scenarios where certain views are partially missing, posing significant challenges to effective clustering. Most existing IMVC approaches face a trade-off: imputation-free methods suffer from information bias and imbalance, while full-imputation methods risk introducing and propagating noise. To overcome these limitations, we propose Energy-Based Deep Incomplete Multi-View Clustering (Energy-DIMC), a novel selective-imputation framework that leverages energy-based models (EBMs) to guide reliable imputations and robust clustering. EBMs assess data compatibility by assigning lower energy to more coherent structures, effectively modeling complex inter-view and inter-sample dependencies. Inspired by EBMs, Energy-DIMC integrates four key components: 1) a …
Icu-Length Of Stay Prediction On Electronic Health Records Using Graph Neural Networks And Homogeneous Similarity Graphs, Ahmad F. Al Musawi, Pratip Rana, Sibtanu Raha, Joshua Braunstein, William C. Sleeman Iv, Rishabh Kapoor, Preetam Ghosh
Icu-Length Of Stay Prediction On Electronic Health Records Using Graph Neural Networks And Homogeneous Similarity Graphs, Ahmad F. Al Musawi, Pratip Rana, Sibtanu Raha, Joshua Braunstein, William C. Sleeman Iv, Rishabh Kapoor, Preetam Ghosh
Computer Science Faculty Publications
Predicting the length of stay (LoS) is important for hospital administration, as it helps allocate proper resources, such as bed management and hospital staffing. Patients' Electronic Health Records (EHRs) contain highly relevant data for LoS prediction; however, their integration and effective use in predictive modeling for accurately estimating LoS remain challenging. To address this, we propose a homogeneous Graph Neural Network (GNN)-based framework for predicting LoS. This method employs a comprehensive data fusion strategy based on the hospital Visit-based Similarity Graph (VSG), which integrates diverse multi-modal clinical features into a coherent, homogeneous graph representation. Next, this VSG is fed into …
Runtime Support For Cpu-Gpu High-Performance Computing On Distributed Memory Platforms, Polykarpos Thomadakis, Nikos Chrisochoides
Runtime Support For Cpu-Gpu High-Performance Computing On Distributed Memory Platforms, Polykarpos Thomadakis, Nikos Chrisochoides
Computer Science Faculty Publications
Hardware heterogeneity is here to stay for high-performance computing. Large-scale systems are currently equipped with multiple GPU accelerators per compute node and are expected to incorporate more specialized hardware. This shift in the computing ecosystem offers many opportunities for performance improvement; however, it also increases the complexity of programming for such architectures. This work introduces a runtime framework that enables effortless programming for heterogeneous systems while efficiently utilizing hardware resources. The framework is integrated within a distributed and scalable runtime system to facilitate performance portability across heterogeneous nodes. Along with the design, this paper describes the implementation and optimizations performed, …
Bayesian Neural Netwok Variational Autoencoder Inverse Mapper (Bnn-Vaim) And Its Application In Compton Form Factors Extraction, Md Fayaz Bin Hossen, Tareq Alghamdi, Manal Almaeen, Yaohang Li
Bayesian Neural Netwok Variational Autoencoder Inverse Mapper (Bnn-Vaim) And Its Application In Compton Form Factors Extraction, Md Fayaz Bin Hossen, Tareq Alghamdi, Manal Almaeen, Yaohang Li
Computer Science Faculty Publications
We extend the Variational Autoencoder Inverse Mapper (VAIM) framework for the inverse problem of extracting Compton Form Factors (CFFs) from deeply virtual exclusive reactions, such as the unpolarized Deeply virtual exclusive scattering (DVCS) cross section. VAIM is an end-to-end deep learning framework to address the solution ambiguity issue in ill-posed inverse problems, which comprises of a forward mapper and a backward mapper to simulate the forward and inverse processes, respectively. In particular, we incorporate Bayesian Neural Network (BNN) into the VAIM architecture (BNN-VAIM) for uncertainty quantification. By sampling the weights and biases distributions of the BNN in the backward mapper …
Understaning Low Vision Graphical Perception Of Bar Charts, Yash Prakash, Akshay Kolgar Nayak, Sampath Jayarathna, Hae-Na Lee, Vikas Ashok
Understaning Low Vision Graphical Perception Of Bar Charts, Yash Prakash, Akshay Kolgar Nayak, Sampath Jayarathna, Hae-Na Lee, Vikas Ashok
Computer Science Faculty Publications
Bar charts are widely used for their simplicity in data representation, prompting numerous studies to explore and model how users interact with and perceive bar chart information. However, these studies have predominantly focused on sighted users, with a few also targeting blind screen-reader users, whereas the graphical perception of low-vision screen magnifier users is still an uncharted research territory. We fill this knowledge gap in this paper by designing four experiments for a laboratory study with 25 low-vision participants to examine their graphical perception while interacting with bar charts. For our investigation, we built a custom screen magnifier-based logger that …
Improving Usability Of Data Charts In Multimodal Documents For Low Vision Users, Yash Prakash, Akshay Kolgar Nayak, Shoaib Mohammed Alyaan, Pathan Aseef Khan, Hae-Na Lee, Vikas Ashok
Improving Usability Of Data Charts In Multimodal Documents For Low Vision Users, Yash Prakash, Akshay Kolgar Nayak, Shoaib Mohammed Alyaan, Pathan Aseef Khan, Hae-Na Lee, Vikas Ashok
Computer Science Faculty Publications
Data chart visualizations and text are often paired in news articles, online blogs, and academic publications to present complex data. While chart visualizations offer graphical summaries of the data, the accompanying text provides essential context and explanation. Associating information from text and charts is straightforward for sighted users but presents significant challenges for individuals with low vision, especially on small-screen devices such as smartphones. The visual nature of charts coupled with the layout of the text inherently makes it difficult for low vision users to mentally associate chart data with text and comprehend the content due to their dependence on …
Exacfs - A Cil Method To Mitigate Catastrophic Forgetting, S. Balasubramanian, Sai Subramaniam M., Sai Sriram Talasu, Manepalli Pranav Phanindra Sai, Yedu P. Krishna, Darshan Gera, Ravi Mukkamala
Exacfs - A Cil Method To Mitigate Catastrophic Forgetting, S. Balasubramanian, Sai Subramaniam M., Sai Sriram Talasu, Manepalli Pranav Phanindra Sai, Yedu P. Krishna, Darshan Gera, Ravi Mukkamala
Computer Science Faculty Publications
Deep neural networks (DNNs) excel at learning from static datasets but struggle with continual learning, where data arrives sequentially. Catastrophic forgetting, the phenomenon of forgetting previously learned knowledge, is a primary challenge. This paper introduces EXponentially Averaged Class-wise Feature Significance (EXACFS) to mitigate this issue in the class incremental learning (CIL) setting. By estimating the significance of model features for each learned class using loss gradients, gradually aging the significance through the incremental tasks and preserving the significant features through a distillation loss, EXACFS effectively balances remembering old knowledge (stability) and learning new knowledge (plasticity). Extensive experiments on CIFAR-100 and …
Learning Optimal Inter-Class Margin Adaptively For Few-Shot Class-Incremental Learning Via Neural Collapse-Based Meta-Learning, Hang Ran, Weijun Li, Lusi Li, Songsong Tian, Xin Ning, Prayag Tiwari
Learning Optimal Inter-Class Margin Adaptively For Few-Shot Class-Incremental Learning Via Neural Collapse-Based Meta-Learning, Hang Ran, Weijun Li, Lusi Li, Songsong Tian, Xin Ning, Prayag Tiwari
Computer Science Faculty Publications
Few-Shot Class-Incremental Learning (FSCIL) aims to learn new classes incrementally with a limited number of samples per class. It faces issues of forgetting previously learned classes and overfitting on few-shot classes. An efficient strategy is to learn features that are discriminative in both base and incremental sessions. Current methods improve discriminability by manually designing inter-class margins based on empirical observations, which can be suboptimal. The emerging Neural Collapse (NC) theory provides a theoretically optimal inter-class margin for classification, serving as a basis for adaptively computing the margin. Yet, it is designed for closed, balanced data, not for sequential or few-shot …
Osfs-Vague: Online Streaming Feature Selection Algorithm Based On A Vague Set, Jie Yang, Zhijun Wang, Guoyin Wang, Yanmin Liu, Yi He, Di Wu
Osfs-Vague: Online Streaming Feature Selection Algorithm Based On A Vague Set, Jie Yang, Zhijun Wang, Guoyin Wang, Yanmin Liu, Yi He, Di Wu
Computer Science Faculty Publications
Online streaming feature selection (OSFS), as an online learning manner to handle streaming features, is critical in addressing high-dimensional data. In real big data-related applications, the patterns and distributions of streaming features constantly change over time due to dynamic data generation environments. However, existing OSFS methods rely on presented and fixed hyperparameters, which undoubtedly lead to poor selection performance when encountering dynamic features. To make up for the existing shortcomings, the authors propose a novel OSFS algorithm based on vague set, named OSFS-Vague. Its main idea is to combine uncertainty and three-way decision theories to improve feature selection from the …
Quantification Of Landside Congestion In Ports: An Analysis Based On Gps Data, Kumushini Thennakoon, Namal Bandaranayake, Senevi Kiridena, Asela K. Kulatunga
Quantification Of Landside Congestion In Ports: An Analysis Based On Gps Data, Kumushini Thennakoon, Namal Bandaranayake, Senevi Kiridena, Asela K. Kulatunga
Computer Science Faculty Publications
Hinterland transport is a critical segment in maritime cross-border logistics, which links the end-users of global supply chains to the maritime segment. Truck-based hinterland transport is known to cause congestion in and around ports. This study aimed to quantify the congestion caused by trucks at the Port of Colombo, which has not been a subject of a systematic study. To this end, the study makes use of GPS data. In addition to revealing heavy congestion within the port, the study also reveals significant variations in congestion during different times of the day with the duration of journeys peaking from 1200hrs …
Deeppatent2: A Large-Scale Benchmarking Corpus For Technical Drawing Understanding, Kehinde Ajayi, Xin Wei, Martin Gryder, Winston Shields, Jian Wu, Shawn M. Jones, Michal Kucer, Diane Oyen
Deeppatent2: A Large-Scale Benchmarking Corpus For Technical Drawing Understanding, Kehinde Ajayi, Xin Wei, Martin Gryder, Winston Shields, Jian Wu, Shawn M. Jones, Michal Kucer, Diane Oyen
Computer Science Faculty Publications
Recent advances in computer vision (CV) and natural language processing have been driven by exploiting big data on practical applications. However, these research fields are still limited by the sheer volume, versatility, and diversity of the available datasets. CV tasks, such as image captioning, which has primarily been carried out on natural images, still struggle to produce accurate and meaningful captions on sketched images often included in scientific and technical documents. The advancement of other tasks such as 3D reconstruction from 2D images requires larger datasets with multiple viewpoints. We introduce DeepPatent2, a large-scale dataset, providing more than 2.7 million …
Online Deep Learning From Doubly-Streaming Data, Heng Lian, John S. Atwood, Bo-Jian Hou, Jian Wu, Yi He
Online Deep Learning From Doubly-Streaming Data, Heng Lian, John S. Atwood, Bo-Jian Hou, Jian Wu, Yi He
Computer Science Faculty Publications
This paper investigates a new online learning problem with doubly-streaming data, where the data streams are described by feature spaces that constantly evolve, with new features emerging and old features fading away. A plausible idea to deal with such data streams is to establish a relationship between the old and new feature spaces, so that an online learner can leverage the knowledge learned from the old features to better the learning performance on the new features. Unfortunately, this idea does not scale up to high-dimensional multimedia data with complex feature interplay, which suffers a tradeoff between onlineness, which biases shallow …
Scholarly Big Data Quality Assessment: A Case Study Of Document Linking And Conflation With S2orc, Jian Wu, Ryan Hiltabrand, Dominik Soós, C. Lee Giles
Scholarly Big Data Quality Assessment: A Case Study Of Document Linking And Conflation With S2orc, Jian Wu, Ryan Hiltabrand, Dominik Soós, C. Lee Giles
Computer Science Faculty Publications
Recently, the Allen Institute for Artificial Intelligence released the Semantic Scholar Open Research Corpus (S2ORC), one of the largest open-access scholarly big datasets with more than 130 million scholarly paper records. S2ORC contains a significant portion of automatically generated metadata. The metadata quality could impact downstream tasks such as citation analysis, citation prediction, and link analysis. In this project, we assess the document linking quality and estimate the document conflation rate for the S2ORC dataset. Using semi-automatically curated ground truth corpora, we estimated that the overall document linking quality is high, with 92.6% of documents correctly linking to six major …
Privacy In Iot Cloud, Aftab Ahmad, Ravi Mukkamala, Karthik Navuluri
Privacy In Iot Cloud, Aftab Ahmad, Ravi Mukkamala, Karthik Navuluri
Computer Science Faculty Publications
We present a framework for privacy preservation in an information cloud of IoT devices. We contend that privacy provisioning should be located in the user device and must protect the user, the information, and the device from breaches in privacy. We elaborate on how the layered privacy model can ensure such privacy provisioning, and justify the device being the provisioning point instead of the cloud alone. We present the point of view that, due to resource limitations of the IoT devices in general, the privacy preserving measures need to be hard-coded in the device technology. We fall short of suggesting …
Results And Challenges In Visualizing Analytic Provenance Of Text Analysis Tasks Using Interaction Logs, Rhema Linder, Alyssa M. Peña, Sampath Jayarathna, Eric D. Ragan
Results And Challenges In Visualizing Analytic Provenance Of Text Analysis Tasks Using Interaction Logs, Rhema Linder, Alyssa M. Peña, Sampath Jayarathna, Eric D. Ragan
Computer Science Faculty Publications
After data analysis, recalling and communicating the steps and rationale followed during the analysis can be difficult. This paper explores the use of interaction logs to generate summaries of an analyst's interest based on interactions with specific data items in a text analysis scenario. Our approach uses data-interaction events as a proxy for user interest in and experience of information. Logging can produce verbose logs that detail all available readable content, so the discussed approach uses topic modeling (LDA) over different time segments to summarize the verbose information and generate visualizations of the history of user interest. Our preliminary results …
Object Reuse And Exchange, Michael L. Nelson, Carl Lagoze, Herbert Van De Sompel, Pete Johnston, Robert Sanderson, Simeon Warner, Jürgen Sieck (Ed.), Michael A. Herzog (Ed.)
Object Reuse And Exchange, Michael L. Nelson, Carl Lagoze, Herbert Van De Sompel, Pete Johnston, Robert Sanderson, Simeon Warner, Jürgen Sieck (Ed.), Michael A. Herzog (Ed.)
Computer Science Faculty Publications
The Open Archives Object Reuse and Exchange (OAI-ORE) project defines standards for the description and exchange of aggregations of Web resources. The OAI-ORE abstract data model is conformant with the Architecture of the World Wide Web and leverages concepts from the Semantic Web, including RDF descriptions and Linked Data. In this paper we provide a brief review of a motivating example and its serialization in Atom.
Synchronization And Multiple Group Server Support For Kepler, K. Maly, M. Zubair, H. Siripuram, S. Zunjarwad, Yannis Manolopoulos (Ed.), Joaquim Filipe (Ed.), Panos Constantopoulos (Ed.), José Cordeiro (Ed.)
Synchronization And Multiple Group Server Support For Kepler, K. Maly, M. Zubair, H. Siripuram, S. Zunjarwad, Yannis Manolopoulos (Ed.), Joaquim Filipe (Ed.), Panos Constantopoulos (Ed.), José Cordeiro (Ed.)
Computer Science Faculty Publications
In the last decade literally thousands of digital libraries have emerged but one of the biggest obstacles for dissemination of information to a user community is that many digital libraries use different, proprietary technologies that inhibit interoperability. Kepler framework addresses interoperability and gives publication control to individual publishers. In Kepler, OAI-PMH is used to support "personal data providers" or "archivelets".". In our vision, individual publishers can be integrated with an institutional repository like Dspace by means of a Kepler Group Digital Library (GDL). The GDL aggregates metadata and full text from archivelets and can act as an OAI-compliant data provider …