Open Access. Powered by Scholars. Published by Universities.®
Graphics and Human Computer Interfaces Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Artificial Intelligence and Robotics (5)
- Databases and Information Systems (5)
- Numerical Analysis and Scientific Computing (2)
- Other Computer Sciences (2)
- Software Engineering (2)
-
- Theory and Algorithms (2)
- Analytical, Diagnostic and Therapeutic Techniques and Equipment (1)
- Anatomy (1)
- Applied Behavior Analysis (1)
- Applied Mathematics (1)
- Archival Science (1)
- Business (1)
- Business Administration, Management, and Operations (1)
- Business Intelligence (1)
- Cataloging and Metadata (1)
- Chemistry (1)
- Cognition and Perception (1)
- Cognitive Psychology (1)
- Collection Development and Management (1)
- Communication (1)
- Communication Sciences and Disorders (1)
- Communication Technology and New Media (1)
- Community-Based Learning (1)
- Community-Based Research (1)
- Comparative Psychology (1)
- Computational Chemistry (1)
- Computational Engineering (1)
- Institution
- Publication
- Publication Type
Articles 1 - 16 of 16
Full-Text Articles in Graphics and Human Computer Interfaces
Streamlined Biomedical Image Processing Pipelines, Jiehyun Kim
Streamlined Biomedical Image Processing Pipelines, Jiehyun Kim
Graduate Doctoral Dissertations
This dissertation focuses on advancing carotid artery analysis through a series of visualizations and deep learning tools for calcified plaque assessment and related biomedical imaging tasks. Accurate plaque evaluation is essential, but current workflows depend on slow, clinician-dependent manual review. To address these limitations, this work introduces the CACTAS framework, a set of tools and methods that enable fast and reliable plaque segmentation for clinicians.
The first study, the CACTAS-Tool, provides a web-based labeling tool that enables clinicians to label plaque directly in three dimensions through a streamlined one-click interface. This tool significantly reduces the effort required to generate high-quality …
Cssa-Fusion: Channel Selective And Spatial Alignment Infrared-Visible Image Fusion, Zhen Li, Zhi Zeng, Zhongrui Xiao, Ming Wen, Zhiyuan Zhang, Yibin Tian
Cssa-Fusion: Channel Selective And Spatial Alignment Infrared-Visible Image Fusion, Zhen Li, Zhi Zeng, Zhongrui Xiao, Ming Wen, Zhiyuan Zhang, Yibin Tian
Research Collection School Of Computing and Information Systems
Infrared-visible image fusion aims to integrate complementary information from two modalities to generate images with enriched semantic content. However, existing methods often neglect two critical aspects: the design of a local–global feature enhancement architecture and spatial alignment. To address these challenges, we propose Channel Selective and Spatial Alignment Fusion (CSSA-Fusion), a novel framework composed of two synergistic modules. The first is a selective channel and redundancy suppression module, which introduces a dual-branch selective channel attention mechanism to jointly capture local saliency and global channel importance for enhanced feature representation, and an informativeness–redundancy separation strategy to suppress redundant information while preserving …
Few-Shot Learning On Graphs: From Meta-Learning To Llm-Empowered Pre-Training And Beyond, Yuan Fang, Yuxia Wu, Xingtong Yu, Shirui Pan
Few-Shot Learning On Graphs: From Meta-Learning To Llm-Empowered Pre-Training And Beyond, Yuan Fang, Yuxia Wu, Xingtong Yu, Shirui Pan
Research Collection School Of Computing and Information Systems
Graph representation learning has become central to many graph-based tasks, driving advancements in various domains such as web search, recommendation systems, and social network analysis. Traditionally, these methods rely on end-to-end supervised learning paradigms that require abundant labeled data, which can be costly and difficult to obtain. To address this limitation, few-shot learning on graphs has emerged as a promising approach, allowing models to generalize with minimal supervision and overcome data scarcity in real-world applications. This tutorial offers an in-depth exploration of recent advancements in few-shot learning for graphs, providing a comparative analysis of state-of-the-art methods and identifying future research …
Real-Time Feedback-Driven Framework For Automated Cybersickness Mitigation, Md Jahirul Islam
Real-Time Feedback-Driven Framework For Automated Cybersickness Mitigation, Md Jahirul Islam
Master's Theses
As technologies are becoming more advanced day by day, the embracement of virtual reality (VR) technology among users is also increasing in daily activities for various purposes, and subsequently, the barrier between the real and virtual world is fading. Despite the versatile uses, cybersickness (CS) is a major problem which is induced among users due to the immersive VR experience. There is a plethora of research findings and methods to measure the users’ CS such as virtual reality sickness questionnaire (VRSQ), simulator sickness questionnaire (SSQ), fast motion scale questionnaire (FMS), and others. Recently, machine learning approaches have also been adopted …
Lecture-Style Tutorial: Towards Graph Foundation Models, Chuan Shi, Cheng Yang, Yuan Fang, Lichao Sun, Philip Yu
Lecture-Style Tutorial: Towards Graph Foundation Models, Chuan Shi, Cheng Yang, Yuan Fang, Lichao Sun, Philip Yu
Research Collection School Of Computing and Information Systems
Emerging as fundamental building blocks for diverse artificial intelligence applications, foundation models have achieved notable success across natural language processing and many other domains. Concurrently, graph machine learning has gradually evolved from shallow methods to deep models to leverage the abundant graph-structured data that constitute an important pillar in the data ecosystem for artificial intelligence. Naturally, the emergence and homogenization capabilities of foundation models have piqued the interest of graph machine learning researchers. This has sparked discussions about developing a next-generation graph learning paradigm, one that is pre-trained on broad graph data and can be adapted to a wide range …
Metaformer Baselines For Vision, Weihao Yu, Chenyang Si, Pan Zhou, Mi Luo, Yichen Zhou, Jiashi Feng, Shuicheng Yan, Xinchao Wang
Metaformer Baselines For Vision, Weihao Yu, Chenyang Si, Pan Zhou, Mi Luo, Yichen Zhou, Jiashi Feng, Shuicheng Yan, Xinchao Wang
Research Collection School Of Computing and Information Systems
Abstract—MetaFormer, the abstracted architecture of Transformer, has been found to play a significant role in achieving competitive performance. In this paper, we further explore the capacity of MetaFormer, again, by migrating our focus away from the token mixer design: we introduce several baseline models under MetaFormer using the most basic or common mixers, and demonstrate their gratifying performance. We summarize our observations as follows: (1) MetaFormer ensures solid lower bound of performance. By merely adopting identity mapping as the token mixer, the MetaFormer model, termed IdentityFormer, achieves >80% accuracy on ImageNet-1K. (2) MetaFormer works well with arbitrary token mixers. When …
Chatgpt As Metamorphosis Designer For The Future Of Artificial Intelligence (Ai): A Conceptual Investigation, Amarjit Kumar Singh (Library Assistant), Dr. Pankaj Mathur (Deputy Librarian)
Chatgpt As Metamorphosis Designer For The Future Of Artificial Intelligence (Ai): A Conceptual Investigation, Amarjit Kumar Singh (Library Assistant), Dr. Pankaj Mathur (Deputy Librarian)
Library Philosophy and Practice (e-journal)
Abstract
Purpose: The purpose of this research paper is to explore ChatGPT’s potential as an innovative designer tool for the future development of artificial intelligence. Specifically, this conceptual investigation aims to analyze ChatGPT’s capabilities as a tool for designing and developing near about human intelligent systems for futuristic used and developed in the field of Artificial Intelligence (AI). Also with the helps of this paper, researchers are analyzed the strengths and weaknesses of ChatGPT as a tool, and identify possible areas for improvement in its development and implementation. This investigation focused on the various features and functions of ChatGPT that …
Enhancing Facial Emotion Recognition Using Image Processing With Cnn, Sourabh Deokar
Enhancing Facial Emotion Recognition Using Image Processing With Cnn, Sourabh Deokar
Master's Projects
Facial expression recognition (FER) has been a challenging task in computer vision for decades. With recent advancements in deep learning, convolutional neural networks (CNNs) have shown promising results in this field. However, the accuracy of FER using CNNs heavily relies on the quality of the input images and the size of the dataset. Moreover, even in pictures of the same person with the same expression, brightness, backdrop, and stance might change. These variations are emphasized when comparing pictures of individuals with varying ethnic backgrounds and facial features, which makes it challenging for deep-learning models to classify. In this paper, we …
Cross-Modal Food Retrieval: Learning A Joint Embedding Of Food Images And Recipes With Semantic Consistency And Attention Mechanism, Hao Wang, Doyen Sahoo, Chenghao Liu, Ke Shu, Palakorn Achananuparp, Ee-Peng Lim, Steven C. H. Hoi
Cross-Modal Food Retrieval: Learning A Joint Embedding Of Food Images And Recipes With Semantic Consistency And Attention Mechanism, Hao Wang, Doyen Sahoo, Chenghao Liu, Ke Shu, Palakorn Achananuparp, Ee-Peng Lim, Steven C. H. Hoi
Research Collection School Of Computing and Information Systems
Food retrieval is an important task to perform analysis of food-related information, where we are interested in retrieving relevant information about the queried food item such as ingredients, cooking instructions, etc. In this paper, we investigate cross-modal retrieval between food images and cooking recipes. The goal is to learn an embedding of images and recipes in a common feature space, such that the corresponding image-recipe embeddings lie close to one another. Two major challenges in addressing this problem are 1) large intra-variance and small inter-variance across cross-modal food data; and 2) difficulties in obtaining discriminative recipe representations. To address these …
A Large-Scale Benchmark For Food Image Segmentation, Xiongwei Wu, Xin Fu, Ying Liu, Ee-Peng Lim, Steven C. H. Hoi, Qianru Sun
A Large-Scale Benchmark For Food Image Segmentation, Xiongwei Wu, Xin Fu, Ying Liu, Ee-Peng Lim, Steven C. H. Hoi, Qianru Sun
Research Collection School Of Computing and Information Systems
Food image segmentation is a critical and indispensible task for developing health-related applications such as estimating food calories and nutrients. Existing food image segmentation models are underperforming due to two reasons: (1) there is a lack of high quality food image datasets with fine-grained ingredient labels and pixel-wise location masks—the existing datasets either carry coarse ingredient labels or are small in size; and (2) the complex appearance of food makes it difficult to localize and recognize ingredients in food images, e.g., the ingredients may overlap one another in the same image, and the identical ingredient may appear distinctly in different …
Cross-Modal Food Retrieval: Learning A Joint Embedding Of Food Images And Recipes With Semantic Consistency And Attention Mechanism;, Hao Wang, Doyen Sahoo, Chenghao Liu, Ke Shu, Achananuparp Palakorn, Ee Peng Lim, Steven Hoi
Cross-Modal Food Retrieval: Learning A Joint Embedding Of Food Images And Recipes With Semantic Consistency And Attention Mechanism;, Hao Wang, Doyen Sahoo, Chenghao Liu, Ke Shu, Achananuparp Palakorn, Ee Peng Lim, Steven Hoi
Research Collection School Of Computing and Information Systems
Food retrieval is an important task to perform analysis of food-related information, where we are interested in retrieving relevant information about the queried food item such as ingredients, cooking instructions, etc. In this paper, we investigate cross-modal retrieval between food images and cooking recipes. The goal is to learn an embedding of images and recipes in a common feature space, such that the corresponding image-recipe embeddings lie close to one another. Two major challenges in addressing this problem are 1) large intra-variance and small inter-variance across cross-modal food data; and 2) difficulties in obtaining discriminative recipe representations. To address these …
Deepdrawing: A Deep Learning Approach To Graph Drawing, Yong Wang, Zhihua Jin, Qianwen Wang, Weiwei Cui, Tengfei Ma, Huamin Qu
Deepdrawing: A Deep Learning Approach To Graph Drawing, Yong Wang, Zhihua Jin, Qianwen Wang, Weiwei Cui, Tengfei Ma, Huamin Qu
Research Collection School Of Computing and Information Systems
Node-link diagrams are widely used to facilitate network explorations. However, when using a graph drawing technique to visualize networks, users often need to tune different algorithm-specific parameters iteratively by comparing the corresponding drawing results in order to achieve a desired visual effect. This trial and error process is often tedious and time-consuming, especially for non-expert users. Inspired by the powerful data modelling and prediction capabilities of deep learning techniques, we explore the possibility of applying deep learning techniques to graph drawing. Specifically, we propose using a graph-LSTM-based approach to directly map network structures to graph drawings. Given a set of …
Improved Generalisation Bounds For Deep Learning Through L∞ Covering Numbers, Antoine Ledent, Yunwen Lei, Marius Kloft
Improved Generalisation Bounds For Deep Learning Through L∞ Covering Numbers, Antoine Ledent, Yunwen Lei, Marius Kloft
Research Collection School Of Computing and Information Systems
Using proof techniques involving L∞ covering numbers, we show generalisation error bounds for deep learning with two main improvements over the state of the art. First, our bounds have no explicit dependence on the number of classes except for logarithmic factors. This holds even when formulating the bounds in terms of the L 2 norm of the weight matrices, while previous bounds exhibit at least a square-root dependence on the number of classes in this case. Second, we adapt the Rademacher analysis of DNNs to incorporate weight sharing—a task of fundamental theoretical importance which was previously attempted only under very …
Sliced Wasserstein Generative Models, Jiqing Wu, Zhiwu Huang, Dinesh Acharya, Wen Li, Janine Thoma, Danda Pani Paudel, Luc Van Gool
Sliced Wasserstein Generative Models, Jiqing Wu, Zhiwu Huang, Dinesh Acharya, Wen Li, Janine Thoma, Danda Pani Paudel, Luc Van Gool
Research Collection School Of Computing and Information Systems
In generative modeling, the Wasserstein distance (WD) has emerged as a useful metric to measure the discrepancy between generated and real data distributions. Unfortunately, it is challenging to approximate the WD of high-dimensional distributions. In contrast, the sliced Wasserstein distance (SWD) factorizes high-dimensional distributions into their multiple one-dimensional marginal distributions and is thus easier to approximate. In this paper, we introduce novel approximations of the primal and dual SWD. Instead of using a large number of random projections, as it is done by conventional SWD approximation methods, we propose to approximate SWDs with a small number of parameterized orthogonal projections …
Improving Asynchronous Advantage Actor Critic With A More Intelligent Exploration Strategy, James B. Holliday
Improving Asynchronous Advantage Actor Critic With A More Intelligent Exploration Strategy, James B. Holliday
Graduate Theses and Dissertations
We propose a simple and efficient modification to the Asynchronous Advantage Actor Critic (A3C)
algorithm that improves training. In 2016 Google’s DeepMind set a new standard for state-of-theart
reinforcement learning performance with the introduction of the A3C algorithm. The goal of
this research is to show that A3C can be improved by the use of a new novel exploration strategy we
call “Follow then Forage Exploration” (FFE). FFE forces the agents to follow the best known path
at the beginning of a training episode and then later in the episode the agent is forced to “forage”
and explores randomly. In …
A Continuous Space Generative Model, Erzen Komoni
A Continuous Space Generative Model, Erzen Komoni
Graduate Theses and Dissertations
Generative models are a class of machine learning models capable of producing digital images with plausibly realistic properties. They are useful in such applications as visualizing designs, rendering game scenes, and improving images at higher magnifications. Unfortunately, existing generative models generate only images with a discrete predetermined resolution. This paper presents the Continuous Space Generative Model (CSGM), a novel generative model capable of generating images as a continuous function, rather than as a discrete set of pixel values. Like generative adversarial networks, CSGM trains by alternating between generative and discriminative steps. But unlike generative adversarial networks, CSGM uses only one …