Open Access. Powered by Scholars. Published by Universities.®
Artificial Intelligence and Robotics Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- Singapore Management University (176)
- Technological University Dublin (28)
- Old Dominion University (20)
- University of Arkansas, Fayetteville (20)
- University of Dayton (16)
-
- City University of New York (CUNY) (14)
- California Polytechnic State University, San Luis Obispo (11)
- San Jose State University (11)
- Dartmouth College (8)
- University of Malaya (8)
- Clemson University (7)
- Embry-Riddle Aeronautical University (7)
- University of Texas at Arlington (6)
- University of Nebraska - Lincoln (5)
- Rochester Institute of Technology (4)
- University of Kentucky (4)
- Central Washington University (3)
- Michigan Technological University (3)
- Montclair State University (3)
- University of Denver (3)
- University of New Mexico (3)
- California State University, San Bernardino (2)
- Dakota State University (2)
- Fort Hays State University (2)
- Georgia Southern University (2)
- Illinois Math and Science Academy (2)
- LSU New Orleans (2)
- Missouri State University (2)
- New Jersey Institute of Technology (2)
- Southern Adventist University (2)
- Keyword
-
- Artificial intelligence (18)
- Computer vision (17)
- Machine Learning (14)
- Machine learning (13)
- Deep learning (12)
-
- Artificial Intelligence (11)
- AI (8)
- Robotics (7)
- Virtual reality (7)
- Visualization (7)
- Augmented reality (6)
- Eye tracking (6)
- HCI (6)
- Image classification (6)
- Automation (5)
- Codes (5)
- Deep Learning (5)
- Feature extraction (5)
- Mental workload (5)
- Personality (5)
- Reinforcement learning (5)
- Accessibility (4)
- Artificial Intelligence (AI) (4)
- Classification (4)
- Computer Science (4)
- Computer Vision (4)
- Human Computer Interaction (4)
- Large Language Models (4)
- Pattern recognition (4)
- Semantics (4)
- Publication Year
- Publication
-
- Research Collection School Of Computing and Information Systems (171)
- H-Workload 2017: Models and Applications (Works in Progress) (14)
- Conference papers (12)
- MAICS: The Modern Artificial Intelligence and Cognitive Science Conference (12)
- Graduate Theses and Dissertations (10)
-
- Publications and Research (10)
- Computer Science Faculty Publications (9)
- Computer Science and Computer Engineering Undergraduate Honors Theses (9)
- Master's Theses (7)
- Dartmouth College Master’s Theses (6)
- Master's Projects (6)
- All Dissertations (5)
- Student Works (2020-2029) (5)
- College of Engineering Summer Undergraduate Research Program (4)
- Dissertations and Theses Collection (Open Access) (4)
- Frameless (4)
- SWITCH (4)
- Theses and Dissertations--Computer Science (4)
- All Master's Theses (3)
- Department of Computer Science Faculty Scholarship and Creative Works (3)
- Dissertations, Master's Theses and Master's Reports (3)
- International Journal of Aviation, Aeronautics, and Aerospace (3)
- Student Works (2000-2009) (3)
- Theses and Dissertations (3)
- All Theses (2)
- College of Graduate Studies: Theses & Dissertations (2)
- Computer Science ETDs (2)
- Computer Science Working Papers (2)
- Computer Science and Engineering Dissertations - Archive (2)
- Dartmouth College Ph.D Dissertations (2)
- Publication Type
- File Type
Articles 211 - 240 of 414
Full-Text Articles in Artificial Intelligence and Robotics
Skeleton-Based Hand Gesture Recognition Using Data-Level Fusion, Oluwaleke Yusuf
Skeleton-Based Hand Gesture Recognition Using Data-Level Fusion, Oluwaleke Yusuf
Theses and Dissertations
Hand Gesture Recognition (HGR) is a form of perceptual computing that allows artificial systems to capture and interpret human gestures. HGR has applications in human-machine interaction, virtual reality, augmented reality, and human behavior analysis. The human hand can assume a near-infinite number of poses and orientations to form myriad gestures, thus increasing the difficulty of the HGR task.
The hand skeleton of connected joints effectively describes the hand’s geometric shape and thus contains richer semantic gesture information while eliminating noise from individual differences in physical hand characteristics. The efficacy and computational efficiency of skeleton-based HGR frameworks can be significantly enhanced …
Reinforcement Learning Enhanced Pichunter For Interactive Search, Zhixin Ma, Jiaxin Wu, Weixiong Loo, Chong-Wah Ngo
Reinforcement Learning Enhanced Pichunter For Interactive Search, Zhixin Ma, Jiaxin Wu, Weixiong Loo, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
With the tremendous increase in video data size, search performance could be impacted significantly. Specifically, in an interactive system, a real-time system allows a user to browse, search and refine a query. Without a speedy system quickly, the main ingredient to engage a user to stay focused, an interactive system becomes less effective even with a sophisticated deep learning system. This paper addresses this challenge by leveraging approximate search, Bayesian inference, and reinforcement learning. For approximate search, we apply a hierarchical navigable small world, which is an efficient approximate nearest neighbor search algorithm. To quickly prune the search scope, we …
The Basil Technique: Bias Adaptive Statistical Inference Learning Agents For Learning From Human Feedback, Jonathan Indigo Watson
The Basil Technique: Bias Adaptive Statistical Inference Learning Agents For Learning From Human Feedback, Jonathan Indigo Watson
Theses and Dissertations--Computer Science
We introduce a novel approach for learning behaviors using human-provided feedback that is subject to systematic bias. Our method, known as BASIL, models the feedback signal as a combination of a heuristic evaluation of an action's utility and a probabilistically-drawn bias value, characterized by unknown parameters. We present both the general framework for our technique and specific algorithms for biases drawn from a normal distribution. We evaluate our approach across various environments and tasks, comparing it to interactive and non-interactive machine learning methods, including deep learning techniques, using human trainers and a synthetic oracle with feedback distorted to varying degrees. …
Digital Transformation, Applications, And Vulnerabilities In Maritime And Shipbuilding Ecosystems, Rafael Diaz, Katherine Smith
Digital Transformation, Applications, And Vulnerabilities In Maritime And Shipbuilding Ecosystems, Rafael Diaz, Katherine Smith
VMASC Publications
The evolution of maritime and shipbuilding supply chains toward digital ecosystems increases operational complexity and needs reliable communication and coordination. As labor and suppliers shift to digital platforms, interconnection, information transparency, and decentralized choices become ubiquitous. In this sense, Industry 4.0 enables "smart digitalization" in these environments. Many applications exist in two distinct but interrelated areas related to shipbuilding design and shipyard operational performance. New digital tools, such as virtual prototypes and augmented reality, begin to be used in the design phases, during the commissioning/quality control activities, and for training workers and crews. An application relates to using Virtual Prototypes …
Video Sign Language Recognition Using Pose Extraction And Deep Learning Models, Shayla Luong
Video Sign Language Recognition Using Pose Extraction And Deep Learning Models, Shayla Luong
Master's Projects
Sign language recognition (SLR) has long been a studied subject and research field within the Computer Vision domain. Appearance-based and pose-based approaches are two ways to tackle SLR tasks. Various models from traditional to current state-of-the-art including HOG-based features, Convolutional Neural Network, Recurrent Neural Network, Transformer, and Graph Convolutional Network have been utilized to tackle the area of SLR. While classifying alphabet letters in sign language has shown high accuracy rates, recognizing words presents its set of difficulties including the large vocabulary size, the subtleties in body motions and hand orientations, and regional dialects and variations. The emergence of deep …
Enhancing Facial Emotion Recognition Using Image Processing With Cnn, Sourabh Deokar
Enhancing Facial Emotion Recognition Using Image Processing With Cnn, Sourabh Deokar
Master's Projects
Facial expression recognition (FER) has been a challenging task in computer vision for decades. With recent advancements in deep learning, convolutional neural networks (CNNs) have shown promising results in this field. However, the accuracy of FER using CNNs heavily relies on the quality of the input images and the size of the dataset. Moreover, even in pictures of the same person with the same expression, brightness, backdrop, and stance might change. These variations are emphasized when comparing pictures of individuals with varying ethnic backgrounds and facial features, which makes it challenging for deep-learning models to classify. In this paper, we …
Directional Speaker Poster, Eugene Ng, Bryan Wong, Ruhaan Das
Directional Speaker Poster, Eugene Ng, Bryan Wong, Ruhaan Das
Student Works
Changi Airport is set to expand with a new terminal, Terminal 5. Currently, many of the airport's processes are manual, requiring a high dependence on staff. This proposal aims to incorporate automation and AI for a smoother passenger experience.
Transfer Of Personality Through Text Style, Michael O'Mahony, Robert Ross
Transfer Of Personality Through Text Style, Michael O'Mahony, Robert Ross
Other resources
The style of generated text is how something is said rather than what is said. We hypothesize that changing the style of generated text can change the perceived personality of the text generation agent. Dialogue systems that aim to imitate a human agent can appear to have a consistent personality through a consistent, controllable style of conversation. Some recent work on the style of generated text [1] performs impressively in the small number of domains selected for their experiments using transformer and LSTM-based models. Lin et al. [1] used weak supervised learning as their data set lacks parallel data. The …
Vr Computing Lab: An Immersive Classroom For Computing Learning, Shawn Pang, Kyong Jin Shim, Yi Meng Lau, Swapna Gottipati
Vr Computing Lab: An Immersive Classroom For Computing Learning, Shawn Pang, Kyong Jin Shim, Yi Meng Lau, Swapna Gottipati
Research Collection School Of Computing and Information Systems
In recent years, virtual reality (VR) is gaining popularity amongst educators and learners. If a picture is worth a thousand words, a VR session is worth a trillion words. VR technology completely immerses users with an experience that transports them into a simulated world. Universities across the United States, United Kingdom, and other countries have already started using VR for higher education in areas such as medicine, business, architecture, vocational training, social work, virtual field trips, virtual campuses, helping students with special needs, and many more. In this paper, we propose a novel VR platform learning framework which maps elements …
Prompting For Multimodal Hateful Meme Classification, Rui Cao, Roy Ka-Wei Lee, Wen-Haw Chong, Jing Jiang
Prompting For Multimodal Hateful Meme Classification, Rui Cao, Roy Ka-Wei Lee, Wen-Haw Chong, Jing Jiang
Research Collection School Of Computing and Information Systems
Hateful meme classification is a challenging multimodal task that requires complex reasoning and contextual background knowledge. Ideally, we could leverage an explicit external knowledge base to supplement contextual and cultural information in hateful memes. However, there is no known explicit external knowledge base that could provide such hate speech contextual information. To address this gap, we propose PromptHate, a simple yet effective prompt-based model that prompts pre-trained language models (PLMs) for hateful meme classification. Specifically, we construct simple prompts and provide a few in-context examples to exploit the implicit knowledge in the pretrained RoBERTa language model for hateful meme classification. …
Movie Reviews Sentiment Analysis Using Bert, Gibson Nkhata
Movie Reviews Sentiment Analysis Using Bert, Gibson Nkhata
Graduate Theses and Dissertations
Sentiment analysis (SA) or opinion mining is analysis of emotions and opinions from texts. It is one of the active research areas in Natural Language Processing (NLP). Various approaches have been deployed in the literature to address the problem. These techniques devise complex and sophisticated frameworks in order to attain optimal accuracy with their focus on polarity classification or binary classification. In this paper, we aim to fine-tune BERT in a simple but robust approach for movie reviews sentiment analysis to provide better accuracy than state-of-the-art (SOTA) methods. We start by conducting sentiment classification for every review, followed by computing …
Enabling The Human Perception Of A Working Camera In Web Conferences Via Its Movement, Anish Shrestha
Enabling The Human Perception Of A Working Camera In Web Conferences Via Its Movement, Anish Shrestha
LSU Master's Theses
In recent years, video conferencing has seen a significant increase in its usage due to the COVID-19 pandemic. When casting user’s video to other participants, the videoconference applications (e.g. Zoom, FaceTime, Skype, etc.) mainly leverage 1) webcam’s LED-light indicator, 2) user’s video feedback in the software and 3) the software’s video on/off icons to remind the user whether the camera is being used. However, these methods all impose the responsibility on the user itself to check the camera status, and there have been numerous cases reported when users expose their privacy inadvertently due to not realizing that their camera is …
Pixel-Wise Energy-Biased Abstention Learning For Anomaly Segmentation On Complex Urban Driving Scenes, Yu Tian, Yuyuan Liu, Guansong Pang, Fengbei Liu, Yuanhong Chen, Gustavo Carneiro
Pixel-Wise Energy-Biased Abstention Learning For Anomaly Segmentation On Complex Urban Driving Scenes, Yu Tian, Yuyuan Liu, Guansong Pang, Fengbei Liu, Yuanhong Chen, Gustavo Carneiro
Research Collection School Of Computing and Information Systems
State-of-the-art (SOTA) anomaly segmentation approaches on complex urban driving scenes explore pixel-wise classification uncertainty learned from outlier exposure, or external reconstruction models. However, previous uncertainty approaches that directly associate high uncertainty to anomaly may sometimes lead to incorrect anomaly predictions, and external reconstruction models tend to be too inefficient for real-time self-driving embedded systems. In this paper, we propose a new anomaly segmentation method, named pixel-wise energy-biased abstention learning (PEBAL), that explores pixel-wise abstention learning (AL) with a model that learns an adaptive pixel-level anomaly class, and an energy-based model (EBM) that learns inlier pixel distribution. More specifically, PEBAL is …
Dualformer: Local-Global Stratified Transformer For Efficient Video Recognition, Yuxuan Liang, Pan Zhou, Roger Zimmermann, Shuicheng Yan
Dualformer: Local-Global Stratified Transformer For Efficient Video Recognition, Yuxuan Liang, Pan Zhou, Roger Zimmermann, Shuicheng Yan
Research Collection School Of Computing and Information Systems
While transformers have shown great potential on video recognition with their strong capability of capturing long-range dependencies, they often suffer high computational costs induced by the self-attention to the huge number of 3D tokens. In this paper, we present a new transformer architecture termed DualFormer, which can efficiently perform space-time attention for video recognition. Concretely, DualFormer stratifies the full space-time attention into dual cascaded levels, i.e., to first learn fine-grained local interactions among nearby 3D tokens, and then to capture coarse-grained global dependencies between the query token and global pyramid contexts. Different from existing methods that apply space-time factorization or …
Interactive Video Corpus Moment Retrieval Using Reinforcement Learning, Zhixin Ma, Chong-Wah Ngo
Interactive Video Corpus Moment Retrieval Using Reinforcement Learning, Zhixin Ma, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
Known-item video search is effective with human-in-the-loop to interactively investigate the search result and refine the initial query. Nevertheless, when the first few pages of results are swamped with visually similar items, or the search target is hidden deep in the ranked list, finding the know-item target usually requires a long duration of browsing and result inspection. This paper tackles the problem by reinforcement learning, aiming to reach a search target within a few rounds of interaction by long-term learning from user feedbacks. Specifically, the system interactively plans for navigation path based on feedback and recommends a potential target that …
Long-Term Leap Attention, Short-Term Periodic Shift For Video Classification, Hao Zhang, Lechao Cheng, Yanbin Hao, Chong-Wah Ngo
Long-Term Leap Attention, Short-Term Periodic Shift For Video Classification, Hao Zhang, Lechao Cheng, Yanbin Hao, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
Video transformer naturally incurs a heavier computation burden than a static vision transformer, as the former processes �� times longer sequence than the latter under the current attention of quadratic complexity (�� 2�� 2 ). The existing works treat the temporal axis as a simple extension of spatial axes, focusing on shortening the spatio-temporal sequence by either generic pooling or local windowing without utilizing temporal redundancy. However, videos naturally contain redundant information between neighboring frames; thereby, we could potentially suppress attention on visually similar frames in a dilated manner. Based on this hypothesis, we propose the LAPS, a long-term “Leap …
Wave-Vit: Unifying Wavelet And Transformers For Visual Representation Learning, Ting Yao, Yingwei Pan, Yehao Li, Chong-Wah Ngo, Tao Mei
Wave-Vit: Unifying Wavelet And Transformers For Visual Representation Learning, Ting Yao, Yingwei Pan, Yehao Li, Chong-Wah Ngo, Tao Mei
Research Collection School Of Computing and Information Systems
Multi-scale Vision Transformer (ViT) has emerged as a powerful backbone for computer vision tasks, while the self-attention computation in Transformer scales quadratically w.r.t. the input patch number. Thus, existing solutions commonly employ down-sampling operations (e.g., average pooling) over keys/values to dramatically reduce the computational cost. In this work, we argue that such over-aggressive down-sampling design is not invertible and inevitably causes information dropping especially for high-frequency components in objects (e.g., texture details). Motivated by the wavelet theory, we construct a new Wavelet Vision Transformer (Wave-ViT) that formulates the invertible down-sampling with wavelet transforms and self-attention learning in a unified way. …
Dynamic Temporal Filtering In Video Models, Fuchen Long, Zhaofan Qiu, Yingwei Pan, Ting Yao, Chong-Wah Ngo, Tao Mei
Dynamic Temporal Filtering In Video Models, Fuchen Long, Zhaofan Qiu, Yingwei Pan, Ting Yao, Chong-Wah Ngo, Tao Mei
Research Collection School Of Computing and Information Systems
Video temporal dynamics is conventionally modeled with 3D spatial-temporal kernel or its factorized version comprised of 2D spatial kernel and 1D temporal kernel. The modeling power, nevertheless, is limited by the fixed window size and static weights of a kernel along the temporal dimension. The pre-determined kernel size severely limits the temporal receptive fields and the fixed weights treat each spatial location across frames equally, resulting in sub-optimal solution for longrange temporal modeling in natural scenes. In this paper, we present a new recipe of temporal feature learning, namely Dynamic Temporal Filter (DTF), that novelly performs spatial-aware temporal modeling in …
Cvfnet: Real-Time 3d Object Detection By Learning Cross View Features, Jiaqi Gu, Zhiyu Xiang, Pan Zhao, Tingming Bai, Lingxuan Wang, Xijun Zhao, Zhiyuan Zhang
Cvfnet: Real-Time 3d Object Detection By Learning Cross View Features, Jiaqi Gu, Zhiyu Xiang, Pan Zhao, Tingming Bai, Lingxuan Wang, Xijun Zhao, Zhiyuan Zhang
Research Collection School Of Computing and Information Systems
In recent years 3D object detection from LiDAR point clouds has made great progress thanks to the development of deep learning technologies. Although voxel or point based methods are popular in 3D object detection, they usually involve time-consuming operations such as 3D convolutions on voxels or ball query among points, making the resulting network inappropriate for time critical applications. On the other hand, 2D view-based methods feature high computing efficiency while usually obtaining inferior performance than the voxel or point based methods. In this work, we present a real-time view-based single stage 3D object detector, namely CVFNet to fulfill this …
Fair, Equitable, And Just: A Socio-Technical Approach To Online Safety, Daricia Wilkinson
Fair, Equitable, And Just: A Socio-Technical Approach To Online Safety, Daricia Wilkinson
All Dissertations
Socio-technical systems have been revolutionary in reshaping how people maintain relationships, learn about new opportunities, engage in meaningful discourse, and even express grief and frustrations. At the same time, these systems have been central in the proliferation of harmful behaviors online as internet users are confronted with serious and pervasive threats at alarming rates. Although researchers and companies have attempted to develop tools to mitigate threats, the perception of dominant (often Western) frameworks as the standard for the implementation of safety mechanisms fails to account for imbalances, inequalities, and injustices in non-Western civilizations like the Caribbean. Therefore, in this dissertation …
Towards Improving System Performance In Large Scale Multi-Agent Systems With Selfish Agents, Rajiv Ranjan Kumar
Towards Improving System Performance In Large Scale Multi-Agent Systems With Selfish Agents, Rajiv Ranjan Kumar
Dissertations and Theses Collection (Open Access)
Intelligent agents are becoming increasingly prevalent in a wide variety of domains including but not limited to transportation, safety and security. To better utilize the intelligence, there has been increasing focus on frameworks and methods for coordinating these intelligent agents. This thesis is specifically targeted at providing solution approaches for improving large scale multi-agent systems with selfish intelligent agents. In such systems, the performance of an agent depends on not just his/her own efforts, but also on other agent’s decisions. The complexity of interactions among multiple agents, coupled with the large scale nature of the problem domains and the uncertainties …
Analysis Of Digital Image Segmentation Algorithms, Khalilov Sirojiddin
Analysis Of Digital Image Segmentation Algorithms, Khalilov Sirojiddin
Karakalpak Scientific Journal
Ushbu maqolada zamonaviy axborot-kommunikatsiya texnologiyalaridan foydalanishni kengaytirish maqsadida raqamli tasvirni qayta ishlash usullari va algoritmlari tahlil qilinadi. Maqolada, shuningdek, raqamli tasvirni qayta ishlash, tasvirni segmentatsiyalash usullari, WaterShed, MeanShift, FloodFill, GrabCut algoritmlarining afzalliklari va kamchiliklari o'rganiladi.
Cross-Lingual Adaptation For Recipe Retrieval With Mixup, Bin Zhu, Chong-Wah Ngo, Jingjing Chen, Wing-Kwong Chan
Cross-Lingual Adaptation For Recipe Retrieval With Mixup, Bin Zhu, Chong-Wah Ngo, Jingjing Chen, Wing-Kwong Chan
Research Collection School Of Computing and Information Systems
Cross-modal recipe retrieval has attracted research attention in recent years, thanks to the availability of large-scale paired data for training. Nevertheless, obtaining adequate recipe-image pairs covering the majority of cuisines for supervised learning is difficult if not impossible. By transferring knowledge learnt from a data-rich cuisine to a data-scarce cuisine, domain adaptation sheds light on this practical problem. Nevertheless, existing works assume recipes in source and target domains are mostly originated from the same cuisine and written in the same language. This paper studies unsupervised domain adaptation for image-to-recipe retrieval, where recipes in source and target domains are in different …
Reinforcement Learning-Based Interactive Video Search, Zhixin Ma, Jiaxin Wu, Zhijian Hou, Chong-Wah Ngo
Reinforcement Learning-Based Interactive Video Search, Zhixin Ma, Jiaxin Wu, Zhijian Hou, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
Despite the rapid progress in text-to-video search due to the advancement of cross-modal representation learning, the existing techniques still fall short in helping users to rapidly identify the search targets. Particularly, in the situation that a system suggests a long list of similar candidates, the user needs to painstakingly inspect every search result. The experience is frustrated with repeated watching of similar clips, and more frustratingly, the search targets may be overlooked due to mental tiredness. This paper explores reinforcement learning-based (RL) searching to relieve the user from the burden of brute force inspection. Specifically, the system maintains a graph …
Group Contextualization For Video Recognition, Yanbin Hao, Hao Zhang, Chong-Wah Ngo, Xiangnan He
Group Contextualization For Video Recognition, Yanbin Hao, Hao Zhang, Chong-Wah Ngo, Xiangnan He
Research Collection School Of Computing and Information Systems
Learning discriminative representation from the complex spatio-temporal dynamic space is essential for video recognition. On top of those stylized spatio-temporal computational units, further refining the learnt feature with axial contexts is demonstrated to be promising in achieving this goal. However, previous works generally focus on utilizing a single kind of contexts to calibrate entire feature channels and could hardly apply to deal with diverse video activities. The problem can be tackled by using pair-wise spatio-temporal attentions to recompute feature response with cross-axis contexts at the expense of heavy computations. In this paper, we propose an efficient feature refinement method that …
Mlp-3d: A Mlp-Like 3d Architecture With Grouped Time Mixing, Zhaofan Qiu, Ting Yao, Chong-Wah Ngo, Tao Mei
Mlp-3d: A Mlp-Like 3d Architecture With Grouped Time Mixing, Zhaofan Qiu, Ting Yao, Chong-Wah Ngo, Tao Mei
Research Collection School Of Computing and Information Systems
Convolutional Neural Networks (CNNs) have been re-garded as the go-to models for visual recognition. More re-cently, convolution-free networks, based on multi-head self-attention (MSA) or multi-layer perceptrons (MLPs), become more and more popular. Nevertheless, it is not trivial when utilizing these newly-minted networks for video recognition due to the large variations and complexities in video data. In this paper, we present MLP-3D networks, a novel MLP-like 3D architecture for video recognition. Specifically, the architecture consists of MLP-3D blocks, where each block contains one MLP applied across tokens (i.e., token-mixing MLP) and one MLP applied independently to each token (i.e., channel MLP). …
Multimodal Zero-Shot Hateful Meme Detection, Jiawen Zhu, Roy Ka-Wei Lee, Wen Haw Chong
Multimodal Zero-Shot Hateful Meme Detection, Jiawen Zhu, Roy Ka-Wei Lee, Wen Haw Chong
Research Collection School Of Computing and Information Systems
Facebook has recently launched the hateful meme detection challenge, which garnered much attention in academic and industry research communities. Researchers have proposed multimodal deep learning classification methods to perform hateful meme detection. While the proposed methods have yielded promising results, these classification methods are mostly supervised and heavily rely on labeled data that are not always available in the real-world setting. Therefore, this paper explores and aims to perform hateful meme detection in a zero-shot setting. Working towards this goal, we propose Target-Aware Multimodal Enhancement (TAME), which is a novel deep generative framework that can improve existing hateful meme classification …
High-Resolution Face Swapping Via Latent Semantics Disentanglement, Yangyang Xu, Bailin Deng, Junle Wang, Yanqing Jing, Jia Pan, Shengfeng He
High-Resolution Face Swapping Via Latent Semantics Disentanglement, Yangyang Xu, Bailin Deng, Junle Wang, Yanqing Jing, Jia Pan, Shengfeng He
Research Collection School Of Computing and Information Systems
We present a novel high-resolution face swapping method using the inherent prior knowledge of a pre-trained GAN model. Although previous research can leverage generative priors to produce high-resolution results, their quality can suffer from the entangled semantics of the latent space. We explicitly disentangle the latent semantics by utilizing the progressive nature of the generator, deriving structure at-tributes from the shallow layers and appearance attributes from the deeper ones. Identity and pose information within the structure attributes are further separated by introducing a landmark-driven structure transfer latent direction. The disentangled latent code produces rich generative features that incorporate feature blending …
Gauging The State-Of-The-Art For Foresight Weight Pruning On Neural Networks, Noah James
Gauging The State-Of-The-Art For Foresight Weight Pruning On Neural Networks, Noah James
Computer Science and Computer Engineering Undergraduate Honors Theses
The state-of-the-art for pruning neural networks is ambiguous due to poor experimental practices in the field. Newly developed approaches rarely compare to each other, and when they do, their comparisons are lackluster or contain errors. In the interest of stabilizing the field of pruning, this paper initiates a dive into reproducing prominent pruning algorithms across several architectures and datasets. As a first step towards this goal, this paper shows results for foresight weight pruning across 6 baseline pruning strategies, 5 modern pruning strategies, random pruning, and one legacy method (Optimal Brain Damage). All strategies are evaluated on 3 different architectures …
Analysis Of Gpu Memory Vulnerabilities, Jarrett Hoover
Analysis Of Gpu Memory Vulnerabilities, Jarrett Hoover
Computer Science and Computer Engineering Undergraduate Honors Theses
Graphics processing units (GPUs) have become a widely used technology for various purposes. While their intended use is accelerating graphics rendering, their parallel computing capabilities have expanded their use into other areas. They are used in computer gaming, deep learning for artificial intelligence and mining cryptocurrencies. Their rise in popularity led to research involving several security aspects, including this paper’s focus, memory vulnerabilities. Research documented many vulnerabilities, including GPUs not implementing address space layout randomization, not zeroing out memory after deallocation, and not initializing newly allocated memory. These vulnerabilities can lead to a victim’s sensitive data being leaked to an …