Proceedings Of The 13th International Workshop On Graph Computation Models (Gcm 2022),
2022
Singapore Management University
Proceedings Of The 13th International Workshop On Graph Computation Models (Gcm 2022), Reiko Heckel, Christopher M. Poskitt
Research Collection School Of Computing and Information Systems
This volume contains the proceedings of the Thirteenth International Workshop on Graph Computation Models (GCM 2022) , which was held in Nantes, France on 6th July 2022 as part of the STAF federation of conferences. Graphs are common mathematical structures that are visual and intuitive. They constitute a natural and seamless way for system modelling in science, engineering, and beyond, including computer science, biology, and business process modelling. Graph computation models constitute a class of very high-level models where graphs are first-class citizens. The aim of the International GCM Workshop series is to bring together researchers interested in all aspects …
Analysis Of Digital Image Segmentation Algorithms,
2022
Nurafshan branch of Tashkent University of Information Technologies named after Muhammad al-Khwarizmi
Analysis Of Digital Image Segmentation Algorithms, Khalilov Sirojiddin
Karakalpak Scientific Journal
Ushbu maqolada zamonaviy axborot-kommunikatsiya texnologiyalaridan foydalanishni kengaytirish maqsadida raqamli tasvirni qayta ishlash usullari va algoritmlari tahlil qilinadi. Maqolada, shuningdek, raqamli tasvirni qayta ishlash, tasvirni segmentatsiyalash usullari, WaterShed, MeanShift, FloodFill, GrabCut algoritmlarining afzalliklari va kamchiliklari o'rganiladi.
Spotlight Report #6: Proffering Machine-Readable Personal Privacy Research Agreements: Pilot Project Findings For Ieee P7012 Wg,
2022
Internet Safety Labs
Spotlight Report #6: Proffering Machine-Readable Personal Privacy Research Agreements: Pilot Project Findings For Ieee P7012 Wg, Noreen Y. Whysel, Lisa Levasseur
Publications and Research
What if people had the ability to assert their own legally binding permissions for data collection, use, sharing, and retention by the technologies they use? The IEEE P7012 has been working on an interoperability specification for machine-readable personal privacy terms to support this ability since 2018. The premise behind the work of IEEE P7012 is that people need technology that works on their behalf—i.e. software agents that assert the individual’s permissions and preferences in a machine-readable format.
Thanks to a grant from the IEEE Technical Activities Board Committee on Standards (TAB CoS), we were able to explore the attitudes of …
Designing Narrative-Based Interfaces For Collective Action: A Case Study Using Amazon, Climate Change, And Consumer Behavior,
2022
Dartmouth College
Designing Narrative-Based Interfaces For Collective Action: A Case Study Using Amazon, Climate Change, And Consumer Behavior, Catherine Parnell
Dartmouth College Undergraduate Theses
Climate change is the most pressing issue facing future generations. Amongst expanses of the population there is a lack of collective action on environmental issues, as there is a large gap between awareness and behavior change. This study suggests persuasive design that utilizes a narrative framing as a solution to reduce barriers to engaging in issues of collective action. Through extensive need-finding studies to understand target users, this thesis uses online-shopping via Amazon as a context for arguing that narrative can support actionable change in behavior. The technical artifact resulting from this research is a developed chrome extension and web …
Metaformer Is Actually What You Need For Vision,
2022
Singapore Management University
Metaformer Is Actually What You Need For Vision, Weihao Yu, Mi Luo, Pan Zhou, Chenyang Si, Yichen Zhou, Xinchao Wang, Jiashi Feng, Shuicheng Yan
Research Collection School Of Computing and Information Systems
Transformers have shown great potential in computer vision tasks. A common belief is their attention-based token mixer module contributes most to their competence. However, recent works show the attention-based module in transformers can be replaced by spatial MLPs and the resulted models still perform quite well. Based on this observation, we hypothesize that the general architecture of the transformers, instead of the specific token mixer module, is more essential to the model's performance. To verify this, we deliberately replace the attention module in transformers with an embarrassingly simple spatial pooling operator to conduct only basic token mixing. Surprisingly, we observe …
Cross-Lingual Adaptation For Recipe Retrieval With Mixup,
2022
Singapore Management University
Cross-Lingual Adaptation For Recipe Retrieval With Mixup, Bin Zhu, Chong-Wah Ngo, Jingjing Chen, Wing-Kwong Chan
Research Collection School Of Computing and Information Systems
Cross-modal recipe retrieval has attracted research attention in recent years, thanks to the availability of large-scale paired data for training. Nevertheless, obtaining adequate recipe-image pairs covering the majority of cuisines for supervised learning is difficult if not impossible. By transferring knowledge learnt from a data-rich cuisine to a data-scarce cuisine, domain adaptation sheds light on this practical problem. Nevertheless, existing works assume recipes in source and target domains are mostly originated from the same cuisine and written in the same language. This paper studies unsupervised domain adaptation for image-to-recipe retrieval, where recipes in source and target domains are in different …
Reinforcement Learning-Based Interactive Video Search,
2022
Singapore Management University
Reinforcement Learning-Based Interactive Video Search, Zhixin Ma, Jiaxin Wu, Zhijian Hou, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
Despite the rapid progress in text-to-video search due to the advancement of cross-modal representation learning, the existing techniques still fall short in helping users to rapidly identify the search targets. Particularly, in the situation that a system suggests a long list of similar candidates, the user needs to painstakingly inspect every search result. The experience is frustrated with repeated watching of similar clips, and more frustratingly, the search targets may be overlooked due to mental tiredness. This paper explores reinforcement learning-based (RL) searching to relieve the user from the burden of brute force inspection. Specifically, the system maintains a graph …
Group Contextualization For Video Recognition,
2022
University of Science and Technology of China
Group Contextualization For Video Recognition, Yanbin Hao, Hao Zhang, Chong-Wah Ngo, Xiangnan He
Research Collection School Of Computing and Information Systems
Learning discriminative representation from the complex spatio-temporal dynamic space is essential for video recognition. On top of those stylized spatio-temporal computational units, further refining the learnt feature with axial contexts is demonstrated to be promising in achieving this goal. However, previous works generally focus on utilizing a single kind of contexts to calibrate entire feature channels and could hardly apply to deal with diverse video activities. The problem can be tackled by using pair-wise spatio-temporal attentions to recompute feature response with cross-axis contexts at the expense of heavy computations. In this paper, we propose an efficient feature refinement method that …
Mlp-3d: A Mlp-Like 3d Architecture With Grouped Time Mixing,
2022
Singapore Management University
Mlp-3d: A Mlp-Like 3d Architecture With Grouped Time Mixing, Zhaofan Qiu, Ting Yao, Chong-Wah Ngo, Tao Mei
Research Collection School Of Computing and Information Systems
Convolutional Neural Networks (CNNs) have been re-garded as the go-to models for visual recognition. More re-cently, convolution-free networks, based on multi-head self-attention (MSA) or multi-layer perceptrons (MLPs), become more and more popular. Nevertheless, it is not trivial when utilizing these newly-minted networks for video recognition due to the large variations and complexities in video data. In this paper, we present MLP-3D networks, a novel MLP-like 3D architecture for video recognition. Specifically, the architecture consists of MLP-3D blocks, where each block contains one MLP applied across tokens (i.e., token-mixing MLP) and one MLP applied independently to each token (i.e., channel MLP). …
Class Re-Activation Maps For Weakly-Supervised Semantic Segmentation,
2022
Singapore Management University
Class Re-Activation Maps For Weakly-Supervised Semantic Segmentation, Zhaozheng Chen, Tan Wang, Xiongwei Wu, Xian-Sheng Hua, Hanwang Zhang, Qianru Sun
Research Collection School Of Computing and Information Systems
Extracting class activation maps (CAM) is arguably the most standard step of generating pseudo masks for weakly supervised semantic segmentation (WSSS). Yet, we find that the crux of the unsatisfactory pseudo masks is the binary cross-entropy loss (BCE) widely used in CAM. Specifically, due to the sum-over-class pooling nature of BCE, each pixel in CAM may be responsive to multiple classes co-occurring in the same receptive field. To this end, we introduce an embarrassingly simple yet surprisingly effective method: Reactivating the converged CAM with BCE by using softmax crossentropy loss (SCE), dubbed ReCAM. Given an image, we use CAM to …
Revisiting Local Descriptor For Improved Few-Shot Classification,
2022
Singapore Management University
Revisiting Local Descriptor For Improved Few-Shot Classification, Jun He, Richang Hong, Xueliang Liu, Mingliang Xu, Qianru Sun
Research Collection School Of Computing and Information Systems
Few-shot classification studies the problem of quickly adapting a deep learner to understanding novel classes based on few support images. In this context, recent research efforts have been aimed at designing more and more complex classifiers that measure similarities between query and support images but left the importance of feature embeddings seldom explored. We show that the reliance on sophisticated classifiers is not necessary, and a simple classifier applied directly to improved feature embeddings can instead outperform most of the leading methods in the literature. To this end, we present a new method, named DCAP, for few-shot classification, in which …
Multimodal Zero-Shot Hateful Meme Detection,
2022
Singapore Management University
Multimodal Zero-Shot Hateful Meme Detection, Jiawen Zhu, Roy Ka-Wei Lee, Wen Haw Chong
Research Collection School Of Computing and Information Systems
Facebook has recently launched the hateful meme detection challenge, which garnered much attention in academic and industry research communities. Researchers have proposed multimodal deep learning classification methods to perform hateful meme detection. While the proposed methods have yielded promising results, these classification methods are mostly supervised and heavily rely on labeled data that are not always available in the real-world setting. Therefore, this paper explores and aims to perform hateful meme detection in a zero-shot setting. Working towards this goal, we propose Target-Aware Multimodal Enhancement (TAME), which is a novel deep generative framework that can improve existing hateful meme classification …
Faithful Extreme Rescaling Via Generative Prior Reciprocated Invertible Representations,
2022
Singapore Management University
Faithful Extreme Rescaling Via Generative Prior Reciprocated Invertible Representations, Zhixuan Zhong, Liangyu Chai, Yang Zhou, Bailin Deng, Jia Pan, Shengfeng He
Research Collection School Of Computing and Information Systems
This paper presents a Generative prior ReciprocAted Invertible rescaling Network (GRAIN) for generating faithful high-resolution (HR) images from low-resolution (LR) invertible images with an extreme upscaling factor (64×). Previous researches have leveraged the prior knowledge of a pretrained GAN model to generate high-quality upscaling results. However, they fail to produce pixel-accurate results due to the highly ambiguous extreme mapping process. We remedy this problem by introducing a reciprocated invertible image rescaling process, in which high-resolution information can be delicately embedded into an invertible low-resolution image and generative prior for a faithful HR reconstruction. In particular, the invertible LR features not …
A Simple Data Mixing Prior For Improving Self-Supervised Learning,
2022
Singapore Management University
A Simple Data Mixing Prior For Improving Self-Supervised Learning, Sucheng Ren, Huiyu Wang, Zhengqi Gao, Shengfeng He, Alan Yuille, Yuyin Zhou, Cihang Xie
Research Collection School Of Computing and Information Systems
Data mixing (e.g., Mixup, Cutmix, ResizeMix) is an essential component for advancing recognition models. In this paper, we focus on studying its effectiveness in the self-supervised setting. By noticing the mixed images that share the same source images are intrinsically related to each other, we hereby propose SDMP, short for Simple Data Mixing Prior, to capture this straightforward yet essential prior, and position such mixed images as additional positive pairs to facilitate self-supervised representation learning. Our experiments verify that the proposed SDMP enables data mixing to help a set of self-supervised learning frameworks (e.g., MoCo) achieve better accuracy and out-of-distribution …
High-Resolution Face Swapping Via Latent Semantics Disentanglement,
2022
Singapore Management University
High-Resolution Face Swapping Via Latent Semantics Disentanglement, Yangyang Xu, Bailin Deng, Junle Wang, Yanqing Jing, Jia Pan, Shengfeng He
Research Collection School Of Computing and Information Systems
We present a novel high-resolution face swapping method using the inherent prior knowledge of a pre-trained GAN model. Although previous research can leverage generative priors to produce high-resolution results, their quality can suffer from the entangled semantics of the latent space. We explicitly disentangle the latent semantics by utilizing the progressive nature of the generator, deriving structure at-tributes from the shallow layers and appearance attributes from the deeper ones. Identity and pose information within the structure attributes are further separated by introducing a landmark-driven structure transfer latent direction. The disentangled latent code produces rich generative features that incorporate feature blending …
Multi-View Scheduling Of Onboard Live Video Analytics To Minimize Frame Processing Latency,
2022
Singapore Management University
Multi-View Scheduling Of Onboard Live Video Analytics To Minimize Frame Processing Latency, Shengzhong Liu, Tianshi Wang, Hongpeng Guo, Xinzhe Fu, Philip David, Maggie Wigness, Archan Misra, Tarek Abdelzaher
Research Collection School Of Computing and Information Systems
This paper presents a real-time multi-view scheduling framework for DNN-based live video analytics at the edge to minimize frame processing latency. The work is motivated by applications where a higher frame rate is important, not to miss actions of interest. Examples include defense, border security, and intruder detection applications where sensors (in this paper, cameras) are deployed to monitor key roads, chokepoints, or passageways to identify events of interest (and intervene in real-time). Supporting a higher frame rate entails lowering frame processing latency. We assume that multiple cameras are deployed with partially overlapping views. Each camera has access to limited …
Flavor-Videos: Enhancing The Flavor Perception Of Food While Eating With Videos,
2022
Singapore Management University
Flavor-Videos: Enhancing The Flavor Perception Of Food While Eating With Videos, Meetha Nesam James, Nimesha Ranasinghe, Anthony Tang, Lora Oehlberg
Research Collection School Of Computing and Information Systems
People are typically involved in different activities while eating, particularly when eating alone, such as watching television or playing games on their phones. Previous research in Human-Food Interaction (HFI) has primarily focused on studying people’s motivation and analyzing of the media content watched while eating. However, their impact on human behavioral and cognitive processes, particularly flavor perception and its attributes, remains underexplored. We present a user study to investigate the influence of six types of videos, including mukbang – a new food video genre, on flavor perceptions (taste sensations, liking, and emotions) while eating plain white rice. Our findings revealed …
A Bidirectional Formulation For Walk On Spheres,
2022
Dartmouth College
A Bidirectional Formulation For Walk On Spheres, Yang Qi
Dartmouth College Master’s Theses
Poisson’s equations and Laplace’s equations are important linear partial differential equations (PDEs)
widely used in many applications. Conventional methods for solving PDEs numerically often need to
discretize the space first, making them less efficient for complex shapes. The random walk on spheres
method (WoS) is a grid-free Monte-Carlo method for solving PDEs that does not need to discrete the
space. We draw analogies between WoS and classical rendering algorithms, and find that the WoS
algorithm is conceptually identical to forward path tracing.
We show that solving the Poisson’s equation is equivalent to solving the Green’s function for every
pair of …
Artist-Configurable Node-Based Approach To Generate Procedural Brush Stroke Textures For Digital Painting,
2022
California Polytechnic State University, San Luis Obispo
Artist-Configurable Node-Based Approach To Generate Procedural Brush Stroke Textures For Digital Painting, Keavon Chambers
Master's Theses
Digital painting is the field of software designed to provide artists a virtual medium to emulate the experience and results of physical drawing. Several hardware and software components come together to form a whole workflow, ranging from the physical input devices, to the stroking process, to the texture content authorship. This thesis explores an artist-friendly approach to synthesize the textures that give life to digital brush strokes.
Most painting software provides a limited library of predefined brush textures. They aim to offer styles approximating physical media like paintbrushes, pencils, markers, and airbrushes. Often these are static bitmap textures that are …
Out-Of-Core Gpu Path Tracing On Large Instanced Scenes Via Geometry Streaming,
2022
California Polytechnic State University, San Luis Obispo
Out-Of-Core Gpu Path Tracing On Large Instanced Scenes Via Geometry Streaming, Jeremy Berchtold
Master's Theses
We present a technique for out-of-core GPU path tracing of arbitrarily large scenes that is compatible with hardware-accelerated ray-tracing. Our technique improves upon previous works by subdividing the scene spatially into streamable chunks that are loaded using a priority system that maximizes ray throughput and minimizes GPU memory usage. This allows for arbitrarily large scaling of scene complexity. Our system required under 19 minutes to render a solid color version of Disney's Moana Island scene (39.3 million instances, 261.1 million unique quads, and 82.4 billion instanced quads at a resolution of 1024x429 and 1024spp on an RTX 5000 (24GB memory …
