Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- Singapore Management University (942)
- University of Dayton (114)
- Old Dominion University (99)
- Air Force Institute of Technology (98)
- California Polytechnic State University, San Luis Obispo (96)
-
- University of Arkansas, Fayetteville (89)
- University of Nebraska - Lincoln (51)
- City University of New York (CUNY) (48)
- Technological University Dublin (48)
- University of Malaya (43)
- San Jose State University (37)
- Dartmouth College (35)
- Embry-Riddle Aeronautical University (24)
- Clemson University (23)
- Purdue University (23)
- Rochester Institute of Technology (23)
- The University of Akron (22)
- Chapman University (20)
- Edith Cowan University (20)
- University of Kentucky (18)
- Michigan Technological University (16)
- University of Central Florida (15)
- Southern Adventist University (13)
- California State University, San Bernardino (12)
- Kennesaw State University (12)
- St. Mary's University (12)
- Nova Southeastern University (11)
- University of Minnesota Morris Digital Well (11)
- Louisiana State University (10)
- University of Nevada, Las Vegas (10)
- Keyword
-
- Virtual reality (62)
- Visualization (46)
- Computer graphics (38)
- Computer vision (37)
- Accessibility (36)
-
- Human-computer interaction (35)
- Augmented reality (33)
- Usability (31)
- Machine learning (29)
- Machine Learning (26)
- Artificial intelligence (25)
- Computer Science (25)
- Data visualization (25)
- Deep learning (24)
- Virtual Reality (23)
- HCI (22)
- Computer science (21)
- Eye tracking (20)
- Human computer interaction (20)
- User experience (20)
- Design (19)
- Deep Learning (16)
- Education (16)
- Feature extraction (15)
- Graph Neural Networks (15)
- Graphics (15)
- VR (15)
- Artificial Intelligence (14)
- Gamification (14)
- Image processing (14)
- Publication Year
- Publication
-
- Research Collection School Of Computing and Information Systems (916)
- Computer Science Faculty Publications (135)
- Theses and Dissertations (98)
- Master's Theses (50)
- Graduate Theses and Dissertations (43)
-
- Student Works (2000-2009) (33)
- Computer Science and Computer Engineering Undergraduate Honors Theses (31)
- 3-D Printed Model Structural Files (29)
- Publications and Research (28)
- Dartmouth College Master’s Theses (24)
- Williams Honors College, Honors Research Projects (22)
- Master's Projects (20)
- Conference papers (19)
- Frameless (19)
- Computer Science and Software Engineering (18)
- All Dissertations (17)
- Dissertations and Theses Collection (Open Access) (16)
- Dissertations, Master's Theses and Master's Reports (16)
- H-Workload 2017: Models and Applications (Works in Progress) (15)
- Computer Engineering (14)
- Electronic Theses and Dissertations (14)
- Theses : Honours (14)
- Honors Theses (13)
- MAICS: The Modern Artificial Intelligence and Cognitive Science Conference (12)
- AFIT Patents (11)
- CCAC Theses and Dissertations (11)
- Engineering Faculty Articles and Research (11)
- Scholarly Horizons: University of Minnesota, Morris Undergraduate Journal (10)
- Student Works (2020-2029) (10)
- Inquiry: The University of Arkansas Undergraduate Research Journal (9)
- Publication Type
- File Type
Articles 511 - 540 of 2372
Full-Text Articles in Computer Sciences
Leveraging Artificial Intelligence For Team Cognition In Human-Ai Teams, Beau Schelble
Leveraging Artificial Intelligence For Team Cognition In Human-Ai Teams, Beau Schelble
All Dissertations
Advances in artificial intelligence (AI) technologies have enabled AI to be applied across a wide variety of new fields like cryptography, art, and data analysis. Several of these fields are social in nature, including decision-making and teaming, which introduces a new set of challenges for AI research. While each of these fields has its unique challenges, the area of human-AI teaming is beset with many that center around the expectations and abilities of AI teammates. One such challenge is understanding team cognition in these human-AI teams and AI teammates' ability to contribute towards, support, and encourage it. Team cognition is …
Video Sentiment Analysis For Child Safety, Yee Sen Tan, Nicole Anne Huiying Teo, Ezekiel En Zhe Ghe, Jolie Zhi Yi Fong, Zhaoxia Wang
Video Sentiment Analysis For Child Safety, Yee Sen Tan, Nicole Anne Huiying Teo, Ezekiel En Zhe Ghe, Jolie Zhi Yi Fong, Zhaoxia Wang
Research Collection School Of Computing and Information Systems
The proliferation of online video content underscores the critical need for effective sentiment analysis, particularly in safeguarding children from potentially harmful material. This research addresses this concern by presenting a multimodal analysis method for assessing video sentiment, categorizing it as either positive (child-friendly) or negative (potentially harmful). This method leverages three key components: text analysis, facial expression analysis, and audio analysis, including music mood analysis, resulting in a comprehensive sentiment assessment. Our evaluation results validate the effectiveness of this approach, making significant contributions to the field of video sentiment analysis and bolstering child safety measures. This research serves as a …
Unifying Text, Tables, And Images For Multimodal Question Answering, Haohao Luo, Ying Shen, Yang Deng
Unifying Text, Tables, And Images For Multimodal Question Answering, Haohao Luo, Ying Shen, Yang Deng
Research Collection School Of Computing and Information Systems
Multimodal question answering (MMQA), which aims to derive the answer from multiple knowledge modalities (e.g., text, tables, and images), has received increasing attention due to its board applications. Current approaches to MMQA often rely on single-modal or bi-modal QA models, which limits their ability to effectively integrate information across all modalities and leverage the power of pre-trained language models. To address these limitations, we propose a novel framework called UniMMQA, which unifies three different input modalities into a text-to-text format by employing position-enhanced table linearization and diversified image captioning techniques. Additionally, we enhance cross-modal reasoning by incorporating a multimodal rationale …
Developing Detection And Mapping Of Roads Within Various Forms Of Media Using Opencv, Jordan C. Lyle
Developing Detection And Mapping Of Roads Within Various Forms Of Media Using Opencv, Jordan C. Lyle
Computer Science and Computer Engineering Undergraduate Honors Theses
OpenCV, and Computer Vision in general, has been a Computer Science topic that has interested me for a long time while completing my Bachelor’s degree at the University of Arkansas. As a result of this, I ended up choosing to utilize OpenCV in order to complete the task of detecting road-lines and mapping roads when given a wide variety of images. The purpose of my Honors research and this thesis is to detail the process of creating an algorithm to detect the road-lines such that the results are effective and instantaneous, as well as detail how Computer Vision can be …
Self-Supervised Pseudo Multi-Class Pre-Training For Unsupervised Anomaly Detection And Segmentation In Medical Images, Yu Tian, Fengbei Liu, Guansong Pang, Yuanhong Chen, Yuyuan Liu, Johan W. Verjans, Rajvinder Singh, Gustavo Carneiro
Self-Supervised Pseudo Multi-Class Pre-Training For Unsupervised Anomaly Detection And Segmentation In Medical Images, Yu Tian, Fengbei Liu, Guansong Pang, Yuanhong Chen, Yuyuan Liu, Johan W. Verjans, Rajvinder Singh, Gustavo Carneiro
Research Collection School Of Computing and Information Systems
Unsupervised anomaly detection (UAD) methods are trained with normal (or healthy) images only, but during testing, they are able to classify normal and abnormal (or disease) images. UAD is an important medical image analysis (MIA) method to be applied in disease screening problems because the training sets available for those problems usually contain only normal images. However, the exclusive reliance on normal images may result in the learning of ineffective low-dimensional image representations that are not sensitive enough to detect and segment unseen abnormal lesions of varying size, appearance, and shape. Pre-training UAD methods with self-supervised learning, based on computer …
Graph Contrastive Learning With Stable And Scalable Spectral Encoding, Deyu Bo, Yuan Fang, Yang Liu, Chuan Shi
Graph Contrastive Learning With Stable And Scalable Spectral Encoding, Deyu Bo, Yuan Fang, Yang Liu, Chuan Shi
Research Collection School Of Computing and Information Systems
Graph contrastive learning (GCL) aims to learn representations by capturing the agreements between different graph views. Traditional GCL methods generate views in the spatial domain, but it has been recently discovered that the spectral domain also plays a vital role in complementing spatial views. However, existing spectral-based graph views either ignore the eigenvectors that encode valuable positional information, or suffer from high complexity when trying to address the instability of spectral features. To tackle these challenges, we first design an informative, stable, and scalable spectral encoder, termed EigenMLP, to learn effective representations from the spectral features. Theoretically, EigenMLP is invariant …
Depwignn: A Depth-Wise Graph Neural Network For Multi-Hop Spatial Reasoning In Text, Shuaiyi Li, Yang Deng, Wai Lam
Depwignn: A Depth-Wise Graph Neural Network For Multi-Hop Spatial Reasoning In Text, Shuaiyi Li, Yang Deng, Wai Lam
Research Collection School Of Computing and Information Systems
Spatial reasoning in text plays a crucial role in various real-world applications. Existing approaches for spatial reasoning typically infer spatial relations from pure text, which overlook the gap between natural language and symbolic structures. Graph neural networks (GNNs) have showcased exceptional proficiency in inducing and aggregating symbolic structures. However, classical GNNs face challenges in handling multi-hop spatial reasoning due to the over-smoothing issue, i.e., the performance decreases substantially as the number of graph layers increases. To cope with these challenges, we propose a novel Depth-Wise Graph Neural Network (DepWiGNN). Specifically, we design a novel node memory scheme and aggregate the …
The Propagation And Execution Of Malware In Images, Piper Hall
The Propagation And Execution Of Malware In Images, Piper Hall
Cybersecurity Undergraduate Research Showcase
Malware has become increasingly prolific and severe in its consequences as information systems mature and users become more reliant on computing in their daily lives. As cybercrime becomes more complex in its strategies, an often-overlooked manner of propagation is through images. In recent years, several high-profile vulnerabilities in image libraries have opened the door for threat actors to steal money and information from unsuspecting users. This paper will explore the mechanisms by which these exploits function and how they can be avoided.
Performative Mixing For Immersive Audio, Brian A. Elizondo
Performative Mixing For Immersive Audio, Brian A. Elizondo
LSU Doctoral Dissertations
Immersive multichannel audio can be produced with specialized setups of loudspeakers, often surrounding the audience. These setups can feature as few as four loudspeakers or more than 300. Performative mixing in these environments requires a bespoke solution offering intuitive gestural control. Beyond the usual faders for gain control, advancements in multichannel sound demand interfaces capable of quickly positioning sounds between channels. The Quad Cartesian Positioner is such a solution in the form of a Eurorack module for surround mixing for use in live or studio performances.
Diffusion/mixing methods for live multichannel immersive music often rely on the repurposing of hardware …
User Feedback On Celebratory Technology Model For Reducing Stigma, Evelyn Lawrie, Daniel Dinh, Sav Avalos, Jack De Bruyn, Spencer Au, Christian Lopez, Ray Tan, Cyrus Fa'amafoe
User Feedback On Celebratory Technology Model For Reducing Stigma, Evelyn Lawrie, Daniel Dinh, Sav Avalos, Jack De Bruyn, Spencer Au, Christian Lopez, Ray Tan, Cyrus Fa'amafoe
Student Scholar Symposium Abstracts and Posters
Social stigma is a complex manifestation that affects humanity, particularly individuals with disabilities and other marginalized groups, including those with physical, cognitive, and emotional conditions. Society often judges these individuals' interactions with the world, and many technologies designed to assist those with disabilities attempt to change their daily interactions and behaviors. Nonetheless, when the emphasis is placed on validating disabled identities, there is a potential for it to be seen as "inspiration porn." This approach might inadvertently reduce inclusivity and do little to challenge negative stereotypes; it can also lead to the objectification of individuals with disabilities. Therefore, this project …
Vrmovian - An Immersive Data Annotation Tool For Visual Analysis Of Human Interactions In Vr, Isaac Browen
Vrmovian - An Immersive Data Annotation Tool For Visual Analysis Of Human Interactions In Vr, Isaac Browen
Student Scholar Symposium Abstracts and Posters
Understanding human behavior in virtual reality (VR) is a key component for developing intelligent systems to enhance human focused VR experiences. The ability to annotate human motion data proves to be a very useful way to analyze and understand human behavior. However, due to the complexity and multi-dimensionality of human activity data, it is necessary to develop software that can display the data in a comprehensible way and can support intuitive data annotation for developing machine learning models able recognize and assist human motions in VR (e.g., remote physical therapy). Although past research has been done to improve VR data …
Bridging Domain Gaps For Cross-Spectrum And Long-Range Face Recognition Using Domain Adaptive Machine Learning, Cedric Armel Nimpa Fondje
Bridging Domain Gaps For Cross-Spectrum And Long-Range Face Recognition Using Domain Adaptive Machine Learning, Cedric Armel Nimpa Fondje
Dissertations and Doctoral Documents, University of Nebraska-Lincoln, 2023–
Face recognition technology has witnessed significant advancements in recent decades, enabling its widespread adoption in various applications such as security, surveillance, and biometrics applications. However, one of the primary challenges faced by existing face recognition systems is their limited performance when presented with images from different modalities or domains( such as infrared to visible, long range to close range, nighttime to daytime, profile to f rontal, etc.) Additionally, advancements in camera sensors, analytics beyond the visible spectrum, and the increasing size of cross-modal datasets have led to a particular interest in cross-modal learning for face recognition in the biometrics and …
Enhancing Search Engine Results: A Comparative Study Of Graph And Timeline Visualizations For Semantic And Temporal Relationship Discovery, Muhammad Shahiq Qureshi
Enhancing Search Engine Results: A Comparative Study Of Graph And Timeline Visualizations For Semantic And Temporal Relationship Discovery, Muhammad Shahiq Qureshi
Electronic Theses and Dissertations
In today’s digital age, search engines have become indispensable tools for finding information among the corpus of billions of webpages. The standard that most search engines follow is to display search results in a list-based format arranged according to a ranking algorithm. Although this format is good for presenting the most relevant results to users, it fails to represent the underlying relations between different results. These relations, among others, can generally be of either a temporal or semantic nature. A user who wants to explore the results that are connected by those relations would have to make a manual effort …
Constructing Holistic Spatio-Temporal Scene Graph For Video Semantic Role Labeling, Yu Zhao, Hao Fei, Yixin Cao, Bobo Li, Meishan Zhang, Jianguo Wei, Min Zhang, Tat-Seng Chua
Constructing Holistic Spatio-Temporal Scene Graph For Video Semantic Role Labeling, Yu Zhao, Hao Fei, Yixin Cao, Bobo Li, Meishan Zhang, Jianguo Wei, Min Zhang, Tat-Seng Chua
Research Collection School Of Computing and Information Systems
As one of the core video semantic understanding tasks, Video Semantic Role Labeling (VidSRL) aims to detect the salient events from given videos, by recognizing the predict-argument event structures and the interrelationships between events. While recent endeavors have put forth methods for VidSRL, they can be mostly subject to two key drawbacks, including the lack of fine-grained spatial scene perception and the insufficiently modeling of video temporality. Towards this end, this work explores a novel holistic spatio-temporal scene graph (namely HostSG) representation based on the existing dynamic scene graph structures, which well model both the fine-grained spatial semantics and temporal …
Npf-200: A Multi-Modal Eye Fixation Dataset And Method For Non-Photorealistic Videos, Ziyu Yang, Sucheng Ren, Zongwei Wu, Nanxuan Zhao, Junle Wang, Jing Qin, Shengfeng He
Npf-200: A Multi-Modal Eye Fixation Dataset And Method For Non-Photorealistic Videos, Ziyu Yang, Sucheng Ren, Zongwei Wu, Nanxuan Zhao, Junle Wang, Jing Qin, Shengfeng He
Research Collection School Of Computing and Information Systems
Non-photorealistic videos are in demand with the wave of the metaverse, but lack of sufficient research studies. This work aims to take a step forward to understand how humans perceive nonphotorealistic videos with eye fixation (i.e., saliency detection), which is critical for enhancing media production, artistic design, and game user experience. To fill in the gap of missing a suitable dataset for this research line, we present NPF-200, the first largescale multi-modal dataset of purely non-photorealistic videos with eye fixations. Our dataset has three characteristics: 1) it contains soundtracks that are essential according to vision and psychological studies; 2) it …
Designing Secure Mental Healthcare Chatbots For Older Adults, Aishwarya Surani
Designing Secure Mental Healthcare Chatbots For Older Adults, Aishwarya Surani
Electronic Theses and Dissertations
The landscape of mental health support has evolved as a result of the rising demand for digital mental healthcare services. Users now have an opportunity to seek mental health support online due to the growth of digital platforms. For those looking for mental health treatments, chatbots have evolved as user-friendly, accessible platforms that provide remote access and convenience. However, for chatbots to be effective, users must divulge personal and sensitive information, such as demographics, insurance information, and a history of mental illness. While chatbots offer services to a variety of demographic users, older adults face unique challenges related to usability, …
Optimizing E-Payment Applications For Older Adults: User-Centered Solutions To Improve Security, Privacy, Usability, And Accessibility, Urvashi Kishnani
Optimizing E-Payment Applications For Older Adults: User-Centered Solutions To Improve Security, Privacy, Usability, And Accessibility, Urvashi Kishnani
Electronic Theses and Dissertations
In an increasingly digital world, older adults are rapidly becoming a vital demographic in the realm of electronic financial transactions. It is imperative to address their unique needs and challenges to ensure their financial well-being. Older adults can be more vulnerable to various online threats, making security and privacy paramount. As they adapt to the digital age, understanding their specific privacy concerns and preferences is crucial for creating trustworthy e-payment systems. Moreover, enhancing the usability of e-payment applications for older adults promotes financial independence and inclusion, contributing to their overall quality of life. By focusing on these critical dimensions, we …
Editanything: Empowering Unparalleled Flexibility In Image Editing And Generation, Shanghua Gao, Zhijie Lin, Xingyu Xie, Pan Zhou, Ming-Ming Cheng, Shuicheng Yan
Editanything: Empowering Unparalleled Flexibility In Image Editing And Generation, Shanghua Gao, Zhijie Lin, Xingyu Xie, Pan Zhou, Ming-Ming Cheng, Shuicheng Yan
Research Collection School Of Computing and Information Systems
Image editing plays a vital role in computer vision field, aiming to realistically manipulate images while ensuring seamless integration. It finds numerous applications across various fields. In this work, we present EditAnything, a novel approach that empowers users with unparalleled flexibility in editing and generating image content. EditAnything introduces an array of advanced features, including crossimage dragging (e.g., try-on), region-interactive editing, controllable layout generation, and virtual character replacement. By harnessing these capabilities, users can engage in interactive and flexible editing, giving captivating outcomes that uphold the integrity of the original image. With its diverse range of tools, EditAnything caters to …
Turn-It-Up: Rendering Resistance For Knobs In Virtual Reality Through Undetectable Pseudo-Haptics, Martin Feick, Andre Zenner, Oscar Ariza, Anthony Tang, Cihan Biyikli, Antonio Kruger
Turn-It-Up: Rendering Resistance For Knobs In Virtual Reality Through Undetectable Pseudo-Haptics, Martin Feick, Andre Zenner, Oscar Ariza, Anthony Tang, Cihan Biyikli, Antonio Kruger
Research Collection School Of Computing and Information Systems
Rendering haptic feedback for interactions with virtual objects is an essential part of effective virtual reality experiences. In this work, we explore providing haptic feedback for rotational manipulations, e.g., through knobs. We propose the use of a Pseudo-Haptic technique alongside a physical proxy knob to simulate various physical resistances. In a psychophysical experiment with 20 participants, we found that designers can introduce unnoticeable offsets between real and virtual rotations of the knob, and we report the corresponding detection thresholds. Based on these, we present the Pseudo-Haptic Resistance technique to convey physical resistance while applying only unnoticeable pseudo-haptic manipulation. Additionally, we …
Voxelhap: A Toolkit For Constructing Proxies Providing Tactile And Kinesthetic Haptic Feedback In Virtual Reality, M. Feick, C. Biyikli, K. Gani, A. Wittig, Anthony Tang, A. Krüger
Voxelhap: A Toolkit For Constructing Proxies Providing Tactile And Kinesthetic Haptic Feedback In Virtual Reality, M. Feick, C. Biyikli, K. Gani, A. Wittig, Anthony Tang, A. Krüger
Research Collection School Of Computing and Information Systems
Experiencing virtual environments is often limited to abstract interactions with objects. Physical proxies allow users to feel virtual objects, but are often inaccessible. We present the VoxelHap toolkit which enables users to construct highly functional proxy objects using Voxels and Plates. Voxels are blocks with special functionalities that form the core of each physical proxy. Plates increase a proxy’s haptic resolution, such as its shape, texture or weight. Beyond providing physical capabilities to realize haptic sensations, VoxelHap utilizes VR illusion techniques to expand its haptic resolution. We evaluated the capabilities of the VoxelHap toolkit through the construction of a range …
Disentangling Multi-View Representations Beyond Inductive Bias, Guanzhou Ke, Yang Yu, Guoqing Chao, Xiaoli Wang, Chenyang Xu, Shengfeng He
Disentangling Multi-View Representations Beyond Inductive Bias, Guanzhou Ke, Yang Yu, Guoqing Chao, Xiaoli Wang, Chenyang Xu, Shengfeng He
Research Collection School Of Computing and Information Systems
Multi-view (or -modality) representation learning aims to understand the relationships between different view representations. Existing methods disentangle multi-view representations into consistent and view-specific representations by introducing strong inductive biases, which can limit their generalization ability. In this paper, we propose a novel multi-view representation disentangling method that aims to go beyond inductive biases, ensuring both interpretability and generalizability of the resulting representations. Our method is based on the observation that discovering multi-view consistency in advance can determine the disentangling information boundary, leading to a decoupled learning objective. We also found that the consistency can be easily extracted by maximizing the …
Revisiting Disentanglement And Fusion On Modality And Context In Conversational Multimodal Emotion Recognition, Bobo Li, Hao Fei, Lizi Liao, Yu Zhao, Chong Teng, Tat-Seng Chua, Donghong Ji, Fei Li
Revisiting Disentanglement And Fusion On Modality And Context In Conversational Multimodal Emotion Recognition, Bobo Li, Hao Fei, Lizi Liao, Yu Zhao, Chong Teng, Tat-Seng Chua, Donghong Ji, Fei Li
Research Collection School Of Computing and Information Systems
It has been a hot research topic to enable machines to understand human emotions in multimodal contexts under dialogue scenarios, which is tasked with multimodal emotion analysis in conversation (MM-ERC). MM-ERC has received consistent attention in recent years, where a diverse range of methods has been proposed for securing better task performance. Most existing works treat MM-ERC as a standard multimodal classification problem and perform multimodal feature disentanglement and fusion for maximizing feature utility. Yet after revisiting the characteristic of MM-ERC, we argue that both the feature multimodality and conversational contextualization should be properly modeled simultaneously during the feature disentanglement …
Cgt-Gan: Clip-Guided Text Gan For Image Captioning, Jiarui Yu, Haoran Li, Yanbin Hao, Bin Zhu, Tong Xu, Xiangnan He
Cgt-Gan: Clip-Guided Text Gan For Image Captioning, Jiarui Yu, Haoran Li, Yanbin Hao, Bin Zhu, Tong Xu, Xiangnan He
Research Collection School Of Computing and Information Systems
The large-scale visual-language pre-trained model, Contrastive Language-Image Pre-training (CLIP), has significantly improved image captioning for scenarios without human-annotated image-caption pairs. Recent advanced CLIP-based image captioning without human annotations follows a text-only training paradigm, i.e., reconstructing text from shared embedding space. Nevertheless, these approaches are limited by the training/inference gap or huge storage requirements for text embeddings. Given that it is trivial to obtain images in the real world, we propose CLIP-guided text GAN (CgT-GAN), which incorporates images into the training process to enable the model to "see" real visual modality. Particularly, we use adversarial training to teach CgT-GAN to mimic …
Pro-Cap: Leveraging A Frozen Vision-Language Model For Hateful Meme Detection, Rui Cao, Ming Shan Hee, Adriel Kuek, Wen Haw Chong, Roy Ka-Wei Lee, Jing Jiang
Pro-Cap: Leveraging A Frozen Vision-Language Model For Hateful Meme Detection, Rui Cao, Ming Shan Hee, Adriel Kuek, Wen Haw Chong, Roy Ka-Wei Lee, Jing Jiang
Research Collection School Of Computing and Information Systems
Hateful meme detection is a challenging multimodal task that requires comprehension of both vision and language, as well as cross-modal interactions. Recent studies have tried to fine-tune pre-trained vision-language models (PVLMs) for this task. However, with increasing model sizes, it becomes important to leverage powerful PVLMs more efficiently, rather than simply fine-tuning them. Recently, researchers have attempted to convert meme images into textual captions and prompt language models for predictions. This approach has shown good performance but suffers from non-informative image captions. Considering the two factors mentioned above, we propose a probing-based captioning approach to leverage PVLMs in a zero-shot …
Matk: The Meme Analytical Tool Kit, Ming Shan Hee, Aditi Kumaresan, Nguyen Khoi Hoang, Nirmalendu Prakash, Rui Cao, Roy Ka-Wei Lee
Matk: The Meme Analytical Tool Kit, Ming Shan Hee, Aditi Kumaresan, Nguyen Khoi Hoang, Nirmalendu Prakash, Rui Cao, Roy Ka-Wei Lee
Research Collection School Of Computing and Information Systems
The rise of social media platforms has brought about a new digital culture called memes. Memes, which combine visuals and text, can strongly influence public opinions on social and cultural issues. As a result, people have become interested in categorizing memes, leading to the development of various datasets and multimodal models that show promising results in this field. However, there is currently a lack of a single library that allows for the reproduction, evaluation, and comparison of these models using fair benchmarks and settings. To fill this gap, we introduce the Meme Analytical Tool Kit (MATK), an open-source toolkit specifically …
Metaformer Baselines For Vision, Weihao Yu, Chenyang Si, Pan Zhou, Mi Luo, Yichen Zhou, Jiashi Feng, Shuicheng Yan, Xinchao Wang
Metaformer Baselines For Vision, Weihao Yu, Chenyang Si, Pan Zhou, Mi Luo, Yichen Zhou, Jiashi Feng, Shuicheng Yan, Xinchao Wang
Research Collection School Of Computing and Information Systems
Abstract—MetaFormer, the abstracted architecture of Transformer, has been found to play a significant role in achieving competitive performance. In this paper, we further explore the capacity of MetaFormer, again, by migrating our focus away from the token mixer design: we introduce several baseline models under MetaFormer using the most basic or common mixers, and demonstrate their gratifying performance. We summarize our observations as follows: (1) MetaFormer ensures solid lower bound of performance. By merely adopting identity mapping as the token mixer, the MetaFormer model, termed IdentityFormer, achieves >80% accuracy on ImageNet-1K. (2) MetaFormer works well with arbitrary token mixers. When …
Demo Abstract: Vgglass - Demonstrating Visual Grounding And Localization Synergy With A Lidar-Enabled Smart-Glass, Darshana Rathnayake, Dulanga Weerakoon, Meeralakshmi Radhakrishnan, Vigneshwaran Subbaraju, Inseok Hwang, Archan Misra
Demo Abstract: Vgglass - Demonstrating Visual Grounding And Localization Synergy With A Lidar-Enabled Smart-Glass, Darshana Rathnayake, Dulanga Weerakoon, Meeralakshmi Radhakrishnan, Vigneshwaran Subbaraju, Inseok Hwang, Archan Misra
Research Collection School Of Computing and Information Systems
This work demonstrates the VGGlass system, which simultaneously interprets human instructions for a target acquisition task and determines the precise 3D positions of both user and the target object. This is achieved by utilizing LiDARs mounted in the infrastructure and a smart glass device worn by the user. Key to our system is the union of LiDAR-based localization termed LiLOC and a multi-modal visual grounding approach termed RealG(2)In-Lite. To demonstrate the system, we use Intel RealSense L515 cameras and a Microsoft HoloLens 2, as the user devices. VGGlass is able to: a) track the user in real-time in a global …
Improving Human-Automation Collaboration In Motion Planning, Torin J. Adamson
Improving Human-Automation Collaboration In Motion Planning, Torin J. Adamson
Computer Science ETDs
Human-automation collaboration is becoming a part of everyday life as AI helps us drive, make decisions, and solve a variety of other tasks. However, safe and effective collaboration systems depend on factors in trust, communication, and more. Existing studies to explore these are typically carried out in laboratory settings, providing robust data under tight environmental control. However, human behavior evolves over time, driven by external factors that cannot be fully captured in single participation sessions. These factors form the "human context", contextualizing the behavioral data for a more complete understanding. In this thesis, video game adaptations upon conventional subject studies …
Evocative And Provocative Image-Making In The Age Of Generative Ai, Julian Kilker
Evocative And Provocative Image-Making In The Age Of Generative Ai, Julian Kilker
Tradition Innovations in Arts, Design, and Media Higher Education
Editorial for inaugural AI-focused special issue of Tradition-Innovations in Arts, Design, and Media Higher Education, published under the auspices of the Alliance for the Arts in Research Universities (a2ru). Discusses three articles by five authors in this issue: (1) Choreographing Shadows: Interdisciplinary collaboration to orchestrate ethical image-making by Mark Burchick and Diana Pasulka; (2) Giving Up Control: Hybrid AI-augmented workflows for image-making by Joshua Vermillion; and (3) Hands are Hard: Unlearning how we talk about machine learning in the arts by Adam Hyland and Oscar Keyes.
Editing this special issue explored several key questions: What does “innovation” mean when …
Deep Video Demoireing Via Compact Invertible Dyadic Decomposition, Yuhui Quan, Haoran Huang, Shengfeng He, Ruotao Xu
Deep Video Demoireing Via Compact Invertible Dyadic Decomposition, Yuhui Quan, Haoran Huang, Shengfeng He, Ruotao Xu
Research Collection School Of Computing and Information Systems
Removing moire patterns from videos recorded on screens or complex textures is known as video demoireing. It is a challenging task as both structures and textures of an image usually exhibit strong periodic patterns, which thus are easily confused with moire patterns and can be significantly erased in the removal process. By interpreting video demoireing as a multi-frame decomposition problem, we propose a compact invertible dyadic network called CIDNet that progressively decouples latent frames and the moire patterns from an input video sequence. Using a dyadic cross-scale coupling structure with coupling layers tailored for multi-scale processing, CIDNet aims at disentangling …