Efficient Unsupervised Video Hashing With Contextual Modeling And Structural Controlling,
2024
University of Science and Technology of China
Efficient Unsupervised Video Hashing With Contextual Modeling And Structural Controlling, Jingru Duan, Yanbin Hao, Bin Zhu, Lechao Cheng, Pengyuan Zhou, Xiang Wang
Research Collection School Of Computing and Information Systems
The most important effect of the video hashing technique is to support fast retrieval, which is benefiting from the high efficiency of binary calculation. Current video hash approaches are thus mainly targeted at learning compact binary codes to represent video content accurately. However, they may overlook the generation efficiency for hash codes, i.e., designing lightweight neural networks. This paper proposes an method, which is not only for computing compact hash codes but also for designing a lightweight deep model. Specifically, we present an MLP-based model, where the video tensor is split into several groups and multiple axial contexts are explored …
Delving Into Multimodal Prompting For Fine-Grained Visual Classification,
2024
Singapore Management University
Delving Into Multimodal Prompting For Fine-Grained Visual Classification, Xin Jiang, Hao Tang, Junyao Gao, Xiaoyu Du, Shengfeng He, Zechao Li
Research Collection School Of Computing and Information Systems
Fine-grained visual classification (FGVC) involves categorizing fine subdivisions within a broader category, which poses challenges due to subtle inter-class discrepancies and large intra-class variations. However, prevailing approaches primarily focus on uni-modal visual concepts. Recent advancements in pre-trained vision-language models have demonstrated remarkable performance in various high-level vision tasks, yet the applicability of such models to FGVC tasks remains uncertain. In this paper, we aim to fully exploit the capabilities of cross-modal description to tackle FGVC tasks and propose a novel multimodal prompting solution, denoted as MP-FGVC, based on the contrastive language-image pertaining (CLIP) model. Our MP-FGVC comprises a multimodal prompts …
Simple Image-Level Classification Improves Open-Vocabulary Object Detection,
2024
Singapore Management University
Simple Image-Level Classification Improves Open-Vocabulary Object Detection, Ruohuan Fang, Guansong Pang, Xiao Bai
Research Collection School Of Computing and Information Systems
Open-Vocabulary Object Detection (OVOD) aims to detect novel objects beyond a given set of base categories on which the detection model is trained. Recent OVOD methods focus on adapting the image-level pre-trained vision-language models (VLMs), such as CLIP, to a region-level object detection task via, eg., region-level knowledge distillation, regional prompt learning, or region-text pre-training, to expand the detection vocabulary. These methods have demonstrated remarkable performance in recognizing regional visual concepts, but they are weak in exploiting the VLMs' powerful global scene understanding ability learned from the billion-scale image-level text descriptions. This limits their capability in detecting hard objects of …
M3sa: Multimodal Sentiment Analysis Based On Multi-Scale Feature Extraction And Multi-Task Learning,
2024
Singapore Management University
M3sa: Multimodal Sentiment Analysis Based On Multi-Scale Feature Extraction And Multi-Task Learning, Changkai Lin, Hongju Cheng, Qiang Rao, Yang Yang
Research Collection School Of Computing and Information Systems
Sentiment analysis plays an indispensable part in human-computer interaction. Multimodal sentiment analysis can overcome the shortcomings of unimodal sentiment analysis by fusing multimodal data. However, how to extracte improved feature representations and how to execute effective modality fusion are two crucial problems in multimodal sentiment analysis. Traditional work uses simple sub-models for feature extraction, and they ignore features of different scales and fuse different modalities of data equally, making it easier to incorporate extraneous information and affect analysis accuracy. In this paper, we propose a Multimodal Sentiment Analysis model based on Multi-scale feature extraction and Multi-task learning (M 3 SA). …
From Canteen Food To Daily Meals: Generalizing Food Recognition To More Practical Scenarios,
2024
Fudan University
From Canteen Food To Daily Meals: Generalizing Food Recognition To More Practical Scenarios, Guoshan Liu, Yang Jiao, Jingjing Chen, Bin Zhu, Yu-Gang Jiang
Research Collection School Of Computing and Information Systems
The precise recognition of food categories plays a pivotal role for intelligent health management, attracting significant research attention in recent years. Prominent benchmarks, such as Food-101 and VIREO Food-172, provide abundant food image resources that catalyze the prosperity of research in this field. Nevertheless, these datasets are well-curated from canteen scenarios and thus deviate from food appearances in daily life. This discrepancy poses great challenges in effectively transferring classifiers trained on these canteen datasets to broader daily-life scenarios encountered by humans. Toward this end, we present two new benchmarks, namely DailyFood-172 and DailyFood-16, specifically designed to curate food images from …
Out-Of-Distribution Detection In Long-Tailed Recognition With Calibrated Outlier Class Learning,
2024
Singapore Management University
Out-Of-Distribution Detection In Long-Tailed Recognition With Calibrated Outlier Class Learning, Wenjun Miao, Guansong Pang, Xiao Bai, Tianqi Li, Jin Zheng
Research Collection School Of Computing and Information Systems
Existing out-of-distribution (OOD) methods have shown great success on balanced datasets but become ineffective in long-tailed recognition (LTR) scenarios where 1) OOD samples are often wrongly classified into head classes and/or 2) tail-class samples are treated as OOD samples. To address these issues, current studies fit a prior distribution of auxiliary/pseudo OOD data to the long-tailed in-distribution (ID) data. However, it is difficult to obtain such an accurate prior distribution given the unknowingness of real OOD samples and heavy class imbalance in LTR. A straightforward solution to avoid the requirement of this prior is to learn an outlier class to …
Vadclip: Adapting Vision-Language Models For Weakly Supervised Video Anomaly Detection,
2024
Singapore Management University
Vadclip: Adapting Vision-Language Models For Weakly Supervised Video Anomaly Detection, Peng Wu, Xuerong Zhou, Guansong Pang, Lingru Zhou, Qingsen Yan, Peng Wang, Yanning Zhang
Research Collection School Of Computing and Information Systems
The recent contrastive language-image pre-training (CLIP) model has shown great success in a wide range of image-level tasks, revealing remarkable ability for learning powerful visual representations with rich semantics. An open and worthwhile problem is efficiently adapting such a strong model to the video domain and designing a robust video anomaly detector. In this work, we propose VadCLIP, a new paradigm for weakly supervised video anomaly detection (WSVAD) by leveraging the frozen CLIP model directly without any pre-training and fine-tuning process. Unlike current works that directly feed extracted features into the weakly supervised classifier for frame-level binary classification, VadCLIP makes …
Learning From Machines: How Negative Feedback From Machines Improves Learning Between Humans,
2024
Zhejiang University
Learning From Machines: How Negative Feedback From Machines Improves Learning Between Humans, Tengjian Zou, Gokhan Ertug, Thomas Roulet
Research Collection Lee Kong Chian School Of Business
Prior studies on learning from failure primarily focus on how individuals learn from failure feedback given by other individuals. It is unclear whether and how the advent of machine feedback may influence individuals’ learning from failures. We suggest that failure feedback provided by machines facilitates learning in two ways. First, it focuses individuals’ attention on their failures, leading them to learn from these failures. Second, it serves as a catalyzer, motivating individuals to learn more from failure feedback given to them by other individuals as well. In addition, this catalyzing effect is stronger if the failure feedback from machines and …
Digitizing Delphi: Educating Audiences Through Virtual Reconstruction,
2024
Purdue University
Digitizing Delphi: Educating Audiences Through Virtual Reconstruction, Kate Koury
The Journal of Purdue Undergraduate Research
Implementing a 3D model into a virtual space allows the general public to engage critically with archaeological processes. There are many unseen decisions that go into reconstructing an ancient temple. Analysis of available materials and techniques, predictions of how objects were used, decisions of what sources to reference, puzzle piecing broken remains together, and even educated guesses used to fill gaps in information often go unobserved by the public. This work will educate users about those choices by allowing the side-by-side comparison of conflicting theories on the reconstruction of the Tholos at Delphi, which is an ideal site because of …
Piecing Together Performance: Collaborative, Participatory Research-Through-Design For Better Diversity In Games,
2024
Chapman University
Piecing Together Performance: Collaborative, Participatory Research-Through-Design For Better Diversity In Games, Daniel L. Gardner, Louanne Boyd, Reginald T. Gardner
Engineering Faculty Articles and Research
Digital games are a multi-billion-dollar industry whose production and consumption extend globally. Representation in games is an increasingly important topic. As those who create and consume the medium grow ever more diverse, it is essential that player or user-experience research, usability, and any consideration of how people interface with their technology is exercised through inclusive and intersectional lenses. Previous research has identified how character configuration interfaces preface white-male defaults [39, 40, 67]. This study relies on 1-on-1 play-interviews where diverse participants attempt to create “themselves” in a series of games and on group design activities to explore how participants may …
Are Academia And Industry Listening To Each Other? A Citation Analysis Of Ux Research Methods Resources,
2024
Minnesota State University, Mankato
Are Academia And Industry Listening To Each Other? A Citation Analysis Of Ux Research Methods Resources, Abigail Bakke, Leonard Dibono
English Department Publications
Technical and Professional Communication (TPC) has been facing concerns of viability, in both its relationship with industry and its ability to build a relevant and valid body of research. TPC’s disconnection with industry may be reflected in its relationship to UX as well, despite both fields’ shared values. To better understand how TPC and User Experience (UX) are relating to each other, we conducted a citation analysis of a sample of SIGDOC papers and a sample of Nielsen Norman Group (NN/g) practitioner articles focused on research methods. The SIGDOC papers tended to cite TPC sources, while the NN/g articles cited …
Dung Dkar Cloak: Exploring Soft Interfaces For Sonic Interactions,
2024
EJ Tech
Dung Dkar Cloak: Exploring Soft Interfaces For Sonic Interactions, Judit Eszter Kárpáti, Esteban De La Torre
Textile Society of America: Symposium Proceedings
The importance of crossmodal interaction within the contemporary cultural, technological and scientific panorama has evidently gained significant attention due to its remarkable advantages in creating a meaningful, interwoven, and integrated experience. The use and recontextualization of textiles in such exploratory quest into the human senses has proven to be critical. Computational science, algorithmic logic and digital devices have always been rooted and closely interwoven with textile crafts and practices. Recent technological advancements have further combined technology and textile, generating interactive textile surfaces, constructing endless possibilities for multisensorial experiences. In this presentation we will examine how we can weave a sensitive …
How Economically Marginalized Adolescents Of Color Negotiate Critical Pedagogy In A Computing Classroom,
2024
Carleton College
How Economically Marginalized Adolescents Of Color Negotiate Critical Pedagogy In A Computing Classroom, Jean Salac, Lena Armstrong, Megumi Kivuva, Jayne Everson, Amy J. Ko
Computer Science Faculty Work
Background and Context: With the growing movement to adopt critical framings of computing, scholars have worked to reframe computing education from the narrow development of programming skills to skills in identifying and resisting oppressive structures in computing. However, we have little guidance on how these framings may manifest in classroom practice. Objectives: To better understand the processes and practice of critical pedagogy in a computing classrooms, we taught a critically conscious computing elective within a summer academic program at a northwest United States university targeted at secondary students (ages 14–18) from low-income backgrounds and would be the first …
Advancing Policy Insights: Opinion Data Analysis And Discourse Structuring Using Llms,
2024
University of Central Florida
Advancing Policy Insights: Opinion Data Analysis And Discourse Structuring Using Llms, Aaditya Bhatia
Graduate Thesis and Dissertation 2023-2024
The growing volume of opinion data presents a significant challenge for policymakers striving to distill public sentiment into actionable decisions. This study aims to explore the capability of large language models (LLMs) to synthesize public opinion data into coherent policy recommendations. We specifically leverage Mistral 7B and Mixtral 8x7B models for text generation and have developed an architecture to process vast amounts of unstructured information, integrate diverse viewpoints, and extract actionable insights aligned with public opinion. Using a retrospective data analysis of the Polis platform debates published by the Computational Democracy Project, this study examines multiple datasets that span local …
Living Datasets: Towards Data-Centric Ai Explainability And Bias Mitigation,
2024
University of Texas at Arlington
Living Datasets: Towards Data-Centric Ai Explainability And Bias Mitigation, Akib Zaman
Computer Science and Engineering Dissertations - Archive
Benchmark datasets are critical to the evolution of AI efforts yet often embed unintended biases that influence the models that drive human-AI interactions. A deeper inspection and awareness of data is needed to understand the biases datasets may contain. In this dissertation, I introduce the Tag-and-Release method, inspired from wildlife research, that treats data as an organism and examines how different environments (i.e., CNNs) select for unique traits or characteristics that ultimately impact data's survival. Using the canonical MNIST handwritten digit dataset as a case study, I describe how the Tag-and-Release method can be used to analyze how dataset imbalance …
A Unified Cross-Modal Interactive System For Assisting Vision Impaired In Human Navigation And Indoor Based Human Robot Interaction,
2024
University of Texas at Arlington
A Unified Cross-Modal Interactive System For Assisting Vision Impaired In Human Navigation And Indoor Based Human Robot Interaction, Harish Ram Nambiappan
Computer Science and Engineering Dissertations - Archive
People who are blind and vision impaired often require assistance in performing various tasks. With new technologies emerging in the recent years, vision impaired people either require assistance in accessing those technologies or in using those technologies to perform different tasks in real life. Previous works have focused on assisting vision impaired people in different scenarios such as navigation, accessing smartphone interfaces etc. With the recent developments in robotics, a new research has emerged where new systems can be developed for vision impaired people to interact with robots to perform various human robot interactive tasks. But with developing new and …
Flexible Attenuation Fields: Tomographic Reconstruction From Heterogeneous Datasets,
2024
University of Kentucky
Flexible Attenuation Fields: Tomographic Reconstruction From Heterogeneous Datasets, Clifford S. Parker
Theses and Dissertations--Computer Science
Traditional reconstruction methods for X-ray computed tomography (CT) are highly constrained in the variety of input datasets they admit. Many of the imaging settings -- the incident energy, field-of-view, effective resolution -- remain fixed across projection images, and the only real variance is in the detector's position and orientation with respect to the scene. In contrast, methods for 3D reconstruction of natural scenes are extremely flexible to the geometric and photometric properties of the input datasets, readily accepting and benefiting from images captured under varying lighting conditions, with different cameras, and at disparate points in time and space. Extending CT …
Escape The Planet: Revolutionizing Game Design With Novel Oop Techniques,
2024
Minnesota State University, Mankato
Escape The Planet: Revolutionizing Game Design With Novel Oop Techniques, Qusai Kamal Fannoun
All Graduate Theses, Dissertations, and Other Capstone Projects
Mobile devices are continuously evolving and greater computing power and graphics capabilities are being introduced every year. As a result, there is an increasing demand for challenging and engaging mobile games that leverage these advanced features. This project explores best design practices using the development of Escape the Planet, which is an intricate maze game for mobile devices in which players navigate using a spaceship that is trapped in a hostile planet’s maze while avoiding obstacles and enemy attacks. The goal is to safely guide the spaceship out of the maze without colliding into walls or taking bullets from defensive …
Poster, Performed: Understanding Public Opinions Of Authorship In Generative Artificial Intelligence Models Via Analogy,
2024
Dartmouth College
Poster, Performed: Understanding Public Opinions Of Authorship In Generative Artificial Intelligence Models Via Analogy, Wylie Z. Kasai
Dartmouth College Master’s Theses
Over the last decade, generative artificial intelligence models have advanced significantly and provided the public with several tools to create new works of art. However, the true authorship of these works has been debated due to their training on web-scraped data. Serving as an analogy to these larger models, Poster, Performed is an interactive artificial intelligence exhibition project that uses image assets submitted by the public to create poster compositions with custom image processing algorithms. During the course of a four-day exhibition, visitors were asked to identify the exhibition’s primary artist from five options: (1) participants who submitted image assets, …
Predicting An Optimal Medication/Prescription Regimen For Patient Discordant Chronic Comorbidities Using Multi-Output Models,
2024
University of Dayton
Predicting An Optimal Medication/Prescription Regimen For Patient Discordant Chronic Comorbidities Using Multi-Output Models, Ichchha Pradeep Sharma, Tam Nguyen, Shruti Ajay Singh, Tom Ongwere
Computer Science Faculty Publications
This paper focuses on addressing the complex healthcare needs of patients struggling with discordant chronic comorbidities (DCCs). Managing these patients within the current healthcare system often proves to be a challenging process, characterized by evolving treatment needs necessitating multiple medical appointments and coordination among different clinical specialists. This makes it difficult for both patients and healthcare providers to set and prioritize medications and understand potential drug interactions. The primary motivation of this research is the need to reduce medication conflict and optimize medication regimens for individuals with DCCs. To achieve this, we allowed patients to specify their health conditions and …
