Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- China Simulation Federation (3880)
- Singapore Management University (1897)
- Old Dominion University (641)
- San Jose State University (277)
- MBZUAI (233)
-
- City University of New York (CUNY) (184)
- Technological University Dublin (157)
- Air Force Institute of Technology (137)
- Chapman University (125)
- California Polytechnic State University, San Luis Obispo (116)
- Chinese Academy of Sciences (113)
- University of Arkansas, Fayetteville (102)
- Lindenwood University (97)
- Edith Cowan University (92)
- Embry-Riddle Aeronautical University (92)
- University of Nebraska - Lincoln (78)
- University of Kentucky (76)
- University of South Florida (71)
- University of Nevada, Las Vegas (63)
- Dartmouth College (62)
- Clemson University (60)
- University of Denver (59)
- University of Michigan Law School (57)
- Utah State University (57)
- The Texas Medical Center Library (54)
- Thomas Jefferson University (54)
- New Jersey Institute of Technology (53)
- University of Malaya (50)
- Purdue University (48)
- Missouri University of Science and Technology (47)
- Keyword
-
- Artificial intelligence (778)
- Machine learning (685)
- Deep learning (435)
- Artificial Intelligence (359)
- Machine Learning (359)
-
- AI (239)
- Deep Learning (201)
- Simulation (160)
- Computer vision (157)
- Reinforcement learning (140)
- Generative AI (134)
- Neural networks (128)
- Large language models (109)
- Natural language processing (108)
- Robotics (97)
- Natural Language Processing (90)
- ChatGPT (89)
- Path planning (89)
- Optimization (82)
- Large Language Models (77)
- Computer Vision (76)
- Classification (71)
- Neural network (67)
- Neural Networks (65)
- Virtual reality (64)
- Reinforcement Learning (63)
- Computer Science (59)
- Cybersecurity (59)
- Genetic algorithm (58)
- Algorithms (57)
- Publication Year
- Publication
-
- Journal of System Simulation (3880)
- Research Collection School Of Computing and Information Systems (1664)
- Master's Projects (248)
- Theses and Dissertations (183)
- Computer Science Faculty Publications (125)
-
- Bulletin of Chinese Academy of Sciences (Chinese Version) (113)
- Faculty Scholarship (108)
- Publications and Research (99)
- Computer Vision Faculty Publications (98)
- Master's Theses (96)
- Conference papers (92)
- Electrical & Computer Engineering Faculty Publications (90)
- Machine Learning Faculty Publications (86)
- Electronic Theses and Dissertations (85)
- Faculty Publications (77)
- Dissertations (70)
- Research outputs 2022 to 2026 (64)
- USF Tampa Graduate Theses and Dissertations (59)
- Dissertations and Theses Collection (Open Access) (57)
- Articles (54)
- Dissertations, Theses, and Capstone Projects (53)
- Theses and Dissertations--Computer Science (48)
- Natural Language Processing Faculty Publications (46)
- Teaching and Generative AI: Pedagogical Possibilities and Productive Tensions (46)
- Graduate Theses and Dissertations (44)
- Open Access Theses & Dissertations (42)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (40)
- Theses (40)
- Electrical & Computer Engineering Theses & Dissertations (39)
- Publications (39)
- Publication Type
- File Type
Articles 2341 - 2370 of 11180
Full-Text Articles in Computer Sciences
High-Fidelity Soh Prediction In Lithium-Ion Batteries Using Hybrid Ml Networks, Shafiyee Islam, Gon Namkoong
High-Fidelity Soh Prediction In Lithium-Ion Batteries Using Hybrid Ml Networks, Shafiyee Islam, Gon Namkoong
Electrical & Computer Engineering Faculty Publications
Accurate and efficient prediction of lithium-ion battery state of health (SOH) is critical for ensuring reliability in electric vehicles, grid storage, and aerospace systems. Traditional SOH estimation methods often struggle with nonlinear degradation behaviors and lack sensitivity to subtle electrochemical signals, limiting their real-world deployment. To address these challenges, this study examines hybrid deep learning models that integrate differential capacity (dQ/dV) analysis to enhance predictive accuracy. Four hybrid architectures - hybrid CNN-LSTM multihead, CNN extractor for LSTM, DNN-LSTM, and DNN Bi-LSTM - were developed and evaluated using the NASA randomized battery usage dataset, offering a realistic benchmark under diverse operational …
Leveraging Large Language Models To Create Learner Personas For Training Design, N. Lovett, M. Yang, K. Herman, B. Li, M. Sosonkina, W. Purwanto, P. Jiang, M. H. Wu
Leveraging Large Language Models To Create Learner Personas For Training Design, N. Lovett, M. Yang, K. Herman, B. Li, M. Sosonkina, W. Purwanto, P. Jiang, M. H. Wu
Electrical & Computer Engineering Faculty Publications
This case examines the innovative use of Large Language Models (LLMs) to generate learner personas for developing learner-centered cybersecurity training materials when direct access to initial learner data is not available. The team developed a nine-stage iterative process for creating and refining AI-generated personas to address this constraint, integrating ethical review, stakeholder feedback, and action research principles. The process expanded upon Kouprie and Visser’s (2009) empathic design framework to ensure cultural responsiveness and mitigate potential biases in LLM outputs. Through multiple refinement cycles, initial generic personas evolved into detailed, context-rich archetypes which informed the development of effective and context-responsive training …
A Novel Intelligent Thermal Feedback Framework For Electric Motor Protection In Embedded Robotic Systems, Mohamed Shili, Salah Hammedi, Hicham Chaoui, Khaled Nouri
A Novel Intelligent Thermal Feedback Framework For Electric Motor Protection In Embedded Robotic Systems, Mohamed Shili, Salah Hammedi, Hicham Chaoui, Khaled Nouri
Electrical & Computer Engineering Faculty Publications
As robotic systems advance in autonomy and sophistication while being used in uncertain environments, the challenge of building reliable and robust electric motors that are embedded into robotic systems has never been a more important engineering problem. Thermal distress caused by extended operation or excessive loading can negatively affect a motor’s performance and efficiency and lead to catastrophic hardware failure. This paper proposes a novel intelligent control framework that includes real-time thermal feedback for hybrid electric motors that are embedded into robotic systems. The framework relies on adaptive control techniques and lightweight machine learning techniques to estimate internal motor temperatures …
The Effect Of Facial Phenotypes On Differential Performance Of Facial Recognition, Evan R. Garrett
The Effect Of Facial Phenotypes On Differential Performance Of Facial Recognition, Evan R. Garrett
Graduate Theses, Dissertations, and Problem Reports (ETD)
Facial recognition technology is utilized in many facets of life. As the use has become more widespread these systems have improved in reliability and performance approaching the level of human accuracy. With these improvements the problem of bias still remains as a persistent problem. Efforts have been made to minimize the bias prevalent in the systems via studies into various demographic factors, creating training datasets that have a more uniform distribution of subjects, and other methods. As facial recognition is one of the most utilized forms of biometric recognition it is vital to analyze potential causes of bias to help …
Fuel Consumption Prediction Using Bayesian Neural Networks, Syarifah Diana Permai, Jurike V. Moniaga, Zener Sukra Lie, Ferry Jie
Fuel Consumption Prediction Using Bayesian Neural Networks, Syarifah Diana Permai, Jurike V. Moniaga, Zener Sukra Lie, Ferry Jie
Research outputs 2022 to 2026
Transportation is one of the necessities of life. Because humans need transportation to move from one location to another. Transportation requires fuel. On the other hand, fuel consumption is important and must be controlled. This is because fuel can come from both renewable and non-renewable energy sources, depending on the type and process of its formation. Several factors influence the fuel efficiency of a car, including the type of engine, vehicle weight, aerodynamics, driving habits, and other vehicle conditions. This research aims to predict car fuel consumption and identify the factors that affect fuel consumption. Several Machine Learning and Statistical …
Embodied Ai For Challenging Rearrangement Tasks In The Context Of Service And Assistive Robots, Mariia Khan
Embodied Ai For Challenging Rearrangement Tasks In The Context Of Service And Assistive Robots, Mariia Khan
Theses: Doctorates and Masters
Embodied AI explores intelligent agents that learn through interaction with their environment, aiming to replicate human-like learning processes. Achieving this requires agents capable of understanding a scene via various sensors, reasoning about their actions, and reacting accordingly. These abilities are necessary for service domestic robots to assist humans in their day-to-day activities. Embodied AI tasks can include but are not limited to: visual exploration, visual navigation, instruction following and embodied question answering, which typically consider static (unchanging) environments, where objects do not move over time. This thesis addresses one of the most challenging Embodied AI tasks – visual room rearrangement, …
The Future Of Ai Regulation In Drug Development: A Comparative Analysis, Gabriela Lenarczyk, Timo Minssen, Nicholson Price, Arti Rai
The Future Of Ai Regulation In Drug Development: A Comparative Analysis, Gabriela Lenarczyk, Timo Minssen, Nicholson Price, Arti Rai
Faculty Scholarship
As artificial intelligence (AI) transforms drug development, regulatory frameworks are evolving to oversee its implementation, particularly at the US Food and Drug Administration (FDA) and the European Medicines Agency (EMA). This paper makes three contributions to understanding emerging regulatory approaches. First, we offer a comparative analysis of how these agencies have responded to AI-driven advances, incorporating new US executive orders and the European Union (EU)’s AI Act. Second, we propose a novel analytical framework to understand regulatory divergence: the FDA’s flexible, dialog-driven model contrasts with the EMA’s structured, risk-tiered approach, reflecting broader institutional and political-economic differences. While the former encourages …
Information Retrieval In The Age Of Generative Ai: A Mismatch That Matters, Alex Zhang
Information Retrieval In The Age Of Generative Ai: A Mismatch That Matters, Alex Zhang
Faculty Scholarship
This short piece explores a widespread and yet underexamined or even overlooked misconception, that is, large language models (LLMs) function like traditional legal research databases. They do not. As a matter of fact, information retrieval from databases functions very differently from LLMs in terms of inputs, retrieval processes, and outputs. These differences have significant implications for transparency, traceability, and overall effectiveness in AI-driven legal research. Without intentional oversight and adaption, these changes could profoundly affect how we develop research skills and a cumulative knowledge base, both of which are essential skills for lifelong learning in the legal field.
This article …
Towards Human Explainable Digital Forensics: Generating Human Interpretable Evidence For Semantic Understanding In Manipulated Images And Text, Yuwei Chen
Electronic Theses & Dissertations (2024 - present)
Detecting and characterizing manipulations in digital media continues to pose a significant challenge within the field of digital forensics. Despite notable advancements, the discipline often remains in a reactive stance against emerging threats. Current state-of-the-art methods, typically evaluated within academic settings, fails to mirror the complexities of real-world disinformation scenarios. These methods generally prioritize high performance based on quantitative metrics, yet they demonstrate a considerable dependency on training data and lack adaptability to new novel attack signatures. With the rapid evolution of attack methodologies, the dependency on highly accurate models that do not generalize or adapt well to unseen threats …
Improving Generalizability In Image Manipulation Detection, Zhenfei Zhang
Improving Generalizability In Image Manipulation Detection, Zhenfei Zhang
Electronic Theses & Dissertations (2024 - present)
Image manipulation detection (IMD) aims to determine whether an image has been tampered with and to identify the manipulated regions. These capabilities have become increasingly important with the rapid advancement of media editing and generation technologies, such as Photoshop and generative AI methods, which underscore the need for robust tools for media authentication. Although current state-of-the-art (SoTA) methods achieve strong results on common manipulation types, such as splicing, copy-move, and removal, they often struggle to generalize to manipulation types not represented in the training data. Consequently, their real-world applicability remains limited, with performance degrading significantly in practical scenarios.
In this …
Artificial Intelligence And Procedural Due Process, Brandon L. Garrett
Artificial Intelligence And Procedural Due Process, Brandon L. Garrett
Faculty Scholarship
Artificial intelligence (AI) violates procedural due process rights if the government uses it to deprive people of life, liberty, and property without adequate notice or an opportunity to be heard. A wide range of government agencies deploy AI systems, including in courts, law enforcement, public benefits administration, and national security. If the government refuses to disclose the reasons why it denied a person bail, public benefits, or immigration status, serious due process concerns arise. If the government delegates such tasks to an AI system, the due process analysis does not change. One asks whether a person received adequate notice and …
Feel Bad To Discard A Fashion Product: How Ai Designers Influence Individuals' Sustainable Consumption, Ha Kyung Lee, Dooyoung Choi
Feel Bad To Discard A Fashion Product: How Ai Designers Influence Individuals' Sustainable Consumption, Ha Kyung Lee, Dooyoung Choi
Educational Leadership & Workforce Development Faculty Publications
This study explores how AI technology in fashion design influences consumers' sustainable consumption behaviors, focusing on emotional attachment to products. By comparing AI-generated and human-designed fashion items, the study examines how designer type impacts negative emotions about discarding products, mediated by emotional attachment. Results from two experimental studies reveal that designer type significantly affects negative emotions toward discarding human-designed items, but emotional attachment was not influenced by designer type in the first study. This lack of difference may be due to personal characteristics that moderate the effect. The second study found that individuals who perceive AI as human-like form stronger …
Towards Dynamic Learner State: Orchestrating Ai Agents And Workplace Performance Via The Model Context Protocol, Mohan Yang, Nolan Lovett, Belle Li, Zhen Hou
Towards Dynamic Learner State: Orchestrating Ai Agents And Workplace Performance Via The Model Context Protocol, Mohan Yang, Nolan Lovett, Belle Li, Zhen Hou
Educational Leadership & Workforce Development Faculty Publications
Current learning and development approaches often struggle to capture dynamic individual capabilities, particularly the skills they acquire informally every day on the job. This dynamic creates a significant gap between what traditional models think people know and their actual performance, leading to an incomplete and often outdated understanding of how ready the workforce truly is, which can hinder organizational adaptability in rapidly evolving environments. This paper proposes a novel dynamic learner-state ecosystem—an AI-driven solution designed to bridge this gap. Our approach leverages specialized AI agents, orchestrated via the Model Context Protocol (MCP), to continuously track and evolve an individual’s multi-dimensional …
Exploring The Impact Of Value Co-Creation Through Ai-Driven Chatbbots On Customer Repeat Purchases, Dooyoung Choi, Jaeha Lee
Exploring The Impact Of Value Co-Creation Through Ai-Driven Chatbbots On Customer Repeat Purchases, Dooyoung Choi, Jaeha Lee
Educational Leadership & Workforce Development Faculty Publications
Drawing on the Stimulus-Organism-Response (S-O-R) framework, this study explores how perceived value co-creation during chatbot interactions influences customer repeat purchase intentions through cognitive, emotional, and social responses to chatbots. A survey of 220 participants revealed that perceived value co-creation significantly affected repeat purchase intentions, with cognitive evaluations, emotional reactions, and social value serving as key mediators. However, the direct effect of value co-creation on purchase intentions was not significant. The findings suggest that while value co-creation enhances consumer engagement, repeat purchases occur only when consumers experience positive cognitive, emotional, and social outcomes. Therefore, it is crucial for retailers to incorporate …
Analysing Nontraditional Students' Chatgpt Interaction, Engagement, Self-Efficacy And Performance: A Mixed-Methods Approach, Mohan Yang, Shiyan Jiang, Belle Li, Kristin Herman, Tian Luo, Shanan Chappell Moots, Nolan Lovett
Analysing Nontraditional Students' Chatgpt Interaction, Engagement, Self-Efficacy And Performance: A Mixed-Methods Approach, Mohan Yang, Shiyan Jiang, Belle Li, Kristin Herman, Tian Luo, Shanan Chappell Moots, Nolan Lovett
STEMPS Faculty Publications
Generative artificial intelligence brings opportunities and unique challenges to nontraditional higher education students, stemming, in part, from the experience of the digital divide. Providing access and practice is critical to bridge this divide and equip students with needed digital competencies. This mixed-methods study investigated how nontraditional higher education students interact with ChatGPT in multiple courses and examined relationships between ChatGPT interactions, engagement, self-efficacy and performance. Data were collected from 73 undergraduate and graduate students through chat logs, course reflections and artefacts, surveys and interviews. ChatGPT interactions were analysed using four metrics: prompt number, depth of knowledge (DoK), prompt relevance and …
Financial Named Entity Recognition: How Far Can Llm Go?, Yi-Te Lu, Yintong Huo
Financial Named Entity Recognition: How Far Can Llm Go?, Yi-Te Lu, Yintong Huo
Research Collection School Of Computing and Information Systems
The surge of large language models (LLMs) has revolutionized the extraction and analysis of crucial information from a growing volume of financial statements, announcements, and business news. Recognition for named entities to construct structured data poses a significant challenge in analyzing financial documents and is a foundational task for intelligent financial analytics. However, how effective are these generic LLMs and their performance under various prompts are yet need a better understanding. To fill in the blank, we present a systematic evaluation of state-of-the-art LLMs and prompting methods in the financial Named Entity Recognition (NER) problem. Specifically, our experimental results highlight …
Pfedrag: A Personalized Federated Retrieval-Augmented Generation System With Depth-Adaptive Tiered Embedding Tuning, Hangyu He, Xin Yuan, Kai Wu, Ren Ping Liu, Wei Ni
Pfedrag: A Personalized Federated Retrieval-Augmented Generation System With Depth-Adaptive Tiered Embedding Tuning, Hangyu He, Xin Yuan, Kai Wu, Ren Ping Liu, Wei Ni
Research outputs 2022 to 2026
Large Language Models (LLMs) can undergo hallucinations in specialized domains, and standard Retrieval-Augmented Generation (RAG) often falters due to general-purpose embeddings ill-suited for domain-specific terminology. Though domain-specific fine-tuning enhances retrieval, centralizing data introduces privacy risks. The use of federated learning (FL) can alleviate this to some extent, but faces challenges of data heterogeneity, poor personalization, and expensive training data generation. We propose pFedRAG, a novel Personalized Federated RAG framework, which enables efficient collaborative fine-tuning of embedding models to address these challenges. The key contribution is a new Depth-Adaptive Tiered Embedding (DATE) architecture, which comprises a Global Shared Layer, combined using …
Toward Embodied Navigation Through Vision And Language, Muraleekrishna Gopinathan
Toward Embodied Navigation Through Vision And Language, Muraleekrishna Gopinathan
Theses: Doctorates and Masters
Embodied AI is a challenging but exciting field in which a robot learns to interact with human-living spaces to perform various tasks. This thesis studies the embodied navigation problem in which a robotic agent navigates in a previously unseen indoor environment based on a challenging task. In particular, the Vision-and-Language Navigation (VLN) task requires a robot to navigate based on a descriptive human-language instruction. This thesis aims to improve VLN agents on four key aspects - their understanding of the environment, training via additional data, correcting navigational errors, and predicting the layout of the environment for better planning.
First, we …
Defending Federated Recommender Systems Against Untargeted Attacks: A Contribution-Aware Robust Aggregation Scheme, Ruicheng Liang, Yuanchun Jiang, Feida Zhu, Ling Cheng, Huiwen Liu
Defending Federated Recommender Systems Against Untargeted Attacks: A Contribution-Aware Robust Aggregation Scheme, Ruicheng Liang, Yuanchun Jiang, Feida Zhu, Ling Cheng, Huiwen Liu
Research Collection School Of Computing and Information Systems
Federated recommender systems (FedRSs) effectively tackle the tradeoff between recommendation accuracy and privacy preservation. However, recent studies have revealed severe vulnerabilities in FedRSs, particularly against untargeted attacks seeking to undermine their overall performance. Defense methods employed in traditional recommender systems are not applicable to FedRSs, and existing robust aggregation schemes for other federated learning-based applications have proven ineffective in FedRSs. Building on the observation that malicious clients contribute negatively to the training process, we design a novel contribution-aware robust aggregation scheme to defend FedRSs against untargeted attacks, named contribution-aware Bayesian knowledge distillation aggregation (ConDA), comprising two key components for the …
Lara : A Light And Anti-Overfitting Retraining Approach For Unsupervised Time Series Anomaly Detection, Feiyi Chen, Zhen Qin, Mengchu Zhou, Yingying Zhang, Shuiguang Deng, Lunting Fan, Guansong Pang, Qingsong Wen
Lara : A Light And Anti-Overfitting Retraining Approach For Unsupervised Time Series Anomaly Detection, Feiyi Chen, Zhen Qin, Mengchu Zhou, Yingying Zhang, Shuiguang Deng, Lunting Fan, Guansong Pang, Qingsong Wen
Research Collection School Of Computing and Information Systems
Most of current anomaly detection models assume that the normal pattern remains the same all the time. However, the normal patterns of web services can change dramatically and frequently over time. The model trained on old-distribution data becomes outdated and ineffective after such changes. Retraining the whole model whenever the pattern is changed is computationally expensive. Further, at the beginning of normal pattern changes, there is not enough observation data from the new distribution. Retraining a large neural network model with limited data is vulnerable to overfitting. Thus, we propose a Light Anti-overfitting Retraining Approach (LARA) based on deep variational …
Fuzzing Drones For Anomaly Detection: A Systematic Literature Review, Vikas Kumar Malviya, Wei Minn, Lwin Khin Shar, Lingxiao Jiang
Fuzzing Drones For Anomaly Detection: A Systematic Literature Review, Vikas Kumar Malviya, Wei Minn, Lwin Khin Shar, Lingxiao Jiang
Research Collection School Of Computing and Information Systems
Drones, also referred to as Unmanned Aerial Vehicles (UAVs), are becoming popular today due to their uses in different fields and recent technological advancements which provide easy control of UAVs via mobile apps. However, UAVs may contain vulnerabilities or software bugs that cause serious safety and security concerns. For example, the communication protocol used by the UAV may contain authentication and authorization vulnerabilities, which may be exploited by attackers to gain remote access over the UAV. Drones must therefore undergo extensive testing before being released or deployed to identify and fix any software bugs or security vulnerabilities. Fuzzing is one …
Measuring Model Alignment For Code Clone Detection Using Causal Interpretation, Shamsa Abid, Xuemeng Cai, Lingxiao Jiang
Measuring Model Alignment For Code Clone Detection Using Causal Interpretation, Shamsa Abid, Xuemeng Cai, Lingxiao Jiang
Research Collection School Of Computing and Information Systems
Deep Neural Network-based models have demonstrated high accuracy for semantic code clone detection. However, the lack of generalization poses a threat to the trustworthiness and reliability of these models. Furthermore, the black-box nature of these models makes interpreting the model’s decisions very challenging. Currently, there is only a limited understanding of the semantic code clone detection behavior of existing models. There is a lack of transparency in understanding how a model identifies semantic code clones and the exact code components influencing its prediction. In this paper, we introduce the use of a causal interpretation framework based on the Neyman-Rubin causal …
A Review Of Chinese Sentiment Analysis: Subjects, Methods, And Trends, Zhaoxia Wang, Donghao Huang, Jingfeng Cui, Xinyue Zhang, Seng-Beng Ho, Erik Cambria
A Review Of Chinese Sentiment Analysis: Subjects, Methods, And Trends, Zhaoxia Wang, Donghao Huang, Jingfeng Cui, Xinyue Zhang, Seng-Beng Ho, Erik Cambria
Research Collection School Of Computing and Information Systems
Sentiment analysis has emerged as a prominent research domain within the realm of natural language processing, garnering increasing attention and a growing body of literature. While numerous literature reviews have examined sentiment analysis techniques, methods, topics and applications, there remains a gap in the literature concerning thematic trends and research methodologies in sentiment analysis, particularly in the context of Chinese text. This study addresses this gap by presenting a comprehensive survey dedicated to the progression of research subjects, methods and trends in sentiment analysis of Chinese text. Employing a framework that combines keyword co-occurrence analysis with a sophisticated community detection …
Llms-Based Augmentation For Domain Adaptation In Long-Tailed Food Datasets, Qing Wang, Chong-Wah Ngo, Ee-Peng Lim, Qianru Sun
Llms-Based Augmentation For Domain Adaptation In Long-Tailed Food Datasets, Qing Wang, Chong-Wah Ngo, Ee-Peng Lim, Qianru Sun
Research Collection School Of Computing and Information Systems
Training a model for food recognition is challenging because the training samples, which are typically crawled from the Internet, are visually different from the pictures captured by users in the free-living environment. In addition to this domain-shift problem, the real-world food datasets tend to be long-tailed distributed and some dishes of different categories exhibit subtle variations that are difficult to distinguish visually. In this paper, we present a framework empowered with large language models (LLMs) to address these challenges in food recognition. We first leverage LLMs to parse food images to generate food titles and ingredients. Then, we project the …
Interactive Video Search With Multi-Modal Llm Video Captioning, Yu-Tong Cheng, Jiaxin Wu, Zhixin Ma, Jiangshan He, Xiao-Yong Wei, Chong-Wah Ngo
Interactive Video Search With Multi-Modal Llm Video Captioning, Yu-Tong Cheng, Jiaxin Wu, Zhixin Ma, Jiangshan He, Xiao-Yong Wei, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
Cross-modal representation learning is essential for interactive text-to-video search tasks. However, the representation learning is limited by the size and quality of video-caption pairs. To improve the search accuracy, we propose to enlarge the size of available video-caption pairs by leveraging multi-model LLM on video captioning. Specifically, we use LLM to generate video captions for a large video collection (i.e., WebVid dataset) and use the generated video-caption pairs to pre-train a text-to-video search model. Additionally, we use LLM to generate fine-grained captions for test video collections to enable text-to-caption retrieval. Furthermore, we build a semantic overview of the retrieved rank …
Weakly-Supervised Semantic Segmentation With Image-Level Labels: From Traditional Models To Foundation Models, Zhaozheng Chen, Qianru Sun
Weakly-Supervised Semantic Segmentation With Image-Level Labels: From Traditional Models To Foundation Models, Zhaozheng Chen, Qianru Sun
Research Collection School Of Computing and Information Systems
The rapid development of deep learning has driven significant progress in image semantic segmentation—a fundamental task in computer vision. Semantic segmentation algorithms often depend on the availability of pixel-level labels (i.e., masks of objects), which are expensive, time consuming, and labor intensive. Weakly supervised semantic segmentation (WSSS) is an effective solution to avoid such labeling. It utilizes only partial or incomplete annotations and provides a cost-effective alternative to fully supervised semantic segmentation. In this article, our focus is on the WSSS with image-level labels, which is the most challenging form of WSSS. Our work has two parts. First, we conduct …
Synthesizing Multi-Person And Rare Pose Images For Human Pose Estimation, Liuqing Zhao, Zichen Tian, Zou Peng, Richang Hong, Qianru Sun
Synthesizing Multi-Person And Rare Pose Images For Human Pose Estimation, Liuqing Zhao, Zichen Tian, Zou Peng, Richang Hong, Qianru Sun
Research Collection School Of Computing and Information Systems
Human pose estimation (HPE) models underperform in recognizing rare poses because they suffer from data imbalance problems (i.e., there are few image samples for rare poses) in their training datasets. From a data perspective, the most intuitive solution is to synthesize data for rare poses. Specifically, the rule-based methods apply manual manipulations (such as Cutout and GridMask) to the existing data, so the limited diversity of the data constrains the model. An alternative method is to learn the underlying data distribution via deep generative models (such as ControlNet and HumanSD) and then sample “new data” from the distribution. This works …
Transforming Urban Dynamics: Harnessing Large Language Models For Smarter Mobility, Hao Xue, Ming Jin, Shirui Pan, Flora Salim, Guansong Pang
Transforming Urban Dynamics: Harnessing Large Language Models For Smarter Mobility, Hao Xue, Ming Jin, Shirui Pan, Flora Salim, Guansong Pang
Research Collection School Of Computing and Information Systems
Artificial intelligence (AI) has the potential to analyze mobility data and make mobility systems smarter by leveraging diverse data sources such as geospatial data, transportation logs, and real-time sensor data to optimize traffic flow, enhance public transportation systems, and support the development of autonomous vehicles. With the newly emerged generative AI paradigm, exemplified by large language models (LLMs), there is great potential to transform the current AI applications in mobility, transportation, and urban domains. This article provides an overview of recent efforts and aims to shed light on the challenges and future opportunities to facilitate the adaptation of LLMs for …
Fedart: A Neural Model Integrating Federated Learning And Adaptive Resonance Theory, Shubham Pateria, Budhitama Subagdja, Ah-Hwee Tan
Fedart: A Neural Model Integrating Federated Learning And Adaptive Resonance Theory, Shubham Pateria, Budhitama Subagdja, Ah-Hwee Tan
Research Collection School Of Computing and Information Systems
Federated Learning (FL) has emerged as a promising paradigm for collaborative model training across distributed clients while preserving data privacy. However, prevailing FL approaches aggregate the clients’ local models into a global model through multi-round iterative parameter averaging. This leads to the undesirable bias of the aggregated model towards certain clients in the presence of heterogeneous data distributions among the clients. Moreover, such approaches are restricted to supervised classification tasks and do not support unsupervised clustering. To address these limitations, we propose a novel one-shot FL approach called Federated Adaptive Resonance Theory (FedART) which leverages self-organizing Adaptive Resonance Theory (ART) …
A Survey Of Multilingual Large Language Models, Libo Qin, Qiguang Chen, Yuhang Zhou, Zhi Chen, Yinghui Li, Lizi Liao, Min Li, Wanxiang Che, Philip S. Yu
A Survey Of Multilingual Large Language Models, Libo Qin, Qiguang Chen, Yuhang Zhou, Zhi Chen, Yinghui Li, Lizi Liao, Min Li, Wanxiang Che, Philip S. Yu
Research Collection School Of Computing and Information Systems
Multilingual large language models (MLLMs) leverage advanced large language models to process and respond to queries across multiple languages, achieving significant success in polyglot tasks. Despite these breakthroughs, a comprehensive survey summarizing existing approaches and recent developments remains absent. To this end, this paper presents a unified and thorough review of the field, highlighting recent progress and emerging trends in MLLM research. The contributions of this paper are as follows. (1) Extensive survey: to our knowledge, this is the pioneering thorough review of multilingual alignment in MLLMs. (2) Unified taxonomy: we provide a unified framework to summarize the current progress …