Lova3 : Learning To Visual Question Answering, Asking And Assessment,
2024
Singapore Management University
Lova3 : Learning To Visual Question Answering, Asking And Assessment, Henry Hengyuan Zhao, Pan Zhou, Difei Gao, Bai Shou, Mike Zheng Shou
Research Collection School Of Computing and Information Systems
Question answering, asking, and assessment are three innate human traits crucial for understanding the world and acquiring knowledge. By enhancing these capabilities, humans can more effectively utilize data, leading to better comprehension and learning outcomes. Current Multimodal Large Language Models (MLLMs) primarily focus on question answering, often neglecting the full potential of questioning and assessment skills. Inspired by the human learning mechanism, we introduce LOVA3 , an innovative framework named “Learning tO Visual question Answering, Asking and Assessment,” designed to equip MLLMs with these additional capabilities. Our approach involves the creation of two supplementary training tasks GenQA and EvalQA, aiming …
Unified Generative And Discriminative Training For Multi-Modal Large Language Models,
2024
Singapore Management University
Unified Generative And Discriminative Training For Multi-Modal Large Language Models, Wei Chow, Juncheng Li, Kaihang Pan, Qifan Yu, Hao Fei, Zhiqi Ge, Shuai Yang, Siliang Teng, Hanwang Zhang, Qianru Sun
Research Collection School Of Computing and Information Systems
In recent times, Vision-Language Models (VLMs) have been trained under two predominant paradigms. Generative training has enabled Multimodal Large Language Models (MLLMs) to tackle various complex tasks, yet issues such as hallucinations and weak object discrimination persist. Discriminative training, exemplified by models like CLIP, excels in zero-shot image-text classification and retrieval, yet struggles with complex scenarios requiring fine-grained semantic differentiation. This paper addresses these challenges by proposing a unified approach that integrates the strengths of both paradigms. Considering interleaved image-text sequences as the general format of input samples, we introduce a structure-induced training strategy that imposes semantic relationships between input …
Automating Maritime Risk Data Collection And Identification Leveraging Large Language Models,
2024
Singapore Management University
Automating Maritime Risk Data Collection And Identification Leveraging Large Language Models, Donghao Huang, Xiuju Fu, Xiaofeng Yin, Haibo Pen, Zhaoxia Wang
Research Collection School Of Computing and Information Systems
Maritime risk research is crucial yet challenging for improving safety, efficiency, and sustainability in maritime operations. This paper presents an innovative method for automating the collection and identification of risk data related to global maritime risks from news sources, addressing the limitations of traditional manual methods. To evaluate the proposed method, different learning-based models, including conventional machine learning approaches and advanced Large Language Models (LLMs) such as GPT-4 and LLaMA-3.1, are comprehensively studied for comparison. In addition, not only do we use popular evaluation metrics to assess the proposed method, but we also introduce a new evaluation metric, called the …
Habit Coach: Customising Rag-Based Chatbots To Support Behavior Change,
2024
Universita di Milano - Bicocca
Habit Coach: Customising Rag-Based Chatbots To Support Behavior Change, Arian Fooroogh Mand Arabi, Cansu Koyuturk, Michael O'Mahony, Raffaella Calati, Dimitri Ognibene
Conference papers
This paper presents the iterative development of Habit Coach, a GPT-based chatbot designed to support users in habit change through personalized interaction. Employing a user-centered design approach, we developed the chatbot using a Retrieval-Augmented Generation (RAG) system, which enables behavior personalization without retraining the underlying language model (GPT-4). The system leverages document retrieval and specialized prompts to tailor interactions, drawing from Cognitive Behavioral Therapy (CBT) and narrative therapy techniques. A key challenge in the development process was the difficulty of translating declarative knowledge into effective interaction behaviors. In the initial phase, the chatbot was provided with declarative knowledge about CBT …
A Systematic Review Of The Effects Of Ai-Assisted Moderation On Individuals And Groups,
2024
GESIS - Leibniz Institute for the Social Sciences, RWTH Aachen
A Systematic Review Of The Effects Of Ai-Assisted Moderation On Individuals And Groups, Zehui Yu, Lukas Otto, Dennis Assenmacher, Claudia Wagner
Human-Machine Communication
This review paper provides a conceptualization of AI-assisted content moderation with various degrees of autonomy and summarizes experimental evidence for how different levels of automation in content moderation and related losses of autonomy affect individuals and groups. Our results show that current research predominantly focuses on individuallevel effects, necessitating a shift toward understanding the impact on groups. The study highlights gaps in exploring different levels of AI-assisted moderation interventions and misalignments of different conceptualizations that make comparing research results difficult. The discussion underscores the prevailing emphasis on harmful content removal and advocates for investigating more constructive moderation techniques, emphasizing the …
Designing Customized Loss Functions For Training Deep Neural Networks,
2024
University of Denver
Designing Customized Loss Functions For Training Deep Neural Networks, Ali Pourramezan Fard
Electronic Theses and Dissertations
This dissertation explores the critical role of loss functions in enhancing the predictive performance of deep machine learning models. Loss functions are an integral element of all the ongoing advances we witness daily in this domain. I design custom loss functions and their impacts on various machine learning tasks, particularly in computer vision.
In the first stage of my research, I aim to improve the prediction performance of deep learning models by providing them with more precise feedback associated with task requirements. This led me to create the concept of assistive loss functions. My first proposed loss function, inspired by …
An Overview Of Generative Ai Initiatives At Minnesota State University, Mankato (So Far),
2024
Minnesota State University, Mankato
An Overview Of Generative Ai Initiatives At Minnesota State University, Mankato (So Far), Evan Rusch, Nat Gustafson-Sundell
Library Services Publications
At Minnesota State University, Mankato, we’ve undertaken several experiments and initiatives focused on Generative Artificial Intelligence. We provided several examples at the Generative AI in Libraries (GAIL) conference. For this presentation, we provided a revised and expanded overview of our initiatives for the Northern Ohio Technical Services Librarians (NOTSL) Fall General Meeting. We explained license-related restrictions on uses of AI. We discussed the limitations of the retrieval-augmented generation tools currently available in the library. We summarized how we’ve tested ChatBots to support licensing and we showed how we’ve tried to use AI to improve data visualization for collections outreach. We …
Autonomous Driving Trajectory Prediction,
2024
University of Nevada, Las Vegas
Autonomous Driving Trajectory Prediction, Carlos Funes
Undergraduate Research Symposium Lightning Talks
Autonomous driving is undoubtedly one of the world's most revolutionary technologies, opening the door to a more secure traffic environment. This innovation has led to vehicles being able to drive by themselves without the necessity of a person behind the wheel, as well as cruise control, lane-keeping assist, and automatic emergency braking. Unfortunately, there is still plenty of work before autonomous driving becomes more popular among drivers. While at UNLV as an undergraduate student/research assistant, one of my goals is to learn how these technologies work to bring ideas into the automotive industry by refining solutions to problems within these …
It's Not As Bad As You Think: Detecting Ai-Generated Voices,
2024
University of Nevada, Las Vegas
It's Not As Bad As You Think: Detecting Ai-Generated Voices, Yong Qin Xu
Undergraduate Research Symposium Lightning Talks
Advances in machine learning have opened up the world to a brand new frontier of fraudulent phone calls which the average person may not be in any way prepared for. From imitations of a loved one's voice to lifelike mimicry of human callers, telephone scams may become harder than ever to anticipate or prevent now that criminals have the help of AI on their side. This is why in my research paper, I aim to analyze and compare two existing methods of detecting the authenticity of human voice recordings in order to demonstrate and explain currently available technology that's capable …
Vision-Language Integration For Enhanced Locomotion Mode Prediction,
2024
Louisiana State University and Agricultural and Mechanical College
Vision-Language Integration For Enhanced Locomotion Mode Prediction, Ehsan Ahmadi
LSU Master's Theses
Wearable exoskeletons offer significant potential in enhancing human mobility in industrial environments. However, their adaptability to dynamic, task-intensive settings presents challenges, especially in accurately predicting locomotion modes such as ladder climbing, stair navigation, low-space movement, and obstacle navigation. This research proposes a multimodal framework that integrates visual data and speech commands to improve locomotion mode prediction in unpredictable environments. Multimodal data was collected using smart glasses, capturing both the user’s perspective (field-of-view, FOV) and voice during locomotion tasks. State-of-the-art models—CLIP, ImageBind, and GPT-4o—process these visual and linguistic inputs to predict locomotion activities. The models were evaluated in zero-shot and fine-tuned …
Dynamic Knowledge Elicitation: Leveraging Student Feedback For Improved Language Model Distillation,
2024
Kennesaw State University
Dynamic Knowledge Elicitation: Leveraging Student Feedback For Improved Language Model Distillation, Reuven Muller
Master's Theses
Large Language Models (LLMs) have significantly advanced the field of natural language processing but remain resource-intensive and impractical for many organizations. Specialist models offer a viable alternative, often developed through Knowledge Distillation (KD) techniques. However, traditional KD methods rely on predefined static datasets to elicit knowledge from the teacher model, failing to dynamically address the weaknesses of the student model during training. This research introduces two novel methods for adaptive knowledge elicitation: Feedback-Driven Question Generation and Agent-Based Targeted Question Generation. These methods iteratively expand the training dataset based on the student model’s performance, leveraging a teacher model to generate targeted …
Artificial Intelligence Foundation Model Risk Identification And Governance Model From Esg Perspective,
2024
AI Safety and Trustworthy Center, Shanghai Artificial Intelligence Laboratory, Shanghai 200232, China
Artificial Intelligence Foundation Model Risk Identification And Governance Model From Esg Perspective, Jincheng Shi, Guoyu Wang, Yingchun Wang
Bulletin of Chinese Academy of Sciences (Chinese Version)
The application ecology of artificial intelligence foundation model is rapidly expanding. The environment, society, and governance are facing new challenges and opportunities. Exploring the construction of a governance framework for the development risks of foundation model has important theoretical value and practical significance for promoting the healthy and sustainable development of artificial intelligence. Based on the theories of ESG and artificial intelligence governance, this study analyzes the development benefits and typical risks of foundation model from the perspective of ESG and then constructs a risk governance framework and implementation strategies for artificial intelligence foundation models. This study shows that a …
Participatory Ethical Regulations: Risk Challenges Of Artificial Intelligence Era And Construction Of Governance Logic,
2024
Department of Sociology, School of Social Sciences, Tsinghua University, Beijing 100084, China; Chinese Association of Development Strategy Studies, Beijing 100190, China
Participatory Ethical Regulations: Risk Challenges Of Artificial Intelligence Era And Construction Of Governance Logic, Chenggang Zhang, Lu Pan
Bulletin of Chinese Academy of Sciences (Chinese Version)
Participatory ethical norms emphasize the involvement of diverse stakeholders, aiming to construct a more comprehensive and balanced ethical governance framework. The rapid development of artificial intelligence (AI) technology is leading society through unprecedented transformations, significantly impacting ethical perspectives, social governance models, and the symbiotic relationship between humans and technology. The participatory ethical norms, characterized by multi-stakeholder participation, interactivity, and openness, represent a crucial pathway for addressing the challenges posed by the rapid development of AI technology. Constructing an AI governance framework based on participatory ethical norms provides solutions for the sustainable, fair, and transparent development of AI from multiple aspects …
Ethical Risks And Challenges Of Chatgpt Applications In Education,
2024
School of Government, University of International Business and Economics, Beijing 100029, China
Ethical Risks And Challenges Of Chatgpt Applications In Education, Jingbo Fan, Hui Liang
Bulletin of Chinese Academy of Sciences (Chinese Version)
ChatGPT is a typical application in the field of natural language processing, with the potential to empower and revolutionize education. It can serve not only as a digital tutor for students but also as a virtual assistant for teachers, driving the transformation of student learning methods and teaching paradigms. Additionally, ChatGPT shows a wide range of applications in the research field. However, while bringing opportunities for educational development, ChatGPT also poses ethical risks and challenges to educational equity. Firstly, ChatGPT may exacerbate the digital divide, leading to unequal educational opportunities. Secondly, it presents risks such as knowledge alienation, algorithmic black-box …
Overview On Autonomous Machine Computing,
2024
Shenzhen Institute of Artificial Intelligence and Robotics for Society, Shenzhen 518129, China
Overview On Autonomous Machine Computing, Shaoshan Liu, Yiming Gan, Yinhe Han
Bulletin of Chinese Academy of Sciences (Chinese Version)
Autonomous machine computing, an innovative blend of algorithms, software, and cutting-edge computing hardware, is poised to be the next major paradigm shift in the global economy, following personal, mobile, and cloud computing. This study delves into the research and commercialization of the robotics industry, underscoring the critical importance of establishing a comprehensive autonomous machine computing ecosystem. This study argues that autonomous machine computing necessitates a complete ecosystem that encompasses applications, programming languages, and the foundational hardware architectures, and presents a comprehensive review of significant research contributions across these areas. Moreover, the study explores the synergy between autonomous machine computing and …
Enlightenment Of Us Nairr To Construction Of Artificial Intelligence Innovation Ecosystem In China,
2024
National Science Library, Chinese Academy of Sciences, Beijing 100190, China
Enlightenment Of Us Nairr To Construction Of Artificial Intelligence Innovation Ecosystem In China, Tian Jiang, Li Qian
Bulletin of Chinese Academy of Sciences (Chinese Version)
The National Artificial Intelligence Research Resource (NAIRR) of the United States is of great significance for addressing the new challenges faced by the innovation of artificial intelligence technology. For China, the experience of NAIRR provides valuable references in resource optimization allocation and efficient utilization, which helps us to overcome resource bottlenecks and promote the rapid development of artificial intelligence technology. This study analyzes the enlightenment of NAIRR to the innovation and development of artificial intelligence in China from two key aspects. The first is the innovation-driven elements, where the integration strategies of NAIRR in computing and storage power, data resource, …
First Amendment Roadblock? Regulating The Misuse Of Generative Ai Technologies: Impersonation And Appropriation Of Likeness Without Permission,
2024
Old Dominion University
First Amendment Roadblock? Regulating The Misuse Of Generative Ai Technologies: Impersonation And Appropriation Of Likeness Without Permission, Muhammad Rabiu
Cybersecurity Undergraduate Research Showcase
This paper initially explores the misuse of generative AI technologies, particularly their role in impersonating or appropriating individuals' likeness without consent. It will then analyze technical and legal mitigation strategies and propose recommendations to address this issue in light of the First Amendment’s constitutional Freedom of Speech provision.
Between Copyright And Computer Science: The Law And Ethics Of Generative Ai,
2024
Georgia Institute of Technology
Between Copyright And Computer Science: The Law And Ethics Of Generative Ai, Devin R. Desai, Mark Riedl
Northwestern Journal of Technology and Intellectual Property
Copyright and computer science continue to intersect and clash, but they can coexist. The advent of new technologies such as digitization of visual and aural creations, sharing technologies, search engines, social media offerings, and more, challenge copyright-based industries and reopen questions about the reach of copyright law. Breakthroughs in artificial intelligence research, especially Large Language Models that leverage copyrighted material as part of training, are the latest examples of the ongoing tension between copyright and computer science. The exuberance, rush-to-market, and edge problem cases created by a few misguided companies now raises challenges to core legal doctrines and may shift …
Machine Learning-Driven Process Analysis And Optimization In Solid-State Welding And Fusion-Based Additive Manufacturing,
2024
Louisiana Tech University
Machine Learning-Driven Process Analysis And Optimization In Solid-State Welding And Fusion-Based Additive Manufacturing, Radif Uddin Ahmed
Master's Theses
In the modern era of advanced manufacturing, optimizing process parameters is pivotal in ensuring the quality and reliability of sophisticated component fabrication. This study presents a novel, data-driven approach to parameter optimization in two cutting-edge manufacturing techniques: Friction Stir Welding (FSW) and Laser Powder Bed Fusion (LPBF). By leveraging machine learning methodologies, this research addresses the critical challenge of efficiently determining optimal process parameters, a task traditionally relying on time-consuming and resource-intensive trial-and-error methods. This study will lead to a robust data-driven framework for process analysis of more advanced manufacturing techniques like the Additive Friction Stir Deposition (AFSD) process. Friction …
Robotic Multi-Object Grasping From A Pile: Techniques And Algorithms For Enhanced Dexterity,
2024
University of South Florida
Robotic Multi-Object Grasping From A Pile: Techniques And Algorithms For Enhanced Dexterity, Tianze Chen
USF Tampa Graduate Theses and Dissertations
As robots become increasingly integrated into real-world applications such as warehousing, fulfillment centers, and manufacturing, the need for efficient and adaptable robotic systems grows. One of the key challenges is enabling robots to grasp multiple objects simultaneously, as this significantly boosts the efficiency of tasks like batch picking, sorting, and object transferring, reducing both time and energy consumption. This dissertation presents a comprehensive multi-object grasping (MOG) pipeline that includes pre-grasp selection, end-pose selection, grasping synergy calculation, and a data-driven model for estimating the number of objects being grasped. Central to this work is the development of the Experience Forest structure, …
