Open Access. Powered by Scholars. Published by Universities.®

Artificial Intelligence and Robotics Commons™

Open Access. Powered by Scholars. Published by Universities.®

11,188 Full-Text Articles 24,563 Authors 5,758,021 Downloads 274 Institutions

All Articles in Artificial Intelligence and Robotics

Faceted Search

11,188 full-text articles. Page 126 of 542.

Lova3 : Learning To Visual Question Answering, Asking And Assessment, Henry Hengyuan ZHAO, Pan ZHOU, Difei GAO, BAI SHOU, Mike Zheng SHOU 2024 Singapore Management University

Lova3 : Learning To Visual Question Answering, Asking And Assessment, Henry Hengyuan Zhao, Pan Zhou, Difei Gao, Bai Shou, Mike Zheng Shou

Research Collection School Of Computing and Information Systems

Question answering, asking, and assessment are three innate human traits crucial for understanding the world and acquiring knowledge. By enhancing these capabilities, humans can more effectively utilize data, leading to better comprehension and learning outcomes. Current Multimodal Large Language Models (MLLMs) primarily focus on question answering, often neglecting the full potential of questioning and assessment skills. Inspired by the human learning mechanism, we introduce LOVA3 , an innovative framework named “Learning tO Visual question Answering, Asking and Assessment,” designed to equip MLLMs with these additional capabilities. Our approach involves the creation of two supplementary training tasks GenQA and EvalQA, aiming …


Unified Generative And Discriminative Training For Multi-Modal Large Language Models, Wei CHOW, Juncheng LI, Kaihang PAN, Qifan YU, Hao FEI, Zhiqi GE, Shuai YANG, Siliang TENG, Hanwang ZHANG, Qianru SUN 2024 Singapore Management University

Unified Generative And Discriminative Training For Multi-Modal Large Language Models, Wei Chow, Juncheng Li, Kaihang Pan, Qifan Yu, Hao Fei, Zhiqi Ge, Shuai Yang, Siliang Teng, Hanwang Zhang, Qianru Sun

Research Collection School Of Computing and Information Systems

In recent times, Vision-Language Models (VLMs) have been trained under two predominant paradigms. Generative training has enabled Multimodal Large Language Models (MLLMs) to tackle various complex tasks, yet issues such as hallucinations and weak object discrimination persist. Discriminative training, exemplified by models like CLIP, excels in zero-shot image-text classification and retrieval, yet struggles with complex scenarios requiring fine-grained semantic differentiation. This paper addresses these challenges by proposing a unified approach that integrates the strengths of both paradigms. Considering interleaved image-text sequences as the general format of input samples, we introduce a structure-induced training strategy that imposes semantic relationships between input …


Automating Maritime Risk Data Collection And Identification Leveraging Large Language Models, Donghao HUANG, Xiuju FU, Xiaofeng YIN, Haibo PEN, Zhaoxia WANG 2024 Singapore Management University

Automating Maritime Risk Data Collection And Identification Leveraging Large Language Models, Donghao Huang, Xiuju Fu, Xiaofeng Yin, Haibo Pen, Zhaoxia Wang

Research Collection School Of Computing and Information Systems

Maritime risk research is crucial yet challenging for improving safety, efficiency, and sustainability in maritime operations. This paper presents an innovative method for automating the collection and identification of risk data related to global maritime risks from news sources, addressing the limitations of traditional manual methods. To evaluate the proposed method, different learning-based models, including conventional machine learning approaches and advanced Large Language Models (LLMs) such as GPT-4 and LLaMA-3.1, are comprehensively studied for comparison. In addition, not only do we use popular evaluation metrics to assess the proposed method, but we also introduce a new evaluation metric, called the …


Habit Coach: Customising Rag-Based Chatbots To Support Behavior Change, Arian Fooroogh Mand Arabi, Cansu Koyuturk, Michael O'Mahony, Raffaella Calati, Dimitri Ognibene 2024 Universita di Milano - Bicocca

Habit Coach: Customising Rag-Based Chatbots To Support Behavior Change, Arian Fooroogh Mand Arabi, Cansu Koyuturk, Michael O'Mahony, Raffaella Calati, Dimitri Ognibene

Conference papers

This paper presents the iterative development of Habit Coach, a GPT-based chatbot designed to support users in habit change through personalized interaction. Employing a user-centered design approach, we developed the chatbot using a Retrieval-Augmented Generation (RAG) system, which enables behavior personalization without retraining the underlying language model (GPT-4). The system leverages document retrieval and specialized prompts to tailor interactions, drawing from Cognitive Behavioral Therapy (CBT) and narrative therapy techniques. A key challenge in the development process was the difficulty of translating declarative knowledge into effective interaction behaviors. In the initial phase, the chatbot was provided with declarative knowledge about CBT …


A Systematic Review Of The Effects Of Ai-Assisted Moderation On Individuals And Groups, Zehui Yu, Lukas Otto, Dennis Assenmacher, Claudia Wagner 2024 GESIS - Leibniz Institute for the Social Sciences, RWTH Aachen

A Systematic Review Of The Effects Of Ai-Assisted Moderation On Individuals And Groups, Zehui Yu, Lukas Otto, Dennis Assenmacher, Claudia Wagner

Human-Machine Communication

This review paper provides a conceptualization of AI-assisted content moderation with various degrees of autonomy and summarizes experimental evidence for how different levels of automation in content moderation and related losses of autonomy affect individuals and groups. Our results show that current research predominantly focuses on individuallevel effects, necessitating a shift toward understanding the impact on groups. The study highlights gaps in exploring different levels of AI-assisted moderation interventions and misalignments of different conceptualizations that make comparing research results difficult. The discussion underscores the prevailing emphasis on harmful content removal and advocates for investigating more constructive moderation techniques, emphasizing the …


Designing Customized Loss Functions For Training Deep Neural Networks, Ali Pourramezan Fard 2024 University of Denver

Designing Customized Loss Functions For Training Deep Neural Networks, Ali Pourramezan Fard

Electronic Theses and Dissertations

This dissertation explores the critical role of loss functions in enhancing the predictive performance of deep machine learning models. Loss functions are an integral element of all the ongoing advances we witness daily in this domain. I design custom loss functions and their impacts on various machine learning tasks, particularly in computer vision.

In the first stage of my research, I aim to improve the prediction performance of deep learning models by providing them with more precise feedback associated with task requirements. This led me to create the concept of assistive loss functions. My first proposed loss function, inspired by …


An Overview Of Generative Ai Initiatives At Minnesota State University, Mankato (So Far), Evan Rusch, Nat Gustafson-Sundell 2024 Minnesota State University, Mankato

An Overview Of Generative Ai Initiatives At Minnesota State University, Mankato (So Far), Evan Rusch, Nat Gustafson-Sundell

Library Services Publications

At Minnesota State University, Mankato, we’ve undertaken several experiments and initiatives focused on Generative Artificial Intelligence. We provided several examples at the Generative AI in Libraries (GAIL) conference. For this presentation, we provided a revised and expanded overview of our initiatives for the Northern Ohio Technical Services Librarians (NOTSL) Fall General Meeting. We explained license-related restrictions on uses of AI. We discussed the limitations of the retrieval-augmented generation tools currently available in the library. We summarized how we’ve tested ChatBots to support licensing and we showed how we’ve tried to use AI to improve data visualization for collections outreach. We …


Autonomous Driving Trajectory Prediction, Carlos Funes 2024 University of Nevada, Las Vegas

Autonomous Driving Trajectory Prediction, Carlos Funes

Undergraduate Research Symposium Lightning Talks

Autonomous driving is undoubtedly one of the world's most revolutionary technologies, opening the door to a more secure traffic environment. This innovation has led to vehicles being able to drive by themselves without the necessity of a person behind the wheel, as well as cruise control, lane-keeping assist, and automatic emergency braking. Unfortunately, there is still plenty of work before autonomous driving becomes more popular among drivers. While at UNLV as an undergraduate student/research assistant, one of my goals is to learn how these technologies work to bring ideas into the automotive industry by refining solutions to problems within these …


It's Not As Bad As You Think: Detecting Ai-Generated Voices, Yong Qin Xu 2024 University of Nevada, Las Vegas

It's Not As Bad As You Think: Detecting Ai-Generated Voices, Yong Qin Xu

Undergraduate Research Symposium Lightning Talks

Advances in machine learning have opened up the world to a brand new frontier of fraudulent phone calls which the average person may not be in any way prepared for. From imitations of a loved one's voice to lifelike mimicry of human callers, telephone scams may become harder than ever to anticipate or prevent now that criminals have the help of AI on their side. This is why in my research paper, I aim to analyze and compare two existing methods of detecting the authenticity of human voice recordings in order to demonstrate and explain currently available technology that's capable …


Vision-Language Integration For Enhanced Locomotion Mode Prediction, Ehsan Ahmadi 2024 Louisiana State University and Agricultural and Mechanical College

Vision-Language Integration For Enhanced Locomotion Mode Prediction, Ehsan Ahmadi

LSU Master's Theses

Wearable exoskeletons offer significant potential in enhancing human mobility in industrial environments. However, their adaptability to dynamic, task-intensive settings presents challenges, especially in accurately predicting locomotion modes such as ladder climbing, stair navigation, low-space movement, and obstacle navigation. This research proposes a multimodal framework that integrates visual data and speech commands to improve locomotion mode prediction in unpredictable environments. Multimodal data was collected using smart glasses, capturing both the user’s perspective (field-of-view, FOV) and voice during locomotion tasks. State-of-the-art models—CLIP, ImageBind, and GPT-4o—process these visual and linguistic inputs to predict locomotion activities. The models were evaluated in zero-shot and fine-tuned …


Dynamic Knowledge Elicitation: Leveraging Student Feedback For Improved Language Model Distillation, Reuven Muller 2024 Kennesaw State University

Dynamic Knowledge Elicitation: Leveraging Student Feedback For Improved Language Model Distillation, Reuven Muller

Master's Theses

Large Language Models (LLMs) have significantly advanced the field of natural language processing but remain resource-intensive and impractical for many organizations. Specialist models offer a viable alternative, often developed through Knowledge Distillation (KD) techniques. However, traditional KD methods rely on predefined static datasets to elicit knowledge from the teacher model, failing to dynamically address the weaknesses of the student model during training. This research introduces two novel methods for adaptive knowledge elicitation: Feedback-Driven Question Generation and Agent-Based Targeted Question Generation. These methods iteratively expand the training dataset based on the student model’s performance, leveraging a teacher model to generate targeted …


Artificial Intelligence Foundation Model Risk Identification And Governance Model From Esg Perspective, Jincheng SHI, Guoyu WANG, Yingchun WANG 2024 AI Safety and Trustworthy Center, Shanghai Artificial Intelligence Laboratory, Shanghai 200232, China

Artificial Intelligence Foundation Model Risk Identification And Governance Model From Esg Perspective, Jincheng Shi, Guoyu Wang, Yingchun Wang

Bulletin of Chinese Academy of Sciences (Chinese Version)

The application ecology of artificial intelligence foundation model is rapidly expanding. The environment, society, and governance are facing new challenges and opportunities. Exploring the construction of a governance framework for the development risks of foundation model has important theoretical value and practical significance for promoting the healthy and sustainable development of artificial intelligence. Based on the theories of ESG and artificial intelligence governance, this study analyzes the development benefits and typical risks of foundation model from the perspective of ESG and then constructs a risk governance framework and implementation strategies for artificial intelligence foundation models. This study shows that a …


Participatory Ethical Regulations: Risk Challenges Of Artificial Intelligence Era And Construction Of Governance Logic, Chenggang ZHANG, Lu PAN 2024 Department of Sociology, School of Social Sciences, Tsinghua University, Beijing 100084, China; Chinese Association of Development Strategy Studies, Beijing 100190, China

Participatory Ethical Regulations: Risk Challenges Of Artificial Intelligence Era And Construction Of Governance Logic, Chenggang Zhang, Lu Pan

Bulletin of Chinese Academy of Sciences (Chinese Version)

Participatory ethical norms emphasize the involvement of diverse stakeholders, aiming to construct a more comprehensive and balanced ethical governance framework. The rapid development of artificial intelligence (AI) technology is leading society through unprecedented transformations, significantly impacting ethical perspectives, social governance models, and the symbiotic relationship between humans and technology. The participatory ethical norms, characterized by multi-stakeholder participation, interactivity, and openness, represent a crucial pathway for addressing the challenges posed by the rapid development of AI technology. Constructing an AI governance framework based on participatory ethical norms provides solutions for the sustainable, fair, and transparent development of AI from multiple aspects …


Ethical Risks And Challenges Of Chatgpt Applications In Education, Jingbo FAN, Hui LIANG 2024 School of Government, University of International Business and Economics, Beijing 100029, China

Ethical Risks And Challenges Of Chatgpt Applications In Education, Jingbo Fan, Hui Liang

Bulletin of Chinese Academy of Sciences (Chinese Version)

ChatGPT is a typical application in the field of natural language processing, with the potential to empower and revolutionize education. It can serve not only as a digital tutor for students but also as a virtual assistant for teachers, driving the transformation of student learning methods and teaching paradigms. Additionally, ChatGPT shows a wide range of applications in the research field. However, while bringing opportunities for educational development, ChatGPT also poses ethical risks and challenges to educational equity. Firstly, ChatGPT may exacerbate the digital divide, leading to unequal educational opportunities. Secondly, it presents risks such as knowledge alienation, algorithmic black-box …


Overview On Autonomous Machine Computing, Shaoshan LIU, Yiming GAN, Yinhe HAN 2024 Shenzhen Institute of Artificial Intelligence and Robotics for Society, Shenzhen 518129, China

Overview On Autonomous Machine Computing, Shaoshan Liu, Yiming Gan, Yinhe Han

Bulletin of Chinese Academy of Sciences (Chinese Version)

Autonomous machine computing, an innovative blend of algorithms, software, and cutting-edge computing hardware, is poised to be the next major paradigm shift in the global economy, following personal, mobile, and cloud computing. This study delves into the research and commercialization of the robotics industry, underscoring the critical importance of establishing a comprehensive autonomous machine computing ecosystem. This study argues that autonomous machine computing necessitates a complete ecosystem that encompasses applications, programming languages, and the foundational hardware architectures, and presents a comprehensive review of significant research contributions across these areas. Moreover, the study explores the synergy between autonomous machine computing and …


Enlightenment Of Us Nairr To Construction Of Artificial Intelligence Innovation Ecosystem In China, Tian JIANG, Li QIAN 2024 National Science Library, Chinese Academy of Sciences, Beijing 100190, China

Enlightenment Of Us Nairr To Construction Of Artificial Intelligence Innovation Ecosystem In China, Tian Jiang, Li Qian

Bulletin of Chinese Academy of Sciences (Chinese Version)

The National Artificial Intelligence Research Resource (NAIRR) of the United States is of great significance for addressing the new challenges faced by the innovation of artificial intelligence technology. For China, the experience of NAIRR provides valuable references in resource optimization allocation and efficient utilization, which helps us to overcome resource bottlenecks and promote the rapid development of artificial intelligence technology. This study analyzes the enlightenment of NAIRR to the innovation and development of artificial intelligence in China from two key aspects. The first is the innovation-driven elements, where the integration strategies of NAIRR in computing and storage power, data resource, …


First Amendment Roadblock? Regulating The Misuse Of Generative Ai Technologies: Impersonation And Appropriation Of Likeness Without Permission, Muhammad Rabiu 2024 Old Dominion University

First Amendment Roadblock? Regulating The Misuse Of Generative Ai Technologies: Impersonation And Appropriation Of Likeness Without Permission, Muhammad Rabiu

Cybersecurity Undergraduate Research Showcase

This paper initially explores the misuse of generative AI technologies, particularly their role in impersonating or appropriating individuals' likeness without consent. It will then analyze technical and legal mitigation strategies and propose recommendations to address this issue in light of the First Amendment’s constitutional Freedom of Speech provision.


Between Copyright And Computer Science: The Law And Ethics Of Generative Ai, Devin R. Desai, Mark Riedl 2024 Georgia Institute of Technology

Between Copyright And Computer Science: The Law And Ethics Of Generative Ai, Devin R. Desai, Mark Riedl

Northwestern Journal of Technology and Intellectual Property

Copyright and computer science continue to intersect and clash, but they can coexist. The advent of new technologies such as digitization of visual and aural creations, sharing technologies, search engines, social media offerings, and more, challenge copyright-based industries and reopen questions about the reach of copyright law. Breakthroughs in artificial intelligence research, especially Large Language Models that leverage copyrighted material as part of training, are the latest examples of the ongoing tension between copyright and computer science. The exuberance, rush-to-market, and edge problem cases created by a few misguided companies now raises challenges to core legal doctrines and may shift …


Machine Learning-Driven Process Analysis And Optimization In Solid-State Welding And Fusion-Based Additive Manufacturing, Radif Uddin Ahmed 2024 Louisiana Tech University

Machine Learning-Driven Process Analysis And Optimization In Solid-State Welding And Fusion-Based Additive Manufacturing, Radif Uddin Ahmed

Master's Theses

In the modern era of advanced manufacturing, optimizing process parameters is pivotal in ensuring the quality and reliability of sophisticated component fabrication. This study presents a novel, data-driven approach to parameter optimization in two cutting-edge manufacturing techniques: Friction Stir Welding (FSW) and Laser Powder Bed Fusion (LPBF). By leveraging machine learning methodologies, this research addresses the critical challenge of efficiently determining optimal process parameters, a task traditionally relying on time-consuming and resource-intensive trial-and-error methods. This study will lead to a robust data-driven framework for process analysis of more advanced manufacturing techniques like the Additive Friction Stir Deposition (AFSD) process. Friction …


Robotic Multi-Object Grasping From A Pile: Techniques And Algorithms For Enhanced Dexterity, Tianze Chen 2024 University of South Florida

Robotic Multi-Object Grasping From A Pile: Techniques And Algorithms For Enhanced Dexterity, Tianze Chen

USF Tampa Graduate Theses and Dissertations

As robots become increasingly integrated into real-world applications such as warehousing, fulfillment centers, and manufacturing, the need for efficient and adaptable robotic systems grows. One of the key challenges is enabling robots to grasp multiple objects simultaneously, as this significantly boosts the efficiency of tasks like batch picking, sorting, and object transferring, reducing both time and energy consumption. This dissertation presents a comprehensive multi-object grasping (MOG) pipeline that includes pre-grasp selection, end-pose selection, grasping synergy calculation, and a data-driven model for estimating the number of objects being grasped. Central to this work is the development of the Experience Forest structure, …


Digital Commons powered by bepress