Open Access. Powered by Scholars. Published by Universities.®

Artificial Intelligence and Robotics Commons™

Open Access. Powered by Scholars. Published by Universities.®

11,192 Full-Text Articles 24,576 Authors 5,758,021 Downloads 274 Institutions

All Articles in Artificial Intelligence and Robotics

Faceted Search

11,192 full-text articles. Page 62 of 543.

Learn To Fly: Enabling Deep Learning Based Perception And Control In Aerial Robotics, Krishna Muvva 2025 University of Nebraska-Lincoln

Learn To Fly: Enabling Deep Learning Based Perception And Control In Aerial Robotics, Krishna Muvva

School of Computing: Dissertations, Theses, and Student Research

Uncrewed Aerial Vehicles (UAVs) are increasingly deployed in dynamic, GPS degraded, and cluttered environments, yet their autonomy remains fundamentally constrained by limitations in onboard perception and real-time control. This dissertation addresses these challenges by proposing a unified framework that co-designs deep learning-based perception and model-based control, organized around three core thrusts: Learn to Track, Learn to Localize, and Learn to Evade.

Learn to Track develops dynamic and adaptive perception control mechanisms that optimize CNN inference for target tracking. A control-aware CNN framework dynamically adjusts inference frequency based on UAV motion, reducing latency while maintaining visual lock. An adaptive CNN with …


The Pastor As Romantic Author: Ai, Preaching, And The Unacknowledged Inheritance Of Authenticity, Daniel Plate, James Hutson 2025 Lindenwood University

The Pastor As Romantic Author: Ai, Preaching, And The Unacknowledged Inheritance Of Authenticity, Daniel Plate, James Hutson

Faculty Scholarship

This article interrogates contemporary reactions to sermons produced with generative technologies through a historical–conceptual lens, arguing that widespread judgments of such outputs as “soulless,” “generic,” or lacking a “beating heart” are best explained by an unacknowledged inheritance from nineteenth-century Romantic expressivism. Rather than treating resistance to machine authorship as a theological verdict on computational incapacity, the study reconstructs how Romanticism centered authorship in sincere self-expression and solitary genius, displacing earlier heraldic expectations that prized fidelity to a received message. Methodologically, the analysis combines intellectual history with discourse analysis of global Christian experiments in synthetic composition (2020–2025), denominational guidance, and media …


Csc36000 - Modern Distributed Computing Assignment, Saptarashmi Bandyopadhyay 2025 CUNY City College

Csc36000 - Modern Distributed Computing Assignment, Saptarashmi Bandyopadhyay

Open Educational Resources

This assignment covers standard performance metrics for Distributed Systems and the basics of Multiprocessing for CSC36000 - Modern Distributed Computing at the City College of New York CUNY. It is an interactive coding assignment intended to be executed in a Python notebook.


Building An Inclusive Ai Chatbot For Diverse Student Communities At Cal Poly: Uplift Ai, Gideon Telahun 2025 California Polytechnic State University, San Luis Obispo

Building An Inclusive Ai Chatbot For Diverse Student Communities At Cal Poly: Uplift Ai, Gideon Telahun

College of Engineering Summer Undergraduate Research Program

This research project will investigate the ability of advanced Large Language Models (LLMs) to identify and assess misinformation across diverse forms of media, including text, images, and video. In an age where misleading content spreads rapidly across digital platforms, evaluating the reliability and integrity of AI systems tasked with fact-checking is critical. We will develop a comprehensive dataset composed of factual and misleading examples drawn from various well-known and reliable fact-checking organizations. Each item will be independently reviewed and transparently labeled to ensure reproducibility. We will then prompt a curated group of state-of-the-art LLMs—including GPT-4, Claude, Gemini, Perplexity, Grok, and …


Leveraging Machine Learning And Causal Inference For Loan Default Prediction, Luca Guida 2025 Embry-Riddle Aeronautical University

Leveraging Machine Learning And Causal Inference For Loan Default Prediction, Luca Guida

Doctoral Dissertations and Master's Theses

This research explores a systematic application of machine learning techniques combined with causal inference to predict loan defaults in peer-to-peer lending. Accurately forecasting loan defaults is crucial for mitigating financial risk and optimizing lending strategies. This analysis is based on multiple datasets of loan applications spanning over a decade, containing detailed financial and credit information about borrowers. Beginning with extensive Exploratory Data Analysis (EDA) coupled with scaling strategies, the research identifies key trends in loan performance across a large number of factors, such as interest rates or borrower creditworthiness, and one objective is to determine from the many available predictors …


A Comprehensive Pipeline For Autonomous Surface Vehicle Navigation In Challenging Aquatic Environments, Mingi Jeong 2025 Dartmouth College

A Comprehensive Pipeline For Autonomous Surface Vehicle Navigation In Challenging Aquatic Environments, Mingi Jeong

Dartmouth College Ph.D Dissertations

The primary objective of this thesis is to develop and validate innovative and robust navigation methods for autonomous surface vehicles (ASVs) in challenging scenarios. These efforts aim to establish a complete autonomy pipeline for robotic decision-making systems, enabling high-level tasks such as environmental monitoring and autonomous transportation with broader impacts. The ocean economy contributes over 1.5 trillion USD annually, supporting diverse cultures and economies through tourism, fisheries, shipping, and renewable energy. The global marine industry handles over 90% of the world’s cargo transportation, underscoring its critical importance. Despite this significance, current maritime navigation relies heavily on human decision-making, which is …


Analysis Of The Status And Thematic Trends Of Ai For Science Research Abroad From 2015 To 2024, Fangyuan WANG, Huiting XU, Jinghua XUE 2025 Shanghai Library(Shanghai Institute of Scientific and Technical Information), Shanghai 200031

Analysis Of The Status And Thematic Trends Of Ai For Science Research Abroad From 2015 To 2024, Fangyuan Wang, Huiting Xu, Jinghua Xue

Journal of Scientific Information Research

[Purpose/significance] This paper analyzes the relevant literature in the field of AI for Science(AI4S)in the WoS core database from 2015 to 2024, and sorts out the research status and development trends in this field, aiming to provide forward-looking insights for the application of AI technology in scientific research.

[Method/process] This paper combines bibliometric analysis with the BERTopic model to analyze the publication trends, publishing countries, core authors, and topic identification and development trends in the field of AI4S.

[Result/conclusion] Through bibliometric analysis, this paper reveals the exponential growth trend of AI4S-related literature, and finds that China ranks first in the …


Trust And Ethics In Ai-Driven E-Commerce: Persuasion Vs. Privacy, Akriti Nepal 2025 Gettysburg College

Trust And Ethics In Ai-Driven E-Commerce: Persuasion Vs. Privacy, Akriti Nepal

Student Publications

This study examines how AI-driven features in e-commerce influence user satisfaction and the role of trust in these interactions. Using a survey-based dataset of 100 consumers, we investigated whether trust moderates or mediates the impact of AI persuasiveness and perceptions of bias, intrusiveness, and preference understanding on satisfaction. Results indicate that AI’s perceived ability to understand user preferences strongly predicts satisfaction, while trust partially mediates the relationship between helpful AI features and urgency messages and user satisfaction. Conversely, trust did not significantly moderate these relationships, and concerns about bias and intrusiveness had minimal impact. Findings suggest that AI-driven satisfaction is …


Harnessing Se Community Knowledge For Developer-Centric Code Intelligence, Chengran YANG 2025 Singapore Management University

Harnessing Se Community Knowledge For Developer-Centric Code Intelligence, Chengran Yang

Dissertations and Theses Collection (Open Access)

The integration of Large Language Models (LLMs), particularly those tailored for programming tasks—referred to as code LLMs—has created novel opportunities to enhance developer productivity. These advanced models automate routine and repetitive coding tasks, such as code generation and debugging, and enable faster prototyping and more efficient problem-solving. Despite these remarkable advantages, the current generation of code LLMs exhibits notable limitations that impact their practical effectiveness in real-world software engineering scenarios. These models frequently produce code that is inefficient or suboptimal in runtime performance, demonstrate opaque reasoning processes, and struggle to adapt effectively to diverse developer contexts and specific requirements. Moreover, …


Memory-Efficient 4-Bit Preconditioned Stochastic Optimization, Jingyang LI, Kuangyu DING, Kim-Chuan TOH, Pan ZHOU 2025 Singapore Management University

Memory-Efficient 4-Bit Preconditioned Stochastic Optimization, Jingyang Li, Kuangyu Ding, Kim-Chuan Toh, Pan Zhou

Research Collection School Of Computing and Information Systems

Preconditioned stochastic optimization algorithms, exemplified by Shampoo, outperform first-order optimizers by offering theoretical convergence benefits and practical gains in large-scale neural network training. However, they incur substantial memory overhead due to the storage demands of non-diagonal preconditioning matrices. To address this, we introduce 4-bit quantization for Shampoo’s preconditioners. We introduce two key methods: First, we apply Cholesky decomposition followed by quantization of the Cholesky factors, reducing memory usage by leveraging their lower triangular structure while better preserving spectral properties to minimize information loss. To our knowledge, this is the first quantization approach applied to Cholesky factors of preconditioners. Second, we …


What Students Really Think: Unpacking Ai Ethics In Educational Assessments Through A Triadic Framework, LIM MING SOON TRISTAN, GOTTIPATI Swapna, Michelle L. F. CHEONG 2025 Singapore Management University

What Students Really Think: Unpacking Ai Ethics In Educational Assessments Through A Triadic Framework, Lim Ming Soon Tristan, Gottipati Swapna, Michelle L. F. Cheong

Research Collection School Of Computing and Information Systems

The rise of AI in educational assessments has significantly enhanced efficiency and accuracy. However, it also introduces critical ethical challenges, including bias in grading, data privacy risks, and accountability gaps. These issues can undermine trust in AI-driven assessments and compromise educational fairness, making a structured ethical framework essential. To address these challenges, this study empirically validates an existing triadic ethical framework for AI-assisted educational assessments, originally proposed by Lim, Gottipati and Cheong (In: Keengwe (ed) Creative AI tools and ethical implications in teaching and learning, IGI Global, 2023), grounded in student perceptions. The framework encompasses three ethical domains—physical, cognitive, and …


Cookingdiffusion: Cooking Procedural Image Generation With Stable Diffusion, Yuan WANG, Bin ZHU, Yanbin HAO, Chong-wah NGO, Yi TAN, Xiang WANG 2025 Singapore Management University

Cookingdiffusion: Cooking Procedural Image Generation With Stable Diffusion, Yuan Wang, Bin Zhu, Yanbin Hao, Chong-Wah Ngo, Yi Tan, Xiang Wang

Research Collection School Of Computing and Information Systems

Recent advancements in text-to-image generation models have excelled in creating diverse and realistic images. This success extends to food imagery, where various conditional inputs like cooking styles, ingredients, and recipes are utilized. However, a yet-unexplored challenge is generating a sequence of procedural images based on cooking steps from a recipe. This could enhance the cooking experience with visual guidance and possibly lead to an intelligent cooking simulation system. To fill this gap, we introduce a novel task called cooking procedural image generation. This task is inherently demanding, as it strives to create photo-realistic images that align with cooking steps while …


Probabilistic Prototype Calibration Of Vision-Language Models For Generalized Few-Shot Semantic Segmentation, Jie LIU, Jiayi SHEN, Pan ZHOU, Jan-Jakob SONKE, Stratis GAVVES 2025 Singapore Management University

Probabilistic Prototype Calibration Of Vision-Language Models For Generalized Few-Shot Semantic Segmentation, Jie Liu, Jiayi Shen, Pan Zhou, Jan-Jakob Sonke, Stratis Gavves

Research Collection School Of Computing and Information Systems

Generalized Few-Shot Semantic Segmentation (GFSS) aims to extend a segmentation model to novel classes with only a few annotated examples while maintaining performance on base classes. Recently, pretrained vision-language models (VLMs) such as CLIP have been leveraged in GFSS to improve generalization on novel classes through multi-modal prototypes learning. However, existing prototype-based methods are inherently deterministic, limiting the adaptability of learned prototypes to diverse samples, particularly for novel classes with scarce annotations. To address this, we propose FewCLIP, a probabilistic prototype calibration framework over multi-modal prototypes from the pretrained CLIP, thus providing more adaptive prototype learning for GFSS. Specifically, FewCLIP …


Visual-Enhanced Multimodal Framework For Flexible Job Shop Scheduling Problem, Peng ZHAO, Zhiguang CAO, Di WANG, Wen SONG, Wei PANG, You ZHOU, Yuan JIANG 2025 Singapore Management University

Visual-Enhanced Multimodal Framework For Flexible Job Shop Scheduling Problem, Peng Zhao, Zhiguang Cao, Di Wang, Wen Song, Wei Pang, You Zhou, Yuan Jiang

Research Collection School Of Computing and Information Systems

Multimodal models leverage complementary information across modalities to enrich feature representations. While visual information shows potential in representing structure for some combinatorial optimization problems (COPs), its application to complex scheduling like the Flexible Job Shop Scheduling Problem (FJSP) remains underexplored. Current learning-based FJSP solvers predominantly rely on handcrafted state features. This dependence can lead to inconsistencies and may not fully capture the problem's intricate dynamics. Crucially, these methods overlook visual modalities. Visual representations offer a distinct advantage by inherently capturing the global topological structure and complex resource interactions within the FJSP state. Unlike localized handcrafted features, this holistic, structural view …


Mitigating Cross-Modal Representation Bias For Multicultural Image-To-Recipe Retrieval, Qing WANG, Chong-wah NGO, Yu CAO, Ee-peng LIM 2025 Singapore Management University

Mitigating Cross-Modal Representation Bias For Multicultural Image-To-Recipe Retrieval, Qing Wang, Chong-Wah Ngo, Yu Cao, Ee-Peng Lim

Research Collection School Of Computing and Information Systems

Existing approaches for image-to-recipe retrieval have the implicit assumption that a food image can fully capture the details textually documented in its recipe. However, a food image only reflects the visual outcome of a cooked dish and not the underlying cooking process. Consequently, learning cross-modal representations to bridge the modality gap between images and recipes tends to ignore subtle, recipe-specific details that are not visually apparent but are crucial for recipe retrieval. Specifically, the representations are biased to capture the dominant visual elements, resulting in difficulty in ranking similar recipes with subtle differences in use of ingredients and cooking methods. …


Boosting Chart-To-Code Generation In Mllm Via Dual Preference-Guided Refinement, Zhihan ZHANG, Yixin CAO, Lizi LIAO 2025 Singapore Management University

Boosting Chart-To-Code Generation In Mllm Via Dual Preference-Guided Refinement, Zhihan Zhang, Yixin Cao, Lizi Liao

Research Collection School Of Computing and Information Systems

Translating chart images into executable plotting scripts-referred to as the chart-to-code generation task-requires Multimodal Large Language Models (MLLMs) to perform fine-grained visual parsing, precise code synthesis, and robust cross-modal reasoning. However, this task is inherently under-constrained: multiple valid code implementations can produce the same visual chart, and evaluation must consider both code correctness and visual fidelity across diverse dimensions. This makes it difficult to learn accurate and generalizable mappings through standard supervised fine-tuning. To address these challenges, we propose a dual preference-guided refinement framework that combines a feedback-driven, dual-modality reward mechanism with iterative preference learning. Our approach introduces a structured …


Art4math: Handwritten Mathematical Expression Recognition Via Multimodal Sketch Grounding, Yang ZHOU, Jin WANG, Yuxiao ZHANG, Kaixiang HUANG, Guodong LU, Jingru YANG, Shengfeng HE 2025 Singapore Management University

Art4math: Handwritten Mathematical Expression Recognition Via Multimodal Sketch Grounding, Yang Zhou, Jin Wang, Yuxiao Zhang, Kaixiang Huang, Guodong Lu, Jingru Yang, Shengfeng He

Research Collection School Of Computing and Information Systems

Handwritten Mathematical Expression Recognition (HMER) remains a challenging task due to the structural complexity of mathematical notation and the ambiguity of handwritten symbols-e.g., ''ρ'' vs. ''p'' or ''B'' vs. ''β''. While stroke-based models offer disambiguation via temporal cues, most existing methods are constrained by coarse modality fusion and a lack of fine-grained cross-modal alignment, further hindered by limited annotated data. We introduce Art for Math (Art4Math), a novel framework that leverages the structural richness of human sketches to enhance HMER through fine-grained, modality-aware learning. Art4Math follows a two-stage training paradigm: Art Grounding (A-Grd) and Math Decoding (M-Dec). In A-Grd, the …


Language-Driven 3d Human Pose Estimation In Multi-Person Scenarios: A New Dataset And Approach, Tingrui SHEN, Bangzhen LIU, Zhirun FAN, Shiting ZHANG, Weifeng PAN, Sun FAN, Dan CAO, Shengfeng HE 2025 Singapore Management University

Language-Driven 3d Human Pose Estimation In Multi-Person Scenarios: A New Dataset And Approach, Tingrui Shen, Bangzhen Liu, Zhirun Fan, Shiting Zhang, Weifeng Pan, Sun Fan, Dan Cao, Shengfeng He

Research Collection School Of Computing and Information Systems

In an NBA game scenario, consider the challenge of locating and analyzing the 3D poses of players performing a user-specified action, such as attempting a shot. Traditional 3D human pose estimation (3DHPE) methods often fall short in such complex, multi-person scenes due to their lack of semantic integration and reliance on isolated pose data. To address these limitations, we introduce Language-Driven 3D Human Pose Estimation (L3DHPE), a novel approach that extends 3DHPE to general multi-person contexts by incorporating detailed language descriptions. We present Panoptic-L3D, the first dataset designed for L3DHPE, featuring 3,838 linguistic annotations for 1,476 individuals across 588 videos, …


Diffusionmat: Alpha Matting As Deterministic Sequential Refinement Learning, Yangyang XU, Shengfeng HE, Wenqi SHAO, Yong DU, Kwan-Yee K. WONG, Yu QIAO, Jun YU, Ping LUO 2025 Singapore Management University

Diffusionmat: Alpha Matting As Deterministic Sequential Refinement Learning, Yangyang Xu, Shengfeng He, Wenqi Shao, Yong Du, Kwan-Yee K. Wong, Yu Qiao, Jun Yu, Ping Luo

Research Collection School Of Computing and Information Systems

In this paper, we introduce DiffusionMat, a novel image matting framework that employs a diffusion model for the transition from coarse to refined alpha mattes. Diverging from conventional methods that utilize trimaps merely as loose guidance for alpha matte prediction, our approach treats image matting as a deterministic sequential refinement learning process. This process begins with the addition of noise to trimaps and iteratively denoises them using a pre-trained diffusion model, which incrementally guides the prediction towards a clean alpha matte. The key innovation of our framework is a correction module that adjusts the output at each denoising step, ensuring …


Fine-Grained Abnormality Prompt Learning For Zero-Shot Anomaly Detection, Jiawen ZHU, Yew‑Soon ONG, Chunhua SHEN, Guansong PANG 2025 Singapore Management University

Fine-Grained Abnormality Prompt Learning For Zero-Shot Anomaly Detection, Jiawen Zhu, Yew‑Soon Ong, Chunhua Shen, Guansong Pang

Research Collection School Of Computing and Information Systems

Current zero-shot anomaly detection (ZSAD) methods show remarkable success in prompting large pre-trained visionlanguage models to detect anomalies in a target dataset without using any dataset-specific training or demonstration. However, these methods often focus on crafting/learning prompts that capture only coarse-grained semantics of abnormality, e.g., high-level semantics like ‘damaged’, ‘imperfect’, or ‘defective’ objects. They therefore have limited capability in recognizing diverse abnormality details that deviate from these general abnormal patterns in various ways. To address this limitation, we propose FAPrompt, a novel framework designed to learn Fine-grained Abnormality Prompts for accurate ZSAD. To this end, a novel Compound Abnormality Prompt …


Digital Commons powered by bepress