Open Access. Powered by Scholars. Published by Universities.®
Artificial Intelligence and Robotics Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- Singapore Management University (176)
- Technological University Dublin (28)
- Old Dominion University (20)
- University of Arkansas, Fayetteville (20)
- University of Dayton (16)
-
- City University of New York (CUNY) (14)
- California Polytechnic State University, San Luis Obispo (11)
- San Jose State University (11)
- Dartmouth College (8)
- University of Malaya (8)
- Clemson University (7)
- Embry-Riddle Aeronautical University (7)
- University of Texas at Arlington (6)
- University of Nebraska - Lincoln (5)
- Rochester Institute of Technology (4)
- University of Kentucky (4)
- Central Washington University (3)
- Michigan Technological University (3)
- Montclair State University (3)
- University of Denver (3)
- University of New Mexico (3)
- California State University, San Bernardino (2)
- Dakota State University (2)
- Fort Hays State University (2)
- Georgia Southern University (2)
- Illinois Math and Science Academy (2)
- LSU New Orleans (2)
- Missouri State University (2)
- New Jersey Institute of Technology (2)
- Southern Adventist University (2)
- Keyword
-
- Artificial intelligence (18)
- Computer vision (17)
- Machine Learning (14)
- Machine learning (13)
- Deep learning (12)
-
- Artificial Intelligence (11)
- AI (8)
- Robotics (7)
- Virtual reality (7)
- Visualization (7)
- Augmented reality (6)
- Eye tracking (6)
- HCI (6)
- Image classification (6)
- Automation (5)
- Codes (5)
- Deep Learning (5)
- Feature extraction (5)
- Mental workload (5)
- Personality (5)
- Reinforcement learning (5)
- Accessibility (4)
- Artificial Intelligence (AI) (4)
- Classification (4)
- Computer Science (4)
- Computer Vision (4)
- Human Computer Interaction (4)
- Large Language Models (4)
- Pattern recognition (4)
- Semantics (4)
- Publication Year
- Publication
-
- Research Collection School Of Computing and Information Systems (171)
- H-Workload 2017: Models and Applications (Works in Progress) (14)
- Conference papers (12)
- MAICS: The Modern Artificial Intelligence and Cognitive Science Conference (12)
- Graduate Theses and Dissertations (10)
-
- Publications and Research (10)
- Computer Science Faculty Publications (9)
- Computer Science and Computer Engineering Undergraduate Honors Theses (9)
- Master's Theses (7)
- Dartmouth College Master’s Theses (6)
- Master's Projects (6)
- All Dissertations (5)
- Student Works (2020-2029) (5)
- College of Engineering Summer Undergraduate Research Program (4)
- Dissertations and Theses Collection (Open Access) (4)
- Frameless (4)
- SWITCH (4)
- Theses and Dissertations--Computer Science (4)
- All Master's Theses (3)
- Department of Computer Science Faculty Scholarship and Creative Works (3)
- Dissertations, Master's Theses and Master's Reports (3)
- International Journal of Aviation, Aeronautics, and Aerospace (3)
- Student Works (2000-2009) (3)
- Theses and Dissertations (3)
- All Theses (2)
- College of Graduate Studies: Theses & Dissertations (2)
- Computer Science ETDs (2)
- Computer Science Working Papers (2)
- Computer Science and Engineering Dissertations - Archive (2)
- Dartmouth College Ph.D Dissertations (2)
- Publication Type
- File Type
Articles 31 - 60 of 414
Full-Text Articles in Artificial Intelligence and Robotics
Dreamcs: Geometry-Aware Text-To-3d Generation With Unpaired 3d Reward Supervision, Xiandong Zou, Ruihao Xia, Hongsong Wang, Pan Zhou
Dreamcs: Geometry-Aware Text-To-3d Generation With Unpaired 3d Reward Supervision, Xiandong Zou, Ruihao Xia, Hongsong Wang, Pan Zhou
Research Collection School Of Computing and Information Systems
While text-to-3D generation has attracted growing interest, existing methods often struggle to produce 3D assets that align well with human preferences. Current preference alignment techniques for 3D content typically rely on hardly-collected preference-paired multi-view 2D images to train 2D reward models, when then guide 3D generation — leading to geometric artifacts, such as the Janus face problem and geometric incompleteness, due to their inherent 2D bias. To address these limitations, we construct 3D-MeshPref, the first large-scale unpaired 3D preference dataset, featuring diverse 3D meshes annotated by a large language model and refined by human evaluators. We then develop RewardCS, the …
From Spatial To Actions: Grounding Vision-Language-Action Model In Spatial Foundation Priors, Zhengshen Zhang, Hao Li, Yalun Dai, Zhengbang Zhu, Lei Zhou, Chenchen Liu, Dong Wang, Francis E. H. Tay, Sijin Chen, Ziwei Liu, Yuxiao Liu, Xinghang Li, Pan Zhou
From Spatial To Actions: Grounding Vision-Language-Action Model In Spatial Foundation Priors, Zhengshen Zhang, Hao Li, Yalun Dai, Zhengbang Zhu, Lei Zhou, Chenchen Liu, Dong Wang, Francis E. H. Tay, Sijin Chen, Ziwei Liu, Yuxiao Liu, Xinghang Li, Pan Zhou
Research Collection School Of Computing and Information Systems
Existing vision-language-action (VLA) models act in 3D real-world but are typically built on 2D encoders, leaving a spatial reasoning gap that limits generalization and adaptability. Recent 3D integration techniques for VLAs either require specialized sensors and transfer poorly across modalities, or inject weak cues that lack geometry and degrade vision-language alignment. In this work, we introduce FALCON (From Spatial to Action), a novel paradigm that injects rich 3D spatial tokens into the action head. FALCON leverages spatial foundation models to deliver strong geometric priors from RGB alone, and includes an Embodied Spatial Model that can optionally fuse depth, or pose …
Who You Explain To Matters: Learning By Explaining To Conversational Agents With Different Pedagogical Roles, Zhengtao Xu, Junti Zhang, Anthony Tang, Yi-Chieh Lee
Who You Explain To Matters: Learning By Explaining To Conversational Agents With Different Pedagogical Roles, Zhengtao Xu, Junti Zhang, Anthony Tang, Yi-Chieh Lee
Research Collection School Of Computing and Information Systems
Conversational agents are increasingly used in education for learning support. An application is “learning by explaining”, where learners explain their understanding to an agent. However, existing research focuses on single roles, leaving it unclear how different pedagogical roles influence learners’ interaction patterns, learning outcomes and experiences. We conducted a between-subjects study (N=96) comparing agents with three pedagogical roles (Tutee, Peer, Challenger) and a control condition while learning an economics concept. We found that different pedagogical roles shaped learning dynamics, including interaction patterns and experiences. Specifically, the Tutee agent elicited the most cognitive investment but led to high pressure. The Peer …
Weakly Supervised Video Anomaly Detection And Localization With Spatio-Temporal Prompts, Peng Wu, Xuerong Zhou, Guansong Pang, Zhiwei Yang, Qingsen Yan, Peng Wang, Yanning Zhang
Weakly Supervised Video Anomaly Detection And Localization With Spatio-Temporal Prompts, Peng Wu, Xuerong Zhou, Guansong Pang, Zhiwei Yang, Qingsen Yan, Peng Wang, Yanning Zhang
Research Collection School Of Computing and Information Systems
Current weakly supervised video anomaly detection (WSVAD) task aims to achieve frame-level anomalous event detection with only coarse video-level annotations available. Existing works typically involve extracting global features from full-resolution video frames and training frame-level classifiers to detect anomalies in the temporal dimension. However, most anomalous events tend to occur in localized spatial regions rather than the entire video frames, which implies existing frame-level feature based works may be misled by the dominant background information and lack the interpretation of the detected anomalies. To address this dilemma, this paper introduces a novel method called STPrompt that learns spatio-temporal prompt embeddings …
Portrait Shadow Removal Via Self-Exemplar Illumination Equalization, Qian Huang, Cheng Xu, Guiqing Li, Ziheng Wu, Shengxin Liu, Shengfeng He
Portrait Shadow Removal Via Self-Exemplar Illumination Equalization, Qian Huang, Cheng Xu, Guiqing Li, Ziheng Wu, Shengxin Liu, Shengfeng He
Research Collection School Of Computing and Information Systems
We introduce the Self-Exemplar Illumination Equalization Network, designed specifically for effective portrait shadow removal. The core idea of our method is that partially shadowed portraits can find ideal exemplars within their non-shadowed facial regions. Rather than directly fusing two distinct classes of facial features, our approach utilizes non-shadowed regions as an illumination indicator to equalize the shadowed regions, generating deshadowed results without boundary-merging artifacts. Our network comprises cascaded Self-Exemplar Illumination Equalization Blocks (SExmBlock), each containing two modules: a self-exemplar feature matching module and a feature-level illumination rectification module. The former identifies and applies internal illumination exemplars to shadowed areas, producing …
Lagrangian Motion Fields For Long-Term Motion Generation, Yifei Yang, Zikai Huang, Chenshu Xu, Shengfeng He
Lagrangian Motion Fields For Long-Term Motion Generation, Yifei Yang, Zikai Huang, Chenshu Xu, Shengfeng He
Research Collection School Of Computing and Information Systems
Long-term motion generation is a challenging task that requires producing coherent and realistic sequences over extended durations. Current methods primarily rely on framewise motion representations, which capture only static spatial details and overlook temporal dynamics. This approach leads to significant redundancy across the temporal dimension, complicating the generation of effective long-term motion. To overcome these limitations, we introduce the novel concept of Lagrangian Motion Fields, specifically designed for long-term motion generation. By treating each joint as a Lagrangian particle with uniform velocity over short intervals, our approach condenses motion representations into a series of "supermotions" (analogous to superpixels). This method …
Shaping The Future: Emerging Technologies And Their Role In Industry 4.0 And Beyond, Liuliu Qin
Shaping The Future: Emerging Technologies And Their Role In Industry 4.0 And Beyond, Liuliu Qin
Information Technology & Decision Sciences Faculty Publications
This paper provides a comprehensive review of emerging technologies driving the transition from Industry 4.0 to Industry 5.0. It examines the foundational concepts and pillars of Industry 4.0 and explores the transformative roles of Artificial Intelligence (AI), Extended Reality (XR), Collaborative Cobots (Cobots), Brain–Computer Interfaces (BCIs), quantum technologies, and next-generation connectivity (5G/6G). By integrating technological, human-centric, and sustainability perspectives, the study outlines how these emerging technologies reshape industrial systems and enable intelligent, adaptive, and inclusive futures.
Trustworthy Multimodal Ai For Medical Imaging: Enhancing Diagnosis, Reasoning, And Human-Agent Interaction In Extended Reality, Jai Prakash Veerla
Trustworthy Multimodal Ai For Medical Imaging: Enhancing Diagnosis, Reasoning, And Human-Agent Interaction In Extended Reality, Jai Prakash Veerla
Computer Science and Engineering Dissertations
The transition from traditional microscopy to digital pathology has digitized diagnostic data, yet clinical workflows remain constrained by two-dimensional screens and passive, opaque analysis tools that fail to capture the spatial complexity of biological systems. While Foundation Models now promise to reason across histology and genomics, a critical disconnect persists between the richness of this data and the limited cognitive bandwidth of clinicians, who currently lack the immersive interfaces and trustworthy agents necessary to utilize it effectively. This dissertation presents a unified framework for "Embodied Agentic AI," establishing a pipeline that augments physician capabilities through immersive visualization, robust security, and …
A Novel Lightweight Framework For Low-Light Image Enhancement Via Gaussian Denoising And Clahe, Daniel Oluwaseun Adesoji
A Novel Lightweight Framework For Low-Light Image Enhancement Via Gaussian Denoising And Clahe, Daniel Oluwaseun Adesoji
Master's Theses or Doctor of Nursing Practice
Low-light image enhancement is a major challenge in digital imaging, especially in medical imaging, surveillance, and autonomous vision systems. Images captured under poor illumination often appear dark, noisy, and low in contrast, which makes it hard to observe important details. Traditional enhancement methods can improve brightness but usually introduce artifacts or increase noise. Although deep learning methods have shown strong performance, they usually require large datasets, and high computational resources. This creates a need for simpler and more efficient enhancement techniques. This study proposes a lightweight framework that incorporates Gaussian denoising with Contrast Limited Adaptive Histogram Equalization (CLAHE) to enhance …
Learning Design To Advance Human-Ai Collaboration In K-12 Education, Wing Sha Chan, Jinhee Kim, Seongryeong Yu, Rita Kay Detrick
Learning Design To Advance Human-Ai Collaboration In K-12 Education, Wing Sha Chan, Jinhee Kim, Seongryeong Yu, Rita Kay Detrick
STEMPS Faculty Publications
This chapter explores key components for designing effective Human-AI Collaboration (HAC) in K–12 education, addressing the current lack of theoretical and conceptual frameworks for structuring and implementing HAC in teaching and learning. It examines four essential areas: curriculum design, student and teacher–AI interaction, learning environments, and the evolution of HAC over time. The chapter introduces the concept of HAC in K–12 contexts, highlighting how humans and AI can leverage each other's strengths through co-evolutionary processes that foster mutual learning and collaboration. It reviews current HAC practices in schools and discusses their contributions to both teaching and learning. Finally, it presents …
Designing Narrative-Based Ai Assistance For Sensemaking In Collaborative Environments: Case Studies In Education And Dementia Care, Dylan Edward Moore
Designing Narrative-Based Ai Assistance For Sensemaking In Collaborative Environments: Case Studies In Education And Dementia Care, Dylan Edward Moore
Dartmouth College Ph.D Dissertations
This thesis addresses a gap in the human-computer interaction literature regarding the design, development, and evaluation of narrative-based AI assistance for collaborative, complex problem solving. I explore this design space through three case studies across the domains of education and dementia care. This work encompasses multi-year industry partnerships and longitudinal fieldwork, user-centered design, dataset curation, model training, and system evaluation.
Specifically, the first case study considers a story-based web platform for teaching AI literacy through peer-generated, personalized narrative scaffolding. Learners on the platform showed significant knowledge gains and other learning-related outcomes. To describe the novel design of this system, I …
Improving Medical Diagnostics With Vision-Language Models: Convex Hull-Based Uncertainty Analysis, Ferhat Ozgur Catak, Murat Kuzlu, Taylor Patrick, Michel Audette
Improving Medical Diagnostics With Vision-Language Models: Convex Hull-Based Uncertainty Analysis, Ferhat Ozgur Catak, Murat Kuzlu, Taylor Patrick, Michel Audette
Engineering Technology Faculty Publications
In recent years, vision-language models (VLMs) have been applied to various fields, including healthcare, education, finance, and manufacturing, with remarkable performance. However, concerns remain regarding VLMs' consistency and uncertainty, particularly in critical applications such as healthcare, which demand a high level of trust and reliability. This paper proposes a novel approach to evaluate uncertainty in VLMs' responses using a convex hull approach on a healthcare application for visual question answering (VQA). For any VLM, temperature refers to a sampling parameter used in probabilistic generation, which controls the randomness of the model's output. The LLM-CXR model is selected as the medical …
Swimming In Uncertainty: Filling Data Gaps And Providing An Educational Platform For Beach Water Quality At Tybee Island, Georgia, Lukas Roberson
Swimming In Uncertainty: Filling Data Gaps And Providing An Educational Platform For Beach Water Quality At Tybee Island, Georgia, Lukas Roberson
College of Graduate Studies: Theses & Dissertations
@font-face {font-family:"Cambria Math"; panose-1:2 4 5 3 5 4 6 3 2 4; mso-font-charset:0; mso-generic-font-family:roman; mso-font-pitch:variable; mso-font-signature:-536870145 1107305727 0 0 415 0;}p.MsoNormal, li.MsoNormal, div.MsoNormal {mso-style-unhide:no; mso-style-qformat:yes; mso-style-parent:""; margin:0in; mso-pagination:widow-orphan; font-size:12.0pt; font-family:"Times New Roman",serif; mso-fareast-font-family:"Times New Roman";}.MsoChpDefault {mso-style-type:export-only; mso-default-props:yes; mso-font-kerning:0pt; mso-ligatures:none;}div.WordSection1 {page:WordSection1;}
Swimming in beaches water contaminated with high levels of bacteria can make you sick. Current monitoring at the public beaches on Tybee Island consists of weekly monitoring and enumeration of fecal indicator bacteria that takes 24 hours for results. If the number of bacteria exceed regulatory limits, a public health advisory is issued, and affected waters are retested until …
A Web-Based Wizard-Of-Oz Platform For Collaborative And Reproducible Human-Robot Interaction Research, Sean O'Connor
A Web-Based Wizard-Of-Oz Platform For Collaborative And Reproducible Human-Robot Interaction Research, Sean O'Connor
Honors Theses
The Wizard-of-Oz (WoZ) technique is widely used in Human-Robot Interaction (HRI) research, but two persistent problems limit its effectiveness: existing tools impose technical barriers that exclude non-engineering domain experts (the Accessibility Problem), and the fragmented landscape of robot-specific implementations makes interaction scripts difficult to port across platforms (the Reproducibility Problem- concerning execution consistency and portability, not third-party replication). Through a literature review, I identified three design principles to address both: a hierarchical specification model, an event-driven execution model, and a plugin architecture that decouples experiment logic from robot-specific implementations. I realized these principles in HRIStudio, an open-source, web-based platform providing …
Mg-Spair: Multi-Grade Sparse-Guided Implicit Representation For Training-Data-Free Image Restoration, Jianmin Liao, Lei Huang, Ronglong Fang, Ashley Prater-Bennette, Lixin Shen, Yuesheng Xu
Mg-Spair: Multi-Grade Sparse-Guided Implicit Representation For Training-Data-Free Image Restoration, Jianmin Liao, Lei Huang, Ronglong Fang, Ashley Prater-Bennette, Lixin Shen, Yuesheng Xu
Mathematics & Statistics Faculty Publications
MG-SpaIR is a training-data-free framework for restoring a clean image from a single observation corrupted by a mixture of blur, downsampling, noise, and missing pixels. Building on implicit neural representations (INRs), we introduce a multi-grade residual hierarchy that progressively refines the reconstruction from low to high spatial frequencies across grades, improving representational fidelity and mitigating spectral limitations. To stabilize reconstruction optimization and suppress INR-induced artifacts, we further propose an explicit sparse proximal regularization (e.g., ℓ0 type) applied directly in the high-resolution image domain, which discourages spurious high-frequency patterns while preserving sharp structures. The resulting optimization is solved efficiently via a …
Cognitive Prosthetic: An Ai-Enabled Multimodal System For Episodic Recall In Knowledge Work, Lawrence Obiuwevwi, Krzystof J. Rechowicz, Vikas Ashok, Sachin Shetty, Sampath Jayarathna
Cognitive Prosthetic: An Ai-Enabled Multimodal System For Episodic Recall In Knowledge Work, Lawrence Obiuwevwi, Krzystof J. Rechowicz, Vikas Ashok, Sachin Shetty, Sampath Jayarathna
Computer Science Faculty Publications
Modern knowledge workplaces increasingly strain human episodic memory as individuals navigate fragmented attention, overlapping meetings, and multimodal information streams. Existing workplace tools provide partial support through note-taking or analytics but rarely integrate cognitive, physiological, and attentional context into retrievable memory representations. This paper presents the Cognitive Prosthetic Multimodal System (CPMS)—an AI-enabled proof-of-concept designed to support episodic recall in knowledge work through structured episodic capture and natural language retrieval. CPMS synchronizes speech transcripts, physiological signals, and gaze behavior into temporally aligned, JSON-based episodic records processed locally for privacy. Beyond data logging, the system includes a web-based retrieval interface that allows users …
Modeling Joint Visual Attention In Naturalistic Dyadic Interactions, Kuushini Thennakoon, Yasasi Abeysinghe, Bhanuka Mahanama, Vikas Ashok, Sampath Jayarathna
Modeling Joint Visual Attention In Naturalistic Dyadic Interactions, Kuushini Thennakoon, Yasasi Abeysinghe, Bhanuka Mahanama, Vikas Ashok, Sampath Jayarathna
Computer Science Faculty Publications
Joint visual attention (JVA) provides important insight into how individuals coordinate attention during social interaction. Egocentric eye tracking enables the study of JVA in natural, multi-user settings. This work presents a multi-stage framework to identify and analyze JVA using egocentric video and gaze data. The approach consists of three steps: spatiotemporal tube-based visual similarity, gaze-guided object detection, and attention pattern analysis using the ambient–focal coefficient K. Results show that object-focused collaborative activities exhibit high JVA, with object detection capturing higher joint attention than visual similarity, whereas conversation-based or independent activities show lower and more fragmented joint attention. Analysis of K …
Lost In Instructions: Study Of Blind Users' Experiences With Diy Manuals And Ai-Rewritten Instructions For Assembly, Operation, And Troubleshooting Of Tangible Products, Monalika Padma Reddy, Aruna Balasubramanian, Jiawei Zhou, Xiaojun Bi, Iv Ramakrishnan, Vikas Ashok
Lost In Instructions: Study Of Blind Users' Experiences With Diy Manuals And Ai-Rewritten Instructions For Assembly, Operation, And Troubleshooting Of Tangible Products, Monalika Padma Reddy, Aruna Balasubramanian, Jiawei Zhou, Xiaojun Bi, Iv Ramakrishnan, Vikas Ashok
Computer Science Faculty Publications
AI tools like ChatGPT and Be-My-AI are increasingly being used by blind individuals. Although prior work has explored their use in some Do-It-Yourself (DIY) tasks by blind individuals, little is known about how they use these tools and the available product-manual resources to assemble, operate, and troubleshoot physical/tangible products – tasks requiring spatial reasoning, structural understanding, and precise execution. We address this knowledge gap via an interview study and a usability study with blind participants, investigating how they leverage AI tools and product manuals for DIY tasks with physical products. Findings show that manuals are essential resources, but product-manual instructions …
Can An Experienced Qualitative Researcher Distinguish Ai From Human Qualitative Content Analysis?, Alexandra T. Lucas, Jianna Ramos, Maria Bajwa, Aaron Calhoun, Mark W. Scerbo, Janice C. Palaganas
Can An Experienced Qualitative Researcher Distinguish Ai From Human Qualitative Content Analysis?, Alexandra T. Lucas, Jianna Ramos, Maria Bajwa, Aaron Calhoun, Mark W. Scerbo, Janice C. Palaganas
Psychology Faculty Publications
Background
Artificial intelligence (AI) has become increasingly embedded in research workflows. Large language models (LLMs) are being used to code segments of text, organise codes into themes and interpret patterns within contexts. Recent comparisons between human and AI analyses demonstrate up to 80% thematic overlap, yet humans consistently exhibit deeper interpretive integration and contextual understanding. This study assesses whether experienced researchers can distinguish between entirely human-generated and AI-generated qualitative content analyses of a simulation debriefing.
Methods
We conducted a qualitative descriptive study comparing human-generated qualitative content analysis (QCA) with ChatGPT-4o-generated QCA using a single focus group transcript on emotion management …
Error-Driven Density Control For Compact Gaussian Splatting Under Sparse Supervision, Abdelrhman Elrawy
Error-Driven Density Control For Compact Gaussian Splatting Under Sparse Supervision, Abdelrhman Elrawy
Theses and Dissertations (Comprehensive)
This thesis studies efficiency and stability challenges in Gaussian-splatting-based reconstruction under sparse supervision. In few-shot novel view synthesis, standard 3D Gaussian Splatting (3DGS) can overfit the limited training views and grow an unnecessarily large number of primitives due to limitations in its Adaptive Density Control (ADC) mechanism. This thesis introduces an error-driven reformulation of ADC that triggers densification using opacity gradients as a lightweight proxy for rendering error, and shows that such aggressive densification must be paired with delayed and conservative pruning to prevent destructive create--destroy cycles. When combined with depth-based geometric regularization, the resulting framework produces substantially more compact …
A Comparative Analysis Of Explainable Ai (Xai) Techniques For Transparent And Reliable Image Classification, Sovon Chakraborty, Shakib Mahmud Dipto, Kevin R. Pilkiewicz, Michael L. Mayo, Pratip Rana
A Comparative Analysis Of Explainable Ai (Xai) Techniques For Transparent And Reliable Image Classification, Sovon Chakraborty, Shakib Mahmud Dipto, Kevin R. Pilkiewicz, Michael L. Mayo, Pratip Rana
Computer Science Faculty Publications
Evaluating the trustworthiness of black-box machine learning models remains a significant methodological challenge. Their lack of transparency and interpretability limits applicability, because stakeholders often seek transparency before trusting the results of black-box machine learning models. Explainable AI (XAI) methods provide for human-understandable justifications and informed decision-making of these black-box architectures. Therefore, it is imperative to select the proper XAI model tailored to specific tasks. In this research, we focus on examining four XAI techniques: PEEK, LRP, GRAD-CAM, and LIME to understand how they perform against each other for image classification tasks. We evaluate the performance, robustness, generalizability, noise stability, and …
Geovig And Purevig: Geometry-Aware Architectures For Efficient Computer Vision, Omar Ismail
Geovig And Purevig: Geometry-Aware Architectures For Efficient Computer Vision, Omar Ismail
Theses and Dissertations (Comprehensive)
Deploying deep learning models for medical image analysis on mobile devices requires a balance between inference latency, memory footprint, and delineating anatomical boundaries with high accuracy. While Convolutional Neural Networks (CNNs) and mobile Vision Transformers (ViTs) offer efficiency, they often struggle to model the irregular, non-local geometric structures inherent in biological tissues without incurring prohibitive computational costs. In this thesis, we introduce GeoViG (Geometric Vision Graph), an architecture that bridges the gap between efficient grid-based processing and explicit Geometric Deep Learning. GeoViG introduces a novel transition from high-resolution pixel grids to low-resolution dynamic graphs via a SpreadEdgePool operator, a geometry-aware …
Cross-Modal Proxy Evolving For Ood Detection With Vision-Language Models, Hao Tang, Yu Liu, Shuanglin Yan, Fei Shen, Shengfeng He, Jing Qin
Cross-Modal Proxy Evolving For Ood Detection With Vision-Language Models, Hao Tang, Yu Liu, Shuanglin Yan, Fei Shen, Shengfeng He, Jing Qin
Research Collection School Of Computing and Information Systems
Reliable zero-shot detection of out-of-distribution (OOD) inputs is critical for deploying vision-language models in open-world settings. However, the lack of labeled negatives in zero-shot OOD detection necessitates proxy signals that remain effective under distribution shift. Existing negative-label methods rely on a fixed set of textual proxies, which (i) sparsely sample the semantic space beyond in-distribution (ID) classes and (ii) remain static while only visual features drift, leading to cross-modal misalignment and unstable predictions. In this paper, we propose CoEvo, a training- and annotation-free test-time framework that performs bidirectional, sample-conditioned adaptation of both textual and visual proxies. Specifically, CoEvo introduces a …
Purified Zero-Shot Sketch-Based Image Retrieval, Yang Zhou, Jingru Yang, Jin Wang, Kaixiang Huang, Guodong Lu, Shengfeng He
Purified Zero-Shot Sketch-Based Image Retrieval, Yang Zhou, Jingru Yang, Jin Wang, Kaixiang Huang, Guodong Lu, Shengfeng He
Research Collection School Of Computing and Information Systems
Sketches, as a new solution in multimedia systems that can replace natural language, are characterized by sparse visual cues such as simple strokes that differ significantly from natural images containing complex elements such as background, foreground, and texture. This misalignment poses substantial challenges for zero-shot sketch-based image retrieval (ZS-SBIR). Prior approaches match sketches to full images and tend to overlook redundant elements in natural images, leading to model distraction and semantic ambiguity. To address this issue, we introduce a distraction-agnostic framework, purified cross-domain matching (PuXIM), which operates on a straightforward principle: masking and matching. We devise a visual-cross-linguistic (VxL) sampler …
Look, Compare And Draw: Differential Query Transformer For Automatic Oil Painting, Lingyu Liu, Yaxiong Wang, Li Zhu, Lizi Liao, Zhedong Zheng
Look, Compare And Draw: Differential Query Transformer For Automatic Oil Painting, Lingyu Liu, Yaxiong Wang, Li Zhu, Lizi Liao, Zhedong Zheng
Research Collection School Of Computing and Information Systems
This work introduces a new approach to automatic oil painting that emphasizes the creation of dynamic and expressive brushstrokes. A pivotal challenge lies in mitigating the duplicate and common-place strokes, which often lead to less aesthetic outcomes. Inspired by the human painting process, i.e., observing, comparing, and drawing, we incorporate differential image analysis into a neural oil painting model, allowing the model to effectively concentrate on the incremental impact of successive brushstrokes. To operationalize this concept, we propose the Differential Query Transformer (DQ-Transformer), a new architecture that leverages differentially derived image representations enriched with positional encoding to guide the stroke …
Developing An Ai-Assisted Grading System Using Large Language Models, Andrei Modiga
Developing An Ai-Assisted Grading System Using Large Language Models, Andrei Modiga
MS in Computer Science Project Reports
We present a grading system that accelerates evaluation of open-ended student work across scanned and digital workflows. The system crops answer regions from PDFs, assigns submissions via OCR on identity regions only, and groups answers by visual semantics using a vision LLM. Instructors review and edit groups, apply rubric items once per group, and export grades from an on-screen table. The solution integrates Ghostscript rasterization, PdfPig page orchestration, SkiaSharp region extraction, Tesseract identity OCR, and GPT-4o Vision for grouping. We detail the architecture, token-budgeted batching strategy, and persistence design, then describe testing results for grouping quality, time-on-task, and usability. The …
My First Conversation With Chatgpt (February 22, 2023): Origins Of A Generative Dialogue, David Smith
My First Conversation With Chatgpt (February 22, 2023): Origins Of A Generative Dialogue, David Smith
Publications and Research
This working paper presents the first recorded interaction between the author and the generative AI system ChatGPT, written on February 22, 2023 during the initial weeks of a faculty sabbatical in Boston. The document preserves a complete and unedited transcript of an exploratory conversation conducted without predetermined research aims, marking the author’s first encounter with a large-language-model conversational interface. Although the exchange includes creative experimentation—including musical and poetic prompts—the discussion remains informal and wide-ranging, and no theoretical framework is articulated at this stage. Rather, this transcript is published as primary-source material documenting the moment of discovery and experimentation that precedes …
Marsanywhere: Dataset And Cross-View Diffusion Model For Satellite-To-Ground View Synthesis With Mars Data, Benjamin T. Hinchliff
Marsanywhere: Dataset And Cross-View Diffusion Model For Satellite-To-Ground View Synthesis With Mars Data, Benjamin T. Hinchliff
Master's Theses
Satellite-to-ground view synthesis aims to create a realistic ground view image from a corresponding satellite view image. This is a well-studied problem for street level imagery, with good results being achieved by using modern image synthesis techniques such as diffusion models. However, despite the public availability of satellite and ground level imagery on Mars, these techniques have yet to be applied to the domain due to difficulties in collating and processing the data into a usable form. We address this deficiency by creating a dataset consisting of ground view panorama imagery from the Perseverance rover, along with associated satellite view …
The Future Is Now: Empowering Society Through Ai Literacy, Jason S. Wrench, Sanae Elmoudden
The Future Is Now: Empowering Society Through Ai Literacy, Jason S. Wrench, Sanae Elmoudden
Milne Open Textbooks
Artificial Intelligence (AI) is no longer a futuristic concept—it is the reality of the present. From the algorithms shaping our social media feeds to the generative tools transforming our workplaces, AI has permeated every aspect of modern life. The Future is Now moves beyond the hype to provide a comprehensive roadmap for understanding, navigating, and shaping this technological revolution.
Demystifying the Machine
This textbook serves as a user-friendly guide to the “black box” of AI. It breaks down complex technical concepts—from machine learning and neural networks to large language models—making them accessible to students across all disciplines. By establishing a …
Examining The Roles Of Embodiment And Theory Of Mind In Shaping User Perceptions Of Llm-Driven Conversational Agents, Elizabeth A. Schlesener
Examining The Roles Of Embodiment And Theory Of Mind In Shaping User Perceptions Of Llm-Driven Conversational Agents, Elizabeth A. Schlesener
All Dissertations
Large Language Models (LLMs) have advanced conversational agents, enabling natural, human-like interactions in domains such as education, programming, and workplace collaboration. Yet, user distrust persists over privacy, accuracy, and bias. As developers work to mitigate these issues and human-AI collaboration expands, reinforcing trust in LLM-driven systems is essential. To address this problem, this dissertation explores the role of anthropomorphic form in LLM-driven conversational agents and its impact on user perception.
According to the familiarity thesis, humans attribute human-like characteristics to nonhuman entities — a process known as anthropomorphism — to better comprehend unfamiliar phenomena, based on the assumption that they …