Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Artificial Intelligence and Robotics (108)
- Engineering (68)
- Computer Engineering (46)
- Databases and Information Systems (41)
- Social and Behavioral Sciences (37)
-
- Electrical and Computer Engineering (28)
- Data Science (24)
- Numerical Analysis and Scientific Computing (23)
- Other Computer Sciences (22)
- Software Engineering (18)
- Medicine and Health Sciences (16)
- Business (14)
- Communication (14)
- Information Security (14)
- Arts and Humanities (12)
- Linguistics (9)
- Social Media (9)
- Theory and Algorithms (9)
- Law (8)
- Library and Information Science (8)
- Computational Linguistics (6)
- Education (6)
- Operations Research, Systems Engineering and Industrial Engineering (6)
- Bioinformatics (5)
- Graphics and Human Computer Interfaces (5)
- Life Sciences (5)
- Public Affairs, Public Policy and Public Administration (5)
- Scholarly Publishing (5)
- Institution
-
- Singapore Management University (47)
- Old Dominion University (30)
- TÜBİTAK (17)
- Technological University Dublin (13)
- Wright State University (9)
-
- Dartmouth College (7)
- New Jersey Institute of Technology (7)
- Boise State University (6)
- City University of New York (CUNY) (6)
- United Arab Emirates University (6)
- Virginia Commonwealth University (6)
- Zayed University (6)
- Brigham Young University (5)
- Loyola University Chicago (5)
- California Polytechnic State University, San Luis Obispo (4)
- San Jose State University (4)
- University of Arkansas Little Rock (4)
- University of Arkansas, Fayetteville (4)
- University of Denver (4)
- Western Michigan University (4)
- Air Force Institute of Technology (3)
- Edith Cowan University (3)
- Mississippi State University (3)
- Missouri University of Science and Technology (3)
- Montclair State University (3)
- Purdue University (3)
- Southern Methodist University (3)
- The Texas Medical Center Library (3)
- University of Kentucky (3)
- University of Nebraska at Omaha (3)
- Publication Year
- Publication
-
- Research Collection School Of Computing and Information Systems (37)
- Theses and Dissertations (18)
- Turkish Journal of Electrical Engineering and Computer Sciences (17)
- Dissertations (15)
- Electronic Theses and Dissertations (10)
-
- Computer Science Faculty Publications (8)
- All Works (6)
- Browse all Theses and Dissertations (6)
- Conference papers (6)
- Master's Theses (6)
- VMASC Publications (6)
- Dissertations and Theses Collection (Open Access) (5)
- Boise State University Theses and Dissertations (4)
- Computer Science Faculty Proceedings & Presentations (3)
- Computer Science Senior Theses (3)
- Computer Science Theses & Dissertations (3)
- Computer Science: Faculty Publications and Other Works (3)
- Dissertations, Theses, and Capstone Projects (3)
- Doctoral Dissertations (3)
- Engineering Management & Systems Engineering Faculty Publications (3)
- Master's Projects (3)
- SMU Data Science Review (3)
- Theses and Dissertations--Computer Science (3)
- Thesis/ Dissertation Defenses (3)
- UNLV Theses, Dissertations, Professional Papers, and Capstones (3)
- Articles (2)
- CGU Faculty Publications and Research (2)
- Computer Science Faculty Research & Creative Works (2)
- Computer Science and Engineering Faculty Publications (2)
- Computer Science and Engineering Theses - Archive (2)
- Publication Type
Articles 61 - 90 of 293
Full-Text Articles in Computer Sciences
Tackling Toxicity And Harassment In Online Environments Through The Use Of Artificial Intelligence, Heba Saleous
Tackling Toxicity And Harassment In Online Environments Through The Use Of Artificial Intelligence, Heba Saleous
Dissertations
With the increase in popularity of online communities, such as social media platforms, online games, and chatroom servers, there is a need to improve chat and content moderation. Platforms have reported an increase in the prevalence of toxic behavior and hate speech. Meanwhile, moderators are reporting difficulties in keeping up with the amount of data to check as well and the type of content they are exposed to, which further harms their own mental health. The main objective of this work is to address the challenges that exist within online communities with the rising prevalence of hate speech. Additionally, some …
Etdsuite: A Toolkit To Mine Electronic Theses And Dissertations To Enrich Scholarly Big Data Using Natural Language Processing And Computer Vision, Muntabir Hasan Choudhury
Etdsuite: A Toolkit To Mine Electronic Theses And Dissertations To Enrich Scholarly Big Data Using Natural Language Processing And Computer Vision, Muntabir Hasan Choudhury
Computer Science Theses & Dissertations
In the past decades, there has been a growing interest in mining scientific documents to obtain domain knowledge automatically. One of the understudied types of scientific documents is Electronic Theses and Dissertations (ETDs), as ETDs have distinct features compared with conference proceedings and journal articles. ETDs usually serve as partial requirements of academic degrees for students pursuing higher education. They are book-length documents (i.e., 100 – 400 pages long), and the topics may shift across chapters, exhibit the significant contribution of a student’s research over the entire degree pursuing period, and have unique metadata schema and page layouts. However, the …
An Artificial Intelligence Report Card For Judicial Review, Zoe E. Niesel
An Artificial Intelligence Report Card For Judicial Review, Zoe E. Niesel
Michigan Journal of Environmental & Administrative Law
The rapid advancement of technology, including artificial intelligence (AI), is creating new challenges for judicial review under the Administrative Procedure Act (APA). In late 2023, federal administrative agencies publicly disclosed over 700 use cases of AI that employ sophisticated techniques like machine learning and natural language processing. While the APA's flexible judicial review framework certainly allows agencies to utilize new technologies, the APA also requires explainability of agency decisions; thus, agencies must be able to articulate the reasoning and methodology behind AI-enabled decisions for the purpose of judicial review. This Article examines APA judicial review as it applies to agency …
Probing Effects Of Contextual Bias On Number Magnitude Estimation, Xuehao Du, Ping Ji, Wei Qin, Lei Wang, Yunshi Lan
Probing Effects Of Contextual Bias On Number Magnitude Estimation, Xuehao Du, Ping Ji, Wei Qin, Lei Wang, Yunshi Lan
Research Collection School Of Computing and Information Systems
The semantic understanding of numbers requires association with context. However, powerful neural networks overfit spurious correlations between context and numbers in training corpus can lead to the occurrence of contextual bias, which may affect the network's accurate estimation of number magnitude when making inferences in real-world data. To investigate the resilience of current methodologies against contextual bias, we introduce a novel out-of- distribution (OOD) numerical question-answering (QA) dataset that features specific correlations between context and numbers in the training data, which are not present in the OOD test data. We evaluate the robustness of different numerical encoding and decoding methods …
A Methodological Framework For Ontology Development, Enrichment, And Application In Natural Language Processing Tasks, Navya Martin Kollapally
A Methodological Framework For Ontology Development, Enrichment, And Application In Natural Language Processing Tasks, Navya Martin Kollapally
Dissertations
Electronic Health Records (EHRs) have been widely used in healthcare to record demographics, vital signs, test results, immunizations, medical imaging reports, differential diagnoses, etc. It is now accepted that non-clinical (e.g., social) factors have a substantial influence on health outcomes. Hence, it is desirable to record these Social and Commercial Determinants of Health (SDoH & CDoH) in an EHR. The "non-text parts" of EHR notes (e.g., data tables) rely on coded terms from underlying ontologies or terminologies to facilitate semantic interoperability. Ontologies help define concepts, the relationships between them, and instances that can be utilized in research.
The first accomplishment …
Democratization Of Custom, High Quality Large Language Models, Pablo Lopez
Democratization Of Custom, High Quality Large Language Models, Pablo Lopez
College of Computing and Digital Media Dissertations
Large Language Models (LLMs) have shown exceptional performance in several natural language processing (NLP) tasks. Customizing LLMs boosts their performance in domain specific tasks but typically requires substantial resources and effort for training, such as supervised fine-tuning. This research proposes methods to achieve significant accuracy improvements given minimal resources, particularly focusing on open-ended question answering with a given piece of context. We utilize an LLM’s self-generated training data to fine-tune the LLM and partial fine-tuning with on-demand GPU to reduce practitioner training costs. The research shows that these methods give significant performance gains in a Retrieval Augmented Generation (RAG) based …
Planning Like Human : A Dual-Process Framework For Dialogue Planning, Tao He, Lizi Liao, Yixin Cao, Yuanxing Liu, Ming Liu, Zerui Chen, Bing Qin
Planning Like Human : A Dual-Process Framework For Dialogue Planning, Tao He, Lizi Liao, Yixin Cao, Yuanxing Liu, Ming Liu, Zerui Chen, Bing Qin
Research Collection School Of Computing and Information Systems
In proactive dialogue, the challenge lies not just in generating responses but in steering conversations toward predetermined goals, a task where Large Language Models (LLMs) typically struggle due to their reactive nature. Traditional approaches to enhance dialogue planning in LLMs, ranging from elaborate prompt engineering to the integration of policy networks, either face efficiency issues or deliver suboptimal performance. Inspired by the dual-process theory in psychology, which identifies two distinct modes of thinking—intuitive (fast) and analytical (slow), we propose the Dual-Process Dialogue Planning (DPDP) framework. DPDP embodies this theory through two complementary planning systems: an instinctive policy model for familiar …
A Comprehensive Dataset For Arabic Word Sense Disambiguation, Sanaa Kaddoura, Reem Nassar
A Comprehensive Dataset For Arabic Word Sense Disambiguation, Sanaa Kaddoura, Reem Nassar
All Works
This data paper introduces a comprehensive dataset tailored for word sense disambiguation tasks, explicitly focusing on a hundred polysemous words frequently employed in Modern Standard Arabic. The dataset encompasses a diverse set of senses for each word, ranging from 3 to 8, resulting in 367 unique senses. Each word sense is accompanied by contextual sentences comprising ten sentence examples that feature the polysemous word in various contexts. The data collection resulted in a dataset of 3670 samples. Significantly, the dataset is in Arabic, which is known for its rich morphology, complex syntax, and extensive polysemy. The data was meticulously collected …
A Survey On Neural Question Generation : Methods, Applications, And Prospects, Shasha Guo, Lizi Liao, Cuiping Li, Tat-Seng Chua
A Survey On Neural Question Generation : Methods, Applications, And Prospects, Shasha Guo, Lizi Liao, Cuiping Li, Tat-Seng Chua
Research Collection School Of Computing and Information Systems
In this survey, we present a detailed examination of the advancements in Neural Question Generation (NQG), a field leveraging neural network techniques to generate relevant questions from diverse inputs like knowledge bases, texts, and images. The survey begins with an overview of NQG’s background, encompassing the task’s problem formulation, prevalent benchmark datasets, established evaluation metrics, and notable applications. It then methodically classifies NQG approaches into three predominant categories: structured NQG, which utilizes organized data sources, unstructured NQG, focusing on more loosely structured inputs like texts or visual content, and hybrid NQG, drawing on diverse input modalities. This classification is followed …
Enrichment Of Turkish Question Answering Systems Using Knowledge Graphs, Okan Çi̇ftçi̇, Fati̇h Soygazi̇, Selma Teki̇r
Enrichment Of Turkish Question Answering Systems Using Knowledge Graphs, Okan Çi̇ftçi̇, Fati̇h Soygazi̇, Selma Teki̇r
Turkish Journal of Electrical Engineering and Computer Sciences
Recent capabilities of large language models (LLMs) have transformed many tasks in Natural Language Processing (NLP), including question answering. The state-of-the-art systems do an excellent job of responding in a relevant, persuasive way but cannot guarantee factuality. Knowledge graphs, representing facts as triplets, can be valuable for avoiding errors and inconsistencies with real-world facts. This work introduces a knowledge graph-based approach to Turkish question answering. The proposed approach aims to develop a methodology capable of drawing inferences from a knowledge graph to answer complex multihop questions. We construct the Beyazperde Movie Knowledge Graph (BPMovieKG) and the Turkish Movie Question Answering …
Who Wrote The Scientific News? Improving The Discernibility Of Llms To Human-Written Scientific News, Dominik Soós
Who Wrote The Scientific News? Improving The Discernibility Of Llms To Human-Written Scientific News, Dominik Soós
Computer Science Theses & Dissertations
Large Language Models (LLMs) have rapidly advanced the field of Natural Language Processing and become powerful tools for generating and evaluating scientific text. Although LLMs have demonstrated promising as evaluators for certain text generation tasks, there is still a gap until they are used as reliable text evaluators for general purposes. In this thesis project, I attempted to fill this gap by examining the discernibility of LLMs from human-written and LLM-generated scientific news. This research demonstrated that although it was relatively straightforward for humans to discern scientific news written by humans from scientific news generated by GPT-3.5 using basic prompts, …
Unveiling The Dynamics Of Crisis Events: Sentiment And Emotion Analysis Via Multi-Task Learning With Attention Mechanism And Subject-Based Intent Prediction, Phyo Yi Win Myint, Siaw Ling Lo, Yuhao Zhang
Unveiling The Dynamics Of Crisis Events: Sentiment And Emotion Analysis Via Multi-Task Learning With Attention Mechanism And Subject-Based Intent Prediction, Phyo Yi Win Myint, Siaw Ling Lo, Yuhao Zhang
Research Collection School Of Computing and Information Systems
In the age of rapid internet expansion, social media platforms like Twitter have become crucial for sharing information, expressing emotions, and revealing intentions during crisis situations. They offer crisis responders a means to assess public sentiment, attitudes, intentions, and emotional shifts by monitoring crisis-related tweets. To enhance sentiment and emotion classification, we adopt a transformer-based multi-task learning (MTL) approach with attention mechanism, enabling simultaneous handling of both tasks, and capitalizing on task interdependencies. Incorporating attention mechanism allows the model to concentrate on important words that strongly convey sentiment and emotion. We compare three baseline models, and our findings show that …
Study On Data Mining Of Hydrogen Energy Policy In China Based On Natural Language Processing Technology, Dongling Huang, Yuan Liu, Xiaoshuai Yuan, Guozhong Jin, Yuanhang Cai, Li Liu, Heng Cao, Wanjun Li, Rui Cai
Study On Data Mining Of Hydrogen Energy Policy In China Based On Natural Language Processing Technology, Dongling Huang, Yuan Liu, Xiaoshuai Yuan, Guozhong Jin, Yuanhang Cai, Li Liu, Heng Cao, Wanjun Li, Rui Cai
Bulletin of Chinese Academy of Sciences (Chinese Version)
Report to the 20th National Congress of the CPC emphasized the importance of “working actively and prudently towards the goals of reaching peak carbon emissions and carbon neutrality”, as well as “speeding up the planning and development of a system for new energy sources”. As a green and low-carbon secondary energy source, hydrogen energy has multiple applications in promoting the large-scale and efficient use of renewable energy as well as energy substitution in the field of transportation. It can also accelerate decarbonization in industry, and as such, is an indispensable part of building a new energy system, reaching peak carbon …
Securing The Inbox: Advancing Cyber Resilience With Fine-Tuned Bert, Fatima Rashed Al Saedi
Securing The Inbox: Advancing Cyber Resilience With Fine-Tuned Bert, Fatima Rashed Al Saedi
Thesis/ Dissertation Defenses
In recent years, phishing attacks have persisted as a widespread threat in the contemporary digital environment, presenting substantial risks to individuals and organizations. Cybercriminals are devising increasingly sophisticated strategies to deceive users through malicious emails. In response to this challenge, this research focuses on developing a new tool for detecting phishing emails utilizing the BERT algorithm. The tool aims to enhance email security by accurately identifying deceptive emails and protecting users from potential cyber threats. The primary objective of this study is to investigate how leveraging the BERT algorithm can improve the detection of phishing emails compared to traditional methods. …
Securing The Inbox: Advancing Phishing Email Detection With Fine-Tuned Bert, Fatima Rashed Al Saedi
Securing The Inbox: Advancing Phishing Email Detection With Fine-Tuned Bert, Fatima Rashed Al Saedi
Theses
In recent years, phishing attacks have persisted as a widespread threat in the contemporary digital environment, presenting substantial risks to individuals and organizations. Cybercriminals are devising increasingly sophisticated strategies to deceive users through malicious emails. In response to this challenge, this research focuses on developing a new tool for detecting phishing emails utilizing the BERT algorithm. The tool aims to enhance email security by accurately identifying deceptive emails and protecting users from potential cyber threats. The primary objective of this study is to investigate how leveraging the BERT algorithm can improve the detection of phishing emails compared to traditional methods. …
Actively Learn From Llms With Uncertainty Propagation For Generalized Category Discovery, Jinggui Liang, Lizi Liao, Hao Fei, Bobo Li, Jing Jiang
Actively Learn From Llms With Uncertainty Propagation For Generalized Category Discovery, Jinggui Liang, Lizi Liao, Hao Fei, Bobo Li, Jing Jiang
Research Collection School Of Computing and Information Systems
Generalized category discovery faces a key issue: the lack of supervision for new and unseen data categories. Traditional methods typically combine supervised pretraining with self-supervised learning to create models, and then employ clustering for category identification. However, these approaches tend to become overly tailored to known categories, failing to fully resolve the core issue. Hence, we propose to integrate the feedback from LLMs into an active learning paradigm. Specifically, our method innovatively employs uncertainty propagation to select data samples from high-uncertainty regions, which are then labeled using LLMs through a comparison-based prompting scheme. This not only eases the labeling task …
Sgsh : Stimulate Large Language Models With Skeleton Heuristics For Knowledge Base Question Generation, Shasha Guo, Lizi Liao, Jing Zhang, Yanling Wang, Cuiping Li, Hong Chen
Sgsh : Stimulate Large Language Models With Skeleton Heuristics For Knowledge Base Question Generation, Shasha Guo, Lizi Liao, Jing Zhang, Yanling Wang, Cuiping Li, Hong Chen
Research Collection School Of Computing and Information Systems
Knowledge base question generation (KBQG) aims to generate natural language questions from a set of triplet facts extracted from KB. Existing methods have significantly boosted the performance of KBQG via pre-trained language models (PLMs) thanks to the richly endowed semantic knowledge. With the advance of pre-training techniques, large language models (LLMs) (e.g., GPT-3.5) undoubtedly possess much more semantic knowledge. Therefore, how to effectively organize and exploit the abundant knowledge for KBQG becomes the focus of our study. In this work, we propose SGSH — a simple and effective framework to Stimulate GPT-3.5 with Skeleton Heuristics to enhance KBQG. The framework …
Text-To-Sql: A Methodical Review Of Challenges And Models, Ali Buğra Kanburoğlu, Faik Boray Tek
Text-To-Sql: A Methodical Review Of Challenges And Models, Ali Buğra Kanburoğlu, Faik Boray Tek
Turkish Journal of Electrical Engineering and Computer Sciences
This survey focuses on Text-to-SQL, automated translation of natural language queries into SQL queries. Initially, we describe the problem and its main challenges. Then, by following the PRISMA systematic review methodology, we survey the existing Text-to-SQL review papers in the literature. We apply the same method to extract proposed Text-to-SQL models and classify them with respect to used evaluation metrics and benchmarks. We highlight the accuracies achieved by various models on Text-to-SQL datasets and discuss execution-guided evaluation strategies. We present insights into model training times and implementations of different models. We also explore the availability of Text-to-SQL datasets in non-English …
Reviving The Past: Enhancing Language Models With Historical Text Optimization, Heather D. Broome
Reviving The Past: Enhancing Language Models With Historical Text Optimization, Heather D. Broome
Honors Theses
Recent advancements in Natural Language Processing (NLP) have brought attention to the significant potential that exists for widespread applications of Large Language Models (LLMs). As demands and expectations for LLMs rise, ensuring efficiency and accuracy becomes paramount. Addressing these challenges requires more than just optimizing current techniques; it urges novel approaches to NLP as a whole. This study investigates novel data preprocessing methods designed to enhance LLM performance by mitigating inefficiencies rooted in natural language, particularly by simplifying the complexities presented by historical texts. Utilizing the classical text The Odyssey by Homer, two preprocessing techniques are introduced: tokenization of names …
Using Chatgpt To Generate Gendered Language, Shweta Soundararajan, Manuela Nayantara Jeyaraj, Sarah Jane Delany
Using Chatgpt To Generate Gendered Language, Shweta Soundararajan, Manuela Nayantara Jeyaraj, Sarah Jane Delany
Conference papers
Gendered language is the use of words that denote an individual's gender. This can be explicit where the gender is evident in the actual word used, e.g. mother, she, man, but it can also be implicit where social roles or behaviours can signal an individual's gender - for example, expectations that women display communal traits (e.g., affectionate, caring, gentle) and men display agentic traits (e.g., assertive, competitive, decisive). The use of gendered language in NLP systems can perpetuate gender stereotypes and bias. This paper proposes an approach to generating gendered language datasets using ChatGPT which will provide data for data-driven …
Natural Language Processing Analysis Of Online Reviews For Small Business: Extracting Insight From Small Corpora, Benjamin J. Mccloskey, Phillip M. Lacasse, Bruce A. Cox
Natural Language Processing Analysis Of Online Reviews For Small Business: Extracting Insight From Small Corpora, Benjamin J. Mccloskey, Phillip M. Lacasse, Bruce A. Cox
Faculty Publications
Receiving and acting on customer input is essential to sustaining and growing any service organization, particularly a small family business whose livelihood depends on strong relationships with its customers. The competitive advantage offered by advanced analytical approaches for supporting decisions is not trivial, and enterprises across virtually all domains of society are investing heavily in this emerging discipline. Natural Language Processing (NLP) is a subset of computer science that employs computational approaches to analyze human language; it is effective at extracting insight from text data but frequently requires large corpora to train its models, in the scale of thousands or …
A Comprehensive Survey On Pretrained Foundation Models: A History From Bert To Chatgpt, Ce Zhou, Qian Li, Chen Li, Jun Yu, Yixin Liu, Guangjing Wang, Kai Zhang, Cheng Ji, Qiben Yan, Lifang He, Hao Peng, Jianxin Li, Jia Wu, Ziwei Liu, Pengtao Xie, Caiming Xiong, Jian Pei, Philip S. Yu, Lichao Sun
A Comprehensive Survey On Pretrained Foundation Models: A History From Bert To Chatgpt, Ce Zhou, Qian Li, Chen Li, Jun Yu, Yixin Liu, Guangjing Wang, Kai Zhang, Cheng Ji, Qiben Yan, Lifang He, Hao Peng, Jianxin Li, Jia Wu, Ziwei Liu, Pengtao Xie, Caiming Xiong, Jian Pei, Philip S. Yu, Lichao Sun
Computer Science Faculty Research & Creative Works
Pretrained Foundation Models (PFMs) are regarded as the foundation for various downstream tasks across different data modalities. A PFM (e.g., BERT, ChatGPT, GPT-4) is trained on large-scale data, providing a solid parameter initialization for a wide range of downstream applications. In contrast to earlier methods that use convolution and recurrent modules for feature extraction, BERT learns bidirectional encoder representations from Transformers, trained on large datasets as contextual language models. Similarly, the Generative Pretrained Transformer (GPT) method employs Transformers as feature extractors and is trained on large datasets using an autoregressive paradigm. Recently, ChatGPT has demonstrated significant success in large language …
Bert-Based Detection Of Ai-Generated Text For Content Verification, Soham Biren Katlariwala
Bert-Based Detection Of Ai-Generated Text For Content Verification, Soham Biren Katlariwala
2024 REYES Proceedings
With advancements in AI-driven natural language generation, distinguishing between AI-generated and human-written text has become imperative for ensuring content authenticity across industries. This study explores the effectiveness of Bidirectional Encoder Representations from Transformers (BERT) in addressing this classification challenge. Utilizing a diverse dataset and robust preprocessing techniques, BERT achieved a peak F1-score of 0.94364, outperforming traditional models such as Logistic Regression and Support Vector Machines. The results underscore the potential of transformer-based models in addressing real-world con- tent verification problems. Future enhancements include fine-tuning and expanding datasets for greater generalizability.
Federated Learning For Sentiment Analysis In Presence Of Non-Iid Data: Sensitivity Of Deep Learning Models, Davoud Gholamiangonabadi, Katarina Grolinger
Federated Learning For Sentiment Analysis In Presence Of Non-Iid Data: Sensitivity Of Deep Learning Models, Davoud Gholamiangonabadi, Katarina Grolinger
Electrical and Computer Engineering Publications
In sentiment analysis, data are commonly distributed across many devices, and traditional machine learning requires transferring these data to a central location exposing data to security and privacy risks. Federated Learning (FL) avoids this transfer by training a model without requiring the clients/devices to share their local data; however, FL performance drops when data are not Independent and Identically Distributed (non-IID), such as when label distribution or data size vary across clients. Although techniques for non-IID data have been proposed primarily in the image domain, the sensitivity of various deep learning models to non-IID data needs to be examined. Consequently, …
Generative Artificial Intelligence: Basic Terminology And Concepts, Kincaid Brown
Generative Artificial Intelligence: Basic Terminology And Concepts, Kincaid Brown
Law Librarian Scholarship
Generative artificial intelligence (GenAI) has been a hard topic to avoid in the media for more than a year. But what do all of the terms mean and what are areas of concern with GenAI tools?
This column aims to provide a baseline explanation of terminology and concepts that are frequently in the media.
Revolutionizing Campus Communication: Nlp-Powered University Chatbots, Ritu Ramakrishnan, Priyanka Thangamuthu, Austin Nguyen, Jinzhu Gao
Revolutionizing Campus Communication: Nlp-Powered University Chatbots, Ritu Ramakrishnan, Priyanka Thangamuthu, Austin Nguyen, Jinzhu Gao
Pacific Faculty Work
Artificial intelligence (AI) based chatbots leverage programmed software instructions to simulate human speech and user interaction. These versatile tools can be employed in various domains, from managing smart home devices to providing personal virtual assistants. They can also be useful in responding to common queries and can make information easier to access. In response to this need, we developed a specialized chatbot tailored for the academic environment by training an NLP model to answer frequently asked questions (FAQs) the need of searching through the university website. The main goal is to optimize user engagement and streamline information retrieval within a …
Unifying Context With Labeled Property Graph: A Pipeline-Based System For Comprehensive Text Representation In Nlp, Ali Hur, Naeem Janjua, Mohiuddin Ahmed
Unifying Context With Labeled Property Graph: A Pipeline-Based System For Comprehensive Text Representation In Nlp, Ali Hur, Naeem Janjua, Mohiuddin Ahmed
Research outputs 2022 to 2026
Extracting valuable insights from vast amounts of unstructured digital text presents significant challenges across diverse domains. This research addresses this challenge by proposing a novel pipeline-based system that generates domain-agnostic and task-agnostic text representations. The proposed approach leverages labeled property graphs (LPG) to encode contextual information, facilitating the integration of diverse linguistic elements into a unified representation. The proposed system enables efficient graph-based querying and manipulation by addressing the crucial aspect of comprehensive context modeling and fine-grained semantics. The effectiveness of the proposed system is demonstrated through the implementation of NLP components that operate on LPG-based representations. Additionally, the proposed …
Integrating Generative Artificial Intelligence With Systems Architecting Diagram Creation: Advancement, Challenges, Opportunities And Future Perspectives, Cansu Yalim, Holly H. Handley
Integrating Generative Artificial Intelligence With Systems Architecting Diagram Creation: Advancement, Challenges, Opportunities And Future Perspectives, Cansu Yalim, Holly H. Handley
Engineering Management & Systems Engineering Faculty Publications
Generative AI (GenAI) serves as a powerful tool that can create a wide range of content, including but not limited to text, speech, images, code, videos, and 3D models. ChatGPT stands out as a particularly appealing Generative Pretrained Transformer (GPT) model that offers supplementary capabilities through GPTs and plugins. These extensions enable users to engage with the chatbot and improve its functionality, surpassing mere content generation. Our study delves into the potential of ChatGPT, specifically GPT-4, to expedite the creation of diagrams to support the system architecting process. To this end, we explored the use of ChatGPT's Diagrams Show Me …
Embedding Software Engineering In Mixed Methods: Computationally Enhanced Risk Communication, Ann Marie Reinhold, Madison H. Munro, Elizabeth A. Shanahan, Ross J. Gore, Barry C. Ezell, Clemente I. Izurieta
Embedding Software Engineering In Mixed Methods: Computationally Enhanced Risk Communication, Ann Marie Reinhold, Madison H. Munro, Elizabeth A. Shanahan, Ross J. Gore, Barry C. Ezell, Clemente I. Izurieta
VMASC Publications
Mixed methods research ameliorates many convergent research challenges within the contemporary sociotechnical landscape. We suggest the integration of software engineering in mixed methods studies is a critical step to address some of the remaining and persistent challenges. One such research challenge where software engineering is particularly well suited is in hazard preparedness—in particular, the creation of risk communication messages to mitigate or prevent harm. Computationally enhanced risk communication is convergent research that integrates software engineering and social science research for the benefit of protecting humans and infrastructure. To this end, we developed a mixed methods framework for the efficient construction …
A Comparison Of Lexical Tokenization Methods, Nathan Culmer
A Comparison Of Lexical Tokenization Methods, Nathan Culmer
Williams Honors College, Honors Research Projects
The purpose of this project was to compare tokenization methods, or methods of breaking up a text into meaningful parts for use in natural language processing. The effectiveness of several commonly used tokenization methods were investigated, including morpheme tokenization, which takes into account the linguistic features of the language. In addition, I proposed and implemented a new technique to consider the capitalization pattern of a word in the tokenization process, in order to allow this process to include more natural language features. The effectiveness of these methods was compared by using them in a sentiment analysis model for various datasets, …