Open Access. Powered by Scholars. Published by Universities.®

Natural language processing

Discipline
Institution
Publication Year
Publication
Publication Type

Articles 31 - 60 of 108

Full-Text Articles in Artificial Intelligence and Robotics

Dc-Instruct : An Effective Framework For Generative Multi-Intent Spoken Language Understanding, Bowen Xing, Lizi Liao, Minlie Huang Nov 2024

Dc-Instruct : An Effective Framework For Generative Multi-Intent Spoken Language Understanding, Bowen Xing, Lizi Liao, Minlie Huang

Research Collection School Of Computing and Information Systems

In the realm of multi-intent spoken language understanding, recent advancements have leveraged the potential of prompt learning frameworks. However, critical gaps exist in these frameworks: the lack of explicit modeling of dual-task dependencies and the oversight of task-specific semantic differences among utterances. To address these shortcomings, we propose DC-Instruct, a novel generative framework based on Dual-task Inter-dependent Instructions (DII) and Supervised Contrastive Instructions (SCI). Specifically, DII guides large language models (LLMs) to generate labels for one task based on the other task’s labels, thereby explicitly capturing dual-task inter-dependencies. Moreover, SCI leverages utterance semantics differences by guiding LLMs to determine whether …


Etdsuite: A Toolkit To Mine Electronic Theses And Dissertations To Enrich Scholarly Big Data Using Natural Language Processing And Computer Vision, Muntabir Hasan Choudhury Oct 2024

Etdsuite: A Toolkit To Mine Electronic Theses And Dissertations To Enrich Scholarly Big Data Using Natural Language Processing And Computer Vision, Muntabir Hasan Choudhury

Computer Science Theses & Dissertations

In the past decades, there has been a growing interest in mining scientific documents to obtain domain knowledge automatically. One of the understudied types of scientific documents is Electronic Theses and Dissertations (ETDs), as ETDs have distinct features compared with conference proceedings and journal articles. ETDs usually serve as partial requirements of academic degrees for students pursuing higher education. They are book-length documents (i.e., 100 – 400 pages long), and the topics may shift across chapters, exhibit the significant contribution of a student’s research over the entire degree pursuing period, and have unique metadata schema and page layouts. However, the …


An Artificial Intelligence Report Card For Judicial Review, Zoe E. Niesel Sep 2024

An Artificial Intelligence Report Card For Judicial Review, Zoe E. Niesel

Michigan Journal of Environmental & Administrative Law

The rapid advancement of technology, including artificial intelligence (AI), is creating new challenges for judicial review under the Administrative Procedure Act (APA). In late 2023, federal administrative agencies publicly disclosed over 700 use cases of AI that employ sophisticated techniques like machine learning and natural language processing. While the APA's flexible judicial review framework certainly allows agencies to utilize new technologies, the APA also requires explainability of agency decisions; thus, agencies must be able to articulate the reasoning and methodology behind AI-enabled decisions for the purpose of judicial review. This Article examines APA judicial review as it applies to agency …


Democratization Of Custom, High Quality Large Language Models, Pablo Lopez Aug 2024

Democratization Of Custom, High Quality Large Language Models, Pablo Lopez

College of Computing and Digital Media Dissertations

Large Language Models (LLMs) have shown exceptional performance in several natural language processing (NLP) tasks. Customizing LLMs boosts their performance in domain specific tasks but typically requires substantial resources and effort for training, such as supervised fine-tuning. This research proposes methods to achieve significant accuracy improvements given minimal resources, particularly focusing on open-ended question answering with a given piece of context. We utilize an LLM’s self-generated training data to fine-tune the LLM and partial fine-tuning with on-demand GPU to reduce practitioner training costs. The research shows that these methods give significant performance gains in a Retrieval Augmented Generation (RAG) based …


Planning Like Human : A Dual-Process Framework For Dialogue Planning, Tao He, Lizi Liao, Yixin Cao, Yuanxing Liu, Ming Liu, Zerui Chen, Bing Qin Aug 2024

Planning Like Human : A Dual-Process Framework For Dialogue Planning, Tao He, Lizi Liao, Yixin Cao, Yuanxing Liu, Ming Liu, Zerui Chen, Bing Qin

Research Collection School Of Computing and Information Systems

In proactive dialogue, the challenge lies not just in generating responses but in steering conversations toward predetermined goals, a task where Large Language Models (LLMs) typically struggle due to their reactive nature. Traditional approaches to enhance dialogue planning in LLMs, ranging from elaborate prompt engineering to the integration of policy networks, either face efficiency issues or deliver suboptimal performance. Inspired by the dual-process theory in psychology, which identifies two distinct modes of thinking—intuitive (fast) and analytical (slow), we propose the Dual-Process Dialogue Planning (DPDP) framework. DPDP embodies this theory through two complementary planning systems: an instinctive policy model for familiar …


A Survey On Neural Question Generation : Methods, Applications, And Prospects, Shasha Guo, Lizi Liao, Cuiping Li, Tat-Seng Chua Aug 2024

A Survey On Neural Question Generation : Methods, Applications, And Prospects, Shasha Guo, Lizi Liao, Cuiping Li, Tat-Seng Chua

Research Collection School Of Computing and Information Systems

In this survey, we present a detailed examination of the advancements in Neural Question Generation (NQG), a field leveraging neural network techniques to generate relevant questions from diverse inputs like knowledge bases, texts, and images. The survey begins with an overview of NQG’s background, encompassing the task’s problem formulation, prevalent benchmark datasets, established evaluation metrics, and notable applications. It then methodically classifies NQG approaches into three predominant categories: structured NQG, which utilizes organized data sources, unstructured NQG, focusing on more loosely structured inputs like texts or visual content, and hybrid NQG, drawing on diverse input modalities. This classification is followed …


Who Wrote The Scientific News? Improving The Discernibility Of Llms To Human-Written Scientific News, Dominik Soós Jul 2024

Who Wrote The Scientific News? Improving The Discernibility Of Llms To Human-Written Scientific News, Dominik Soós

Computer Science Theses & Dissertations

Large Language Models (LLMs) have rapidly advanced the field of Natural Language Processing and become powerful tools for generating and evaluating scientific text. Although LLMs have demonstrated promising as evaluators for certain text generation tasks, there is still a gap until they are used as reliable text evaluators for general purposes. In this thesis project, I attempted to fill this gap by examining the discernibility of LLMs from human-written and LLM-generated scientific news. This research demonstrated that although it was relatively straightforward for humans to discern scientific news written by humans from scientific news generated by GPT-3.5 using basic prompts, …


Study On Data Mining Of Hydrogen Energy Policy In China Based On Natural Language Processing Technology, Dongling Huang, Yuan Liu, Xiaoshuai Yuan, Guozhong Jin, Yuanhang Cai, Li Liu, Heng Cao, Wanjun Li, Rui Cai Jun 2024

Study On Data Mining Of Hydrogen Energy Policy In China Based On Natural Language Processing Technology, Dongling Huang, Yuan Liu, Xiaoshuai Yuan, Guozhong Jin, Yuanhang Cai, Li Liu, Heng Cao, Wanjun Li, Rui Cai

Bulletin of Chinese Academy of Sciences (Chinese Version)

Report to the 20th National Congress of the CPC emphasized the importance of “working actively and prudently towards the goals of reaching peak carbon emissions and carbon neutrality”, as well as “speeding up the planning and development of a system for new energy sources”. As a green and low-carbon secondary energy source, hydrogen energy has multiple applications in promoting the large-scale and efficient use of renewable energy as well as energy substitution in the field of transportation. It can also accelerate decarbonization in industry, and as such, is an indispensable part of building a new energy system, reaching peak carbon …


Sgsh : Stimulate Large Language Models With Skeleton Heuristics For Knowledge Base Question Generation, Shasha Guo, Lizi Liao, Jing Zhang, Yanling Wang, Cuiping Li, Hong Chen Jun 2024

Sgsh : Stimulate Large Language Models With Skeleton Heuristics For Knowledge Base Question Generation, Shasha Guo, Lizi Liao, Jing Zhang, Yanling Wang, Cuiping Li, Hong Chen

Research Collection School Of Computing and Information Systems

Knowledge base question generation (KBQG) aims to generate natural language questions from a set of triplet facts extracted from KB. Existing methods have significantly boosted the performance of KBQG via pre-trained language models (PLMs) thanks to the richly endowed semantic knowledge. With the advance of pre-training techniques, large language models (LLMs) (e.g., GPT-3.5) undoubtedly possess much more semantic knowledge. Therefore, how to effectively organize and exploit the abundant knowledge for KBQG becomes the focus of our study. In this work, we propose SGSH — a simple and effective framework to Stimulate GPT-3.5 with Skeleton Heuristics to enhance KBQG. The framework …


Actively Learn From Llms With Uncertainty Propagation For Generalized Category Discovery, Jinggui Liang, Lizi Liao, Hao Fei, Bobo Li, Jing Jiang Jun 2024

Actively Learn From Llms With Uncertainty Propagation For Generalized Category Discovery, Jinggui Liang, Lizi Liao, Hao Fei, Bobo Li, Jing Jiang

Research Collection School Of Computing and Information Systems

Generalized category discovery faces a key issue: the lack of supervision for new and unseen data categories. Traditional methods typically combine supervised pretraining with self-supervised learning to create models, and then employ clustering for category identification. However, these approaches tend to become overly tailored to known categories, failing to fully resolve the core issue. Hence, we propose to integrate the feedback from LLMs into an active learning paradigm. Specifically, our method innovatively employs uncertainty propagation to select data samples from high-uncertainty regions, which are then labeled using LLMs through a comparison-based prompting scheme. This not only eases the labeling task …


Using Chatgpt To Generate Gendered Language, Shweta Soundararajan, Manuela Nayantara Jeyaraj, Sarah Jane Delany Mar 2024

Using Chatgpt To Generate Gendered Language, Shweta Soundararajan, Manuela Nayantara Jeyaraj, Sarah Jane Delany

Conference papers

Gendered language is the use of words that denote an individual's gender. This can be explicit where the gender is evident in the actual word used, e.g. mother, she, man, but it can also be implicit where social roles or behaviours can signal an individual's gender - for example, expectations that women display communal traits (e.g., affectionate, caring, gentle) and men display agentic traits (e.g., assertive, competitive, decisive). The use of gendered language in NLP systems can perpetuate gender stereotypes and bias. This paper proposes an approach to generating gendered language datasets using ChatGPT which will provide data for data-driven …


Natural Language Processing Analysis Of Online Reviews For Small Business: Extracting Insight From Small Corpora, Benjamin J. Mccloskey, Phillip M. Lacasse, Bruce A. Cox Jan 2024

Natural Language Processing Analysis Of Online Reviews For Small Business: Extracting Insight From Small Corpora, Benjamin J. Mccloskey, Phillip M. Lacasse, Bruce A. Cox

Faculty Publications

Receiving and acting on customer input is essential to sustaining and growing any service organization, particularly a small family business whose livelihood depends on strong relationships with its customers. The competitive advantage offered by advanced analytical approaches for supporting decisions is not trivial, and enterprises across virtually all domains of society are investing heavily in this emerging discipline. Natural Language Processing (NLP) is a subset of computer science that employs computational approaches to analyze human language; it is effective at extracting insight from text data but frequently requires large corpora to train its models, in the scale of thousands or …


Bert-Based Detection Of Ai-Generated Text For Content Verification, Soham Biren Katlariwala Jan 2024

Bert-Based Detection Of Ai-Generated Text For Content Verification, Soham Biren Katlariwala

2024 REYES Proceedings

With advancements in AI-driven natural language generation, distinguishing between AI-generated and human-written text has become imperative for ensuring content authenticity across industries. This study explores the effectiveness of Bidirectional Encoder Representations from Transformers (BERT) in addressing this classification challenge. Utilizing a diverse dataset and robust preprocessing techniques, BERT achieved a peak F1-score of 0.94364, outperforming traditional models such as Logistic Regression and Support Vector Machines. The results underscore the potential of transformer-based models in addressing real-world con- tent verification problems. Future enhancements include fine-tuning and expanding datasets for greater generalizability.


Federated Learning For Sentiment Analysis In Presence Of Non-Iid Data: Sensitivity Of Deep Learning Models, Davoud Gholamiangonabadi, Katarina Grolinger Jan 2024

Federated Learning For Sentiment Analysis In Presence Of Non-Iid Data: Sensitivity Of Deep Learning Models, Davoud Gholamiangonabadi, Katarina Grolinger

Electrical and Computer Engineering Publications

In sentiment analysis, data are commonly distributed across many devices, and traditional machine learning requires transferring these data to a central location exposing data to security and privacy risks. Federated Learning (FL) avoids this transfer by training a model without requiring the clients/devices to share their local data; however, FL performance drops when data are not Independent and Identically Distributed (non-IID), such as when label distribution or data size vary across clients. Although techniques for non-IID data have been proposed primarily in the image domain, the sensitivity of various deep learning models to non-IID data needs to be examined. Consequently, …


A Prototype Of A Conversational Virtual University Support Agent Powered By A Large Language Model That Addresses Inquiries About Policies In The Student Handbook, Joseph Benjamin R. Ilagan, Jose Ramon Ilagan Jan 2024

A Prototype Of A Conversational Virtual University Support Agent Powered By A Large Language Model That Addresses Inquiries About Policies In The Student Handbook, Joseph Benjamin R. Ilagan, Jose Ramon Ilagan

Quantitative Methods and Information Technology Faculty Publications

Universities gain a competitive advantage by deliberately improving overall service, student, faculty, and staff experience, leading to attractiveness, retention, and improved outcomes. Quality services are achieved partly by addressing employee satisfaction, specifically in the work environment. This paper presents a prototype study of a virtual university support agent, a system grounded in a Large Language Model (LLM) engineered to address inquiries from university students, faculty and staff related to the student handbook. The study investigates the integration of generative artificial intelligence and natural conversation properties inherent in LLMs to overcome customer service shortcomings identified in previous chatbot applications. The LLMs' …


Exploratory Prompting Of Large Language Models To Act As Co-Pilots For Augmenting Business Process Work In Document Classification, Jose Ramon Ilagan, Joseph Benjamin R. Ilagan, Claire Louisse Basallo, Zachary Matthew Alabastro Jan 2024

Exploratory Prompting Of Large Language Models To Act As Co-Pilots For Augmenting Business Process Work In Document Classification, Jose Ramon Ilagan, Joseph Benjamin R. Ilagan, Claire Louisse Basallo, Zachary Matthew Alabastro

Quantitative Methods and Information Technology Faculty Publications

Businesses deal with different types of documents containing unstructured documents. The data in these documents must be converted into digital forms other automated systems could only process. One generic use case is document classification, which usually involves manual transformation due to human understanding needed in the process. These documents go beyond those generated through regular business transactions and operations and also include web-based content such as online news, blogs, e-mails, and various digital libraries. Recent developments in robotic process automation (RPA) and artificial intelligence (AI) aim to automate the otherwise expensive, time-consuming, and repetitive manual steps. Through more powerful natural …


Towards Dynamic Context Detection From Voice Commands And Conversations With Smart Assistants In Smart Homes, Jeniya Sultana Jan 2024

Towards Dynamic Context Detection From Voice Commands And Conversations With Smart Assistants In Smart Homes, Jeniya Sultana

Graduate Theses/Dissertations

Voice-enabled interactions have become increasingly popular with the rise of voice assistants. Identifying contexts or meanings from voice commands and conversations with smart assistants can contribute to the autonomous control of smart home devices and appliances. To improve automation, there is a growing need for efficient context detection that eliminates the need to memorize voice commands. To address this need, I followed a two-step approach in my research. In the first step, I developed a unique context recognition model using a transformer, an attention mechanism, and a fully connected neural network. I trained this model on a conversational dataset of …


Generative Artificial Intelligence: Basic Terminology And Concepts, Kincaid Brown Jan 2024

Generative Artificial Intelligence: Basic Terminology And Concepts, Kincaid Brown

Law Librarian Scholarship

Generative artificial intelligence (GenAI) has been a hard topic to avoid in the media for more than a year. But what do all of the terms mean and what are areas of concern with GenAI tools?

This column aims to provide a baseline explanation of terminology and concepts that are frequently in the media.


Integrating Generative Artificial Intelligence With Systems Architecting Diagram Creation: Advancement, Challenges, Opportunities And Future Perspectives, Cansu Yalim, Holly H. Handley Jan 2024

Integrating Generative Artificial Intelligence With Systems Architecting Diagram Creation: Advancement, Challenges, Opportunities And Future Perspectives, Cansu Yalim, Holly H. Handley

Engineering Management & Systems Engineering Faculty Publications

Generative AI (GenAI) serves as a powerful tool that can create a wide range of content, including but not limited to text, speech, images, code, videos, and 3D models. ChatGPT stands out as a particularly appealing Generative Pretrained Transformer (GPT) model that offers supplementary capabilities through GPTs and plugins. These extensions enable users to engage with the chatbot and improve its functionality, surpassing mere content generation. Our study delves into the potential of ChatGPT, specifically GPT-4, to expedite the creation of diagrams to support the system architecting process. To this end, we explored the use of ChatGPT's Diagrams Show Me …


A Comparison Of Lexical Tokenization Methods, Nathan Culmer Jan 2024

A Comparison Of Lexical Tokenization Methods, Nathan Culmer

Williams Honors College, Honors Research Projects

The purpose of this project was to compare tokenization methods, or methods of breaking up a text into meaningful parts for use in natural language processing. The effectiveness of several commonly used tokenization methods were investigated, including morpheme tokenization, which takes into account the linguistic features of the language. In addition, I proposed and implemented a new technique to consider the capitalization pattern of a word in the tokenization process, in order to allow this process to include more natural language features. The effectiveness of these methods was compared by using them in a sentiment analysis model for various datasets, …


Conflict Profiles And Team Outcomes In Cross-Disciplinary Teams: An Integrated Latent Profile Analysis And Natural Language Processing Approach, Francisco Cima, Pilar Pazos Jan 2024

Conflict Profiles And Team Outcomes In Cross-Disciplinary Teams: An Integrated Latent Profile Analysis And Natural Language Processing Approach, Francisco Cima, Pilar Pazos

Engineering Management & Systems Engineering Faculty Publications

Team conflict is a naturally emerging phenomenon resulting from individuals' interactions during project execution. Cross-disciplinary teams can experience higher levels of conflict than single-discipline teams because of the increased diversity of knowledge and perspectives. Research has shown that team conflict can emerge from different types of disagreements (cognitive and interpersonal), which have different implications for team functioning. Past empirical research has focused on the impact of both conflict types independent from each other while overlooking their combined effects. This work examines the conflict profiles resulting from the combined levels of interpersonal and cognitive disagreements and their association with team outcomes. …


Uncertainty Quantification In Large Language Models Through Convex Hull Analysis, Ferhat Ozgur Catak, Murat Kuzlu Jan 2024

Uncertainty Quantification In Large Language Models Through Convex Hull Analysis, Ferhat Ozgur Catak, Murat Kuzlu

Engineering Technology Faculty Publications

Uncertainty quantification approaches have been more critical in large language models (LLMs), particularly high-risk applications requiring reliable outputs. However, traditional methods for uncertainty quantification, such as probabilistic models and ensemble techniques, face challenges when applied to the complex and high-dimensional nature of LLM-generated outputs. This study proposes a novel geometric approach to uncertainty quantification using convex hull analysis. The proposed method leverages the spatial properties of response embeddings to measure the dispersion and variability of model outputs. The prompts are categorized into three types, i.e., ’easy’, ’moderate’, and ’confusing’, to generate multiple responses using different LLMs at varying temperature settings. …


A Survey On Few-Shot Class-Incremental Learning, Songsong Tian, Lusi Li, Weijun Li, Hang Ran, Xin Ning, Prayag Tiwari Jan 2024

A Survey On Few-Shot Class-Incremental Learning, Songsong Tian, Lusi Li, Weijun Li, Hang Ran, Xin Ning, Prayag Tiwari

Computer Science Faculty Publications

Large deep learning models are impressive, but they struggle when real-time data is not available. Few-shot class-incremental learning (FSCIL) poses a significant challenge for deep neural networks to learn new tasks from just a few labeled samples without forgetting the previously learned ones. This setup can easily leads to catastrophic forgetting and overfitting problems, severely affecting model performance. Studying FSCIL helps overcome deep learning model limitations on data volume and acquisition time, while improving practicality and adaptability of machine learning models. This paper provides a comprehensive survey on FSCIL. Unlike previous surveys, we aim to synthesize few-shot learning and incremental …


A Chinese Power Text Classification Algorithm Based On Deep Active Learning, Song Deng, Qianliang Li, Renjie Dai, Siming Wei, Di Wu, Yi He, Xindong Wu Jan 2024

A Chinese Power Text Classification Algorithm Based On Deep Active Learning, Song Deng, Qianliang Li, Renjie Dai, Siming Wei, Di Wu, Yi He, Xindong Wu

Computer Science Faculty Publications

The construction of knowledge graph is beneficial for grid production, electrical safety protection, fault diagnosis and traceability in an observable and controllable way. Highly-precision text classification algorithm is crucial to build a professional knowledge graph in power system. Unfortunately, there are a large number of poorly described and specialized texts in the power business system, and the amount of data containing valid labels in these texts is low. This will bring great challenges to improve the precision of text classification models. To offset the gap, we propose a classification algorithm for Chinese text in the power system based on deep …


Assessing The Effectiveness Of A Chatbot Workshop As Experiential Teaching And Learning Tool To Engage Undergraduate Students, Kyong Jin Shim, Thomas Menkhoff, Ying Qian Teo, Clement Shi Qi Ong Dec 2023

Assessing The Effectiveness Of A Chatbot Workshop As Experiential Teaching And Learning Tool To Engage Undergraduate Students, Kyong Jin Shim, Thomas Menkhoff, Ying Qian Teo, Clement Shi Qi Ong

Research Collection School Of Computing and Information Systems

In this paper, we empirically examine and assess the effectiveness of a chatbot workshop as experiential teaching and learning tool to engage undergraduate students enrolled in an elective course “Doing Business with A.I.” in the Lee Kong Chian School of Business (LKCSB) at Singapore Management University. The chatbot workshop provides non-STEM students with an opportunity to acquire basic skills to build a chatbot prototype using the ‘Dialogflow’ program. The workshop and the experiential learning activity are designed to impart conversation and user-centric design know how and know why to students. A key didactical aspect which informs the design and flow …


N-Shot Benchmarking Of Whisper On Diverse Arabic Speech Recognition, Bashar Talafha, Abdul Waheed, Muhammad Abdul-Mageed Aug 2023

N-Shot Benchmarking Of Whisper On Diverse Arabic Speech Recognition, Bashar Talafha, Abdul Waheed, Muhammad Abdul-Mageed

Natural Language Processing Faculty Publications

Whisper, the recently developed multilingual weakly supervised model, is reported to perform well on multiple speech recognition benchmarks in both monolingual and multilingual settings. However, it is not clear how Whisper would fare under diverse conditions even on languages it was evaluated on such as Arabic. In this work, we address this gap by comprehensively evaluating Whisper on several varieties of Arabic speech for the ASR task. Our evaluation covers most publicly available Arabic speech data and is performed under n-shot (zero-, few-, and full) finetuning. We also investigate the robustness of Whisper under completely novel conditions, such as in …


Ocr Post-Processing Using Large Language Models, Mahdi Hajiali Aug 2023

Ocr Post-Processing Using Large Language Models, Mahdi Hajiali

UNLV Theses, Dissertations, Professional Papers, and Capstones

Optical Character Recognition (OCR) technology transforms textual visuals into an electronically readable, non-graphical format of the text. This allows the editing and other text manipulation of the content by language technology software such as machine translation, text comprehension, query-answering systems, and search engines. While Optical Character Recognition (OCR) systems continually progress towards greater precision, several complications persist when dealing with low-resolution source images or those with multicolored backgrounds. Consequently, the text derived from OCR necessitates additional refinement to optimize accuracy, beneficial for various subsequent applications. It is recognized that the character accuracy of OCR-generated text may influence certain natural language …


Multi-Head Attention Graph Convolutional Network Model: End-To-End Entity And Relation Joint Extraction Based On Multi-Head Attention Graph Convolutional Network, Zhihua Tao, Chunping Ouyang, Yongbin Liu, Tonglee Chung, Yixin Cao Jun 2023

Multi-Head Attention Graph Convolutional Network Model: End-To-End Entity And Relation Joint Extraction Based On Multi-Head Attention Graph Convolutional Network, Zhihua Tao, Chunping Ouyang, Yongbin Liu, Tonglee Chung, Yixin Cao

Research Collection School Of Computing and Information Systems

At present, the entity and relation joint extraction task has attracted more and more scholars' attention in the field of natural language processing (NLP). However, most of their methods rely on NLP tools to construct dependency trees to obtain sentence structure information. The adjacency matrix constructed by the dependency tree can convey syntactic information. Dependency trees obtained through NLP tools are too dependent on the tools and may not be very accurate in contextual semantic description. At the same time, a large amount of irrelevant information will cause redundancy. This paper presents a novel end-to-end entity and relation joint extraction …


Law Smells: Defining And Detecting Problematic Patterns In Legal Drafting, Corinna Coupette, Dirk Hartung, Janis Beckedorf, Maximilian Böther, Daniel Martin Katz Jun 2023

Law Smells: Defining And Detecting Problematic Patterns In Legal Drafting, Corinna Coupette, Dirk Hartung, Janis Beckedorf, Maximilian Böther, Daniel Martin Katz

Research Collection Yong Pung How School Of Law

Building on the computer science concept of code smells, we initiate the study of law smells, i.e., patterns in legal texts that pose threats to the comprehensibility and maintainability of the law. With five intuitive law smells as running examples—namely, duplicated phrase, long element, large reference tree, ambiguous syntax, and natural language obsession—, we develop a comprehensive law smell taxonomy. This taxonomy classifies law smells by when they can be detected, which aspects of law they relate to, and how they can be discovered. We introduce textbased and graph-based methods to identify instances of law smells, confirming their utility in …


Law Smells: Defining And Detecting Problematic Patterns In Legal Drafting, Corinna Coupette, Dirk Hartung, Janis Beckedorf, Maximilian Bother, Daniel Martin Katz Jun 2023

Law Smells: Defining And Detecting Problematic Patterns In Legal Drafting, Corinna Coupette, Dirk Hartung, Janis Beckedorf, Maximilian Bother, Daniel Martin Katz

Research Collection Yong Pung How School Of Law

Building on the computer science concept of code smells, we initiate the study of law smells, i.e., patterns in legal texts that pose threats to the comprehensibility and maintainability of the law. With five intuitive law smells as running examples—namely, duplicated phrase, long element, large reference tree, ambiguous syntax, and natural language obsession—, we develop a comprehensive law smell taxonomy. This taxonomy classifies law smells by when they can be detected, which aspects of law they relate to, and how they can be discovered. We introduce text-based and graph-based methods to identify instances of law smells, confirming their utility in …