Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Computer Sciences (24)
- Life Sciences (20)
- Medicine and Health Sciences (20)
- Bioinformatics (19)
- Biomedical Informatics (18)
-
- Artificial Intelligence and Robotics (15)
- Social and Behavioral Sciences (12)
- Engineering (9)
- Linguistics (7)
- Business (4)
- Communication (4)
- Statistics and Probability (4)
- Computational Engineering (3)
- Computational Linguistics (3)
- Databases and Information Systems (3)
- Electrical and Computer Engineering (3)
- Medical Specialties (3)
- Numerical Analysis and Scientific Computing (3)
- Other Computer Sciences (3)
- Political Science (3)
- Software Engineering (3)
- Statistical Models (3)
- Systems and Communications (3)
- American Politics (2)
- Analytical, Diagnostic and Therapeutic Techniques and Equipment (2)
- Applied Mathematics (2)
- Business Intelligence (2)
- Cognitive Psychology (2)
- Institution
-
- The Texas Medical Center Library (17)
- Southern Methodist University (4)
- Kennesaw State University (3)
- Technological University Dublin (3)
- Virginia Commonwealth University (3)
-
- California Polytechnic State University, San Luis Obispo (2)
- Chapman University (2)
- City University of New York (CUNY) (2)
- Dartmouth College (2)
- The University of Southern Mississippi (2)
- University of Arkansas, Fayetteville (2)
- University of Kentucky (2)
- Bellarmine University (1)
- Clemson University (1)
- DePaul University (1)
- Florida Institute of Technology (1)
- Karbala International Journal of Modern Science (1)
- LSU New Orleans (1)
- Mississippi State University (1)
- New Jersey Institute of Technology (1)
- San Jose State University (1)
- Universitas Negeri Malang (1)
- University of Mary Washington (1)
- Washington University in St. Louis (1)
- Publication
-
- Faculty, Staff and Student Publications (17)
- Dissertations (4)
- SMU Data Science Review (4)
- Theses and Dissertations (4)
- Dartmouth College Undergraduate Theses (2)
-
- Data Science Undergraduate Honors Theses (2)
- Theses and Dissertations--Computer Science (2)
- All Dissertations (1)
- Articles (1)
- Capstone Projects (1)
- College of Computing and Digital Media Dissertations (1)
- College of Engineering Summer Undergraduate Research Program (1)
- Computational and Data Sciences (MS) Theses (1)
- Departmental Honors & Graduate Capstone Projects (1)
- Dissertations, Theses, and Capstone Projects (1)
- Doctor of Data Science and Analytics Dissertations (1)
- Engineering Faculty Articles and Research (1)
- Karbala International Journal of Modern Science (1)
- Knowledge Engineering and Data Science (1)
- LSU New Orleans Theses and Dissertations (1)
- Library Philosophy and Practice (e-journal) (1)
- Master's Theses (1)
- Open Educational Resources (1)
- Other Resources (1)
- Other resources (1)
- Published and Grey Literature from PhD Candidates (1)
- Scholarship@WashULaw (1)
- Undergraduate Theses (1)
- Publication Type
Articles 31 - 56 of 56
Full-Text Articles in Data Science
Identifying Features And Predicting Consumer Helpfulness Of Product Reviews, Triston Hudgins, Shijo Joseph, Douglas Yip, Gaston Besanson
Identifying Features And Predicting Consumer Helpfulness Of Product Reviews, Triston Hudgins, Shijo Joseph, Douglas Yip, Gaston Besanson
SMU Data Science Review
Major corporations utilize data from online platforms to make user product or service recommendations. Companies like Netflix, Amazon, Yelp, and Spotify rely on purchasing trends, user reviews, and helpfulness votes to make content recommendations. This strategy can increase user engagement on a company's platform. However, misleading and/or spam reviews significantly hinder the success of these recommendation strategies. The rise of social media has made it increasingly difficult to distinguish between authentic content and advertising, leading to a burst of deceptive reviews across the marketplace. The helpfulness of the review is subjective to a voting system. As such, this study aims …
Chatgpt As Metamorphosis Designer For The Future Of Artificial Intelligence (Ai): A Conceptual Investigation, Amarjit Kumar Singh (Library Assistant), Dr. Pankaj Mathur (Deputy Librarian)
Chatgpt As Metamorphosis Designer For The Future Of Artificial Intelligence (Ai): A Conceptual Investigation, Amarjit Kumar Singh (Library Assistant), Dr. Pankaj Mathur (Deputy Librarian)
Library Philosophy and Practice (e-journal)
Abstract
Purpose: The purpose of this research paper is to explore ChatGPT’s potential as an innovative designer tool for the future development of artificial intelligence. Specifically, this conceptual investigation aims to analyze ChatGPT’s capabilities as a tool for designing and developing near about human intelligent systems for futuristic used and developed in the field of Artificial Intelligence (AI). Also with the helps of this paper, researchers are analyzed the strengths and weaknesses of ChatGPT as a tool, and identify possible areas for improvement in its development and implementation. This investigation focused on the various features and functions of ChatGPT that …
Practical Ai Value Alignment Using Stories, Md Sultan Al Nahian
Practical Ai Value Alignment Using Stories, Md Sultan Al Nahian
Theses and Dissertations--Computer Science
As more machine learning agents interact with humans, it is increasingly a prospect that an agent trained to perform a task optimally - using only a measure of task performance as feedback--can violate societal norms for acceptable behavior or cause harm. Consequently, it becomes necessary to prioritize task performance and ensure that AI actions do not have detrimental effects. Value alignment is a property of intelligent agents, wherein they solely pursue goals and activities that are non-harmful and beneficial to humans. Current approaches to value alignment largely depend on imitation learning or learning from demonstration methods. However, the dynamic nature …
Use Gpt-J Prompt Generation With Roberta For Ner Models On Diagnosis Extraction Of Periodontal Diagnosis From Electronic Dental Records, Yao-Shun Chuang, Xiaoqian Jiang, Chun-Teh Lee, Ryan Brandon, Duong Tran, Oluwabunmi Tokede, Muhammad F Walji
Use Gpt-J Prompt Generation With Roberta For Ner Models On Diagnosis Extraction Of Periodontal Diagnosis From Electronic Dental Records, Yao-Shun Chuang, Xiaoqian Jiang, Chun-Teh Lee, Ryan Brandon, Duong Tran, Oluwabunmi Tokede, Muhammad F Walji
Faculty, Staff and Student Publications
This study explored the usability of prompt generation on named entity recognition (NER) tasks and the performance in different settings of the prompt. The prompt generation by GPT-J models was utilized to directly test the gold standard as well as to generate the seed and further fed to the RoBERTa model with the spaCy package. In the direct test, a lower ratio of negative examples with higher numbers of examples in prompt achieved the best results with a F1 score of 0.72. The performance revealed consistency, 0.92-0.97 in the F1 score, in all settings after training with the RoBERTa model. …
Contextual Variation Of Clinical Notes Induced By Ehr Migration, Kurt Miller, Sungrim Moon, Sunyang Fu, Hongfang Liu
Contextual Variation Of Clinical Notes Induced By Ehr Migration, Kurt Miller, Sungrim Moon, Sunyang Fu, Hongfang Liu
Faculty, Staff and Student Publications
The structure and semantics of clinical notes vary considerably across different Electronic Health Record (EHR) systems, sites, and institutions. Such heterogeneity hampers the portability of natural language processing (NLP) models in extracting information from the text for clinical research or practice. In this study, we evaluate the contextual variation of clinical notes by measuring the semantic and syntactic similarity of the notes of two sets of physicians comprising four medical specialties across EHR migrations at two Mayo Clinic sites. We find significant semantic and syntactic variation imposed by the context of the EHR system and between medical specialties whereas only …
Text Classification Of Cancer Clinical Trial Eligibility Criteria, Yumeng Yang, Soumya Jayaraj, Ethan Ludmir, Kirk Roberts
Text Classification Of Cancer Clinical Trial Eligibility Criteria, Yumeng Yang, Soumya Jayaraj, Ethan Ludmir, Kirk Roberts
Faculty, Staff and Student Publications
Automatic identification of clinical trials for which a patient is eligible is complicated by the fact that trial eligibility are stated in natural language. A potential solution to this problem is to employ text classification methods for common types of eligibility criteria. In this study, we focus on seven common exclusion criteria in cancer trials: prior malignancy, human immunodeficiency virus, hepatitis B, hepatitis C, psychiatric illness, drug/substance abuse, and autoimmune illness. Our dataset consists of 764 phase III cancer trials with these exclusions annotated at the trial level. We experiment with common transformer models as well as a new pre-trained …
Transfer Of Personality Through Text Style, Michael O'Mahony, Robert Ross
Transfer Of Personality Through Text Style, Michael O'Mahony, Robert Ross
Other resources
The style of generated text is how something is said rather than what is said. We hypothesize that changing the style of generated text can change the perceived personality of the text generation agent. Dialogue systems that aim to imitate a human agent can appear to have a consistent personality through a consistent, controllable style of conversation. Some recent work on the style of generated text [1] performs impressively in the small number of domains selected for their experiments using transformer and LSTM-based models. Lin et al. [1] used weak supervised learning as their data set lacks parallel data. The …
Assessment Of Electronic Health Record For Cancer Research And Patient Care Through A Scoping Review Of Cancer Natural Language Processing, Liwei Wang, Sunyang Fu, Andrew Wen, Xiaoyang Ruan, Huan He, Sijia Liu, Sungrim Moon, Michelle Mai, Irbaz B Riaz, Nan Wang, Ping Yang, Hua Xu, Jeremy L Warner, Hongfang Liu
Assessment Of Electronic Health Record For Cancer Research And Patient Care Through A Scoping Review Of Cancer Natural Language Processing, Liwei Wang, Sunyang Fu, Andrew Wen, Xiaoyang Ruan, Huan He, Sijia Liu, Sungrim Moon, Michelle Mai, Irbaz B Riaz, Nan Wang, Ping Yang, Hua Xu, Jeremy L Warner, Hongfang Liu
Faculty, Staff and Student Publications
Purpose: The advancement of natural language processing (NLP) has promoted the use of detailed textual data in electronic health records (EHRs) to support cancer research and to facilitate patient care. In this review, we aim to assess EHR for cancer research and patient care by using the Minimal Common Oncology Data Elements (mCODE), which is a community-driven effort to define a minimal set of data elements for cancer research and practice. Specifically, we aim to assess the alignment of NLP-extracted data elements with mCODE and review existing NLP methodologies for extracting said data elements.
Methods: Published literature studies were searched …
A Large-Scale Sentiment Analysis Of Tweets Pertaining To The 2020 Us Presidential Election, Rao Hamza Ali, Gabriela Pinto, Evelyn Lawrie, Erik J. Linstead
A Large-Scale Sentiment Analysis Of Tweets Pertaining To The 2020 Us Presidential Election, Rao Hamza Ali, Gabriela Pinto, Evelyn Lawrie, Erik J. Linstead
Engineering Faculty Articles and Research
We capture the public sentiment towards candidates in the 2020 US Presidential Elections, by analyzing 7.6 million tweets sent out between October 31st and November 9th, 2020. We apply a novel approach to first identify tweets and user accounts in our database that were later deleted or suspended from Twitter. This approach allows us to observe the sentiment held for each presidential candidate across various groups of users and tweets: accessible tweets and accounts, deleted tweets and accounts, and suspended or inaccessible tweets and accounts. We compare the sentiment scores calculated for these groups and provide key insights into the …
Legislative Language For Success, Sanjana Gundala
Legislative Language For Success, Sanjana Gundala
Master's Theses
Legislative committee meetings are an integral part of the lawmaking process for local and state bills. The testimony presented during these meetings is a large factor in the outcome of the proposed bill. This research uses Natural Language Processing and Machine Learning techniques to analyze testimonies from California Legislative committee meetings from 2015-2016 in order to identify what aspects of a testimony makes it successful. A testimony is considered successful if the alignment of the testimony matches the bill outcome (alignment is "For" and the bill passes or alignment is "Against" and the bill fails). The process of finding what …
Directional Pairwise Class Confusion Bias And Its Mitigation, Sudhashree Sayenju, Ramazan Aygun Phd, Jonathan Boardman, Duleep Prasanna Rathgamage Don, Yifan Zhang Phd, Bill Franks, Sereres Johnston Phd, George Lee, Dan Sullivan, Girish Modgil Phd
Directional Pairwise Class Confusion Bias And Its Mitigation, Sudhashree Sayenju, Ramazan Aygun Phd, Jonathan Boardman, Duleep Prasanna Rathgamage Don, Yifan Zhang Phd, Bill Franks, Sereres Johnston Phd, George Lee, Dan Sullivan, Girish Modgil Phd
Published and Grey Literature from PhD Candidates
Recent advances in Natural Language Processing have led to powerful and sophisticated models like BERT (Bidirectional Encoder Representations from Transformers) that have bias. These models are mostly trained on text corpora that deviate in important ways from the text encountered by a chatbot in a problem-specific context. While a lot of research in the past has focused on measuring and mitigating bias with respect to protected attributes (stereotyping like gender, race, ethnicity, etc.), there is lack of research in model bias with respect to classification labels. We investigate whether a classification model hugely favors one class with respect to another. …
A Study On Developing Novel Methods For Relation Extraction, Darshini Mahendran
A Study On Developing Novel Methods For Relation Extraction, Darshini Mahendran
Theses and Dissertations
Relation Extraction (RE) is a task of Natural Language Processing (NLP) to detect and classify the relations between two entities. Relation extraction in the biomedical and scientific literature domain is challenging as text can contain multiple pairs of entities in the same instance. During the course of this research, we developed an RE framework (RelEx), which consists of five main RE paradigms: rule-based, machine learning-based, Convolutional Neural Network (CNN)-based, Bidirectional Encoder Representations from Transformers (BERT)-based, and Graph Convolutional Networks (GCNs)-based approaches. RelEx's rule-based approach uses co-location information of the entities to determine whether a relation exists between a selected entity …
The Detection Of Sexual Harassment And Chat Predators Using Artificial Neural Network, Noor Amer Hamzah, Ban N. Dhannoon
The Detection Of Sexual Harassment And Chat Predators Using Artificial Neural Network, Noor Amer Hamzah, Ban N. Dhannoon
Karbala International Journal of Modern Science
The vast increase in using social media sites like Twitter and Facebook led to frequent sexual_harassment on the Internet, which is considered a major societal problem. This paper aims to detect sexual_harassment and cyber_predators in early phase. We used deeplearning like Bidirectionally-long-short-term memory. Word representations are carefully reviewed in text specific to mapping to real number vectors. The chat sexual predators Detection_approach with the proposed_model. The best results obtained by the performance measured with F0.5-score were the result is_0.927 with proposed_models. The accuracy measured is_97.27% in the proposed_model. The comments sexual_harassment Detection_approach the result is_0.925 F0.5-score, and accuracy measured is_99.12%.
Lexical Complexity Prediction With Assembly Models, Aadil Islam
Lexical Complexity Prediction With Assembly Models, Aadil Islam
Dartmouth College Undergraduate Theses
Tuning the complexity of one's writing is essential to presenting ideas in a logical, intuitive manner to audiences. This paper describes a system submitted by team BigGreen to LCP 2021 for predicting the lexical complexity of English words in a given context. We assemble a feature engineering-based model and a deep neural network model with an underlying Transformer architecture based on BERT. While BERT itself performs competitively, our feature engineering-based model helps in extreme cases, eg. separating instances of easy and neutral difficulty. Our handcrafted features comprise a breadth of lexical, semantic, syntactic, and novel phonetic measures. Visualizations of BERT …
Fine-Grained Detection Of Hate Speech Using Bertoxic, Yakoob Khan
Fine-Grained Detection Of Hate Speech Using Bertoxic, Yakoob Khan
Dartmouth College Undergraduate Theses
This thesis describes our approach towards the fine-grained detection of hate speech using deep learning. We leverage the transformer encoder architecture to propose BERToxic, a system that fine-tunes a pre-trained BERT model to locate toxic text spans in a given text and utilizes additional post-processing steps to refine the prediction boundaries. The post-processing steps involve (1) labeling character offsets between consecutive toxic tokens as toxic and (2) assigning a toxic label to words that have at least one token labeled as toxic. Through experiments, we show that these two post-processing steps improve the performance of our model by 4.16% on …
Automated Analysis Of Rfps Using Natural Language Processing (Nlp) For The Technology Domain, Sterling Beason, William Hinton, Yousri A. Salamah, Jordan Salsman
Automated Analysis Of Rfps Using Natural Language Processing (Nlp) For The Technology Domain, Sterling Beason, William Hinton, Yousri A. Salamah, Jordan Salsman
SMU Data Science Review
Much progress has been made in text analysis, specifically within the statistical domain of Term Frequency (TF) and Inverse Document Frequency (IDF). However, there is much room for improvement especially within the area of discovering Emerging Trends. Emerging Trend Detection Systems (ETDS) depend on ingesting a collection of textual data and TF/IDF to identify new or up-trending topics within the Corpus. However, the tremendous rate of change and the amount of digital information presents a challenge that makes it almost impossible for a human expert to spot emerging trends without relying on an automated ETD system. Since the U.S. Government …
Semantic Classification Of Multidialectal Arabic Social Media, Tom Rishel
Semantic Classification Of Multidialectal Arabic Social Media, Tom Rishel
Dissertations
Arabic is one of the most widely used languages in the world, but due in part to its morphological and syntactic richness, resources for automated processing of Arabic are relatively rare. Arabic takes three primary forms: Classical Arabic as seen in the Qur’an and other classical texts; Modern Standard Arabic (MSA) as seen in newspapers, formal documents, and other written text intended for widespread distribution; and dialectal Arabic as used in common speech and informal communication. Social media posts are often written in informal language and may include non-standard spellings, abbreviations, emoticons, hashtags, and emojis. Dialectal Arabic is commonly used …
Generalized And Transferable Patient Language Representation For Phenotyping With Limited Data, Yuqi Si, Elmer V Bernstam, Kirk Roberts
Generalized And Transferable Patient Language Representation For Phenotyping With Limited Data, Yuqi Si, Elmer V Bernstam, Kirk Roberts
Faculty, Staff and Student Publications
The paradigm of representation learning through transfer learning has the potential to greatly enhance clinical natural language processing. In this work, we propose a multi-task pre-training and fine-tuning approach for learning generalized and transferable patient representations from medical language. The model is first pre-trained with different but related high-prevalence phenotypes and further fine-tuned on downstream target tasks. Our main contribution focuses on the impact this technique can have on low-prevalence phenotypes, a challenging task due to the dearth of data. We validate the representation from pre-training, and fine-tune the multi-task pre-trained models on low-prevalence phenotypes including 38 circulatory diseases, 23 …
Improving Space Efficiency Of Deep Neural Networks, Aliakbar Panahi
Improving Space Efficiency Of Deep Neural Networks, Aliakbar Panahi
Theses and Dissertations
Language models employ a very large number of trainable parameters. Despite being highly overparameterized, these networks often achieve good out-of-sample test performance on the original task and easily fine-tune to related tasks. Recent observations involving, for example, intrinsic dimension of the objective landscape and the lottery ticket hypothesis, indicate that often training actively involves only a small fraction of the parameter space. Thus, a question remains how large a parameter space needs to be in the first place — the evidence from recent work on model compression, parameter sharing, factorized representations, and knowledge distillation increasingly shows that models can be …
Ensemble Labeling Towards Scientific Information Extraction (Elsie), Erin Murphy
Ensemble Labeling Towards Scientific Information Extraction (Elsie), Erin Murphy
College of Computing and Digital Media Dissertations
Extracting scientific facts from unstructured text is difficult due to challenges specific to the ambiguity of the language, the complexity of the scientific named entities and relations to be extracted. This problem is well illustrated through the extraction of polymer names and their properties. Even in the cases where the property is a temperature, identifying the polymer name associated with the temperature may require expertise due to the use of acronyms, synonyms, complicated naming conventions and by the fact that new polymer names are being “introduced” to the vernacular as polymer science advances. While there exist domain-specific machine learning toolkits …
Toxic Language Detection Using Robust Filters, Deepti Kunupudi, Shantanu Godbole, Pankaj Kumar, Suhas Pai
Toxic Language Detection Using Robust Filters, Deepti Kunupudi, Shantanu Godbole, Pankaj Kumar, Suhas Pai
SMU Data Science Review
Social networks sometimes become a medium for threats, insults, and other types of cyberbullying. A large number of people are involved in online social networks. Hence, the protection of network users from anti-social behavior is a critical activity [19]. One of the significant tasks of such activity is the detection of toxic language. Abusive/Toxic language in user-generated online content has become an issue of increasing importance in recent years. Most current commercial methods use blacklists and regular expressions; however, these measures fall short when contending with more subtle, lesser-known examples of hate speech, profanity, or swearing[6]. Abusive language classification has …
A Study Of Information Bots And Knowledge Bots, Amartya Hatua
A Study Of Information Bots And Knowledge Bots, Amartya Hatua
Dissertations
In this dissertation, a study of different aspects of information bots and knowledge bots is done. The research contributes to a better understanding of the various characteristics of information bots as well as the different patterns and factors responsible for the information diffusion in a social network. This research also shows how these factors can be used to predict information diffusion for a particular topic in a social network. The second part of the research is focused on strategies for improving the knowledge base of knowledge bots, where two different approaches are studied. In the first approach, knowledge is transferred …
Understanding Spatial Language In Radiology: Representation Framework, Annotation, And Spatial Relation Extraction From Chest X-Ray Reports Using Deep Learning, Surabhi Datta, Yuqi Si, Laritza Rodriguez, Sonya E Shooshan, Dina Demner-Fushman, Kirk Roberts
Understanding Spatial Language In Radiology: Representation Framework, Annotation, And Spatial Relation Extraction From Chest X-Ray Reports Using Deep Learning, Surabhi Datta, Yuqi Si, Laritza Rodriguez, Sonya E Shooshan, Dina Demner-Fushman, Kirk Roberts
Faculty, Staff and Student Publications
Radiology reports contain a radiologist's interpretations of images, and these images frequently describe spatial relations. Important radiographic findings are mostly described in reference to an anatomical location through spatial prepositions. Such spatial relationships are also linked to various differential diagnoses and often described through uncertainty phrases. Structured representation of this clinically significant spatial information has the potential to be used in a variety of downstream clinical informatics applications. Our focus is to extract these spatial representations from the reports. For this, we first define a representation framework based on the Spatial Role Labeling (SpRL) scheme, which we refer to as …
Deep Learning In Clinical Natural Language Processing: A Methodical Review, Stephen Wu, Kirk Roberts, Surabhi Datta, Jingcheng Du, Zongcheng Ji, Yuqi Si, Sarvesh Soni, Qiong Wang, Qiang Wei, Yang Xiang, Bo Zhao, Hua Xu
Deep Learning In Clinical Natural Language Processing: A Methodical Review, Stephen Wu, Kirk Roberts, Surabhi Datta, Jingcheng Du, Zongcheng Ji, Yuqi Si, Sarvesh Soni, Qiong Wang, Qiang Wei, Yang Xiang, Bo Zhao, Hua Xu
Faculty, Staff and Student Publications
OBJECTIVE: This article methodically reviews the literature on deep learning (DL) for natural language processing (NLP) in the clinical domain, providing quantitative analysis to answer 3 research questions concerning methods, scope, and context of current research.
MATERIALS AND METHODS: We searched MEDLINE, EMBASE, Scopus, the Association for Computing Machinery Digital Library, and the Association for Computational Linguistics Anthology for articles using DL-based approaches to NLP problems in electronic health records. After screening 1,737 articles, we collected data on 25 variables across 212 papers.
RESULTS: DL in clinical NLP publications more than doubled each year, through 2018. Recurrent neural networks (60.8%) …
Enhancing Clinical Concept Extraction With Contextual Embeddings, Yuqi Si, Jingqi Wang, Hua Xu, Kirk Roberts
Enhancing Clinical Concept Extraction With Contextual Embeddings, Yuqi Si, Jingqi Wang, Hua Xu, Kirk Roberts
Faculty, Staff and Student Publications
OBJECTIVE: Neural network-based representations ("embeddings") have dramatically advanced natural language processing (NLP) tasks, including clinical NLP tasks such as concept extraction. Recently, however, more advanced embedding methods and representations (eg, ELMo, BERT) have further pushed the state of the art in NLP, yet there are no common best practices for how to integrate these representations into clinical tasks. The purpose of this study, then, is to explore the space of possible options in utilizing these new models for clinical concept extraction, including comparing these to traditional word embedding methods (word2vec, GloVe, fastText).
MATERIALS AND METHODS: Both off-the-shelf, open-domain embeddings and …
A Frame-Based Nlp System For Cancer-Related Information Extraction, Yuqi Si, Kirk Roberts
A Frame-Based Nlp System For Cancer-Related Information Extraction, Yuqi Si, Kirk Roberts
Faculty, Staff and Student Publications
We propose a frame-based natural language processing (NLP) method that extracts cancer-related information from clinical narratives. We focus on three frames: cancer diagnosis, cancer therapeutic procedure, and tumor description. We utilize a deep learning-based approach, bidirectional Long Short-term Memory (LSTM) Conditional Random Field (CRF), which uses both character and word embeddings. The system consists of two constituent sequence classifiers: a frame identification (lexical unit) classifier and a frame element classifier. The classifier achieves an F