Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Computer Sciences (24)
- Medicine and Health Sciences (14)
- Artificial Intelligence and Robotics (11)
- Biomedical Informatics (8)
- Social and Behavioral Sciences (8)
-
- Bioinformatics (7)
- Life Sciences (7)
- Communication (5)
- Statistics and Probability (5)
- Engineering (4)
- Linguistics (3)
- Mental and Social Health (3)
- Other Computer Sciences (3)
- Social Media (3)
- Sociology (3)
- Business (2)
- Categorical Data Analysis (2)
- Computational Linguistics (2)
- Computer Engineering (2)
- Education (2)
- Electrical and Computer Engineering (2)
- Law (2)
- Mathematics (2)
- Other Mathematics (2)
- Political Science (2)
- Software Engineering (2)
- Translational Medical Research (2)
- Allergy and Immunology (1)
- Institution
-
- Southern Methodist University (9)
- The Texas Medical Center Library (7)
- New Jersey Institute of Technology (4)
- Virginia Commonwealth University (3)
- Air Force Institute of Technology (2)
-
- Dartmouth College (2)
- University of Kentucky (2)
- Binghamton University (1)
- California Polytechnic State University, San Luis Obispo (1)
- California State University, San Bernardino (1)
- Harrisburg University of Science and Technology (1)
- Indian Statistical Institute (1)
- Kennesaw State University (1)
- Loyola University Chicago (1)
- Minnesota State University, Mankato (1)
- Mississippi State University (1)
- Missouri State University (1)
- Old Dominion University (1)
- Rochester Institute of Technology (1)
- Seattle Pacific University (1)
- University of Denver (1)
- University of Malaya (1)
- West Virginia University (1)
- Western University (1)
- Publication
-
- SMU Data Science Review (9)
- Faculty, Staff and Student Publications (7)
- Theses and Dissertations (5)
- Dissertations (4)
- All Graduate Theses, Dissertations, and Other Capstone Projects (1)
-
- Articles (1)
- College of Engineering Summer Undergraduate Research Program (1)
- Computer Science Senior Theses (1)
- Computer Science: Faculty Publications and Other Works (1)
- Dartmouth College Master’s Theses (1)
- Doctoral Theses (1)
- Electrical and Computer Engineering Publications (1)
- Electronic Theses and Dissertations (1)
- Electronic Theses, Projects, and Dissertations (1)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (1)
- Graduate Theses/Dissertations (1)
- Harrisburg University Other Works (1)
- Honors Projects (1)
- Master of Science in Computer Science Theses (1)
- Modeling, Simulation and Visualization Student Capstone Conference (1)
- Northeast Journal of Complex Systems (NEJCS) (1)
- Student Works (2020-2029) (1)
- Theses and Dissertations--Computer Science (1)
- Theses and Dissertations--Epidemiology and Biostatistics (1)
- Wright Center for Clinical and Translational Research Works (1)
- Publication Type
Articles 1 - 30 of 46
Full-Text Articles in Data Science
Nlp Bias And African American English, Kenya Roy, Faizan Javed
Nlp Bias And African American English, Kenya Roy, Faizan Javed
SMU Data Science Review
African American English (AAE), also referred to as African American Vernacular English (AAVE), is widely used on social media, but most sentiment analysis tools are trained only on Standard American English (SAE). This mismatch can cause models to misclassify dialectal expressions—especially by labeling neutral or positive AAE as negative or toxic. These errors matter, since Natural Language Processing (NLP) systems are now central to content moderation and brand monitoring. This research will evaluate the VADER, RoBERTa, GPT-OSS, and Gemma’s handling of AAE in comparison to SAE using the TwitterAAE corpus, a public dataset of tweets with estimated AAVE usage. The …
Automating Cardiff Model Data Capture In Emergency Departments: Ambient Nlp Integration With Oracle-Cerner Fhir Systems, Simi Augustine, Marco A. Lopez, Jacquelyn Cheun, Chris Papesh
Automating Cardiff Model Data Capture In Emergency Departments: Ambient Nlp Integration With Oracle-Cerner Fhir Systems, Simi Augustine, Marco A. Lopez, Jacquelyn Cheun, Chris Papesh
SMU Data Science Review
Violence and overdose events in Las Vegas occur at rates above the national average, with fewer than half of violent injuries reported to law enforcement [2,7]. The Cardiff Model offers a proven framework for standardized data collection and sharing between hospitals and public safety partners, yet many implementations still rely on manual entry. We propose an ambient triage pipeline integrated with Oracle-Cerner electronic health record systems to listen to nurse–patient dialogue, convert speech to text, extract Cardiff fields, and write standards-based FHIR Bundles for analytics. Using SMART on FHIR standards and Cerner Millennium APIs, the study evaluates whether ambient capture …
Effective Wordle Heuristics, Ronald I. Greenberg
Effective Wordle Heuristics, Ronald I. Greenberg
Computer Science: Faculty Publications and Other Works
While previous researchers have performed an exhaustive search to determine an optimal Wordle strategy, that computation is very time consuming and produced a strategy using words that are unfamiliar to most people. With Wordle solutions being gradually eliminated (with a new puzzle each day and no reuse), an improved strategy could be generated each day, but the computation time makes a daily exhaustive search impractical. This paper shows that simple heuristics allow for fast generation of effective strategies and that little is lost by guessing only words that are possible solution words rather than more obscure words.
Previsit Ai: A Retrieval-Augmented Generation For Patient Readiness In Clinical Encounters, Rolande Umuhoza
Previsit Ai: A Retrieval-Augmented Generation For Patient Readiness In Clinical Encounters, Rolande Umuhoza
All Graduate Theses, Dissertations, and Other Capstone Projects
With healthcare systems under growing pressure from rising patient volumes and shrinking consultation windows, improving how patients communicate with physicians has become essential to delivering quality care. Yet patients routinely arrive at appointments unable to clearly describe their symptoms, recall their medical history, or articulate concerns, contributing to miscommunication, diagnostic inefficiency, and pre-visit anxiety. This study introduces PreVisit AI, a conversational system designed to address this gap through structured, knowledge-based patient preparation. The system is built on a Retrieval-Augmented Generation (RAG) architecture combining HuggingFace sentence embeddings (all-MiniLM-L6-v2), a Chroma vector store, and Google’s Gemini language model over a curated seven-document …
Advancing Hate Speech Detection: Binary And Multiclass Approaches From Traditional Methods To Parameter-Efficient And Ontology-Guided Language Models, Mahmoud Abusaqer
Advancing Hate Speech Detection: Binary And Multiclass Approaches From Traditional Methods To Parameter-Efficient And Ontology-Guided Language Models, Mahmoud Abusaqer
Graduate Theses/Dissertations
The widespread proliferation of hate speech on social media platforms poses significant challenges for content moderation and user safety, requiring automated systems that are simultaneously accurate, efficient, and capable of fine-grained distinctions. This thesis investigates hate speech detection through five published manuscripts organized into two complementary threads: binary detection (hateful vs. non-hateful) and multiclass detection across demographic targeting categories. The binary thread progresses from a broad 38-model baseline spanning traditional machine learning, deep learning, and transformer architectures (where RoBERTa reaches 91.48% accuracy and CatBoost remains competitive at 88.60%) to parameter-efficient adaptation, in which Low-Rank Adaptation (LoRA) of large language models …
Harnessing Graphs For Knowledge Representation In Natural Language Processing, Uras Varolgunes
Harnessing Graphs For Knowledge Representation In Natural Language Processing, Uras Varolgunes
Dissertations
This work proposes innovative methods for integrating domain-specific knowledge into natural language processing tasks through the use of graphs, aiming to enhance the performance of models across various domains, including finance and healthcare. Several novel approaches are proposed that fuse graph structures with modern deep learning techniques, addressing the challenges of missing word embeddings, label prediction, and graph representation learning for large language models.
First, a powerful embedding method built on top of the recent advances in latent graph learning is introduced to address the critical problem of word embedding imputation. Second, a graph-enhanced label attention model designed for medical …
Domain Obedient Deep Learning, Soumadeep Saha
Domain Obedient Deep Learning, Soumadeep Saha
Doctoral Theses
Deep learning, a family of data-driven artificial intelligence techniques, has shown immense promise in a plethora of applications, and it has even outpaced experts in several domains. However, unlike symbolic approaches to learning, these methods fall short when it comes to abiding by and learning from pre-existing established principles. This is a significant deficit for deployment in critical applications such as robotics, medicine, industrial automation, etc. For a decision system to be considered for adoption in such fields, it must demonstrate the ability to adhere to specified constraints, an ability missing in deep learning-based approaches. Exploring this problem serves as …
Leveraging Large Language Models For Knowledge-Free Weak Supervision In Clinical Natural Language Processing, Enshuo Hsu, Kirk Roberts
Leveraging Large Language Models For Knowledge-Free Weak Supervision In Clinical Natural Language Processing, Enshuo Hsu, Kirk Roberts
Faculty, Staff and Student Publications
The performance of deep learning-based natural language processing systems is based on large amounts of labeled training data which, in the clinical domain, are not easily available or affordable. Weak supervision and in-context learning offer partial solutions to this issue, particularly using large language models (LLMs), but their performance still trails traditional supervised methods with moderate amounts of gold-standard data. In particular, inferencing with LLMs is computationally heavy. We propose an approach leveraging fine-tuning LLMs and weak supervision with virtually no domain knowledge that still achieves consistently dominant performance. Using a prompt-based approach, the LLM is used to generate weakly-labeled …
Computational Bridges: Enhancing Natural Language Processing Of Swahili., Joyce Murungi
Computational Bridges: Enhancing Natural Language Processing Of Swahili., Joyce Murungi
Harrisburg University Other Works
Swahili remains significantly underrepresented in natural language processing (NLP) despite being one of the most widely spoken languages in Africa. Computational Bridges: Enhancing Natural Language Processing of Swahili addresses this gap through computational linguistics, corpus creation, and large-scale analysis of Swahili syntax and lexical structure. Central to this study is GUMZO, a novel corpus developed from spontaneous conversational data collected from YouTube videos, television panel discussions, political speeches, religious discourse, and unscripted broadcasts. Unlike many existing datasets that rely on formal or translated text, GUMZO captures authentic language use and provides a stronger foundation for NLP research involving low-resource languages. …
A Methodological Framework For Ontology Development, Enrichment, And Application In Natural Language Processing Tasks, Navya Martin Kollapally
A Methodological Framework For Ontology Development, Enrichment, And Application In Natural Language Processing Tasks, Navya Martin Kollapally
Dissertations
Electronic Health Records (EHRs) have been widely used in healthcare to record demographics, vital signs, test results, immunizations, medical imaging reports, differential diagnoses, etc. It is now accepted that non-clinical (e.g., social) factors have a substantial influence on health outcomes. Hence, it is desirable to record these Social and Commercial Determinants of Health (SDoH & CDoH) in an EHR. The "non-text parts" of EHR notes (e.g., data tables) rely on coded terms from underlying ontologies or terminologies to facilitate semantic interoperability. Ontologies help define concepts, the relationships between them, and instances that can be utilized in research.
The first accomplishment …
Natural Language Processing Of Clinical Notes Enables Early Inborn Error Of Immunity Risk Ascertainment, Kirk Roberts, Aaron T Chin, Klaus Loewy, Lisa Pompeii, Harold Shin, Nicholas L Rider
Natural Language Processing Of Clinical Notes Enables Early Inborn Error Of Immunity Risk Ascertainment, Kirk Roberts, Aaron T Chin, Klaus Loewy, Lisa Pompeii, Harold Shin, Nicholas L Rider
Faculty, Staff and Student Publications
BACKGROUND: There are now approximately 450 discrete inborn errors of immunity (IEI) described; however, diagnostic rates remain suboptimal. Use of structured health record data has proven useful for patient detection but may be augmented by natural language processing (NLP). Here we present a machine learning model that can distinguish patients from controls significantly in advance of ultimate diagnosis date.
OBJECTIVE: We sought to create an NLP machine learning algorithm that could identify IEI patients early during the disease course and shorten the diagnostic odyssey.
METHODS: Our approach involved extracting a large corpus of IEI patient clinical-note text from a major …
Spec: A Soft Prompt-Based Calibration On Performance Variability Of Large Language Model In Clinical Notes Summarization, Yu-Neng Chuang, Ruixiang Tang, Xiaoqian Jiang, Xia Hu
Spec: A Soft Prompt-Based Calibration On Performance Variability Of Large Language Model In Clinical Notes Summarization, Yu-Neng Chuang, Ruixiang Tang, Xiaoqian Jiang, Xia Hu
Faculty, Staff and Student Publications
Electronic health records (EHRs) store an extensive array of patient information, encompassing medical histories, diagnoses, treatments, and test outcomes. These records are crucial for enabling healthcare providers to make well-informed decisions regarding patient care. Summarizing clinical notes further assists healthcare professionals in pinpointing potential health risks and making better-informed decisions. This process contributes to reducing errors and enhancing patient outcomes by ensuring providers have access to the most pertinent and current patient data. Recent research has shown that incorporating instruction prompts with large language models (LLMs) substantially boosts the efficacy of summarization tasks. However, we show that this approach also …
Federated Learning For Sentiment Analysis In Presence Of Non-Iid Data: Sensitivity Of Deep Learning Models, Davoud Gholamiangonabadi, Katarina Grolinger
Federated Learning For Sentiment Analysis In Presence Of Non-Iid Data: Sensitivity Of Deep Learning Models, Davoud Gholamiangonabadi, Katarina Grolinger
Electrical and Computer Engineering Publications
In sentiment analysis, data are commonly distributed across many devices, and traditional machine learning requires transferring these data to a central location exposing data to security and privacy risks. Federated Learning (FL) avoids this transfer by training a model without requiring the clients/devices to share their local data; however, FL performance drops when data are not Independent and Identically Distributed (non-IID), such as when label distribution or data size vary across clients. Although techniques for non-IID data have been proposed primarily in the image domain, the sensitivity of various deep learning models to non-IID data needs to be examined. Consequently, …
The Enact Network Is Acting On Housing Instability And The Unhoused Using The Open Health Natural Language Processing Toolkit, Daniel R Harris, Sunyang Fu, Andrew Wen, Alexandria Corbeau, Darren Henderson, Jordan Hilsman, David Oniani, Yanshan Wang
The Enact Network Is Acting On Housing Instability And The Unhoused Using The Open Health Natural Language Processing Toolkit, Daniel R Harris, Sunyang Fu, Andrew Wen, Alexandria Corbeau, Darren Henderson, Jordan Hilsman, David Oniani, Yanshan Wang
Faculty, Staff and Student Publications
Housing is an environmental social determinant of health that is linked to mortality and clinical outcomes. We developed a lexicon of housing-related concepts and rule-based natural language processing methods for identifying these housing-related concepts within clinical text. We piloted our methods on several test cohorts: a synthetic cohort generated by ChatGPT for initial infrastructure testing, a cohort with substance use disorders (SUD), and a cohort diagnosed with problems related to housing and economic circumstances (HEC). Our methods successfully identified housing concepts in our ChatGPT notes (recall = 1.0, precision = 1.0), our SUD population (recall = 0.9798, precision = 0.9898), …
Review Classification Using Natural Language Processing And Deep Learning, Brian Nazareth
Review Classification Using Natural Language Processing And Deep Learning, Brian Nazareth
Electronic Theses, Projects, and Dissertations
Sentiment Analysis is an ongoing research in the field of Natural Language Processing (NLP). In this project, I will evaluate my testing against an Amazon Reviews Dataset, which contains more than 100 thousand reviews from customers. This project classifies the reviews using three methods – using a sentiment score by comparing the words of the reviews based on every positive and negative word that appears in the text with the Opinion Lexicon dataset, by considering the text’s variating sentiment polarity scores with a Python library called TextBlob, and with the help of neural network training. I have created a neural …
Dei: Exploring Academic Reflections Using Natural Language Processing To Create A Roadmap Of Student Success And Foster Inclusive Engineering Education, Rajvir H. Vyas, Nidhi Raviprasad
Dei: Exploring Academic Reflections Using Natural Language Processing To Create A Roadmap Of Student Success And Foster Inclusive Engineering Education, Rajvir H. Vyas, Nidhi Raviprasad
College of Engineering Summer Undergraduate Research Program
Every year, the College of Engineering (CENG) students and faculty reach out to admitted students through “Text-a-Thon” programs to answer their questions about being a student at Cal Poly. In order to improve CENG outreach efforts, we analyzed these text conversations to predict the likelihood of an admitted student accepting an offer of admission from Cal Poly. Through our research, we discovered key factors that play a role in a student committing to Cal Poly through data-based insights. Additionally, we successfully used a human-on-the-loop system to help create Machine Learning (ML) models that predict satisfaction of response by way of …
Weakly Supervised Spatial Relation Extraction From Radiology Reports, Surabhi Datta, Kirk Roberts
Weakly Supervised Spatial Relation Extraction From Radiology Reports, Surabhi Datta, Kirk Roberts
Faculty, Staff and Student Publications
OBJECTIVE: Weak supervision holds significant promise to improve clinical natural language processing by leveraging domain resources and expertise instead of large manually annotated datasets alone. Here, our objective is to evaluate a weak supervision approach to extract spatial information from radiology reports.
MATERIALS AND METHODS: Our weak supervision approach is based on data programming that uses rules (or labeling functions) relying on domain-specific dictionaries and radiology language characteristics to generate weak labels. The labels correspond to different spatial relations that are critical to understanding radiology reports. These weak labels are then used to fine-tune a pretrained Bidirectional Encoder Representations from …
Acquisition Of A Lexicon For Family History Information: Bidirectional Encoder Representations From Transformers-Assisted Sublanguage Analysis, Liwei Wang, Huan He, Andrew Wen, Sungrim Moon, Sunyang Fu, Kevin J Peterson, Xuguang Ai, Sijia Liu, Ramakanth Kavuluru, Hongfang Liu
Acquisition Of A Lexicon For Family History Information: Bidirectional Encoder Representations From Transformers-Assisted Sublanguage Analysis, Liwei Wang, Huan He, Andrew Wen, Sungrim Moon, Sunyang Fu, Kevin J Peterson, Xuguang Ai, Sijia Liu, Ramakanth Kavuluru, Hongfang Liu
Faculty, Staff and Student Publications
BACKGROUND: A patient's family history (FH) information significantly influences downstream clinical care. Despite this importance, there is no standardized method to capture FH information in electronic health records and a substantial portion of FH information is frequently embedded in clinical notes. This renders FH information difficult to use in downstream data analytics or clinical decision support applications. To address this issue, a natural language processing system capable of extracting and normalizing FH information can be used.
OBJECTIVE: In this study, we aimed to construct an FH lexical resource for information extraction and normalization.
METHODS: We exploited a transformer-based method to …
An Analysis Of Text-Based Machine Learning Models For Vulnerability Detection, Kollin Ryne Napier
An Analysis Of Text-Based Machine Learning Models For Vulnerability Detection, Kollin Ryne Napier
Theses and Dissertations
With an increase in complexity of software, developers rely more on reuse and dependencies in their source code via code snippets. As a result, it is becoming harder to identify and mitigate vulnerabilities. Although traditional analysis tools are still utilized, machine learning models are being adopted to expand efforts and combat such threats. Given the possibilities towards usage of such models, research in this area has introduced various approaches which vary in usability and prediction. In generalizing models to a more natural language approach, researchers have opted to train models on source code to identify existing and potential vulnerabilities. Exploratory …
Behind Derogatory Migrants' Terms For Venezuelan Migrants: Xenophobia And Sexism Identification With Twitter Data And Nlp, Joseph Martínez, Melissa Miller-Felton, Jose Padilla, Erika Frydenlund
Behind Derogatory Migrants' Terms For Venezuelan Migrants: Xenophobia And Sexism Identification With Twitter Data And Nlp, Joseph Martínez, Melissa Miller-Felton, Jose Padilla, Erika Frydenlund
Modeling, Simulation and Visualization Student Capstone Conference
The sudden arrival of many migrants can present new challenges for host communities and create negative attitudes that reflect that tension. In the case of Colombia, with the influx of over 2.5 million Venezuelan migrants, such tensions arose. Our research objective is to investigate how those sentiments arise in social media. We focused on monitoring derogatory terms for Venezuelans, specifically veneco and veneca. Using a dataset of 5.7 million tweets from Colombian users between 2015 and 2021, we determined the proportion of tweets containing those terms. We observed a high prevalence of xenophobic and defamatory language correlated with the …
Using Nlp To Model U.S. Supreme Court Cases, Katherine Lockard, Robert Slater, Brandon Sucrese
Using Nlp To Model U.S. Supreme Court Cases, Katherine Lockard, Robert Slater, Brandon Sucrese
SMU Data Science Review
The advantages of employing text analysis to uncover policy positions, generate legal predictions, and inform or evaluate reform practices are multifold. Given the far-reaching effects of legislation at all levels of society these insights and their continued improvement are impactful. This research explores the use of natural language processing (NLP) and machine learning to predictively model U.S. Supreme Court case outcomes based on textual case facts. The final model achieved an F1-score of .324 and an AUC of .68. This suggests that the model can distinguish between the two target classes; however, further research is needed before machine learning models …
Content-Based Unsupervised Fake News Detection On Ukraine-Russia War, Yucheol Shin, Yvan Sojdehei, Limin Zheng, Brad Blanchard
Content-Based Unsupervised Fake News Detection On Ukraine-Russia War, Yucheol Shin, Yvan Sojdehei, Limin Zheng, Brad Blanchard
SMU Data Science Review
The Ukrainian-Russian war has garnered significant attention worldwide, with fake news obstructing the formation of public opinion and disseminating false information. This scholarly paper explores the use of unsupervised learning methods and the Bidirectional Encoder Representations from Transformers (BERT) to detect fake news in news articles from various sources. BERT topic modeling is applied to cluster news articles by their respective topics, followed by summarization to measure the similarity scores. The hypothesis posits that topics with larger variances are more likely to contain fake news. The proposed method was evaluated using a dataset of approximately 1000 labeled news articles related …
Professor Text: University Fundraising Optimization, Braden Anderson, Connor Dobbs, Hien Lam, John Santerre
Professor Text: University Fundraising Optimization, Braden Anderson, Connor Dobbs, Hien Lam, John Santerre
SMU Data Science Review
University fundraising campaigns are a unique type of cause-related marketing with its own challenges and opportunities. Campaigns like this typically last an extended period, such as five or more years, and goals exist beyond the dollar amount raised. These supplemental goals, such as awareness among potential future donators or brand reputation within the local community, are important to consider and strategize. There can also be unique limitations, such as requiring advertising specifically on recent large gifts or endowment programs. This research explores how machine learning techniques such as natural language processing can be used to optimize a fundraising campaign strategy, …
Beyond News Values On Twitter: Predicting Factors That Drive User Engagement In News, Zhiyan Zhong
Beyond News Values On Twitter: Predicting Factors That Drive User Engagement In News, Zhiyan Zhong
Dartmouth College Master’s Theses
When deciding on what news stories to cover, traditional journalism determines news values by following several elements of newsworthiness, such as impact, timeliness, and prominence. However, these guidelines do not always seem to correspond with the success of content on social media. As people are increasingly turning to social media for news, our research aims to understand and predict factors that drive user engagement for news on social media. In this study, we analyze news content published on Twitter, and examine a diverse set of characteristics like metrics retrieved from the Twitter API and semantics by natural language processing, including …
A Systematic Approach To Configuring Metamap For Optimal Performance, Xia Jing, Akash Indani, Nina Hubig, Hua Min, Yang Gong, James J Cimino, Dean F Sittig, Lior Rennert, David Robinson, Paul Biondich, Adam Wright, Christian Nøhr, Timothy Law, Arild Faxvaag, Ronald Gimbel
A Systematic Approach To Configuring Metamap For Optimal Performance, Xia Jing, Akash Indani, Nina Hubig, Hua Min, Yang Gong, James J Cimino, Dean F Sittig, Lior Rennert, David Robinson, Paul Biondich, Adam Wright, Christian Nøhr, Timothy Law, Arild Faxvaag, Ronald Gimbel
Faculty, Staff and Student Publications
BACKGROUND: MetaMap is a valuable tool for processing biomedical texts to identify concepts. Although MetaMap is highly configurative, configuration decisions are not straightforward.
OBJECTIVE: To develop a systematic, data-driven methodology for configuring MetaMap for optimal performance.
METHODS: MetaMap, the word2vec model, and the phrase model were used to build a pipeline. For unsupervised training, the phrase and word2vec models used abstracts related to clinical decision support as input. During testing, MetaMap was configured with the default option, one behavior option, and two behavior options. For each configuration, cosine and soft cosine similarity scores between identified entities and gold-standard terms were …
Analyzing Fluctuation Of Topics And Public Sentiment Through Social Media Data, Haoyue Liu
Analyzing Fluctuation Of Topics And Public Sentiment Through Social Media Data, Haoyue Liu
Dissertations
Over the past decade years, Internet users were expending rapidly in the world. They form various online social networks through such Internet platforms as Twitter, Facebook and Instagram. These platforms provide a fast way that helps their users receive and disseminate information and express personal opinions in virtual space. When dealing with massive and chaotic social media data, how to accurately determine what events or concepts users are discussing is an interesting and important problem.
This dissertation work mainly consists of two parts. First, this research pays attention to mining the hidden topics and user interest trend by analyzing real-world …
Leveraging Context Patterns For Medical Entity Classification, Garrett Johnston
Leveraging Context Patterns For Medical Entity Classification, Garrett Johnston
Computer Science Senior Theses
The ability of patients to understand health-related text is important for optimal health outcomes. A system that can automatically annotate medical entities could help patients better understand health-related text. Such a system would also accelerate manual data annotation for this low-resource domain as well as assist in down- stream medical NLP tasks such as finding textual similarity, identifying conflicting medical advice, and aspect-based sentiment analysis. In this work, we investigate a state-of-the-art entity set expansion model, BootstrapNet, for the task of medical entity classification on a new dataset of medical advice text. We also propose EP SBERT, a simple model …
Innovative Heuristics To Improve The Latent Dirichlet Allocation Methodology For Textual Analysis And A New Modernized Topic Modeling Approach, Jamie T. Zimmerman
Innovative Heuristics To Improve The Latent Dirichlet Allocation Methodology For Textual Analysis And A New Modernized Topic Modeling Approach, Jamie T. Zimmerman
Theses and Dissertations
Natural Language Processing is a complex method of data mining the vast trove of documents created and made available every day. Topic modeling seeks to identify the topics within textual corpora with limited human input into the process to speed analysis. Current topic modeling techniques used in Natural Language Processing have limitations in the pre-processing steps. This dissertation studies topic modeling techniques, those limitations in the pre-processing, and introduces new algorithms to gain improvements from existing topic modeling techniques while being competitive with computational complexity. This research introduces four contributions to the field of Natural Language Processing and topic modeling. …
Developing A Natural Language Processing Approach For Analyzing Student Ideas In Calculus-Based Introductory Physics, Jon M. Geiger, Lisa M. Goodhew, Tor Ole B. Odden
Developing A Natural Language Processing Approach For Analyzing Student Ideas In Calculus-Based Introductory Physics, Jon M. Geiger, Lisa M. Goodhew, Tor Ole B. Odden
Honors Projects
Research characterizing common student ideas about particular physics topics has made a significant impact on university-level physics teaching by providing knowledge that supports instructors to target their instruction and by informing curriculum development. This work utilizes a Natural Language Processing algorithm (Latent Dirichlet Allocation, or LDA) to categorize student ideas, with the goal of significantly expediting the process of categorizing student ideas. We preliminarily test the LDA approach by applying the algorithm to a collection of introductory physics student responses to a conceptual question about circuits, specifically attending to whether it is useful for characterizing conceptual resources, or student ideas …
Toward Suicidal Ideation Detection With Lexical Network Features And Machine Learning, Ulya Bayram, William Lee, Daniel Santel, Ali Minai, Peggy Clark, Tracy Glauser, John Pestian
Toward Suicidal Ideation Detection With Lexical Network Features And Machine Learning, Ulya Bayram, William Lee, Daniel Santel, Ali Minai, Peggy Clark, Tracy Glauser, John Pestian
Northeast Journal of Complex Systems (NEJCS)
In this study, we introduce a new network feature for detecting suicidal ideation from clinical texts and conduct various additional experiments to enrich the state of knowledge. We evaluate statistical features with and without stopwords, use lexical networks for feature extraction and classification, and compare the results with standard machine learning methods using a logistic classifier, a neural network, and a deep learning method. We utilize three text collections. The first two contain transcriptions of interviews conducted by experts with suicidal (n=161 patients that experienced severe ideation) and control subjects (n=153). The third collection consists of interviews conducted by experts …