Open Access. Powered by Scholars. Published by Universities.®
Library and Information Science Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- Old Dominion University (91)
- Walden University (37)
- Hunan Provincial Institute of Scientific and Technology Information (36)
- University of Nebraska - Lincoln (36)
- San Jose State University (33)
-
- Singapore Management University (31)
- University of Rhode Island (18)
- University of Malaya (17)
- Chapman University (16)
- Nova Southeastern University (16)
- University of South Carolina (16)
- Purdue University (15)
- University of Nevada, Las Vegas (15)
- Kennesaw State University (12)
- University of Kentucky (12)
- City University of New York (CUNY) (11)
- Virginia Commonwealth University (10)
- University of Denver (8)
- University of South Florida (8)
- Minnesota State University, Mankato (7)
- University at Albany, State University of New York (7)
- James Madison University (6)
- Brigham Young University (5)
- California State University, San Bernardino (5)
- East Texas A&M University (4)
- Rochester Institute of Technology (4)
- Seattle Pacific University (4)
- Syracuse University (4)
- University of Texas at El Paso (4)
- Wayne State University (4)
- Keyword
-
- Artificial intelligence (43)
- Digital libraries (25)
- Machine learning (24)
- Web archives (22)
- AI (21)
-
- Computer science (21)
- Library science (21)
- Digital preservation (20)
- Libraries (17)
- Web archiving (17)
- Data representation (15)
- Robert Hooke (15)
- Scientific imaging (15)
- Archives (14)
- Artificial Intelligence (13)
- Metadata (13)
- Information retrieval (12)
- Cybersecurity (11)
- Academic libraries (10)
- Text mining (10)
- Lib_pub (9)
- Technology (9)
- Natural language processing (8)
- Neural networks (8)
- Open access (8)
- Privacy (8)
- Search engines (8)
- Security (8)
- Social media (8)
- Collection management (7)
- Publication Year
- Publication
-
- Computer Science Faculty Publications (53)
- Walden Dissertations and Doctoral Studies (37)
- Journal of Scientific Information Research (36)
- Library Philosophy and Practice (e-journal) (24)
- Student Works (2000-2009) (17)
-
- Library Impact Statements (16)
- CCAC Theses and Dissertations (15)
- Copyright, Fair Use, Scholarly Communication, etc. (15)
- Computer Science Presentations (14)
- Computer Science Theses & Dissertations (14)
- Publications and Research (11)
- Information Science Faculty Publications (10)
- Research Collection Library (10)
- FORCE 2026 (9)
- Research Collection School Of Computing and Information Systems (8)
- Section 5: Imaging at the Nano Scale (8)
- USF Tampa Graduate Theses and Dissertations (8)
- Faculty Publications (7)
- University of Nebraska-Lincoln Libraries: Faculty Publications (7)
- VCU Libraries Faculty and Staff Publications (7)
- Handouts (6)
- Library Articles and Research (6)
- Library Presentations, Posters, and Audiovisual Materials (6)
- Library Services Publications (6)
- University of Nebraska-Lincoln Libraries: Presentations (6)
- Inaugural CSU IR Conference, 2015 (5)
- Libraries (5)
- Library Faculty Presentations (5)
- UNLV Theses, Dissertations, Professional Papers, and Capstones (5)
- Legacy Theses & Dissertations (2009 - 2024) (4)
- Publication Type
- File Type
Articles 61 - 90 of 585
Full-Text Articles in Library and Information Science
In Memoriam - Nora Sabelli: Master Orchestrator Of Grant Programs And Mentor For Advancing The Interdisciplinary Learning Sciences Field, Eric Hamilton, Jeremy Roschelle, Roy Pea, Barbara Means, Louis Gomez, Kim Gomez, Nancy Butler Songer
In Memoriam - Nora Sabelli: Master Orchestrator Of Grant Programs And Mentor For Advancing The Interdisciplinary Learning Sciences Field, Eric Hamilton, Jeremy Roschelle, Roy Pea, Barbara Means, Louis Gomez, Kim Gomez, Nancy Butler Songer
Education Division Scholarship
On Friday, September 6, 2024, the learning sciences field lost a giant in Dr. Nora Sabelli, 87 years old, a personal mentor to many researchers and an inspiration to so many learning scientists and STEM leaders. Nora’s first professional career was as a computational chemist, and later she became a passionate leader in research for improving STEM education. Nora’s time as a senior program officer at the National Science Foundation’s (NSF) Education and Human Resources (EHR) directorate was legendary; she was a force of nature who reshaped funding priorities for stronger science and a stronger connection of science to education …
The Year Of Ai: Raising Campus Awareness Through Art, Exhibits, And Community Engagement, Essraa Nawar
The Year Of Ai: Raising Campus Awareness Through Art, Exhibits, And Community Engagement, Essraa Nawar
Library Articles and Research
This poster highlights the Leatherby Libraries’ leadership in advancing AI literacy through creative, inclusive, and interdisciplinary approaches. As part of Chapman University’s “Year of AI,” the library launched initiatives such as Beyond the Lens and AI: The Next Chapter, blending art, ethics, and education to inspire campus-wide engagement. Through collaboration with IS&T, Town & Gown, and academic departments, the library positioned itself as a hub for ethical dialogue and innovation. The poster shares replicable models for how libraries can foster AI awareness through community partnerships, exhibitions, and experiential learning.
Making The Most Of Artificial Intelligence And Large Language Models To Support Collection Development In Health Sciences Libraries, Ivan Portillo, David Carson
Making The Most Of Artificial Intelligence And Large Language Models To Support Collection Development In Health Sciences Libraries, Ivan Portillo, David Carson
Library Articles and Research
This project investigated the potential of generative AI models in aiding health sciences librarians with collection development. Researchers at Chapman University’s Harry and Diane Rinker Health Science campus evaluated four generative AI models—ChatGPT 4.0, Google Gemini, Perplexity, and Microsoft Copilot—over six months starting in March 2024. Two prompts were used: one to generate recent eBook titles in specific health sciences fields and another to identify subject gaps in the existing collection. The first prompt revealed inconsistencies across models, with Copilot and Perplexity providing sources but also inaccuracies. The second prompt yielded more useful results, with all models offering helpful analysis …
Ai 101: What It Can (And Can't) Do For You, April Sheppard
Ai 101: What It Can (And Can't) Do For You, April Sheppard
Staff and Faculty Scholarship
In this presentation, April defines AI, describes how it works, reviews some pros and cons, and finally discusses what AI can actually accomplish in its current iteration.
Relational Database Schema To Support Research Profiling Studies, Natural Language Processing, And Bibliometric Analysis, Darnelle Melvin
Relational Database Schema To Support Research Profiling Studies, Natural Language Processing, And Bibliometric Analysis, Darnelle Melvin
Library Faculty Research
In this paper, a relational database schema is introduced that supports rapid prototyping, data preprocessing, and warehousing tasks associated with research profiling studies, natural language processing, and bibliometric analysis. Python scripts are leveraged for the seamless retrieval and processing of data from Semantic Scholar. This schema is tailored to efficiently analyze entities such as authors, their scientific papers, referenced papers, and cited papers. Adhering to the relational model, this schema offers a standardized approach to data storage and detailed information retrieval for scientific papers. Enhancing knowledge discovery in scientific databases, this schema provides researchers with a powerful platform for robust …
Analysis And Research On The Guiding Role Of Xi Jinping Thought On Socialism With Chinese Characteristics For A New Era In The Discipline Of Information Resources Management, Sanhong Deng, Yiqin Zhang, Hao Wang
Analysis And Research On The Guiding Role Of Xi Jinping Thought On Socialism With Chinese Characteristics For A New Era In The Discipline Of Information Resources Management, Sanhong Deng, Yiqin Zhang, Hao Wang
Journal of Scientific Information Research
[Purpose/significance]This paper explores the guiding role of Xi Jinping Thought on Socialism with Chinese Characteristics for a New Era in the development of the Information Resource Management discipline with Chinese characteristics, providing significant insights for the innovative advancement of China's Information Resource Management discipline and strengthening the discourse power of Chinese social sciences. [Method/process]This paper systematically reviews the core elements of the development philosophy of the Information Resource Management discipline within Xi Jinping Thought on Socialism with Chinese Characteristics for a New Era from a holistic perspective,elucidates the logical system of the development of the discipline from the diverse perspectives …
Interaction Mechanism Between Health Anxiety And Information Seeking Behavior From The Perspective Of Phenomenology, Yanfeng Zhang, Minqian Yu
Interaction Mechanism Between Health Anxiety And Information Seeking Behavior From The Perspective Of Phenomenology, Yanfeng Zhang, Minqian Yu
Journal of Scientific Information Research
[Purpose/significance]To analyze the evolution characteristics of health anxiety before and after information search behavior from the perspective of phenomenological graph analysis, and to explain the internal mechanism of the interaction between health anxiety and information search behavior. [Method/process]By using the phenomenological qualitative research method, the interactive mechanism between health anxiety and information search behavior was deeply explored. Based on the I-PACE theoretical model framework, the model elements of users' health anxiety and information search behavior were analyzed from the four dimensions of "Person-Affect-Cognition-Execution". To construct a mechanistic relationship model between health anxiety and information search behavior. [Result/conclusion]The research results revealed …
Research On Automated Generation And Evaluation Of Patent Claimsbased On Gpt-4, Junhua Li, Qian Yuan, Xiang Yan, Changhong Lv
Research On Automated Generation And Evaluation Of Patent Claimsbased On Gpt-4, Junhua Li, Qian Yuan, Xiang Yan, Changhong Lv
Journal of Scientific Information Research
[Purpose/significance]This study aims to automatically generate claims using the GPT-4 model, in order to reduce the writing difficulty for inventor and improve the work efficiency and quality. [Method/process]The article constructs Prompts suitable for automatically generating patent claims and implements four prompting strategies: ZeroShot, Exact-Drafting, Stepwise-Claim, and Exact-Step Claim. By inputting patent specifications and technical disclosure documents into the GPT-4 model and using Prompts to guide its output, the automated generation of patent claims is achieved. The ROUGE and BERTScore evaluation metrics were used to assess the quality of the text, and the generated text was analyzed in comparison with the …
Research On Emerging Technology Topic Identification Based On Bertopic, Dakun Wang, Bolin Hua
Research On Emerging Technology Topic Identification Based On Bertopic, Dakun Wang, Bolin Hua
Journal of Scientific Information Research
[Purpose/significance]Identifying and foreseeing emerging technologies, bring technological first-mover advantages to enterprises and governments, and grasp technological development trends in a timely manner. [Method/process]This study uses BERTopic's topic modeling method to obtain domain topic distribution, and merges paper and patent topics based on the cosine similarity of topic vectors to identify emerging topics. [Result/conclusion]Using the BERTopic topic modeling method combined with index evaluation can effectively identify emerging topics and emerging terms.Taking the field of new energy vehicles as an example to carry out empirical research, using two methods: divided verification period and data verification method, 12 of the 16 identified topics …
Extending A Pretrained Language Model (Bert) Using An Ontological Perspective To Classify Students' Scientific Expertise Level From Written Responses, Heqiao Wang, Kevin C. Haudek, Amanda D. Manzanares, Chelsie L. Romulo, Emily A. Royse, Caterina B. Azzarello
Extending A Pretrained Language Model (Bert) Using An Ontological Perspective To Classify Students' Scientific Expertise Level From Written Responses, Heqiao Wang, Kevin C. Haudek, Amanda D. Manzanares, Chelsie L. Romulo, Emily A. Royse, Caterina B. Azzarello
Human Movement Studies & Special Education Faculty Publications
The complex and interdisciplinary nature of scientific concepts presents formidable challenges for students in developing their knowledge-in-use skills. The utilization of computerized analysis for evaluating students' contextualized constructed responses offers a potential avenue for educators to develop personalized and scalable interventions, thus supporting the current teaching and learning of science. While prior research in artificial intelligence has demonstrated the effectiveness of algorithms, including Bidirectional Encoder Representations from Transformers (BERT), in tasks like automated classifications of constructed responses, these efforts have predominantly leaned towards text-level features, often overlooking the exploration of conceptual ideas embedded in students' responses from a cognitive perspective. …
Application Of Large Language Model Methods In Scientific And Technical Intelligence Practice, Bolin Hua, Yingze Wang
Application Of Large Language Model Methods In Scientific And Technical Intelligence Practice, Bolin Hua, Yingze Wang
Journal of Scientific Information Research
[Purpose/significance]With the strong ability to process large-scale datasets and outstanding performance in various natural language processing tasks, large language models (LLMs) have excelled across multiple industries.Since scientific and technical intelligence primarily relies on textual data, LLMs are naturally well-suited for this field, ushering in a new wave of transformative changes. [Method /process]This article discusses the advantages of LLMs from five perspectives: low-dimensional dense vector representations of text, large-scale pre-trained models,fine-tuning and prompt learning, high-quality large-scale training data, and human alignment techniques. [Result/conclusion]LLMs have extensive applications in tasks such as intelligence identification, intelligence tracking, intelligence evaluation, and intelligence prediction, resulting in …
How Do Selected Biomedical And Health Sciences Journals React To Submissions Of Artificial Intelligence (Ai) Assisted Manuscripts?, Misa Mi, Lin Wu, Yingting Zhang, Wendy Wu
How Do Selected Biomedical And Health Sciences Journals React To Submissions Of Artificial Intelligence (Ai) Assisted Manuscripts?, Misa Mi, Lin Wu, Yingting Zhang, Wendy Wu
Library Scholarly Publications
Background and Objectives: Generative artificial intelligence (GenAI) increasingly impacts research and scholarly communication. Given the evolving application of ChatGPT and other AI tools in scholarly communications, health sciences librarians must become cognizant of any existing journal publishing guidelines for AI-created or assisted manuscripts. The study aims to examine how scholarly biomedical and health sciences journals and publishers respond to submissions of these manuscripts and what requirements or policies have been put in place to guide and instruct authors on AI use.
Methods: We first retrieved and consolidated a list of journals representing disciplines in biomedical and health sciences from four …
Method Entity And Relation Extraction Based On Automatically Generated Syntactic Templates: A Case Study Of Csdn Artificial Intelligence Blog, Kuiliang Li, Bolin Huang
Method Entity And Relation Extraction Based On Automatically Generated Syntactic Templates: A Case Study Of Csdn Artificial Intelligence Blog, Kuiliang Li, Bolin Huang
Journal of Scientific Information Research
[Purpose/significance]There are many relationships between method entities and application scenarios, problems,organizations and other entities. Extracting these entity relationships helps to capture the development trend of technology and promote the improvement of innovation ability.[Method/process]This paper discusses a method for extracting method entities and relations based on automatically generated syntactic templates. By designing a new adaptive template, the method improves flexibility and adaptability, reducing dependence on large-scale labeled data. Using a small number of seed triples, the method iteratively generates syntactic templates and extracts method entities and relations for the CSDN artificial intelligence topic blog. It also improves the extraction quality using …
Generative Ai And Finding The Law, Paul D. Callister
Generative Ai And Finding The Law, Paul D. Callister
Faculty Works
Legal information science requires, among other things, principles and theories. The article states six principles or considerations that any discussion of generative AI large language models and their role in finding the law must include. The article concludes that law librarianship will increasingly become legal information science and require new paradigms. In addition to the six principles, the article applies ecological holistic media theory to understand the relationship of the legal community’s cognitive authority, institutions, techné (technology, medium and method), geopolitical factors, and the past and future to understand the changes in this information milieu. The article also explains generative …
Evaluating The Evaluators: The Role Of Benchmarks In Legal Ai, Jonathan A. Franklin, Sean Harrington, Christine Hye Won Park
Evaluating The Evaluators: The Role Of Benchmarks In Legal Ai, Jonathan A. Franklin, Sean Harrington, Christine Hye Won Park
Other Faculty Publications
No abstract provided.
A Bibliometric Analysis Of Ai-Driven Healthcare Literature Containing Kos Keywords: Trends, Themes, And Gaps, Julaine Clunis, Eric Asare
A Bibliometric Analysis Of Ai-Driven Healthcare Literature Containing Kos Keywords: Trends, Themes, And Gaps, Julaine Clunis, Eric Asare
STEMPS Faculty Publications
As artificial intelligence (AI) becomes increasingly embedded in healthcare applications, concerns have emerged around the trustworthiness, interpretability, and context-awareness of these systems. Knowledge Organization Systems (KOS) hold considerable potential to address these challenges by supporting semantic standardization, explainability, and domain alignment. This study presents a bibliometric analysis of scholarly publications referencing both AI and healthcare concepts to examine how KOS are positioned within this evolving discourse. The findings indicate that while early literature frequently and explicitly referenced KOS—such as ontologies, controlled vocabularies, and classification systems—their visibility has declined relative to newer paradigms such as machine learning and large language models. …
Comprehensive Benchmarking Of Several Machine Learning And Bayesian Models For Early-Stage Diabetes Risk Prediction: A Large-Scale Comparative Study, Md. Iqbal Hossain, Najila Alam Porno
Comprehensive Benchmarking Of Several Machine Learning And Bayesian Models For Early-Stage Diabetes Risk Prediction: A Large-Scale Comparative Study, Md. Iqbal Hossain, Najila Alam Porno
Mathematics & Statistics Faculty Publications
Diabetes remains a critical global health challenge, with early detection is crucial for effective management. This study presents a comprehensive benchmarking analysis of 14 diverse machine learning and Bayesian models for early-stage diabetes risk prediction using clinical data [2] from Sylhet, Bangladesh. This research evaluated traditional methods (Logistic Regression, Decision Trees), ensemble techniques (Random Forest, XGBoost, LightGBM), Bayesian approaches (BART, Bayesian Logistic Regression), and advanced neural architectures (Deep Belief Networks) using both 70-30 train-test splits and 10-fold cross-validation. The results demonstrate that ensemble methods consistently outperformed other approaches, with Random Forest(RF) achieving the highest cross-validated AUC (0.9951) and accuracy (0.9699). …
Leveraging Transformer-Based Ocr Model With Generative Data Augmentation For Engineering Document Recognition, Wael Khallouli, Mohammad Shahab Uddin, Andres Sousa-Poza, Jiang Li, Samuel Kovacic
Leveraging Transformer-Based Ocr Model With Generative Data Augmentation For Engineering Document Recognition, Wael Khallouli, Mohammad Shahab Uddin, Andres Sousa-Poza, Jiang Li, Samuel Kovacic
Engineering Management & Systems Engineering Faculty Publications
The long-standing practice of document-based engineering has resulted in the accumulation of a large number of engineering documents across various industries. Engineering documents, such as 2D drawings, continue to play a significant role in exchanging information and sharing knowledge across multiple engineering processes. However, these documents are often stored in non-digitized formats, such as paper and portable document format (PDF) files, making automation difficult. As digital engineering transforms processes in many industries, digitizing engineering documents presents a crucial challenge that requires advanced methods. This research addresses the problem of automatically extracting textual content from non-digitized legacy engineering documents. We introduced …
Copyright And Artificial Intelligence, Part 2: Copyrightability
Copyright And Artificial Intelligence, Part 2: Copyrightability
Copyright, Fair Use, Scholarly Communication, etc.
This report by the United States Copyright Office addresses the legal and policy issues related to artificial intelligence (AI) and copyright as outlined in the Office’s August 2023 Notice of Inquiry (NOI).
The report will be published in several parts each one addressing a different topic. This part addresses the copyrightability of works created using generative AI. The first part, published in 2024, addresses the topic of digital replicas—the use of digital technology to realistically replicate an individual’s voice or appearance. A subsequent part will turn to the training of AI models on copyrighted works, licensing considerations, and allocation of …
Not Here, Go There: Analyzing Redirection Patterns On The Web, Kritika Garg, Sawood Alam, Dietrich Ayala, Michele C. Weigle, Michael L. Nelson
Not Here, Go There: Analyzing Redirection Patterns On The Web, Kritika Garg, Sawood Alam, Dietrich Ayala, Michele C. Weigle, Michael L. Nelson
Computer Science Faculty Publications
URI redirections are integral to web management, supporting structural changes, SEO optimization, and security. However, their complexities affect usability, SEO performance, and digital preservation. This study analyzed 11 million unique redirecting URIs, following redirections up to 10 hops per URI, to uncover patterns and implications of redirection practices. Our findings revealed that 50% of the URIs terminated successfully, while 50% resulted in errors, including 0.06% exceeding 10 hops. Canonical redirects, such as HTTP to HTTPS transitions, were prevalent, reflecting adherence to SEO best practices. Non-canonical redirects, often involving domain or path changes, highlighted significant web migrations, rebranding, and security risks. …
Github Repository Complexity Leads To Diminished Web Archive Availability, David Calano, Michael Nelson, Michele Weigle
Github Repository Complexity Leads To Diminished Web Archive Availability, David Calano, Michael Nelson, Michele Weigle
Computer Science Faculty Publications
Software is often developed using versioned controlled software, such as Git, and hosted on centralized Web hosts, such as GitHub and GitLab. These Web hosted software repositories are made available to users in the form of traditional HTML Web pages for each source file and directory, as well as a presentational home page and various descriptive pages. We examined more than 12,000 Web hosted Git repository project home pages, primarily from GitHub, to measure how well their presentational components are preserved in the Internet Archive, as well as the source trees of the collected GitHub repositories to assess the extent …
Coming Back Differently: An Exploratory Case Study Of Near Death Experiences Of Webpages, Lesley Frew, Michael L. Nelson, Michele Weigle
Coming Back Differently: An Exploratory Case Study Of Near Death Experiences Of Webpages, Lesley Frew, Michael L. Nelson, Michele Weigle
Computer Science Faculty Publications
In this case study, we use web archives to analyze 8,824 webpages that were taken offline and subsequently put back online, thus experiencing a “near death experience.” We enumerate the stages of a webpage’s near death experience, including the change from a successful HTTP status code to non-successful and back, the intermediate stage with markers such as an under construction banner, and an analysis of how the pages came back differently.
From Philosophy To Nlu: Evolving Definitions With Research Hypotheses, Jian Wu, Sarah Rajtmajer
From Philosophy To Nlu: Evolving Definitions With Research Hypotheses, Jian Wu, Sarah Rajtmajer
Computer Science Faculty Publications
Over the past decades, alongside advancements in natural language processing, significant attention has been paid to training models to automatically extract, understand, test, and generate hypotheses in open and scientific domains. However, interpretations of the term hypothesis for various natural language understanding (NLU) tasks have migrated from traditional definitions in the natural, social, and formal sciences. Even within NLU, we observe differences defining hypotheses across literature. In this paper, we overview and delineate various definitions of hypothesis. Especially, we discern the nuances of definitions across recently published NLU tasks. We highlight the importance of well-structured and well-defined hypotheses, particularly as …
From Philosophy To Nlu: Evolving Definitions Of Research Hypotheses, Jian Wu, Sarah Rajtmajer
From Philosophy To Nlu: Evolving Definitions Of Research Hypotheses, Jian Wu, Sarah Rajtmajer
Computer Science Faculty Publications
Over the past decades, alongside advancements in natural language processing, significant attention has been paid to training models to automatically extract, understand, test, and generate hypotheses in open and scientific domains. However, interpretations of the term hypothesis for various natural language understanding (NLU) tasks have migrated from traditional definitions in the natural, social, and formal sciences. Even within NLU, we observe differences defining hypotheses across literature. In this paper, we overview and delineate various definitions of hypothesis. Especially, we discern the nuances of definitions across recently published NLU tasks. We highlight the importance of well-structured and well-defined hypotheses, particularly as …
Uso De Herramientas De Inteligencia Artificial Generativa Por Parte De Docentes En Una Escuela Del Suroeste De Puerto Rico, Glorimar N. Rodríguez Guiliani
Uso De Herramientas De Inteligencia Artificial Generativa Por Parte De Docentes En Una Escuela Del Suroeste De Puerto Rico, Glorimar N. Rodríguez Guiliani
Theses and Dissertations
Esta disertación aplicada fue diseñada para investigar el nivel de conocimiento, uso y dificultades que enfrentan los docentes de quinto a duodécimo grado en una escuela privada del suroeste de Puerto Rico respecto a tecnologías emergentes las cuales presentan desafíos significativos para los docentes y los estudiantes. Se exploró la utilización de herramientas de inteligencia artificial generativa (GenAI) como ChatGPT dentro y fuera del aula para actividades pedagógicas y administrativas.
Los hallazgos revelaron una notable carencia en el conocimiento docente sobre el uso y habilidades de la inteligencia artificial. Se identificó, también, una deficiencia en la capacidad de los docentes …
The Invisible Influencer In Information Infrastructure, Herbert Van De Sompel, Michael L. Nelson
The Invisible Influencer In Information Infrastructure, Herbert Van De Sompel, Michael L. Nelson
Computer Science Faculty Publications
The UPS Prototype was a proof-of-concept web portal built in preparation for the Universal Preprint Service Meeting held in October 1999 in Santa Fe, New Mexico. The portal provided search functionality for a set of metadata records that had been aggregated from a range of repositories that hosted preprints, working papers, and technical reports. Every search result was overlaid with a dynamically generated menu, called an SFX-menu, that provided a selection of value-adding links for the described scholarly work. The meeting eventually led to the Open Archives Initiative and its Protocol for Metadata Harvesting (OAI-PMH), which remains widely used in …
An Analytical Review Of Preprocessing Techniques In Bengali Natural Language Processing, Sovon Chakraborty, Protiva Das, Shakib Mahmud Dipto, Md Aktaruzzaman Pramanik, Jannatun Noor
An Analytical Review Of Preprocessing Techniques In Bengali Natural Language Processing, Sovon Chakraborty, Protiva Das, Shakib Mahmud Dipto, Md Aktaruzzaman Pramanik, Jannatun Noor
Computer Science Faculty Publications
Research in Bengali Natural Language Processing (BNLP) is rapidly expanding. Despite being one of the most widely spoken languages in the world, BNLP research remains insufficient, particularly in Bengali speech recognition. The languages rich morphology, agglutinative structure, and diverse dialects make text and speech processing especially challenging. However, these challenges can be addressed with effective preprocessing techniques. Various organizations in Bangladesh and West Bengal are integrating Natural Language Processing (NLP) into their services, but without a thorough understanding of preprocessing, these implementations remain incomplete. Applying proper preprocessing techniques to the Bengali language will serve as a foundation for developing robust …
A Bibliographic And Topic Modeling Analysis Of The P-Adic Theory Literature Using Latent Dirichlet Allocation, Humberto Llinás, Ismael Gutiérrez, Anselmo Torresblanca, Javier De La Hoz, Brian Llinás
A Bibliographic And Topic Modeling Analysis Of The P-Adic Theory Literature Using Latent Dirichlet Allocation, Humberto Llinás, Ismael Gutiérrez, Anselmo Torresblanca, Javier De La Hoz, Brian Llinás
Computer Science Faculty Publications
P-adic analysis, introduced by Kurt Hensel in the early 20th century, has developed into a fundamental area of mathematical research with broad applications in number theory, algebraic geometry, and mathematical physics. This study aims to examine the thematic evolution and scholarly impact of p-adic research through a comprehensive topic modeling and bibliometric analysis. Using classical bibliometric techniques (e.g., performance analysis, co-authorship, and co-citation networks) combined with Latent Dirichlet Allocation (LDA), we analyzed 7388 peer-reviewed documents published between 1965 and 2024. The computational workflow was conducted using R (version 4.4.1) and VOSviewer (version 1.6.20), which enabled the identification of 20 distinct …
Can Llms Beat Humans On Discerning Human-Written And Llm-Generated Science News, Dominik Soós, Meng Jiang, Jian Wu
Can Llms Beat Humans On Discerning Human-Written And Llm-Generated Science News, Dominik Soós, Meng Jiang, Jian Wu
Computer Science Faculty Publications
Science news is increasingly important in connecting scientists and the public by sharing discoveries and innovations. With the rise of large language models (LLMs), there is potential to automate science news creation, but concerns exist about the quality of LLM-generated news versus human-written news. This paper explores whether LLMs can outperform humans in distinguishing between human-written and LLM-generated news. Inspired by the Chain-of-Thought prompting method, we designed a simple yet effective variant called Guided Few-shot (GFS), which encodes the characteristics of news of two types with examples. Our experiments indicated that GFS with just a single example effectively boosted the …
Humans Vs. Llms On Open Domain Scientific Claim Verification: A Baseline Study, Benjamin Curtis, Stefania Dzhaman, Matthew Maisonave, Jian Wu
Humans Vs. Llms On Open Domain Scientific Claim Verification: A Baseline Study, Benjamin Curtis, Stefania Dzhaman, Matthew Maisonave, Jian Wu
Computer Science Faculty Publications
Verifying scientific claims is challenging for the general public because most people lack domain knowledge. Manual verification by subject domain experts is accurate, but it is obviously not scalable to meet the rising number of scientific claims on the Web. Whether the emerging large language models and large reasoning models can be used for scientific claim verification, and how their performances compare to humans, are still research questions. To this end, we developed a new benchmark MSVEC2 that consists of 138 claims from credible fact verification websites and science news outlets. Two tasks were given to both human and LLM …