Open Access. Powered by Scholars. Published by Universities.®

Scholarly Publishing Commons

Open Access. Powered by Scholars. Published by Universities.®

Articles 1 - 16 of 16

Full-Text Articles in Scholarly Publishing

Exploring Marshall–Olkin Models Through Bibliometric And Topic Modeling Approaches Uses Latent Dirichlet Allocation (1981-2025): A Study Based On Scopus Data, Humberto Llinás, Brian Llinás, Carlos López, Daniela Nuñez Jan 2026

Exploring Marshall–Olkin Models Through Bibliometric And Topic Modeling Approaches Uses Latent Dirichlet Allocation (1981-2025): A Study Based On Scopus Data, Humberto Llinás, Brian Llinás, Carlos López, Daniela Nuñez

Computer Science Faculty Publications

The Marshall–Olkin family of distributions has gained increasing attention in fields such as reliability engineering, survival analysis, financial risk modeling, and actuarial science because of its flexibility in modeling dependence among events and its wide range of extensions. Despite its growing relevance, a systematic understanding of how research on Marshall–Olkin models has evolved over time is still limited. This study addresses this gap by combining bibliometric techniques with topic modeling to analyze the structure and evolution of the scientific literature on Marshall–Olkin models. The analysis includes all 266 peer-reviewed publications on Marshall–Olkin models indexed in Scopus between 1981 and 2025. …


Open Scholarly Information Systems: Status Quo, Challenges, Opportunities, Hannah Bast, Guillaume Cabanac, Paolo Manghi, Jian Wu, Marcel R. Ackermann Jan 2026

Open Scholarly Information Systems: Status Quo, Challenges, Opportunities, Hannah Bast, Guillaume Cabanac, Paolo Manghi, Jian Wu, Marcel R. Ackermann

Computer Science Faculty Publications

Over the past 30 years, a rich ecosystem of scholarly information systems has developed that openly provide their services to the scientific community. These systems include aggregators of bibliographic metadata (e.g., DBLP, OpenCitations, OpenAIRE Graph, OpenAlex, ORKG, Semantic Scholar, CiteSeerX, and CORE); publication, data, and software repositories (e.g., Arxiv.org, Figshare, Zenodo, Software Heritage, and Dataverse); and PID authorities (e.g., ORCID, ROR, Crossref, and DataCite). This interdisciplinary Dagstuhl Seminar "Open Scholarly Information Systems: Status Quo, Challenges, Opportunities" (25381) was the first of its kind to bring together practitioners from this ecosystem, as well as researchers investigating related questions or relying on …


An Analytical Review Of Preprocessing Techniques In Bengali Natural Language Processing, Sovon Chakraborty, Protiva Das, Shakib Mahmud Dipto, Md Aktaruzzaman Pramanik, Jannatun Noor Jan 2025

An Analytical Review Of Preprocessing Techniques In Bengali Natural Language Processing, Sovon Chakraborty, Protiva Das, Shakib Mahmud Dipto, Md Aktaruzzaman Pramanik, Jannatun Noor

Computer Science Faculty Publications

Research in Bengali Natural Language Processing (BNLP) is rapidly expanding. Despite being one of the most widely spoken languages in the world, BNLP research remains insufficient, particularly in Bengali speech recognition. The languages rich morphology, agglutinative structure, and diverse dialects make text and speech processing especially challenging. However, these challenges can be addressed with effective preprocessing techniques. Various organizations in Bangladesh and West Bengal are integrating Natural Language Processing (NLP) into their services, but without a thorough understanding of preprocessing, these implementations remain incomplete. Applying proper preprocessing techniques to the Bengali language will serve as a foundation for developing robust …


A Bibliographic And Topic Modeling Analysis Of The P-Adic Theory Literature Using Latent Dirichlet Allocation, Humberto Llinás, Ismael Gutiérrez, Anselmo Torresblanca, Javier De La Hoz, Brian Llinás Jan 2025

A Bibliographic And Topic Modeling Analysis Of The P-Adic Theory Literature Using Latent Dirichlet Allocation, Humberto Llinás, Ismael Gutiérrez, Anselmo Torresblanca, Javier De La Hoz, Brian Llinás

Computer Science Faculty Publications

P-adic analysis, introduced by Kurt Hensel in the early 20th century, has developed into a fundamental area of mathematical research with broad applications in number theory, algebraic geometry, and mathematical physics. This study aims to examine the thematic evolution and scholarly impact of p-adic research through a comprehensive topic modeling and bibliometric analysis. Using classical bibliometric techniques (e.g., performance analysis, co-authorship, and co-citation networks) combined with Latent Dirichlet Allocation (LDA), we analyzed 7388 peer-reviewed documents published between 1965 and 2024. The computational workflow was conducted using R (version 4.4.1) and VOSviewer (version 1.6.20), which enabled the identification of 20 distinct …


From Philosophy To Nlu: Evolving Definitions With Research Hypotheses, Jian Wu, Sarah Rajtmajer Jan 2025

From Philosophy To Nlu: Evolving Definitions With Research Hypotheses, Jian Wu, Sarah Rajtmajer

Computer Science Faculty Publications

Over the past decades, alongside advancements in natural language processing, significant attention has been paid to training models to automatically extract, understand, test, and generate hypotheses in open and scientific domains. However, interpretations of the term hypothesis for various natural language understanding (NLU) tasks have migrated from traditional definitions in the natural, social, and formal sciences. Even within NLU, we observe differences defining hypotheses across literature. In this paper, we overview and delineate various definitions of hypothesis. Especially, we discern the nuances of definitions across recently published NLU tasks. We highlight the importance of well-structured and well-defined hypotheses, particularly as …


Can Llms Beat Humans On Discerning Human-Written And Llm-Generated Science News, Dominik Soós, Meng Jiang, Jian Wu Jan 2025

Can Llms Beat Humans On Discerning Human-Written And Llm-Generated Science News, Dominik Soós, Meng Jiang, Jian Wu

Computer Science Faculty Publications

Science news is increasingly important in connecting scientists and the public by sharing discoveries and innovations. With the rise of large language models (LLMs), there is potential to automate science news creation, but concerns exist about the quality of LLM-generated news versus human-written news. This paper explores whether LLMs can outperform humans in distinguishing between human-written and LLM-generated news. Inspired by the Chain-of-Thought prompting method, we designed a simple yet effective variant called Guided Few-shot (GFS), which encodes the characteristics of news of two types with examples. Our experiments indicated that GFS with just a single example effectively boosted the …


From Philosophy To Nlu: Evolving Definitions Of Research Hypotheses, Jian Wu, Sarah Rajtmajer Jan 2025

From Philosophy To Nlu: Evolving Definitions Of Research Hypotheses, Jian Wu, Sarah Rajtmajer

Computer Science Faculty Publications

Over the past decades, alongside advancements in natural language processing, significant attention has been paid to training models to automatically extract, understand, test, and generate hypotheses in open and scientific domains. However, interpretations of the term hypothesis for various natural language understanding (NLU) tasks have migrated from traditional definitions in the natural, social, and formal sciences. Even within NLU, we observe differences defining hypotheses across literature. In this paper, we overview and delineate various definitions of hypothesis. Especially, we discern the nuances of definitions across recently published NLU tasks. We highlight the importance of well-structured and well-defined hypotheses, particularly as …


Can Large Language Models Discern Evidence For Scientific Hypotheses? Case Studies In The Social Sciences, Sai Koneru, Jian Wu, Sarah Rajtmajer Jan 2024

Can Large Language Models Discern Evidence For Scientific Hypotheses? Case Studies In The Social Sciences, Sai Koneru, Jian Wu, Sarah Rajtmajer

Computer Science Faculty Publications

Hypothesis formulation and testing are central to empirical research. A strong hypothesis is a best guess based on existing evidence and informed by a comprehensive view of relevant literature. However, with exponential increase in the number of scientific articles published annually, manual aggregation and synthesis of evidence related to a given hypothesis is a challenge. Our work explores the ability of current large language models (LLMs) to discern evidence in support or refute of specific hypotheses based on the text of scientific abstracts. We share a novel dataset for the task of scientific hypothesis evidencing using community-driven annotations of studies …


Short: Can Citations Tell Us About A Paper's Reproducibility? A Case Study Of Machine Learning Papers, Rochana R. Obadage, Sarah M. Rajtmajer, Jian Wu Jan 2024

Short: Can Citations Tell Us About A Paper's Reproducibility? A Case Study Of Machine Learning Papers, Rochana R. Obadage, Sarah M. Rajtmajer, Jian Wu

Computer Science Faculty Publications

The iterative character of work in machine learning (ML) and artificial intelligence (AI) and reliance on comparisons against benchmark datasets emphasize the importance of reproducibility in that literature. Yet, resource constraints and inadequate documentation can make running replications particularly challenging. Our work explores the potential of using downstream citation contexts as a signal of reproducibility. We introduce a sentiment analysis framework applied to citation contexts from papers involved in Machine Learning Reproducibility Challenges in order to interpret the positive or negative outcomes of reproduction attempts. Our contributions include training classifiers for reproducibility-related contexts and sentiment analysis, and exploring correlations between …


Retrogressive Document Manipulation Of Us Federal Environmental Websites, Lesley Frew, Michael L. Nelson, Michele C. Weigle Jan 2024

Retrogressive Document Manipulation Of Us Federal Environmental Websites, Lesley Frew, Michael L. Nelson, Michele C. Weigle

Computer Science Faculty Publications

Changes made to webpages can affect their retrievability. Often this is done with the intention of increasing the page's search engine ranking to improve overall access to information on the page. The Environmental Data and Governance Initiative (EDGI) created a dataset that describes changes on US federal environmental webpages between 2016 and 2020. EDGI noted that many environmental terms were deleted from the pages, but without user data, claims that page retrievability and public information access were lowered are only anecdotal. The Open Resource for Click Analysis in Search (ORCAS) dataset was created during the same time frame, from 2017 …


Building Datasets To Support Information Extraction And Structure Parsing From Electronic Theses And Dissertations, William A. Ingram, Jian Wu, Sampanna Yashwant Kahu, Javaid Akbar Manzoor, Bipasha Banerjee, Aman Ahuja, Muntabir Hasan Choudhury, Lamia Salsabil, Winston Shields, Edward A. Fox Jan 2024

Building Datasets To Support Information Extraction And Structure Parsing From Electronic Theses And Dissertations, William A. Ingram, Jian Wu, Sampanna Yashwant Kahu, Javaid Akbar Manzoor, Bipasha Banerjee, Aman Ahuja, Muntabir Hasan Choudhury, Lamia Salsabil, Winston Shields, Edward A. Fox

Computer Science Faculty Publications

Despite the millions of electronic theses and dissertations (ETDs) publicly available online, digital library services for ETDs have not evolved past simple search and browse at the metadata level. We need better digital library services that allow users to discover and explore the content buried in these long documents. Recent advances in machine learning have shown promising results for decomposing documents into their constituent parts, but these models and techniques require data for training and evaluation. In this article, we present high-quality datasets to train, evaluate, and compare machine learning methods in tasks that are specifically suited to identify and …


D-Lib Magazine Pioneered Web-Based Scholarly Communication, Michael L. Nelson, Herbert Van De Sompel Jan 2022

D-Lib Magazine Pioneered Web-Based Scholarly Communication, Michael L. Nelson, Herbert Van De Sompel

Computer Science Faculty Publications

The web began with a vision of, as stated by Tim Berners-Lee in 1991, “that much academic information should be freely available to anyone”. For many years, the development of the web and the development of digital libraries and other scholarly communications infrastructure proceeded in tandem. A milestone occurred in July, 1995, when the first issue of D-Lib Magazine was published as an online, HTML-only, open access magazine, serving as the focal point for the then emerging digital library research community. In 2017 it ceased publication, in part due to the maturity of the community it served as well as …


Extractive Research Slide Generation Using Windowed Labeling Ranking, Athar Sefid, Prasenjit Mitra, Jian Wu, C. Lee Giles Jan 2021

Extractive Research Slide Generation Using Windowed Labeling Ranking, Athar Sefid, Prasenjit Mitra, Jian Wu, C. Lee Giles

Computer Science Faculty Publications

Presentation slides generated from original research papers provide an efficient form to present research innovations. Manually generating presentation slides is labor-intensive. We propose a method to automatically generates slides for scientific articles based on a corpus of 5000 paper-slide pairs compiled from conference proceedings websites. The sentence labeling module of our method is based on SummaRuNNer, a neural sequence model for extractive summarization. Instead of ranking sentences based on semantic similarities in the whole document, our algorithm measures the importance and novelty of sentences by combining semantic and lexical features within a sentence window. Our method outperforms several baseline methods …


Acknowledgement Entity Recognition In Cord-19 Papers, Jian Wu, Pei Wang, Xin Wei, Sarah Rajtmajer, C. Lee Giles, Christopher Griffin Jan 2020

Acknowledgement Entity Recognition In Cord-19 Papers, Jian Wu, Pei Wang, Xin Wei, Sarah Rajtmajer, C. Lee Giles, Christopher Griffin

Computer Science Faculty Publications

Acknowledgements are ubiquitous in scholarly papers. Existing acknowledgement entity recognition methods assume all named entities are acknowledged. Here, we examine the nuances between acknowledged and named entities by analyzing sentence structure. We develop an acknowledgement extraction system, AckExtract based on open-source text mining software and evaluate our method using manually labeled data. AckExtract uses the PDF of a scholarly paper as input and outputs acknowledgement entities. Results show an overall performance of F1=0.92. We built a supplementary database by linking CORD-19 papers with acknowledgement entities extracted by AckExtract including persons and organizations and find that only up to …


Opening Books And The National Corpus Of Graduate Research, William A. Ingram, Edward A. Fox, Jian Wu Jan 2020

Opening Books And The National Corpus Of Graduate Research, William A. Ingram, Edward A. Fox, Jian Wu

Computer Science Faculty Publications

Virginia Tech University Libraries, in collaboration with Virginia Tech Department of Computer Science and Old Dominion University Department of Computer Science, request $505,214 in grant funding for a 3-year project, the goal of which is to bring computational access to book-length documents, demonstrating that with Electronic Theses and Dissertations (ETDs). The project is motivated by the following library and community needs. (1) Despite huge volumes of book-length documents in digital libraries, there is a lack of models offering effective and efficient computational access to these long documents. (2) Nationwide open access services for ETDs generally function at the metadata level. …


Automatic Slide Generation For Scientific Papers, Athar Sefid, Jian Wu, Prasenjit Mitra, C. Lee Giles Jan 2019

Automatic Slide Generation For Scientific Papers, Athar Sefid, Jian Wu, Prasenjit Mitra, C. Lee Giles

Computer Science Faculty Publications

We describe our approach for automatically generating presentation slides for scientific papers using deep neural networks. Such slides can help authors have a starting point for their slide generation process. Extractive summarization techniques are applied to rank and select important sentences from the original document. Previous work identified important sentences based only on a limited number of features that were extracted from the position and structure of sentences in the paper. Our method extends previous work by (1) extracting a more comprehensive list of surface features, (2) considering semantic or meaning of the sentence, and (3) using context around the …