Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Databases and Information Systems (90)
- Engineering (60)
- Numerical Analysis and Scientific Computing (47)
- Computer Engineering (46)
- Artificial Intelligence and Robotics (36)
-
- Social and Behavioral Sciences (32)
- Software Engineering (24)
- Business (22)
- Electrical and Computer Engineering (22)
- Mathematics (19)
- Data Science (17)
- Systems Architecture (17)
- Theory and Algorithms (17)
- Life Sciences (16)
- Statistics and Probability (16)
- Logic and Foundations (13)
- Communication (11)
- Information Security (11)
- Medicine and Health Sciences (11)
- Operations Research, Systems Engineering and Industrial Engineering (11)
- Other Computer Sciences (11)
- Bioinformatics (9)
- Management Information Systems (9)
- Social Media (9)
- Systems Science (8)
- Education (7)
- Library and Information Science (6)
- OS and Networks (6)
- Institution
-
- Singapore Management University (70)
- Portland State University (28)
- New Jersey Institute of Technology (16)
- TÜBİTAK (16)
- University at Albany, State University of New York (9)
-
- Zayed University (9)
- Air Force Institute of Technology (8)
- China Simulation Federation (8)
- Louisiana State University (8)
- Old Dominion University (8)
- Institute of Business Administration (7)
- Louisiana Tech University (7)
- University of Nebraska - Lincoln (7)
- Edith Cowan University (6)
- Technological University Dublin (6)
- University of Texas at Arlington (6)
- Brigham Young University (4)
- Kennesaw State University (4)
- MBZUAI (4)
- University of Louisville (4)
- University of Nevada, Las Vegas (4)
- Western Kentucky University (4)
- Central Washington University (3)
- Claremont Colleges (3)
- Clemson University (3)
- Embry-Riddle Aeronautical University (3)
- Nova Southeastern University (3)
- Purdue University (3)
- San Jose State University (3)
- University of Kentucky (3)
- Publication Year
- Publication
-
- Research Collection School Of Computing and Information Systems (64)
- Complex Systems Faculty Publications and Presentations (22)
- Dissertations (16)
- Turkish Journal of Electrical Engineering and Computer Sciences (16)
- Theses and Dissertations (11)
-
- All Works (9)
- Legacy Theses & Dissertations (2009 - 2024) (9)
- Doctoral Dissertations (8)
- Journal of System Simulation (8)
- International Conference on Information and Communication Technologies (7)
- Computer Science Faculty Publications (6)
- Electronic Theses and Dissertations (6)
- Computer Science Faculty Publications and Presentations (5)
- Faculty Publications (5)
- LSU Doctoral Dissertations (5)
- Computer Science and Engineering Theses - Archive (4)
- Dissertations and Theses Collection (Open Access) (4)
- Faculty Articles (4)
- Theses (4)
- All Faculty Scholarship for the College of the Sciences (3)
- CCAC Theses and Dissertations (3)
- CGU Faculty Publications and Research (3)
- Conference papers (3)
- LSU Master's Theses (3)
- Machine Learning Faculty Publications (3)
- Masters Theses & Specialist Projects (3)
- Theses and Dissertations--Computer Science (3)
- UNLV Theses, Dissertations, Professional Papers, and Capstones (3)
- All Dissertations (2)
- All Graduate Theses, Dissertations, and Other Capstone Projects (2)
- Publication Type
Articles 1 - 30 of 337
Full-Text Articles in Computer Sciences
A Data-Driven Framework For Mitigating Breast Cancer Overdiagnosis: From Estimation To Risk-Adjusted Computer-Aided Diagnosis, William M. Brown Jr.
A Data-Driven Framework For Mitigating Breast Cancer Overdiagnosis: From Estimation To Risk-Adjusted Computer-Aided Diagnosis, William M. Brown Jr.
LSU Doctoral Dissertations
In Computer-Aided Diagnosis (CAD) of cancer, standard cost metrics (false-positives and false-negatives) fundamentally fail to account for overdiagnosis. Overdiagnosis is a critical scenario where a disease is correctly detected (true-positive) but is biologically indolent and would never have caused the patient harm or symptoms. While widely recognized in the medical community as a major healthcare crisis driving stressful and invasive overtreatment, overdiagnosis remains severely under-researched within computer science and engineering. This dissertation addresses this interdisciplinary gap by defining the three key computational challenges of overdiagnosis: (i) accurate estimation, (ii) harm quantification, and (iii) algorithmic mitigation. To overcome the estimation challenge, …
Predict Social Economic Outcomes By Transferred Knowledge With Satellite Imagery, Yang Tang, Shih-Fen Cheng, Yunqiang Zhu, Yichen Yang, Zhiqiang Zou
Predict Social Economic Outcomes By Transferred Knowledge With Satellite Imagery, Yang Tang, Shih-Fen Cheng, Yunqiang Zhu, Yichen Yang, Zhiqiang Zou
Research Collection School Of Computing and Information Systems
Traditional deep learning methods and econometric models have played a crucial role in the field of data mining, particularly in the prediction of socioeconomic outcomes. However, socio-economic information is unable to be directly extracted from remote sensing data. So, in this paper, we propose a method to leverage transfer learning to predict socioeconomic indicators (outcomes) through satellite imagery. Specifically, we use road network types as a proxy for socioeconomic factors, which is more effective and stable than using nightlight. We have extracted eleven distinct road topological features to generate reasonable road network types. Given the unique characteristics of road networks, …
Freqllm: Frequency-Aware Large Language Models For Time Series Forecasting, Shunan Wang, Min Gao, Zongwei Wang, Yibing Bai, Feng Jiang, Guansong Pang
Freqllm: Frequency-Aware Large Language Models For Time Series Forecasting, Shunan Wang, Min Gao, Zongwei Wang, Yibing Bai, Feng Jiang, Guansong Pang
Research Collection School Of Computing and Information Systems
Large Language Models (LLMs) have recently shown promise in Time Series Forecasting (TSF) by effectively capturing intricate time-domain dependencies. However, our preliminary experiments reveal that standard LLM-based approaches often fail to capture global correlations, limiting predictive performance. We found that embedding frequency-domain signals smooths weight distributions and enhances structured correlations by clearly separating global trends (low-frequency components) from local variations (high-frequency components). Building on these insights, we propose FreqLLM, a novel framework that integrates frequency-domain semantic alignment into LLMs to refine prompts for improved time series analysis. By bridging the gap between frequency signals and textual embeddings, FreqLLM effectively captures …
Towards Explainable Ai On Graph Neural Networks: Xaig, Jiaxing Zhang
Towards Explainable Ai On Graph Neural Networks: Xaig, Jiaxing Zhang
Dissertations
In the evolving landscape of artificial intelligence (AI), Graph Neural Networks (GNNs) have garnered growing prominence for their adeptness in processing graph-structured data. Despite this, the interpretability of their predictions often remains elusive. The demand for transparency and explainability in complex prediction models has reached unprecedented levels. To address this, post-hoc instance-level explanation techniques have emerged, aiming to unveil the rationale behind GNN predictions. These techniques endeavor to unearth substructures that elucidate the predictive behavior of trained GNNs.
This dissertation embarks on an exploration of Explainable AI (XAI) technologies within the realm of GNNs. Amid the challenges posed by the …
A Human-In-The-Loop Framework For Scalable And Interpretable Event Triaging In Large-Scale Systems, Ibrahim Khaled Al-Agha
A Human-In-The-Loop Framework For Scalable And Interpretable Event Triaging In Large-Scale Systems, Ibrahim Khaled Al-Agha
Doctoral Dissertations
This dissertation presents a comprehensive and scalable framework for real-time fault detection and event triage in industrial systems, addressing critical challenges such as class imbalance, ambiguous feature boundaries, and the prioritization of complex, high-dimensional event data. The proposed framework integrates advanced methodologies, including micro-batch processing, retrospective divergence-based event detection (DB-RED), association rule mining (ARM), clustering, and Dempster-Shafer Theory (DST) for conflict resolution. Together, these components enable the systematic stratification of events into actionable priority levels, ensuring robust and interpretable decision-making in real-time environments. DB-RED forms the cornerstone of the framework, leveraging KL-divergence and PE-divergence metrics to detect subtle and transient …
Applying Machine Learning And Optimization Algorithms To Perform Feature Selection, Shizhao Yu
Applying Machine Learning And Optimization Algorithms To Perform Feature Selection, Shizhao Yu
Theses and Dissertations (Comprehensive)
The objective of feature selection in the realms of machine learning and data mining is integral, serving as an efficient mechanism to eradicate redundant or irrelevant features, and subsequently augmenting the performance of predictive models. In the contemporary landscape of big data, with the escalating dimensionality of datasets, the efficacy of traditional feature selection methodologies is compromised, due to their computational complexity and ineptitude in addressing the curse of dimensionality. This thesis posits a pioneering feature selection framework that amalgamates machine learning with advanced optimization algorithms. The methodology employs a Support Vector Machine (SVM), in conjunction with a cutting-edge metaheuristic …
A Threat Assessment Method In Uncertain Dynamic Environments, Mei Yang, Bingkun Wang, Zhongjie Zhang, Yan Zeng, Jian Huang
A Threat Assessment Method In Uncertain Dynamic Environments, Mei Yang, Bingkun Wang, Zhongjie Zhang, Yan Zeng, Jian Huang
Journal of System Simulation
Abstract: A threat assessment method based on priori information and dynamic observation results is studied for the existence of dynamic uncertainty in complex war systems. The data mining is applied to obtain prior knowledge on the battlefield situation and construct an equipment-related confidence matrix. The sensor model is constructed to dynamically update the number of blue-side entities under the current situation by using the Bayesian method and considering both intelligence and observation results. The threat evaluation indicators and their weights are determined, and the TOPSIS method is used to finish the threat assessment. This method can well describe the complex …
Developing Linguistic Patterns To Mitigate Inherent Human Bias In Offensive Language Detection, Toygar Tanyel, Besher Alkurdi, Serkan Ayvaz
Developing Linguistic Patterns To Mitigate Inherent Human Bias In Offensive Language Detection, Toygar Tanyel, Besher Alkurdi, Serkan Ayvaz
Turkish Journal of Electrical Engineering and Computer Sciences
With the proliferation of social media, there has been a sharp increase in offensive content, particularly targeting vulnerable groups, exacerbating social problems such as hatred, racism, and sexism. Detecting offensive language use is crucial to prevent offensive language from being widely shared on social media. However, the accurate detection of irony, implication, and various forms of hate speech on social media remains a challenge. Natural language-based deep learning models require extensive training with large, comprehensive, and labeled datasets. Unfortunately, manually creating such datasets is both costly and error-prone. Additionally, the presence of human-bias in offensive language datasets is a major …
Latent Representation Learning For Geospatial Entities, Ween Jiann Lee, Hady Wirawan Lauw
Latent Representation Learning For Geospatial Entities, Ween Jiann Lee, Hady Wirawan Lauw
Research Collection School Of Computing and Information Systems
Representation learning has been instrumental in the success of machine learning, offering compact and performant data representations for diverse downstream tasks. In the spatial domain, it has been pivotal in extracting latent patterns from various data types, including points, polylines, polygons, and networked structures. However, existing approaches often fall short of explicitly capturing both semantic and spatial information, relying on proxies and synthetic features. This article presents GeoNN, a novel graph neural network-based model designed to learn spatially-aware embeddings for geospatial entities. GeoNN leverages edge features generated from geodesic functions, dynamically selecting relevant features based on relative locations. It introduces …
A Multimodal Foundation Agent For Financial Trading : Tool-Augmented, Diversified, And Generalist, Wentao Zhang, Lingxuan Zhao, Haochong Xia, Shuo Sun, Jiaze Sun, Molei Qin, Xinyi Li, Yuqing Zhao, Yilei Zhao, Xinyu Cai, Longtao Zheng, Xinrun Wang, Bo An
A Multimodal Foundation Agent For Financial Trading : Tool-Augmented, Diversified, And Generalist, Wentao Zhang, Lingxuan Zhao, Haochong Xia, Shuo Sun, Jiaze Sun, Molei Qin, Xinyi Li, Yuqing Zhao, Yilei Zhao, Xinyu Cai, Longtao Zheng, Xinrun Wang, Bo An
Research Collection School Of Computing and Information Systems
Financial trading is a crucial component of the markets, informed by a multimodal information landscape encompassing news, prices, and Kline charts, and encompasses diverse tasks such as quantitative trading and high-frequency trading with various assets. While advanced AI techniques like deep learning and reinforcement learning are extensively utilized in finance, their application in financial trading tasks often faces challenges due to inadequate handling of multimodal data and limited generalizability across various tasks. To address these challenges, we present FinAgent, a multimodal foundational agent with tool augmentation for financial trading. FinAgent's market intelligence module processes a diverse range of data-numerical, textual, …
Study On Data Mining Of Hydrogen Energy Policy In China Based On Natural Language Processing Technology, Dongling Huang, Yuan Liu, Xiaoshuai Yuan, Guozhong Jin, Yuanhang Cai, Li Liu, Heng Cao, Wanjun Li, Rui Cai
Study On Data Mining Of Hydrogen Energy Policy In China Based On Natural Language Processing Technology, Dongling Huang, Yuan Liu, Xiaoshuai Yuan, Guozhong Jin, Yuanhang Cai, Li Liu, Heng Cao, Wanjun Li, Rui Cai
Bulletin of Chinese Academy of Sciences (Chinese Version)
Report to the 20th National Congress of the CPC emphasized the importance of “working actively and prudently towards the goals of reaching peak carbon emissions and carbon neutrality”, as well as “speeding up the planning and development of a system for new energy sources”. As a green and low-carbon secondary energy source, hydrogen energy has multiple applications in promoting the large-scale and efficient use of renewable energy as well as energy substitution in the field of transportation. It can also accelerate decarbonization in industry, and as such, is an indispensable part of building a new energy system, reaching peak carbon …
High-Dimensional Data Analysis Using Parameter Free Algorithm Data Point Positioning Analysis, S. M. F. D. Syed Mustapha
High-Dimensional Data Analysis Using Parameter Free Algorithm Data Point Positioning Analysis, S. M. F. D. Syed Mustapha
All Works
Clustering is an effective statistical data analysis technique; it has several applications, including data mining, pattern recognition, image analysis, bioinformatics, and machine learning. Clustering helps to partition data into groups of objects with distinct characteristics. Most of the methods for clustering use manually selected parameters to find the clusters from the dataset. Consequently, it can be very challenging and time-consuming to extract the optimal parameters for clustering a dataset. Moreover, some clustering methods are inadequate for locating clusters in high-dimensional data. To address these concerns systematically, this paper introduces a novel selection-free clustering technique named data point positioning analysis (DPPA). …
Dynamic Storytelling Algorithms Using Contextual Aspects Of A Large Language Model, Alireza Pasha Nouri
Dynamic Storytelling Algorithms Using Contextual Aspects Of A Large Language Model, Alireza Pasha Nouri
Open Access Theses & Dissertations
Storytelling is a set of algorithms used to create narratives by connecting documents in a sequencethat accurately reflects the evolution of events and entities within a particular topic or theme. Early storytelling algorithms face challenges in encoding the progression and interconnections of information between consecutive texts, given that the conventional approaches rely primarily on connecting document pairs based on content overlap. They often neglect critical linguistic features, such as word contexts, semantics, the roles words play across different documents, and attention to the historical contexts of the underlying documents. Many existing storytelling models frequently produce story chains that, while connected …
Using Data Mining To Analyze Job Reviews, Nicholas Bornkamp, Tony Breitzman
Using Data Mining To Analyze Job Reviews, Nicholas Bornkamp, Tony Breitzman
STEM Student Research Symposium Posters
Job review websites like Glassdoor are not always clear on how well the company operates, especially as viewed from differing levels of employment. For instance, a middle or upper manager from Amazon may have an overall positive review of the company with minor issues about it, but someone who works in the warehouse may have a mixed experience. To solve this issue and determine any correlation between employee level and their review, data mining techniques were utilized such as website scraping and neural network training to develop a model that analyzes employee reviews.
Spatio-Temporal Association Rule Mining Of Traffic Congestion In A Large-Scale Road Network Based On Trajectory Data, Qifan Zhou, Haixu Liu, Zhipeng Dong, Yin Xu
Spatio-Temporal Association Rule Mining Of Traffic Congestion In A Large-Scale Road Network Based On Trajectory Data, Qifan Zhou, Haixu Liu, Zhipeng Dong, Yin Xu
Journal of System Simulation
Abstract: A K neighbor-RElim (KNR) algorithm and a sequential KNbr-RElim (SKNR) algorithm are proposed to mine traffic congestion association rules and congestion propagation spatio-temporal association rules by vehicle trajectory data in a large-scale road network. The KNR algorithm extends the spatial topology constraint based on the RElim algorithm. The KNR can be used to mine the road links prone to congestion from the large-scale trajectory dataset in a large-scale road network and quantify the strength of association for congested road links. The SKNR algorithm expands the time dimension in the form of sliding window and can be applied for mining …
Mobilytics: Mobility Analytics Framework For Transferring Semantic Knowledge, Shreya Ghosh, Soumya K. Ghosh, Sajal K. Das, Prasenjit Mitra
Mobilytics: Mobility Analytics Framework For Transferring Semantic Knowledge, Shreya Ghosh, Soumya K. Ghosh, Sajal K. Das, Prasenjit Mitra
Computer Science Faculty Research & Creative Works
The proliferation of sensor-equipped smartphones has led to the generation of vast amounts of GPS data, such as timestamped location points, enabling a range of location-based services. However, deciphering the spatio-temporal dynamics of mobility to understand the underlying motivations behind travel patterns presents a significant challenge. his paper focuses on how individuals' GPS traces (latitude, longitude, timestamp) interpret the connection and correlations among different entities such as people, locations or point-of-interests (POIs), and semantic contexts (trip-purpose). We introduce a mobility analytics framework, named Mobilytics designed to identify trip purposes from individual GPS traces by leveraging a “mobility knowledge graph” (MKG) …
Chatting With Ai: Deciphering Developer Conversations With Chatgpt, Esteban Parra Rodriguez, Suad Mohamed, Abdullah Parvin
Chatting With Ai: Deciphering Developer Conversations With Chatgpt, Esteban Parra Rodriguez, Suad Mohamed, Abdullah Parvin
Funded Scholarship
Large Language Models (LLMs) have been widely adopted and are becoming ubiquitous and integral to software development. However, we have little knowledge as to how these tools are being used by software developers beyond anecdotal evidence and word-of-mouth reports. In this work, we present a study toward understanding how developers engage with and utilize LLMs by reporting the results of an empirical study identifying patterns in the conversation that developers have with LLMs. We identified a total of 19 topics describing the purpose of the developers in their conversations with LLMs. Our findings reveal that developers use LLMs to facilitate …
A Systemic Mapping Study On Intrusion Response Systems, Adel Rezapour, Mohammad Ghasemigol, Daniel Takabi
A Systemic Mapping Study On Intrusion Response Systems, Adel Rezapour, Mohammad Ghasemigol, Daniel Takabi
School of Cybersecurity Faculty Publications
With the increasing frequency and sophistication of network attacks, network administrators are facing tremendous challenges in making fast and optimum decisions during critical situations. The ability to effectively respond to intrusions requires solving a multi-objective decision-making problem. While several research studies have been conducted to address this issue, the development of a reliable and automated Intrusion Response System (IRS) remains unattainable. This paper provides a Systematic Mapping Study (SMS) for IRS, aiming to investigate the existing studies, their limitations, and future directions in this field. A novel semi-automated research methodology is developed to identify and summarize related works. The innovative …
Conceptthread: Visualizing Threaded Concepts In Mooc Videos, Zhiguang Zhou, Li Ye, Lihong Cai, Lei Wang, Yigang Wang, Yongheng Wang, Wei Chen, Yong Wang
Conceptthread: Visualizing Threaded Concepts In Mooc Videos, Zhiguang Zhou, Li Ye, Lihong Cai, Lei Wang, Yigang Wang, Yongheng Wang, Wei Chen, Yong Wang
Research Collection School Of Computing and Information Systems
Massive Open Online Courses (MOOCs) platforms are becoming increasingly popular in recent years. Online learners need to watch the whole course video on MOOC platforms to learn the underlying new knowledge, which is often tedious and time-consuming due to the lack of a quick overview of the covered knowledge and their structures. In this paper, we propose ConceptThread , a visual analytics approach to effectively show the concepts and the relations among them to facilitate effective online learning. Specifically, given that the majority of MOOC videos contain slides, we first leverage video processing and speech analysis techniques, including shot recognition, …
Short: Can Citations Tell Us About A Paper's Reproducibility? A Case Study Of Machine Learning Papers, Rochana R. Obadage, Sarah M. Rajtmajer, Jian Wu
Short: Can Citations Tell Us About A Paper's Reproducibility? A Case Study Of Machine Learning Papers, Rochana R. Obadage, Sarah M. Rajtmajer, Jian Wu
Computer Science Faculty Publications
The iterative character of work in machine learning (ML) and artificial intelligence (AI) and reliance on comparisons against benchmark datasets emphasize the importance of reproducibility in that literature. Yet, resource constraints and inadequate documentation can make running replications particularly challenging. Our work explores the potential of using downstream citation contexts as a signal of reproducibility. We introduce a sentiment analysis framework applied to citation contexts from papers involved in Machine Learning Reproducibility Challenges in order to interpret the positive or negative outcomes of reproduction attempts. Our contributions include training classifiers for reproducibility-related contexts and sentiment analysis, and exploring correlations between …
Evaluating Social Media Reach Via Mainstream Media Discourse, Himarsha R. Jayanetti
Evaluating Social Media Reach Via Mainstream Media Discourse, Himarsha R. Jayanetti
Computer Science Faculty Publications
This study examines the intersection between social media and mainstream television (TV) news with an aim to understand how social media content amplifies its impact through TV broadcasts. While many studies emphasize social media as a primary platform for information dissemination, they often underestimate its total influence by focusing solely on interactions within the platform. This research examines instances where social media posts gain prominence on TV broadcasts, reaching new audiences and prompting public discourse. By using TV news closed captions, on-screen text recognition, and social media logo detection, we analyze how social media is referenced in TV news. Our …
Homln-Sd: Substructure Discovery In Homogeneous Multilayer Networks, Arshdeep Singh
Homln-Sd: Substructure Discovery In Homogeneous Multilayer Networks, Arshdeep Singh
Computer Science and Engineering Theses - Archive
Substructure discovery is a process in data analysis and data mining that involves identifying and extracting meaningful patterns, structures, or components within a larger dataset. These substructures can be of various types, such as frequent patterns, motifs, or any other relevant features within the data. The growth of the internet and the proliferation of mobile devices have led to the generation of enormous amounts of data. Companies like Facebook and Twitter can generate large datasets from user interactions on their websites, such as connections between users and user generated content. Moreover, advances in processing power and storage capacity have made …
On The Effect Of Emotion Identification From Limited Translated Text Samples Using Computational Intelligence, Madiha Tahir, Zahid Halim, Muhmmad Waqas, Shanshan Tu
On The Effect Of Emotion Identification From Limited Translated Text Samples Using Computational Intelligence, Madiha Tahir, Zahid Halim, Muhmmad Waqas, Shanshan Tu
Research outputs 2022 to 2026
Emotion identification from text data has recently gained focus of the research community. This has multiple utilities in an assortment of domains. Many times, the original text is written in a different language and the end-user translates it to her native language using online utilities. Therefore, this paper presents a framework to detect emotions on translated text data in four different languages. The source language is English, whereas the four target languages include Chinese, French, German, and Spanish. Computational intelligence (CI) techniques are applied to extract features, dimensionality reduction, and classification of data into five basic classes of emotions. Results …
Predictive Analysis Of Students’ Learning Performance Using Data Mining Techniques: A Comparative Study Of Feature Selection Methods, S. M. F. D. Syed Mustapha
Predictive Analysis Of Students’ Learning Performance Using Data Mining Techniques: A Comparative Study Of Feature Selection Methods, S. M. F. D. Syed Mustapha
All Works
The utilization of data mining techniques for the prompt prediction of academic success has gained significant importance in the current era. There is an increasing interest in utilizing these methodologies to forecast the academic performance of students, thereby facilitating educators to intervene and furnish suitable assistance when required. The purpose of this study was to determine the optimal methods for feature engineering and selection in the context of regression and classification tasks. This study compared the Boruta algorithm and Lasso regression for regression, and Recursive Feature Elimination (RFE) and Random Forest Importance (RFI) for classification. According to the findings, Gradient …
Cannabidiol Tweet Miner: A Framework For Identifying Misinformation In Cbd Tweets., Jason Turner
Cannabidiol Tweet Miner: A Framework For Identifying Misinformation In Cbd Tweets., Jason Turner
Electronic Theses and Dissertations
As regulations surrounding cannabis continue to develop, the demand for cannabis-based products is on the rise. Despite not producing the psychoactive effects commonly associated with THC, products containing cannabidiol (CBD) have gained immense popularity in recent years as a potential treatment option for a range of conditions, particularly those associated with pain or sleep disorders. However, due to current federal policies, these products have yet to undergo comprehensive safety and efficacy testing. Fortunately, utilizing advanced natural language processing (NLP) techniques, data harvested from social networks have been employed to investigate various social trends within healthcare, such as disease tracking and …
Bertnet: Harvesting Knowledge Graphs With Arbitrary Relations From Pretrained Language Models, Shibo Hao, Bowen Tan, Kaiwen Tang, Bin Ni, Xiyan Shao, Hengzhe Zhang, Eric P. Xing, Zhiting Hu
Bertnet: Harvesting Knowledge Graphs With Arbitrary Relations From Pretrained Language Models, Shibo Hao, Bowen Tan, Kaiwen Tang, Bin Ni, Xiyan Shao, Hengzhe Zhang, Eric P. Xing, Zhiting Hu
Machine Learning Faculty Publications
It is crucial to automatically construct knowledge graphs (KGs) of diverse new relations to support knowledge discovery and broad applications. Previous KG construction methods, based on either crowdsourcing or text mining, are often limited to a small predefined set of relations due to manual cost or restrictions in text corpus. Recent research proposed to use pretrained language models (LMs) as implicit knowledge bases that accept knowledge queries with prompts. Yet, the implicit knowledge lacks many desirable properties of a full-scale symbolic KG, such as easy access, navigation, editing, and quality assurance. In this paper, we propose a new approach of …
Understanding Masked Autoencoders Via Hierarchical Latent Variable Models, Lingjing Kong, Martin Q. Ma, Guangyi Chen, Eric P. Xing, Yuejie Chi, Louis Philippe Morency, Kun Zhang
Understanding Masked Autoencoders Via Hierarchical Latent Variable Models, Lingjing Kong, Martin Q. Ma, Guangyi Chen, Eric P. Xing, Yuejie Chi, Louis Philippe Morency, Kun Zhang
Machine Learning Faculty Publications
Masked autoencoder (MAE), a simple and effective self-supervised learning framework based on the reconstruction of masked image regions, has recently achieved prominent success in a variety of vision tasks. Despite the emergence of intriguing empirical observations on MAE, a theoretically principled understanding is still lacking. In this work, we formally characterize and justify existing empirical in-sights and provide theoretical guarantees of MAE. We formulate the underlying data-generating process as a hierarchical latent variable model, and show that under reasonable assumptions, MAE provably identifies a set of latent variables in the hierarchical model, explaining why MAE can extract high-level information from …
Ptmtorrent: A Dataset For Mining Open-Source Pre-Trained Model Packages, Wenxin Jiang, Nicholas Synovic, Purvish Jajal, Taylor R. Schorlemmer, Arav Tewari, Bhavesh Pareek, George K. Thiruvathukal, James C. Davis
Ptmtorrent: A Dataset For Mining Open-Source Pre-Trained Model Packages, Wenxin Jiang, Nicholas Synovic, Purvish Jajal, Taylor R. Schorlemmer, Arav Tewari, Bhavesh Pareek, George K. Thiruvathukal, James C. Davis
Computer Science: Faculty Publications and Other Works
Due to the cost of developing and training deep learning models from scratch, machine learning engineers have begun to reuse pre-trained models (PTMs) and fine-tune them for downstream tasks. PTM registries known as “model hubs” support engineers in distributing and reusing deep learning models. PTM packages include pre-trained weights, documentation, model architectures, datasets, and metadata. Mining the information in PTM packages will enable the discovery of engineering phenomena and tools to support software engineers. However, accessing this information is difficult — there are many PTM registries, and both the registries and the individual packages may have rate limiting for accessing …
Towards Understanding The Open Source Interest In Gender-Related Github Projects, Rita Garcia, Christoph Treude, Wendy La
Towards Understanding The Open Source Interest In Gender-Related Github Projects, Rita Garcia, Christoph Treude, Wendy La
Research Collection School Of Computing and Information Systems
The open-source community uses the GitHub platform to exchange and share software applications and services of interest. This paper aims to identify the open-source community’s interest in gender-related projects on GitHub. Our findings create research opportunities and identify resources by the open-source community that promote diversity, equity, and inclusion. We use data mining to identify GitHub projects that focus on gender-related topics. We apply quantitative and qualitative methodologies to examine the projects’ attributes and to classify them within a gender social structure and a gender bias taxonomy. We aim to understand the open-source community’s efforts and interests in gender topics …
Analyzing Syntactic Constructs Of Java Programs With Machine Learning, Francisco Ortin, Guillermo Facundo, Miguel Garcia
Analyzing Syntactic Constructs Of Java Programs With Machine Learning, Francisco Ortin, Guillermo Facundo, Miguel Garcia
Department of Computer Science Publications
The massive number of open-source projects in public repositories has notably increased in the last years. Such repositories represent valuable information to be mined for different purposes, such as documenting recurrent syntactic constructs, analyzing the particular constructs used by experts and beginners, using them to teach programming and to detect bad programming practices, and building programming tools such as decompilers, Integrated Development Environments or Intelligent Tutoring Systems. An inherent problem of source code is that its syntactic information is represented with tree structures, while traditional machine learning algorithms use -dimensional datasets. Therefore, we present a feature engineering process to translate …