Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences Commons™

Open Access. Powered by Scholars. Published by Universities.®

Databases and Information Systems

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 1801 - 1830 of 7256

Full-Text Articles in Computer Sciences

Structurally Enriched Entity Mention Embedding From Semi-Structured Textual Content, Lee Hsun Hsieh, Yang Yin Lee, Ee-Peng Lim Mar 2021

Structurally Enriched Entity Mention Embedding From Semi-Structured Textual Content, Lee Hsun Hsieh, Yang Yin Lee, Ee-Peng Lim

Research Collection School Of Computing and Information Systems

In this research, we propose a novel and effective entity mention embedding framework that learns from semi-structured text corpus with annotated entity mentions without the aid of well-constructed knowledge graph or external semantic information other than the corpus itself. Based on the co-occurrence of words and entity mentions, we enrich the co-occurrence matrix with entity-entity, entity-word, and word-entity relationships as well as the simple structures within the documents. Experimentally, we show that our proposed entity mention embedding benefits from the structural information in link prediction task measured by mean reciprocal rank (MRR) and mean precision@K (MP@K) on two datasets for …


Improving Multi-Hop Knowledge Base Question Answering By Learning Intermediate Supervision Signals, Gaole He, Yunshi Lan, Jing Jiang, Wayne Xin Zhao, Ji Rong Wen Mar 2021

Improving Multi-Hop Knowledge Base Question Answering By Learning Intermediate Supervision Signals, Gaole He, Yunshi Lan, Jing Jiang, Wayne Xin Zhao, Ji Rong Wen

Research Collection School Of Computing and Information Systems

Multi-hop Knowledge Base Question Answering (KBQA) aims to find the answer entities that are multiple hops away in the Knowledge Base (KB) from the entities in the question. A major challenge is the lack of supervision signals at intermediate steps. Therefore, multi-hop KBQA algorithms can only receive the feedback from the final answer, which makes the learning unstable or ineffective. To address this challenge, we propose a novel teacher-student approach for the multi-hop KBQA task. In our approach, the student network aims to find the correct answer to the query, while the teacher network tries to learn intermediate supervision signals …


All The Wiser: Fake News Intervention Using User Reading Preferences, Kuan Chieh Lo, Shih Chieh Dai, Aiping Xiong, Jing Jiang, Lun Wei Ku Mar 2021

All The Wiser: Fake News Intervention Using User Reading Preferences, Kuan Chieh Lo, Shih Chieh Dai, Aiping Xiong, Jing Jiang, Lun Wei Ku

Research Collection School Of Computing and Information Systems

To address the increasingly significant issue of fake news, we develop a news reading platform in which we propose an implicit approach to reduce people's belief in fake news. Specifically, we leverage reinforcement learning to learn an intervention module on top of a recommender system (RS) such that the module is activated to replace RS to recommend news toward the verification once users touch the fake news. To examine the effect of the proposed method, we conduct a comprehensive evaluation with 89 human subjects and check the effective rate of change in belief but without their other limitations. Moreover, 84% …


Bilateral Variational Autoencoder For Collaborative Filtering, Quoc Tuan Truong, Aghiles Salah, Hady W. Lauw Mar 2021

Bilateral Variational Autoencoder For Collaborative Filtering, Quoc Tuan Truong, Aghiles Salah, Hady W. Lauw

Research Collection School Of Computing and Information Systems

Preference data is a form of dyadic data, with measurements associated with pairs of elements arising from two discrete sets of objects. These are users and items, as well as their interactions, e.g., ratings. We are interested in learning representations for both sets of objects, i.e., users and items, to predict unknown pairwise interactions. Motivated by the recent successes of deep latent variable models, we propose Bilateral Variational Autoencoder (BiVAE), which arises from a combination of a generative model of dyadic data with two inference models, user- and item-based, parameterized by neural networks. Interestingly, our model can take the form …


Explainable Recommendation With Comparative Constraints On Product Aspects, Trung-Hoang Le, Hady W. Lauw Mar 2021

Explainable Recommendation With Comparative Constraints On Product Aspects, Trung-Hoang Le, Hady W. Lauw

Research Collection School Of Computing and Information Systems

To aid users in choice-making, explainable recommendation models seek to provide not only accurate recommendations but also accompanying explanations that help to make sense of those recommendations. Most of the previous approaches rely on evaluative explanations, assessing the quality of an individual item along some aspects of interest to the user. In this work, we are interested in comparative explanations, the less studied problem of assessing a recommended item in comparison to another reference item.

In particular, we propose to anchor reference items on the previously adopted items in a user's history. Not only do we aim at providing comparative …


How Do Monetary Incentives Influence Giving? An Empirical Investigation Of Matching Subsidies On Kiva, Zhiyuan Gao, Zhiling Guo, Qian Tang Mar 2021

How Do Monetary Incentives Influence Giving? An Empirical Investigation Of Matching Subsidies On Kiva, Zhiyuan Gao, Zhiling Guo, Qian Tang

Research Collection School Of Computing and Information Systems

Matching subsidies, through which third-party institutions provide a dollar-for-dollar match of private contributions made through selected campaigns, have served as effective tools to boost fundraising. We utilize a quasi-experiment on a prosocial crowdfunding platform to examine the effectiveness of matching subsidies in shaping funding outcomes and lender behaviors. Although matching subsidies offer matched loans competitive advantages over unmatched loans, we find that total private contributions made to both matched and unmatched loans increase compared to their prematching counterparts, suggesting a positive spillover effect on unmatched loans. However, matching subsidies lead to decreased private contributions made on the platform after a …


Enhancing Healthcare Professional And Caregiving Staff Informedness With Data Analytics For Chronic Disease Management, Na Liu, Robert John Kauffman Mar 2021

Enhancing Healthcare Professional And Caregiving Staff Informedness With Data Analytics For Chronic Disease Management, Na Liu, Robert John Kauffman

Research Collection School Of Computing and Information Systems

An important area in healthcare to which data analytics can be applied is chronic disease management. The chronic care model is mostly patient-centric, so patients have been considered as the end users of data analytics. The information needs of healthcare providers have been overlooked. Drawing upon the theory of informedness and the transtheoretical model of health behavior change, we use a multicase study approach to investigate the information needs of different caregiving stakeholders in the spectrum of chronic diseases, and how data analytics can be designed to meet the varying needs of professionals and staff to support their informedness.


How Do Users Answer Matlab Questions On Q&A Sites? A Case Study On Stack Overflow And Mathworks, Mahshid Naghashzadeh, Amir Hagshenas, Ashkan Sami, David Lo Mar 2021

How Do Users Answer Matlab Questions On Q&A Sites? A Case Study On Stack Overflow And Mathworks, Mahshid Naghashzadeh, Amir Hagshenas, Ashkan Sami, David Lo

Research Collection School Of Computing and Information Systems

MATLAB is an engineering programming language with various toolboxes that has a dedicated Question and Answer (Q&A) platform on the MathWorks website, which is similar to Stack Overflow (SO). Moreover, some MATLAB users ask their questions on SO. This paper aims to compare these two Q&A platforms to see what kind of questions are asked and how developers answer these questions in each platform. The result of our analysis on 80,382 MATLAB questions on SO and 266,367 questions on MathWorks show that MATLAB questions on topics ranging from the MATLAB software installation to questions related to programming received high votes …


Is The Ground Truth Really Accurate? Dataset Purification For Automated Program Repair, Deheng Yang, Yan Lei, Xiaoguang Mao, David Lo, Huan Xie, Meng Yan Mar 2021

Is The Ground Truth Really Accurate? Dataset Purification For Automated Program Repair, Deheng Yang, Yan Lei, Xiaoguang Mao, David Lo, Huan Xie, Meng Yan

Research Collection School Of Computing and Information Systems

Datasets of real-world bugs shipped with human-written patches are intensively used in the evaluation of existing automated program repair (APR) techniques, wherein the human-written patches always serve as the ground truth, for manual or automated assessment approaches, to evaluate the correctness of test-suite adequate patches. An inaccurate human-written patch tangled with other code changes will pose threats to the reliability of the assessment results. Therefore, the construction of such datasets always requires much manual effort on isolating real bug fixes from bug fixing commits. However, the manual work is time-consuming and prone to mistakes, and little has been known on …


Learning To Assess The Quality Of Stroke Rehabilitation Exercises, Min Hun Lee, Daniel P. Siewiorek, Asim Smailagic, Alexandre Bernardino, Sergi Bermúdez I Badia Mar 2021

Learning To Assess The Quality Of Stroke Rehabilitation Exercises, Min Hun Lee, Daniel P. Siewiorek, Asim Smailagic, Alexandre Bernardino, Sergi Bermúdez I Badia

Research Collection School Of Computing and Information Systems

Due to the limited number of therapists, task-oriented exercises are often prescribed for post-stroke survivors as in-home rehabilitation. During in-home rehabilitation, a patient may become unmotivated or confused to comply prescriptions without the feedback of a therapist. To address this challenge, this paper proposes an automated method that can achieve not only qualitative, but also quantitative assessment of stroke rehabilitation exercises. Specifically, we explored a threshold model that utilizes the outputs of binary classifiers to quantify the correctness of a movements into a performance score. We collected movements of 11 healthy subjects and 15 post-stroke survivors using a Kinect sensor …


Privacy-Preserving Multi-Keyword Searchable Encryption For Distributed Systems, Xueqiao Liu, Guomin Yang, Willy Susilo, Joseph Tonien, Jian Shen Mar 2021

Privacy-Preserving Multi-Keyword Searchable Encryption For Distributed Systems, Xueqiao Liu, Guomin Yang, Willy Susilo, Joseph Tonien, Jian Shen

Research Collection School Of Computing and Information Systems

As cloud storage has been widely adopted in various applications, how to protect data privacy while allowing efficient data search and retrieval in a distributed environment remains a challenging research problem. Existing searchable encryption schemes are still inadequate on desired functionality and security/privacy perspectives. Specifically, supporting multi-keyword search under the multi-user setting, hiding search pattern and access pattern, and resisting keyword guessing attacks (KGA) are the most challenging tasks. In this article, we present a new searchable encryption scheme that addresses the above problems simultaneously, which makes it practical to be adopted in distributed systems. It not only enables multi-keyword …


Deepis: Susceptibility Estimation On Social Networks, Wenwen Xia, Yuchen Li, Jun Wu, Shenghong Li Mar 2021

Deepis: Susceptibility Estimation On Social Networks, Wenwen Xia, Yuchen Li, Jun Wu, Shenghong Li

Research Collection School Of Computing and Information Systems

Influence diffusion estimation is a crucial problem in social network analysis. Most prior works mainly focus on predicting the total influence spread, i.e., the expected number of influenced nodes given an initial set of active nodes (aka. seeds). However, accurate estimation of susceptibility, i.e., the probability of being influenced for each individual, is more appealing and valuable in real-world applications. Previous methods generally adopt Monte Carlo simulation or heuristic rules to estimate the influence, resulting in high computational cost or unsatisfactory estimation error when these methods are used to estimate susceptibility. In this work, we propose to leverage graph neural …


Web Design Attributes Guideline To Reinforce User Trust, Satisfaction, And Loyalty For Malaysian University Students, Ranjen Naidu Vasudiven Feb 2021

Web Design Attributes Guideline To Reinforce User Trust, Satisfaction, And Loyalty For Malaysian University Students, Ranjen Naidu Vasudiven

Student Works (2020-2029)

Web site design is the most crucial factor influencing the online e-commerce business. It determines the customer's trust in the products and owners, increasing sales and recognizing its brand. There is currently no reference to understanding Malaysian university students' user preference in web site design. For this study, we evaluated Malaysian university students' user preferences for website design attributes mainly comprised of (interactivity, navigation typography, colour, and content quality). The analysis of the user preferences on the website attributes, the relationship of the attributes with the Malaysian cultures, and the findings based on the three Malaysian mostly visited bookstore websites …


Google Books, Jody Condit Fagan Feb 2021

Google Books, Jody Condit Fagan

Libraries

Google Books’ (GB) full-text search of more than 40 million books offers significant value for libraries and their patrons. However, Google’s refusal to disclose information about the coverage of GB, as well as observed gaps and inaccuracies in the collection and its metadata, makes it difficult to recommend with confidence for a given research need. While most search and retrieval functions work well, glitches aren’t hard to find, which suggests GB development is focused on user experiences that relate to monetization. Privacy and equity concerns surrounding GB mirror those of other big technology platforms. Still, every librarian should familiarize themselves …


Unsupervised Data Mining Technique For Clustering Library In Indonesia, Robbi Rahim, Joseph Teguh Santoso, Sri Jumini, Gita Widi Bhawika, Daniel Susilo, Danny Wibowo Feb 2021

Unsupervised Data Mining Technique For Clustering Library In Indonesia, Robbi Rahim, Joseph Teguh Santoso, Sri Jumini, Gita Widi Bhawika, Daniel Susilo, Danny Wibowo

Library Philosophy and Practice (e-journal)

Organizing school libraries not only keeps library materials, but helps students and teachers in completing tasks in the teaching process so that national development goals are in order to improve community welfare by producing quality and competitive human resources. The purpose of this study is to analyze the Unsupervised Learning technique in conducting cluster mapping of the number of libraries at education levels in Indonesia. The data source was obtained from the Ministry of Education and Culture which was processed by the Central Statistics Agency (abbreviated as BPS) with url: bps.go.id/. The data consisted of 34 records where the attribute …


Hybrid Cloud Workload Monitoring As A Service, Shreya Kundu Feb 2021

Hybrid Cloud Workload Monitoring As A Service, Shreya Kundu

Master's Projects

Cloud computing and cloud-based hosting has become embedded in our daily lives. It is imperative for cloud providers to make sure all services used by both enterprises and consumers have high availability and elasticity to prevent any downtime, which impacts negatively for any business. To ensure cloud infrastructures are working reliably, cloud monitoring becomes an essential need for both businesses, the provider and the consumer. This thesis project reports on the need of efficient scalable monitoring, enumerating the necessary types of metrics of interest to be collected. Current understanding of various architectures designed to collect, store and process monitoring data …


A New Feature Selection Method Based On Class Association Rule, Sami A. Al-Dhaheri Feb 2021

A New Feature Selection Method Based On Class Association Rule, Sami A. Al-Dhaheri

Dissertations, Theses, and Capstone Projects

Feature selection is a key process for supervised learning algorithms. It involves discarding irrelevant attributes from the training dataset from which the models are derived. One of the vital feature selection approaches is Filtering, which often uses mathematical models to compute the relevance for each feature in the training dataset and then sorts the features into descending order based on their computed scores. However, most Filtering methods face several challenges including, but not limited to, merely considering feature-class correlation when defining a feature’s relevance; additionally, not recommending which subset of features to retain. Leaving this decision to the end-user may …


Accelerating Large-Scale Heterogeneous Interaction Graph Embedding Learning Via Importance Sampling, Yugang Ji, Mingyang Yin, Hongxia Yang, Jingren Zhou, Vincent W. Zheng, Chuan Shi, Yuan Fang Feb 2021

Accelerating Large-Scale Heterogeneous Interaction Graph Embedding Learning Via Importance Sampling, Yugang Ji, Mingyang Yin, Hongxia Yang, Jingren Zhou, Vincent W. Zheng, Chuan Shi, Yuan Fang

Research Collection School Of Computing and Information Systems

In real-world problems, heterogeneous entities are often related to each other through multiple interactions, forming a Heterogeneous Interaction Graph (HIG in short). While modeling HIGs to deal with fundamental tasks, graph neural networks present an attractive opportunity that can make full use of the heterogeneity and rich semantic information by aggregating and propagating information from different types of neighborhoods. However, learning on such complex graphs, often with millions or billions of nodes, edges, and various attributes, could suffer from expensive time cost and high memory consumption. In this paper, we attempt to accelerate representation learning on large-scale HIGs by adopting …


Revman: Revenue-Aware Multi-Task Online Insurance Recommendation, Yu Li, Yi Zhang, Lu Gan, Gengwei Hong, Zimu Zhou, Qiang Li Feb 2021

Revman: Revenue-Aware Multi-Task Online Insurance Recommendation, Yu Li, Yi Zhang, Lu Gan, Gengwei Hong, Zimu Zhou, Qiang Li

Research Collection School Of Computing and Information Systems

Online insurance is a new type of e-commerce with exponential growth. An effective recommendation model that maximizes the total revenue of insurance products listed in multiple customized sales scenarios is crucial for the success of online insurance business. Prior recommendation models are ineffective because they fail to characterize the complex relatedness of insurance products in multiple sales scenarios and maximize the overall conversion rate rather than the total revenue. Even worse, it is impractical to collect training data online for total revenue maximization due to the business logic of online insurance. We propose RevMan, a Revenue-aware Multi-task Network for online …


Multi-Decoder Attention Model With Embedding Glimpse For Solving Vehicle Routing Problems, Liang Xin, Wen Song, Zhiguang Cao, Jie Zhang Feb 2021

Multi-Decoder Attention Model With Embedding Glimpse For Solving Vehicle Routing Problems, Liang Xin, Wen Song, Zhiguang Cao, Jie Zhang

Research Collection School Of Computing and Information Systems

We present a novel deep reinforcement learning method to learn construction heuristics for vehicle routing problems. In specific, we propose a Multi-Decoder Attention Model (MDAM) to train multiple diverse policies, which effectively increases the chance of finding good solutions compared with existing methods that train only one policy. A customized beam search strategy is designed to fully exploit the diversity of MDAM. In addition, we propose an Embedding Glimpse layer in MDAM based on the recursive nature of construction, which can improve the quality of each policy by providing more informative embeddings. Extensive experiments on six different routing problems show …


Differential Training: A Generic Framework To Reduce Label Noises For Android Malware Detection, Jiayun Xu, Yingjiu Li, Robert H. Deng Feb 2021

Differential Training: A Generic Framework To Reduce Label Noises For Android Malware Detection, Jiayun Xu, Yingjiu Li, Robert H. Deng

Research Collection School Of Computing and Information Systems

A common problem in machine learning-based malware detection is that training data may contain noisy labels and it is challenging to make the training data noise-free at a large scale. To address this problem, we propose a generic framework to reduce the noise level of training data for the training of any machine learning-based Android malware detection. Our framework makes use of all intermediate states of two identical deep learning classification models during their training with a given noisy training dataset and generate a noise-detection feature vector for each input sample. Our framework then applies a set of outlier detection …


Evidence Aware Neural Pornographic Text Identification For Child Protection, Kaisong Song, Yangyang Kang, Wei Gao, Zhe Gao, Changlong Sun, Xiaozhong Liu Feb 2021

Evidence Aware Neural Pornographic Text Identification For Child Protection, Kaisong Song, Yangyang Kang, Wei Gao, Zhe Gao, Changlong Sun, Xiaozhong Liu

Research Collection School Of Computing and Information Systems

Identifying pornographic text online is practically useful to protect children from access to such adult content. However, some authors may intentionally avoid using sensitive words in their pornographic texts to take advantage of the lack of human audits. Without prior knowledge guidance, real semantics of such pornographic text is difficult to understand by existing methods due to its high context-sensitivity and heavy usage of figurative language, which brings huge challenges to the porn detection systems used in social media platforms. In this paper, we approach to the problem as a document-level porn identification task by locating and integrating sentence-level evidence …


An Exploratory Study On The Introduction And Removal Of Different Types Of Technical Debt In Deep Learning Frameworks, Jiakun Liu, Qiao Huang, Xin Xia, Emad Shihab, David Lo, Shanping Li Feb 2021

An Exploratory Study On The Introduction And Removal Of Different Types Of Technical Debt In Deep Learning Frameworks, Jiakun Liu, Qiao Huang, Xin Xia, Emad Shihab, David Lo, Shanping Li

Research Collection School Of Computing and Information Systems

To complete tasks faster, developers often have to sacrifice the quality of the software. Such compromised practice results in the increasing burden to developers in future development. The metaphor, technical debt, describes such practice. Prior research has illustrated the negative impact of technical debt, and many researchers investigated how developers deal with a certain type of technical debt. However, few studies focused on the removal of different types of technical debt in practice. To fill this gap, we use the introduction and removal of different types of self-admitted technical debt (i.e., SATD) in 7 deep learning frameworks as an example. …


Evoking Empathy: A Framework For Describing Empathy Tools, Sydney Pratte, Anthony Tang, Lora Oehlberg Feb 2021

Evoking Empathy: A Framework For Describing Empathy Tools, Sydney Pratte, Anthony Tang, Lora Oehlberg

Research Collection School Of Computing and Information Systems

Empathy tools are experiences designed to evoke empathetic responses by placing the user in another’s lived and felt experience. The problem is that designers do not have a common vocabulary to describe empathy tool experiences; consequently, it is difficult to compare/contrast empathy tool designs or to think about their efficacy. To address this problem, we analyzed 26 publications on empathy tools to develop a descriptive framework for designers of empathy tools. Based on our analysis, we found that empathy tools can be described along three dimensions: (i) the amount of agency the tool allows, (ii) the user’s perspective while using …


Visual Analysis Of Discrimination In Machine Learning, Qianwen Wang, Zhenghua Xu, Zhutian Chen, Yong Wang, Shixia Liu, Huamin Qu Feb 2021

Visual Analysis Of Discrimination In Machine Learning, Qianwen Wang, Zhenghua Xu, Zhutian Chen, Yong Wang, Shixia Liu, Huamin Qu

Research Collection School Of Computing and Information Systems

The growing use of automated decision-making in critical applications, such as crime prediction and college admission, has raised questions about fairness in machine learning. How can we decide whether different treatments are reasonable or discriminatory? In this paper, we investigate discrimination in machine learning from a visual analytics perspective and propose an interactive visualization tool, DiscriLens, to support a more comprehensive analysis. To reveal detailed information on algorithmic discrimination, DiscriLens identifies a collection of potentially discriminatory itemsets based on causal modeling and classification rules mining. By combining an extended Euler diagram with a matrix-based visualization, we develop a novel set …


Qlens: Visual Analytics Of Multi-Step Problem-Solving Behaviors For Improving Question Design, Meng Xia, Reshika P. Velumani, Yong Wang, Huamin Qu, Xiaojuan Ma Feb 2021

Qlens: Visual Analytics Of Multi-Step Problem-Solving Behaviors For Improving Question Design, Meng Xia, Reshika P. Velumani, Yong Wang, Huamin Qu, Xiaojuan Ma

Research Collection School Of Computing and Information Systems

With the rapid development of online education in recent years, there has been an increasing number of learning platforms that provide students with multi-step questions to cultivate their problem-solving skills. To guarantee the high quality of such learning materials, question designers need to inspect how students’ problem-solving processes unfold step by step to infer whether students’ problem-solving logic matches their design intent. They also need to compare the behaviors of different groups (e.g., students from different grades) to distribute questions to students with the right level of knowledge. The availability of fine-grained interaction data, such as mouse movement trajectories from …


Learning To Pre-Train Graph Neural Networks, Yuanfu Lu, Xunqiang Jiang, Yuan Fang, Chuan Shi Feb 2021

Learning To Pre-Train Graph Neural Networks, Yuanfu Lu, Xunqiang Jiang, Yuan Fang, Chuan Shi

Research Collection School Of Computing and Information Systems

Graph neural networks (GNNs) have become the de facto standard for representation learning on graphs, which derive effective node representations by recursively aggregating information from graph neighborhoods. While GNNs can be trained from scratch, pre-training GNNs to learn transferable knowledge for downstream tasks has recently been demonstrated to improve the state of the art. However, conventional GNN pre-training methods follow a two-step paradigm: 1) pre-training on abundant unlabeled data and 2) fine-tuning on downstream labeled data, between which there exists a significant gap due to the divergence of optimization objectives in the two steps. In this paper, we conduct an …


Relative And Absolute Location Embedding For Few-Shot Node Classification On Graph, Zemin Liu, Yuan Fang, Chenghao Liu, Steven C. H. Hoi Feb 2021

Relative And Absolute Location Embedding For Few-Shot Node Classification On Graph, Zemin Liu, Yuan Fang, Chenghao Liu, Steven C. H. Hoi

Research Collection School Of Computing and Information Systems

Node classification is an important problem on graphs. While recent advances in graph neural networks achieve promising performance, they require abundant labeled nodes for training. However, in many practical scenarios there often exist novel classes in which only one or a few labeled nodes are available as supervision, known as few-shot node classification. Although meta-learning has been widely used in vision and language domains to address few-shot learning, its adoption on graphs has been limited. In particular, graph nodes in a few-shot task are not independent and relate to each other. To deal with this, we propose a novel model …


Delineating Knowledge Domains In Scientific Domains In Scientific Literature Using Machine Learning (Ml), Abhay Maurya, Smarajit Paul Choudhury Mr., Kshitij Jaiswal Mr. Jan 2021

Delineating Knowledge Domains In Scientific Domains In Scientific Literature Using Machine Learning (Ml), Abhay Maurya, Smarajit Paul Choudhury Mr., Kshitij Jaiswal Mr.

Library Philosophy and Practice (e-journal)

The recent years have witnessed an upsurge in the number of published documents. Organizations are showing an increased interest in text classification for effective use of the information. Manual procedures for text classification can be fruitful for a handful of documents, but the same lack in credibility when the number of documents increases besides being laborious and time-consuming. Text mining techniques facilitate assigning text strings to categories rendering the process of classification fast, accurate, and hence reliable. This paper classifies chemistry documents using machine learning and statistical methods. The procedure of text classification has been described in chronological order like …


Automation Of Crawling Blogosphere Based On Pattern Recognition, Anal Kanti Roy Jan 2021

Automation Of Crawling Blogosphere Based On Pattern Recognition, Anal Kanti Roy

Theses and Dissertations

Social media plays an important role in the propagation and dissemination of ideas and thoughts. Compared to other social media platforms, blogs provide a convenient platform for users to post detailed information, engage in active discussions and share the content on other social media sites, such as Facebook and Twitter. Thus, the blogosphere has been an enormous and ever-growing part of the open-source intelligence. In order to track and monitor online social behavior particularly from blogs, the first challenging part is to mine the vast pool of unstructured data. To scale up this process and cope with the continuously changing …