Open Access. Powered by Scholars. Published by Universities.®

Databases and Information Systems Commons™

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 2101 - 2130 of 7256

Full-Text Articles in Databases and Information Systems

Probabilistic Value Selection For Space Efficient Model, Gunarto Sindoro Njoo, Baihua Zheng, Kuo-Wei Hsu, Wen-Chih Peng Jul 2020

Probabilistic Value Selection For Space Efficient Model, Gunarto Sindoro Njoo, Baihua Zheng, Kuo-Wei Hsu, Wen-Chih Peng

Research Collection School Of Computing and Information Systems

An alternative to current mainstream preprocessing methods is proposed: Value Selection (VS). Unlike the existing methods such as feature selection that removes features and instance selection that eliminates instances, value selection eliminates the values (with respect to each feature) in the dataset with two purposes: reducing the model size and preserving its accuracy. Two probabilistic methods based on information theory's metric are proposed: PVS and P + VS. Extensive experiments on the benchmark datasets with various sizes are elaborated. Those results are compared with the existing preprocessing methods such as feature selection, feature transformation, and instance selection methods. Experiment results …


Next-Term Grade Prediction: A Machine Learning Approach, Audrey Tedja Widjaja, Lei Wang, Nghia Truong Trong, Aldy Gunawan, Ee-Peng Lim Jul 2020

Next-Term Grade Prediction: A Machine Learning Approach, Audrey Tedja Widjaja, Lei Wang, Nghia Truong Trong, Aldy Gunawan, Ee-Peng Lim

Research Collection School Of Computing and Information Systems

As students progress in their university programs, they have to face many course choices. It is important for them to receive guidance based on not only their interest, but also the "predicted" course performance so as to improve learning experience and optimise academic performance. In this paper, we propose the next-term grade prediction task as a useful course selection guidance. We propose a machine learning framework to predict course grades in a specific program term using the historical student-course data. In this framework, we develop the prediction model using Factorization Machine (FM) and Long Short Term Memory combined with FM …


Improving Multimodal Named Entity Recognition Via Entity Span Detection With Unified Multimodal Transformer, Jianfei Yu, Jing Jiang, Li Yang, Rui Xia Jul 2020

Improving Multimodal Named Entity Recognition Via Entity Span Detection With Unified Multimodal Transformer, Jianfei Yu, Jing Jiang, Li Yang, Rui Xia

Research Collection School Of Computing and Information Systems

In this paper, we study Multimodal Named Entity Recognition (MNER) for social media posts. Existing approaches for MNER mainly suffer from two drawbacks: (1) despite generating word-aware visual representations, their word representations are insensitive to the visual context; (2) most of them ignore the bias brought by the visual context. To tackle the first issue, we propose a multimodal interaction module to obtain both image-aware word representations and word-aware visual representations. To alleviate the visual bias, we further propose to leverage purely text-based entity span detection as an auxiliary module, and design a Unified Multimodal Transformer to guide the final …


Biane: Bipartite Attributed Network Embedding, Wentao Huang, Yuchen Li, Yuan Fang, Ju Fan, Hongxia Yang Jul 2020

Biane: Bipartite Attributed Network Embedding, Wentao Huang, Yuchen Li, Yuan Fang, Ju Fan, Hongxia Yang

Research Collection School Of Computing and Information Systems

Network embedding effectively transforms complex network data into a low-dimensional vector space and has shown great performance in many real-world scenarios, such as link prediction, node classification, and similarity search. A plethora of methods have been proposed to learn node representations and achieve encouraging results. Nevertheless, little attention has been paid on the embedding technique for bipartite attributed networks, which is a typical data structure for modeling nodes from two distinct partitions. In this paper, we propose a novel model called BiANE, short for Bipartite Attributed Network Embedding. In particular, BiANE not only models the inter-partition proximity but also models …


Semi-Supervised Co-Clustering On Attributed Heterogeneous Information Networks, Yugang Ji, Chuan Shi, Yuan Fang, Xiangnan Kong, Mingyang Yin Jul 2020

Semi-Supervised Co-Clustering On Attributed Heterogeneous Information Networks, Yugang Ji, Chuan Shi, Yuan Fang, Xiangnan Kong, Mingyang Yin

Research Collection School Of Computing and Information Systems

Node clustering on heterogeneous information networks (HINs) plays an important role in many real-world applications. While previous research mainly clusters same-type nodes independently via exploiting structural similarity search, they ignore the correlations of different-type nodes. In this paper, we focus on the problem of co-clustering heterogeneous nodes where the goal is to mine the latent relevance of heterogeneous nodes and simultaneously partition them into the corresponding type-aware clusters. This problem is challenging in two aspects. First, the similarity or relevance of nodes is not only associated with multiple meta-path-based structures but also related to numerical and categorical attributes. Second, clusters …


Covid-19 Pandemic: Role Of Technology In Transforming Business To The New Normal, Fiona Fui-Hoon Nah, Keng Siau Jul 2020

Covid-19 Pandemic: Role Of Technology In Transforming Business To The New Normal, Fiona Fui-Hoon Nah, Keng Siau

Research Collection School Of Computing and Information Systems

COVID-19 has disrupted our lives and the economy. In this paper, we outline approaches in which information technology can be used to implement business strategies to enhance resilience by coping with, adapting to, and recovering from adversity resulting from the COVID-19 pandemic. We discuss how information technology such as digital supply chain, data analytics, artificial intelligence, machine learning, robotics, digital commerce, and Internet of Things can be used to enhance resilience and continuity of business.


Answer Ranking For Product-Related Questions Via Multiple Semantic Relations Modeling, Wenxuan Zhang, Yang Deng, Wai Lam Jul 2020

Answer Ranking For Product-Related Questions Via Multiple Semantic Relations Modeling, Wenxuan Zhang, Yang Deng, Wai Lam

Research Collection School Of Computing and Information Systems

Many E-commerce sites now offer product-specific question answering platforms for users to communicate with each other by posting and answering questions during online shopping. However, the multiple answers provided by ordinary users usually vary diversely in their qualities and thus need to be appropriately ranked for each question to improve user satisfaction. It can be observed that product reviews usually provide useful information for a given question, and thus can assist the ranking process. In this paper, we investigate the answer ranking problem for product-related questions, with the relevant reviews treated as auxiliary information that can be exploited for facilitating …


Bridging Hierarchical And Sequential Context Modeling For Question-Driven Extractive Answer Summarization, Yang Deng, Wenxuan Zhang, Yaliang Li, Min Yang, Wai Lam, Ying Shen Jul 2020

Bridging Hierarchical And Sequential Context Modeling For Question-Driven Extractive Answer Summarization, Yang Deng, Wenxuan Zhang, Yaliang Li, Min Yang, Wai Lam, Ying Shen

Research Collection School Of Computing and Information Systems

Non-factoid question answering (QA) is one of the most extensive yet challenging application and research areas of retrieval-based question answering. In particular, answers to non-factoid questions can often be too lengthy and redundant to comprehend, which leads to the great demand on answer sumamrization in non-factoid QA. However, the multi-level interactions between QA pairs and the interrelation among different answer sentences are usually modeled separately on current answer summarization studies. In this paper, we propose a unified model to bridge hierarchical and sequential context modeling for question-driven extractive answer summarization. Specifically, we design a hierarchical compare-aggregate method to integrate the …


Towards An Optimal Outdoor Advertising Placement: When A Budget Constraint Meets Moving Trajectories, Ping Zhang, Zhifeng Bao, Yuchen Li, Guoliang Li, Yipeng Zhang, Zhiyong Peng Jul 2020

Towards An Optimal Outdoor Advertising Placement: When A Budget Constraint Meets Moving Trajectories, Ping Zhang, Zhifeng Bao, Yuchen Li, Guoliang Li, Yipeng Zhang, Zhiyong Peng

Research Collection School Of Computing and Information Systems

In this article, we propose and study the problem of trajectory-driven influential billboard placement: given a set of billboards U (each with a location and a cost), a database of trajectories T, and a budget L, we find a set of billboards within the budget to influence the largest number of trajectories. One core challenge is to identify and reduce the overlap of the influence from different billboards to the same trajectories, while keeping the budget constraint into consideration. We show that this problem is NP-hard and present an enumeration based algorithm with (1-1/e) approximation ratio. However, the enumeration would …


Introduction To The R-Package: Usdampr, Elliott James Dennis, Bowen Chen Jun 2020

Introduction To The R-Package: Usdampr, Elliott James Dennis, Bowen Chen

Extension Farm and Ranch Management News

Why the Need for the Package? In the 1990’s, concern over growing packer concentration and a hog industry market shock resulted in discontent among producers and packers. As a result, the United States Congress passed the Livestock Mandatory Reporting Act of 1999 (1999 Act) [Pub. L. 106-78, Title IX] which is required to be reauthorized every five years. See here for a full history of the Livestock Mandatory Reporting Background.

Market reports were publicly issued in the form of .txt files with varying frequency from April 2000 to April 2020. Current and historical data were also housed in a USDA-AMS …


Translating Counting Problems Into Computable Language Expressions, Zach Prescott Jun 2020

Translating Counting Problems Into Computable Language Expressions, Zach Prescott

Theses

The realm of automated problem solving is a relatively new field, even in the context of natural language processing. One area where this is often demonstrated is that of creating a program that can solve word problems. The program must understand the problem, perform some processing, and then convey this information to a user in a way that is accessible and understandable. There has been quite a lot of progress in this area with simpler problems. However, when it comes to understanding problems that involve a level of NLP, the results are not conclusive. In this paper, we would like …


Context-Aware And Scale-Insensitive Temporal Repetition Counting, Huaidong Zhang, Xuemiao Xu, Guoqiang Han, Shengfeng He Jun 2020

Context-Aware And Scale-Insensitive Temporal Repetition Counting, Huaidong Zhang, Xuemiao Xu, Guoqiang Han, Shengfeng He

Research Collection School Of Computing and Information Systems

Temporal repetition counting aims to estimate the number of cycles of a given repetitive action. Existing deep learning methods assume repetitive actions are performed in a fixed time-scale, which is invalid for the complex repetitive actions in real life. In this paper, we tailor a context-aware and scale-insensitive framework, to tackle the challenges in repetition counting caused by the unknown and diverse cycle-lengths. Our approach combines two key insights: (1) Cycle lengths from different actions are unpredictable that require large-scale searching, but, once a coarse cycle length is determined, the variety between repetitions can be overcome by regression. (2) Determining …


Mining User-Generated Content Of Mobile Patient Portal: Dimensions Of User Experience, Mohammad Al-Ramahi, Cherie Noteboom Jun 2020

Mining User-Generated Content Of Mobile Patient Portal: Dimensions Of User Experience, Mohammad Al-Ramahi, Cherie Noteboom

Research & Publications

Patient portals are positioned as a central component of patient engagement through the potential to change the physician-patient relationship and enable chronic disease self-management. The incorporation of patient portals provides the promise to deliver excellent quality, at optimized costs, while improving the health of the population. This study extends the existing literature by extracting dimensions related to the Mobile Patient Portal Use. We use a topic modeling approach to systematically analyze users’ feedback from the actual use of a common mobile patient portal, Epic’s MyChart. Comparing results of Latent Dirichlet Allocation analysis with those of human analysis validated the extracted …


Adaptive Loss-Aware Quantization For Multi-Bit Networks, Zhongnan Qu, Zimu Zhou, Yun Cheng, Lothar Thiele Jun 2020

Adaptive Loss-Aware Quantization For Multi-Bit Networks, Zhongnan Qu, Zimu Zhou, Yun Cheng, Lothar Thiele

Research Collection School Of Computing and Information Systems

We investigate the compression of deep neural networks by quantizing their weights and activations into multiple binary bases, known as multi-bit networks (MBNs), which accelerate the inference and reduce the storage for the deployment on low-resource mobile and embedded platforms. We propose Adaptive Loss-aware Quantization (ALQ), a new MBN quantization pipeline that is able to achieve an average bitwidth below one-bit without notable loss in inference accuracy. Unlike previous MBN quantization solutions that train a quantizer by minimizing the error to reconstruct full precision weights, ALQ directly minimizes the quantizationinduced error on the loss function involving neither gradient approximation nor …


Mnemonics Training: Multi-Class Incremental Learning Without Forgetting, Yaoyao Liu, Yuting Su, An-An Liu, Bernt Schiele, Qianru Sun Jun 2020

Mnemonics Training: Multi-Class Incremental Learning Without Forgetting, Yaoyao Liu, Yuting Su, An-An Liu, Bernt Schiele, Qianru Sun

Research Collection School Of Computing and Information Systems

Multi-Class Incremental Learning (MCIL) aims to learn new concepts by incrementally updating a model trained on previous concepts. However, there is an inherent trade-off to effectively learning new concepts without catastrophic forgetting of previous ones. To alleviate this issue, it has been proposed to keep around a few examples of the previous concepts but the effectiveness of this approach heavily depends on the representativeness of these examples. This paper proposes a novel and automatic framework we call mnemonics, where we parameterize exemplars and make them optimizable in an end-to-end manner. We train the framework through bilevel optimizations, i.e., model-level and …


Knowledge Enhanced Neural Fashion Trend Forecasting, Yunshan Ma, Yujuan Ding, Xun Yang, Lizi Liao, Wai Keung Wong, Tat-Seng Chua Jun 2020

Knowledge Enhanced Neural Fashion Trend Forecasting, Yunshan Ma, Yujuan Ding, Xun Yang, Lizi Liao, Wai Keung Wong, Tat-Seng Chua

Research Collection School Of Computing and Information Systems

Fashion trend forecasting is a crucial task for both academia andindustry. Although some efforts have been devoted to tackling this challenging task, they only studied limited fashion elements with highly seasonal or simple patterns, which could hardly reveal thereal fashion trends. Towards insightful fashion trend forecasting,this work focuses on investigating fine-grained fashion element trends for specific user groups. We first contribute a large-scale fashion trend dataset (FIT) collected from Instagram with extracted time series fashion element records and user information. Furthermore, to effectively model the time series data of fashion elements with rather complex patterns, we propose a Knowledge Enhanced …


Ntire 2020 Challenge On Video Quality Mapping: Methods And Results, D. Fuoli, Zhiwu Huang, M. Danelljan, R. Timofte, H. Wang, L. Jin, D. Su, J. Liu, J. Lee, M. Kudelski, L. Bala, D. Hryboy, M. Mozejko, M. Li, S. Li, B. Pang, C. Lu, Li C., He D., Li F. Jun 2020

Ntire 2020 Challenge On Video Quality Mapping: Methods And Results, D. Fuoli, Zhiwu Huang, M. Danelljan, R. Timofte, H. Wang, L. Jin, D. Su, J. Liu, J. Lee, M. Kudelski, L. Bala, D. Hryboy, M. Mozejko, M. Li, S. Li, B. Pang, C. Lu, Li C., He D., Li F.

Research Collection School Of Computing and Information Systems

This paper reviews the NTIRE 2020 challenge on video quality mapping (VQM), which addresses the issues of quality mapping from source video domain to target video domain. The challenge includes both a supervised track (track 1) and a weakly-supervised track (track 2) for two benchmark datasets. In particular, track 1 offers a new Internet video benchmark, requiring algorithms to learn the map from more compressed videos to less compressed videos in a supervised training manner. In track 2, algorithms are required to learn the quality mapping from one device to another when their quality varies substantially and weaklyaligned video pairs …


Secure Server-Aided Data Sharing Clique With Attestation, Yujue Wang, Hwee Hwa Pang, Robert H. Deng, Yong Ding, Qianhong Wu, Bo Qin, Kefeng Fan Jun 2020

Secure Server-Aided Data Sharing Clique With Attestation, Yujue Wang, Hwee Hwa Pang, Robert H. Deng, Yong Ding, Qianhong Wu, Bo Qin, Kefeng Fan

Research Collection School Of Computing and Information Systems

In this paper, we consider the security issues in data sharing cliques via remote server. We present a public key re-encryption scheme with delegated equality test on ciphertexts (PRE-DET). The scheme allows users to share outsourced data on the server without performing decryption-then-encryption procedures, allows new users to dynamically join the clique, allows clique users to attest the message underlying a ciphertext, and enables the server to partition outsourced user data without any further help of users after being delegated. We introduce the PRE-DET framework, propose a concrete construction and formally prove its security against five types of adversaries regarding …


Goods Consumed During Transit In Split Delivery Vehicle Routing Problems: Modeling And Solution, Wenzhe Yang, Di Wang, Wei Pang, Ah-Hwee Tan, You Zhou Jun 2020

Goods Consumed During Transit In Split Delivery Vehicle Routing Problems: Modeling And Solution, Wenzhe Yang, Di Wang, Wei Pang, Ah-Hwee Tan, You Zhou

Research Collection School Of Computing and Information Systems

This article presents the modeling and solution of an extended type of split delivery vehicle routing problem (SDVRP). In SDVRP, the demands of customers need to be met by efficiently routing a given number of capacitated vehicles, wherein each customer may be served multiple times by more than one vehicle. Furthermore, in many real-world scenarios, consumption of vehicles en route is the same as the goods being delivered to customers, such as food, water and fuel in rescue or replenishment missions in harsh environments. Moreover, the consumption may also be in virtual forms, such as time spent in constrained tasks. …


Maximum A Posteriori Estimation For Information Source Detection, Biao Chang, Enhong Chen, Feida Zhu, Qi Liu, Tong Xu Jun 2020

Maximum A Posteriori Estimation For Information Source Detection, Biao Chang, Enhong Chen, Feida Zhu, Qi Liu, Tong Xu

Research Collection School Of Computing and Information Systems

Information source detection is to identify nodes initiating the diffusion process in a network, which has a wide range of applications including epidemic outbreak prevention, Internet virus source identification, and rumor source tracing in social networks. Although it has attracted ever-increasing attention from research community in recent years, existing solutions still suffer from high time complexity and inadequate effectiveness, due to high dynamics of information diffusion and observing just a snapshot of the whole process. To this end, we present a comprehensive study for single information source detection in weighted graphs. Specifically, we first propose a maximum a posteriori (MAP) …


Visual Commonsense Representation Learning Via Causal Inference, Tan Wang, Jianqiang Huang, Hanwang Zhang, Qianru Sun Jun 2020

Visual Commonsense Representation Learning Via Causal Inference, Tan Wang, Jianqiang Huang, Hanwang Zhang, Qianru Sun

Research Collection School Of Computing and Information Systems

We present a novel unsupervised feature representation learning method, Visual Commonsense Region-based Convolutional Neural Network (VC R-CNN), to serve as an improved visual region encoder for high-level tasks such as captioning and VQA. Given a set of detected object regions in an image (e.g., using Faster R-CNN), like any other unsupervised feature learning methods (e.g., word2vec), the proxy training objective of VC R-CNN is to predict the con-textual objects of a region. However, they are fundamentally different: the prediction of VC R-CNN is by using causal intervention: P(Y|do(X)), while others are by using the conventional likelihood: P(Y|X). We extensively apply …


Gpu-Accelerated Subgraph Enumeration On Partitioned Graphs, Wentian Guo, Yuchen Li, Mo Sha, Bingsheng He, Xiaokui Xiao, Kian-Lee Tan Jun 2020

Gpu-Accelerated Subgraph Enumeration On Partitioned Graphs, Wentian Guo, Yuchen Li, Mo Sha, Bingsheng He, Xiaokui Xiao, Kian-Lee Tan

Research Collection School Of Computing and Information Systems

Subgraph enumeration is important for many applications such as network motif discovery and community detection. Recent works utilize graphics processing units (GPUs) to parallelize subgraph enumeration, but they can only handle graphs that fit into the GPU memory. In this paper, we propose a new approach for GPU-accelerated subgraph enumeration that can efficiently scale to large graphs beyond the GPU memory. Our approach divides the graph into partitions, each of which fits into the GPU memory. The GPU processes one partition at a time and searches the matched subgraphs of a given pattern (i.e., instances) within the partition as in …


Hyperbolic Visual Embedding Learning For Zero-Shot Recognition, Shaoteng Liu, Jingjing Chen, Liangming Pan, Chong-Wah Ngo, Tat-Seng Chua, Yu-Gang Jiang Jun 2020

Hyperbolic Visual Embedding Learning For Zero-Shot Recognition, Shaoteng Liu, Jingjing Chen, Liangming Pan, Chong-Wah Ngo, Tat-Seng Chua, Yu-Gang Jiang

Research Collection School Of Computing and Information Systems

This paper proposes a Hyperbolic Visual Embedding Learning Network for zero-shot recognition. The network learns image embeddings in hyperbolic space, which is capable of preserving the hierarchical structure of semantic classes in low dimensions. Comparing with existing zeroshot learning approaches, the network is more robust because the embedding feature in hyperbolic space better represents class hierarchy and thereby avoid misleading resulted from unrelated siblings. Our network outperforms exiting baselines under hierarchical evaluation with an extremely challenging setting, i.e., learning only from 1,000 categories to recognize 20,841 unseen categories. While under flat evaluation, it has competitive performance as state-of-the-art methods but …


An Investigation Into The Optimal Use Of Frequency-Based Weights To Improve The Performance Of Entity Resolution, Bingyi Zhong May 2020

An Investigation Into The Optimal Use Of Frequency-Based Weights To Improve The Performance Of Entity Resolution, Bingyi Zhong

Theses and Dissertations

Using a weight-based match score has been studied as a way to improve the accuracy of record linking (entity resolution) since Fellegi and Sunter first described the idea of probabilistic agreement and disagreement weights in their seminal work “A Theory of Record Linking.” However, the original work only described weight associated with an entity attribute. Later researchers such as Herzog et al suggested the weighting scheme could be extended to apply to frequently-occurring attribute values (frequency-based weights) instead of just the attribute as in the Fellegi-Sunter scheme. However, there has been little definitive research as to how frequency-based weights should …


Improved Chinese Language Processing For An Open Source Search Engine, Xianghong Sun May 2020

Improved Chinese Language Processing For An Open Source Search Engine, Xianghong Sun

Master's Projects

Natural Language Processing (NLP) is the process of computers analyzing on human languages. There are also many areas in NLP. Some of the areas include speech recognition, natural language understanding, and natural language generation.

Information retrieval and natural language processing for Asians languages has its own unique set of challenges not present for Indo-European languages. Some of these are text segmentation, named entity recognition in unsegmented text, and part of speech tagging. In this report, we describe our implementation of and experiments with improving the Chinese language processing sub-component of an open source search engine, Yioop. In particular, we rewrote …


An Inventory Of Existing Neuroprivacy Controls, Dustin Steinhagen, Houssain Kettani May 2020

An Inventory Of Existing Neuroprivacy Controls, Dustin Steinhagen, Houssain Kettani

Research & Publications

Brain-Computer Interfaces (BCIs) facilitate communication between brains and computers. As these devices become increasingly popular outside of the medical context, research interest in brain privacy risks and countermeasures has bloomed. Several neuroprivacy threats have been identified in the literature, including brain malware, personal data being contained in collected brainwaves and the inadequacy of legal regimes with regards to neural data protection. Dozens of controls have been proposed or implemented for protecting neuroprivacy, although it has not been immediately apparent what the landscape of neuroprivacy controls consists of. This paper inventories the implemented and proposed neuroprivacy risk mitigation techniques from open …


Benchmarking Mongodb Multi-Document Transactions In A Sharded Cluster, Tushar Panpaliya May 2020

Benchmarking Mongodb Multi-Document Transactions In A Sharded Cluster, Tushar Panpaliya

Master's Projects

Relational databases like Oracle, MySQL, and Microsoft SQL Server offer trans- action processing as an integral part of their design. These databases have been a primary choice among developers for business-critical workloads that need the highest form of consistency. On the other hand, the distributed nature of NoSQL databases makes them suitable for scenarios needing scalability, faster data access, and flexible schema design. Recent developments in the NoSQL database community show that NoSQL databases have started to incorporate transactions in their drivers to let users work on business-critical scenarios without compromising the power of distributed NoSQL features [1].

MongoDB is …


Server Score, Zachary Buresh May 2020

Server Score, Zachary Buresh

Student Academic Conference

This presentation is in regards to the Android mobile application that I developed in the Kotlin programming language named "Server Score". The app helps waiters/waitresses calculate, track, and predict performance related statistics on the job.


Predictive Modeling Of Asynchronous Event Sequence Data, Jin Shang May 2020

Predictive Modeling Of Asynchronous Event Sequence Data, Jin Shang

LSU Doctoral Dissertations

Large volumes of temporal event data, such as online check-ins and electronic records of hospital admissions, are becoming increasingly available in a wide variety of applications including healthcare analytics, smart cities, and social network analysis. Those temporal events are often asynchronous, interdependent, and exhibiting self-exciting properties. For example, in the patient's diagnosis events, the elevated risk exists for a patient that has been recently at risk. Machine learning that leverages event sequence data can improve the prediction accuracy of future events and provide valuable services. For example, in e-commerce and network traffic diagnosis, the analysis of user activities can be …


Building A Library Search Infrastructure With Elasticsearch, Kim Pham, Fernando Reyes, Jeff Rynhart May 2020

Building A Library Search Infrastructure With Elasticsearch, Kim Pham, Fernando Reyes, Jeff Rynhart

University Libraries: Faculty Scholarship

This article discusses our implementation of an Elastic cluster to address our search, search administration and indexing needs, how it integrates in our technology infrastructure, and finally takes a close look at the way that we built a reusable, dynamic search engine that powers our digital repository search. We cover the lessons learned with our early implementations and how to address them to lay the groundwork for a scalable, networked search environment that can also be applied to alternative search engines such as Solr.