Process Models Discovery And Traces Classification: A Fuzzy-Bpmn Mining Approach.,
2017
University of East London
Process Models Discovery And Traces Classification: A Fuzzy-Bpmn Mining Approach., Kingsley Okoye Dr, Usman Naeem Dr, Syed Islam Dr, Abdel-Rahman H. Tawil Dr, Elyes Lamine Dr
Journal of International Technology and Information Management
The discovery of useful or worthwhile process models must be performed with due regards to the transformation that needs to be achieved. The blend of the data representations (i.e data mining) and process modelling methods, often allied to the field of Process Mining (PM), has proven to be effective in the process analysis of the event logs readily available in many organisations information systems. Moreover, the Process Discovery has been lately seen as the most important and most visible intellectual challenge related to the process mining. The method involves automatic construction of process models from event logs about any domain …
A Novel Density Peak Clustering Algorithm Based On Squared Residual Error,
2017
Singapore Management University
A Novel Density Peak Clustering Algorithm Based On Squared Residual Error, Milan Parmar, Di Wang, Ah-Hwee Tan, Chunyan Miao, Jianhua Jiang, You Zhou
Research Collection School Of Computing and Information Systems
The density peak clustering (DPC) algorithm is designed to quickly identify intricate-shaped clusters with high dimensionality by finding high-density peaks in a non-iterative manner and using only one threshold parameter. However, DPC has certain limitations in processing low-density data points because it only takes the global data density distribution into account. As such, DPC may confine in forming low-density data clusters, or in other words, DPC may fail in detecting anomalies and borderline points. In this paper, we analyze the limitations of DPC and propose a novel density peak clustering algorithm to better handle low-density clustering tasks. Specifically, our algorithm …
Btci: A New Framework For Identifying Congestion Cascades Using Bus Trajectory Data,
2017
Singapore Management University
Btci: A New Framework For Identifying Congestion Cascades Using Bus Trajectory Data, Meng-Fen Chiang, Ee Peng Lim, Wang-Chien Lee, Agus Trisnajaya Kwee
Research Collection School Of Computing and Information Systems
The knowledge of traffic health status is essential to the general public and urban traffic management. To identify congestion cascades, an important phenomenon of traffic health, we propose a Bus Trajectory based Congestion Identification (BTCI) framework that explores the anomalous traffic health status and structure properties of congestion cascades using bus trajectory data. BTCI consists of two main steps, congested segment extraction and congestion cascades identification. The former constructs path speed models from historical vehicle transitions and design a non-parametric Kernel Density Estimation (KDE) function to derive a measure of congestion score. The latter aggregates congested segments (i.e., those with …
Ethics And Bias In Machine Learning: A Technical Study Of What Makes Us “Good”,
2017
CUNY John Jay College of Criminal Justice
Ethics And Bias In Machine Learning: A Technical Study Of What Makes Us “Good”, Ashley Nicole Shadowen
Student Theses
The topic of machine ethics is growing in recognition and energy, but bias in machine learning algorithms outpaces it to date. Bias is a complicated term with good and bad connotations in the field of algorithmic prediction making. Especially in circumstances with legal and ethical consequences, we must study the results of these machines to ensure fairness. This paper attempts to address ethics at the algorithmic level of autonomous machines. There is no one solution to solving machine bias, it depends on the context of the given system and the most reasonable way to avoid biased decisions while maintaining the …
Web Application For Graduate Course Recommendation System,
2017
California State University, San Bernardino
Web Application For Graduate Course Recommendation System, Sayali Dhumal
Electronic Theses, Projects, and Dissertations
The main aim of the course advising system is to build a course recommendation path for students to help them plan courses to successfully graduate on time. The recommendation path displays the list of courses a student can take in each quarter from the first quarter after admission until the graduation quarter. The courses are filtered as per the student’s interest obtained from a questionnaire asked to the student.
The business logic involves building the recommendation algorithm. Also, the application is functionality-tested end-to-end by using nightwatch.js which is built on top of node.js. Test cases are written for every module …
Web Application For Graduate Course Advising System,
2017
California State University, San Bernardino
Web Application For Graduate Course Advising System, Sanjay Karrolla
Electronic Theses, Projects, and Dissertations
The main aim of the course recommendation system is to build a course recommendation path for students to help them plan courses to successfully graduate on time. The Model-View-Controller (MVC) architecture is used to isolate the user interface (UI) design from the business logic. The front-end of the application develops the UI using AngularJS. The front-end design is done by gathering the functionality system requirements -- input controls, navigational components, informational components and containers and usability testing. The back-end of the application involves setting up the database and server-side routing. Server-side routing is done using Express JS.
Scalable Online Kernel Learning,
2017
Singapore Management University
Scalable Online Kernel Learning, Jing Lu
Dissertations and Theses Collection (Open Access)
One critical deficiency of traditional online kernel learning methods is their increasing and unbounded number of support vectors (SV’s), making them inefficient and non-scalable for large-scale applications. Recent studies on budget online learning have attempted to overcome this shortcoming by bounding the number of SV’s. Despite being extensively studied, budget algorithms usually suffer from several drawbacks.
First of all, although existing algorithms attempt to bound the number of SV’s at each iteration, most of them fail to bound the number of SV’s for the final averaged classifier, which is commonly used for online-to-batch conversion. To solve this problem, we propose …
Second-Order Online Active Learning And Its Applications,
2017
Institute of High Performance Computing
Second-Order Online Active Learning And Its Applications, Shuji Hao, Jing Lu, Peilin Zhao, Chi Zhang, Steven C. H. Hoi, Chunyan Miao
Research Collection School Of Computing and Information Systems
The goal of online active learning is to learn predictive models from a sequence of unlabeled data given limited label querybudget. Unlike conventional online learning tasks, online active learning is considerably more challenging because of two reasons.Firstly, it is difficult to design an effective query strategy to decide when is appropriate to query the label of an incoming instance givenlimited query budget. Secondly, it is also challenging to decide how to update the predictive models effectively whenever the true labelof an instance is queried. Most existing approaches for online active learning are often based on a family of first-order online …
Detecting Semantic Uncertainty By Learning Hedge Cues In Sentences Using An Hmm,
2017
Singapore Management University
Detecting Semantic Uncertainty By Learning Hedge Cues In Sentences Using An Hmm, Xiujun Li, Wei Gao, Jude Shavlik
Research Collection School Of Computing and Information Systems
Detecting speculative assertions is essential to distinguish semantically uncertain information from the factual ones in text. This is critical to the trustworthiness of many intelligent systems that are based on information retrieval and natural language processing techniques, such as question answering or information extraction. We empirically explore three fundamental issues of uncertainty detection: (1) the predictive ability of different learning methods on this task; (2) whether using unlabeled data can lead to a more accurate model; and (3) whether closed-domain training or crossdomain training is better. For these purposes, we adopt two statistical learning approaches to this problem: the commonly …
Large Scale Kernel Methods For Online Auc Maximization,
2017
University of Chicago
Large Scale Kernel Methods For Online Auc Maximization, Yi Ding, Chenghao Liu, Peilin Zhao, Steven C. H. Hoi
Research Collection School Of Computing and Information Systems
Learning to optimize AUC performance for classifying label imbalanced data in online scenarios has been extensively studied in recent years. Most of the existing work has attempted to address the problem directly in the original feature space, which may not suitable for non-linearly separable datasets. To solve this issue, some kernel-based learning methods are proposed for non-linearly separable datasets. However, such kernel approaches have been shown to be inefficient and failed to scale well on large scale datasets in practice. Taking this cue, in this work, we explore the use of scalable kernel-based learning techniques as surrogates to existing approaches: …
File-Level Defect Prediction: Unsupervised Vs. Supervised Models,
2017
Singapore Management University
File-Level Defect Prediction: Unsupervised Vs. Supervised Models, Meng Yan, Yicheng Fang, David Lo, Xin Xia, Xiaohong Zhang
Research Collection School Of Computing and Information Systems
Background: Software defect models can help software quality assurance teams to allocate testing or code review resources. A variety of techniques have been used to build defect prediction models, including supervised and unsupervised methods. Recently, Yang et al. [1] surprisingly find that unsupervised models can perform statistically significantly better than supervised models in effort-aware change-level defect prediction. However, little is known about relative performance of unsupervised and supervised models for effort-aware file-level defect prediction. Goal: Inspired by their work, we aim to investigate whether a similar finding holds in effort-aware file-level defect prediction. Method: We replicate Yang et al.'s study …
Power-Efficient And Highly Scalable Parallel Graph Sampling Using Fpgas,
2017
WMU
Power-Efficient And Highly Scalable Parallel Graph Sampling Using Fpgas, Usman Tariq, Umer Cheema, Fahad Saeed
Parallel Computing and Data Science Lab Technical Reports
Energy efficiency is a crucial problem in data centers where big data is generally represented by directed or undirected graphs. Analysis of this big data graph is challenging due to volume and velocity of the data as well as irregular memory access patterns. Graph sampling is one of the most effective ways to reduce the size of graph while maintaining crucial characteristics. In this paper we present design and implementation of an FPGA based graph sampling method which is both time- and energy-efficient. This is in contrast to existing parallel approaches which include memory-distributed clusters, multicore and GPUs. Our …
Benchmarking Single-Image Reflection Removal Algorithms,
2017
Singapore Management University
Benchmarking Single-Image Reflection Removal Algorithms, Renjie Wan, Boxin Shi, Ling-Yu Duan, Ah-Hwee Tan, Alex C. Kot
Research Collection School Of Computing and Information Systems
Removing undesired reflections from a photo taken in front of a glass is of great importance for enhancing the efficiency of visual computing systems. Various approaches have been proposed and shown to be visually plausible on small datasets collected by their authors. A quantitative comparison of existing approaches using the same dataset has never been conducted due to the lack of suitable benchmark data with ground truth. This paper presents the first captured Single-image Reflection Removal dataset `SIR 2 ' with 40 controlled and 100 wild scenes, ground truth of background and reflection. For each controlled scene, we further provide …
Sol: A Library For Scalable Online Learning Algorithms,
2017
University of Science and Technology of China
Sol: A Library For Scalable Online Learning Algorithms, Yue Wu, Steven C. H. Hoi, Chenghao Liu, Jing Lu, Doyen Sahoo, Nenghai Yu
Research Collection School Of Computing and Information Systems
SOL is an open-source library for scalable online learning with high-dimensional data. The library provides a family of regular and sparse online learning algorithms for large-scale classification tasks with high efficiency, scalability, portability, and extensibility. We provide easy-to-use command-line tools, python wrappers and library calls for users and developers, and comprehensive documents for both beginners and advanced users. SOL is not only a machine learning toolbox, but also a comprehensive experimental platform for online learning research. Experiments demonstrate that SOL is highly efficient and scalable for large-scale learning with high-dimensional data.
High-Order Relaxed Multirate Infinitesimal Step Methods For Multiphysics Applications,
2017
Southern Methodist University
High-Order Relaxed Multirate Infinitesimal Step Methods For Multiphysics Applications, Jean Sexton
Mathematics Theses and Dissertations
In this work, we consider numerical methods for integrating multirate ordinary differential equations. We are interested in the development of new multirate methods with good stability properties and improved efficiency over existing methods. We discuss the development of multirate methods, particularly focusing on those that are based on Runge-Kutta theory. We introduce the theory of Generalized Additive Runge-Kutta methods proposed by Sandu and Günther. We also introduce the theory of Recursive Flux Splitting Multirate Methods with Sub-cycling described by Schlegel, as well as the Multirate Infinitesimal Step methods this work is based on. We propose a generic structure called Flexible …
Methods For Real-Time Prediction Of The Mode Of Travel Using Smartphone-Based Gps And Accelerometer Data,
2017
University of Washington - Seattle Campus
Methods For Real-Time Prediction Of The Mode Of Travel Using Smartphone-Based Gps And Accelerometer Data, Bryan D. Martin, Vittorio Addona, Julian Wolfson, Gediminas Adomavicius, Yingling Fan
Faculty Publications
We propose and compare combinations of several methods for classifying transportation activity data from smartphone GPS and accelerometer sensors. We have two main objectives. First, we aim to classify our data as accurately as possible. Second, we aim to reduce the dimensionality of the data as much as possible in order to reduce the computational burden of the classification. We combine dimension reduction and classification algorithms and compare them with a metric that balances accuracy and dimensionality. In doing so, we develop a classification algorithm that accurately classifies five different modes of transportation (i.e., walking, biking, car, bus and rail) …
Micro-Review Synthesis For Multi-Entity Summarization,
2017
Singapore Management University
Micro-Review Synthesis For Multi-Entity Summarization, Thanh-Son Nguyen, Hady W. Lauw, Panayiotis Tsaparas
Research Collection School Of Computing and Information Systems
Location-based social networks (LBSNs), exemplified by Foursquare, are fast gaining popularity. One important feature of LBSNs is micro-review. Upon check-in at a particular venue, a user may leave a short review (up to 200 characters long), also known as a tip. These tips are an important source of information for others to know more about various aspects of an entity (e.g., restaurant), such as food, waiting time, or service. However, a user is often interested not in one particular entity, but rather in several entities collectively, for instance within a neighborhood or a category. In this paper, we address the …
Approximation Algorithms For Effective Team Formation,
2017
CUNY Graduate Center
Approximation Algorithms For Effective Team Formation, George Rabanca
Dissertations, Theses, and Capstone Projects
This dissertation investigates the problem of creating multiple disjoint teams of maximum efficacy from a fixed set of workers. We identify three parameters which directly correlate to the team effectiveness — team expertise, team cohesion and team size — and propose efficient algorithms for optimizing each in various settings. We show that under standard assumptions the problems we explore are not optimally solvable in polynomial time, and thus we focus on developing efficient algorithms with guaranteed worst case approximation bounds. First, we investigate maximizing team expertise in a setting where each worker has different expertise for each job and each …
A Combinatorial Framework For Multiple Rna Interaction Prediction,
2017
CUNY Graduate Center
A Combinatorial Framework For Multiple Rna Interaction Prediction, Syed Ali Ahmed
Dissertations, Theses, and Capstone Projects
The interaction of two RNA molecules involves a complex interplay between folding and binding that warranted recent developments in RNA-RNA interaction algorithms. However, biological mechanisms in which more than two RNAs take part in an interaction also exist.
A typical algorithmic approach to such problems is to find the minimum energy structure. Often the computationally optimal solution does not represent the biologically correct structure of the interaction. In addition, different biological structures may be observed, depending on several factors. Furthermore, scoring techniques often miss critical details about dependencies within different parts of the structure, which typically leads to lower scores …
Morphogenesis And Growth Driven By Selection Of Dynamical Properties,
2017
CUNY Graduate Center
Morphogenesis And Growth Driven By Selection Of Dynamical Properties, Yuri Cantor
Dissertations, Theses, and Capstone Projects
Organisms are understood to be complex adaptive systems that evolved to thrive in hostile environments. Though widely studied, the phenomena of organism development and growth, and their relationship to organism dynamics is not well understood. Indeed, the large number of components, their interconnectivity, and complex system interactions all obscure our ability to see, describe, and understand the functioning of biological organisms.
Here we take a synthetic and computational approach to the problem, abstracting the organism as a cellular automaton. Such systems are discrete digital models of real-world environments, making them more accessible and easier to study then their physical world …
