Representing And Learning Preferences Over Combinatorial Domains,
2021
University of Kentucky
Representing And Learning Preferences Over Combinatorial Domains, Michael Huelsman
Theses and Dissertations--Computer Science
Agents make decisions based on their preferences. Thus, to predict their decisions one has to learn the agent's preferences. A key step in the learning process is selecting a model to represent those preferences. We studied this problem by borrowing techniques from the algorithm selection problem to analyze preference example sets and select the most appropriate preference representation for learning. We approached this problem in multiple steps.
First, we determined which representations to consider. For this problem we developed the notion of preference representation language subsumption, which compares representations based on their expressive power. Subsumption creates a hierarchy of preference …
Understanding The Research And Applications Of Quantum Computing,
2021
The University of Akron
Understanding The Research And Applications Of Quantum Computing, Joshua Foss
Williams Honors College, Honors Research Projects
In-Depth research of current quantum computing understanding and practices. Presentation of possible new and creative applications of quantum computing.
Learning From Multi-Class Imbalanced Big Data With Apache Spark,
2021
Virginia Commonwealth University
Learning From Multi-Class Imbalanced Big Data With Apache Spark, William C. Sleeman Iv
Theses and Dissertations
With data becoming a new form of currency, its analysis has become a top priority in both academia and industry, furthering advancements in high-performance computing and machine learning. However, these large, real-world datasets come with additional complications such as noise and class overlap. Problems are magnified when with multi-class data is presented, especially since many of the popular algorithms were originally designed for binary data. Another challenge arises when the number of examples are not evenly distributed across all classes in a dataset. This often causes classifiers to favor the majority class over the minority classes, leading to undesirable results …
Scaling Up Exact Neural Network Compression By Relu Stability,
2021
The University Of Utah
Scaling Up Exact Neural Network Compression By Relu Stability, Thiago Serra, Xin Yu, Abhinav Kumar, Srikumar Ramalingam
Faculty Conference Papers and Presentations
We can compress a rectifier network while exactly preserving its underlying functionality with respect to a given input domain if some of its neurons are stable. However, current approaches to determine the stability of neurons with Rectified Linear Unit (ReLU) activations require solving or finding a good approximation to multiple discrete optimization problems. In this work, we introduce an algorithm based on solving a single optimization problem to identify all stable neurons. Our approach is on median 183 times faster than the state-of-art method on CIFAR-10, which allows us to explore exact compression on deeper (5 x 100) and wider …
Optimal Construction Of A Layer-Ordered Heap And Its Applications,
2021
University of Montana
Optimal Construction Of A Layer-Ordered Heap And Its Applications, Jake Pennington
Graduate Student Theses, Dissertations, & Professional Papers
The layer-ordered heap (LOH) is a simple data structure used in algorithms that perform optimal top-$k$ on $X+Y$, algorithms with the best known runtime for top-$k$ on $X_1+X_2+\cdots+X_m$, and the fastest method in practice for computing the most abundant isotopologue peaks in a chemical compound. In the analysis of these algorithms, the rank, $\alpha$, has been treated as a constant and $n$, the size of the array, has been treated as the sole parameter. Here, we explore the algorithmic complexity of LOH construction with $\alpha$ as a parameter, introduce a few algorithms for constructing LOHs, analyze their complexity in both …
Ssentiaa: A Self-Supervised Sentiment Analyzer For Classification From Unlabeled Data,
2021
Old Dominion University
Ssentiaa: A Self-Supervised Sentiment Analyzer For Classification From Unlabeled Data, Salim Sazzed, Sampath Jayarathna
Computer Science Faculty Publications
In recent years, supervised machine learning (ML) methods have realized remarkable performance gains for sentiment classification utilizing labeled data. However, labeled data are usually expensive to obtain, thus, not always achievable. When annotated data are unavailable, the unsupervised tools are exercised, which still lag behind the performance of supervised ML methods by a large margin. Therefore, in this work, we focus on improving the performance of sentiment classification from unlabeled data. We present a self-supervised hybrid methodology SSentiA (Self-supervised Sentiment Analyzer) that couples an ML classifier with a lexicon-based method for sentiment classification from unlabeled data. We first introduce LRSentiA …
Systematizing Confidence In Open Research And Evidence (Score),
2021
Old Dominion University
Systematizing Confidence In Open Research And Evidence (Score), Nazanin Alipourfard, Beatrix Arendt, Daniel M. Benjamin, Noam Benkler, Michael Bishop, Mark Burstein, Martin Bush, James Caverlee, Yiling Chen, Chae Clark, Anna Dreber Almenberg, Timothy M. Errington, Fiona Fidler, Nicholas Fox, Aaron Frank, Hannah Fraser, Scott Friedman, Ben Gelman, James Gentile, Jian Wu, Et Al., Score Collaboration
Computer Science Faculty Publications
Assessing the credibility of research claims is a central, continuous, and laborious part of the scientific process. Credibility assessment strategies range from expert judgment to aggregating existing evidence to systematic replication efforts. Such assessments can require substantial time and effort. Research progress could be accelerated if there were rapid, scalable, accurate credibility indicators to guide attention and resource allocation for further assessment. The SCORE program is creating and validating algorithms to provide confidence scores for research claims at scale. To investigate the viability of scalable tools, teams are creating: a database of claims from papers in the social and behavioral …
Efficient Algorithms For Identifying Loop Formation And Computing Θ Value For Solving Minimum Cost Flow Network Problems,
2021
Old Dominion University
Efficient Algorithms For Identifying Loop Formation And Computing Θ Value For Solving Minimum Cost Flow Network Problems, Timothy Michael Chávez, Duc Thai Nguyen
Computational Modeling & Simulation Engineering Faculty Publications
While the minimum cost flow (MCF) problems have been well documented in many publications, due to its broad applications, little or no effort have been devoted to explaining the algorithms for identifying loop formation and computing the value needed to solve MCF network problems. This paper proposes efficient algorithms, and MATLAB computer implementation, for solving MCF problems. Several academic and real-life network problems have been solved to validate the proposed algorithms; the numerical results obtained by the developed MCF code have been compared and matched with the built-in MATLAB function Linprog() (Simplex algorithm) for further validation.
Gene Selection For Cancer Classification: A New Hybrid Filter-C5.0 Approach For Breast Cancer Risk Prediction,
2021
ENSAM-Casablanca Université Hassan II
Gene Selection For Cancer Classification: A New Hybrid Filter-C5.0 Approach For Breast Cancer Risk Prediction, Mohammed Hamim, Ismail El Moudden, Hicham Moutachaouik, Mustapha Hain
Department of Medicine Faculty Publications
Despite the significant progress made in data mining technologies in recent years, breast cancer risk prediction and diagnosis at an early stage using DNA microarray technology still a real challenging task. This challenge comes especially from the high-dimensionality in gene expression data, i.e., an enormous number of genes versus a few tens of subjects (samples). To overcome this problem of data imbalance, a gene selection phase becomes a crucial step for gene expression data analysis. This study proposes a new Decision Tree model-based attributes (genes) selection strategy, which incorporates two stages: fisher-score-based filter technique and the gene selection ability of …
Identifying Football Conflict Using Soft-Set Theory In Indonesia Super League,
2021
Universiti Malaya
Identifying Football Conflict Using Soft-Set Theory In Indonesia Super League, Kukuh Wahyudin Pratama
Student Works (2020-2029)
There are several mathematical formal models that handle conflict situations and the most popular one is a rough set theory. With the ability to handle vagueness from the conflict data set, rough set theory has been successfully used in many research. This research used an alternative approach as a method to handle conflict situation in Indonesia Super League. This method was implemented on the respondents or agents who were involved with football club management, match inspector, organizing committee, referees, supporters and players. The novelty of the proposed approach is discussed in rough set theory that include decision rules. It is …
K-Nearest Neighbors Density-Based Clustering,
2021
Virginia Commonwealth University
K-Nearest Neighbors Density-Based Clustering, Avory C. Bryant
Theses and Dissertations
Traditional density-based clustering approaches rely on a distance-based parameter to define data connectivity and density. However, an appropriate value of this parameter can be difficult to determine as it is highly dependent on the underlying distribution of the data. In particular, distribution parameters affect the scale of inter-group distances (e.g., variance); this dependence leads to a well-known inability to simultaneously detect clusters at varying levels of density. In this work, connectivity and density are defined according to the rank-order induced by the distance metric (i.e., invariant to the expected scale of the distances). Connectivity by k-nearest neighbors and density by …
Deep Unsupervised Anomaly Detection,
2021
National University of Singapore
Deep Unsupervised Anomaly Detection, Tangqing Li, Zheng Wang, Siying Liu, Wen-Yan Lin
Research Collection School Of Computing and Information Systems
This paper proposes a novel method to detect anomalies in large datasets under a fully unsupervised setting. The key idea behind our algorithm is to learn the representation underlying normal data. To this end, we leverage the latest clustering technique suitable for handling high dimensional data. This hypothesis provides a reliable starting point for normal data selection. We train an autoencoder from the normal data subset, and iterate between hypothesizing normal candidate subset based on clustering and representation learning. The reconstruction error from the learned autoencoder serves as a scoring function to assess the normality of the data. Experimental results …
Infinite-Duration All-Pay Bidding Games,
2021
Singapore Management University
Infinite-Duration All-Pay Bidding Games, Guy Avni, Ismäel Jecker, Dorde Zikelic
Research Collection School Of Computing and Information Systems
In a two-player zero-sum graph game the players move a token throughout a graph to produce an infinite path, which determines the winner or payoff of the game. Traditionally, the players alternate turns in moving the token. In bidding games, however, the players have budgets, and in each turn, we hold an "auction" (bidding) to determine which player moves the token: both players simultaneously submit bids and the higher bidder moves the token. The bidding mechanisms differ in their payment schemes. Bidding games were largely studied with variants of first-price bidding in which only the higher bidder pays his bid. …
Explainable Feature- And Decision-Level Fusion,
2021
Michigan Technological University
Explainable Feature- And Decision-Level Fusion, Siva Krishna Kakula
Dissertations, Master's Theses and Master's Reports
Information fusion is the process of aggregating knowledge from multiple data sources to produce more consistent, accurate, and useful information than any one individual source can provide. In general, there are three primary sources of data/information: humans, algorithms, and sensors. Typically, objective data---e.g., measurements---arise from sensors. Using these data sources, applications such as computer vision and remote sensing have long been applying fusion at different "levels" (signal, feature, decision, etc.). Furthermore, the daily advancement in engineering technologies like smart cars, which operate in complex and dynamic environments using multiple sensors, are raising both the demand for and complexity of fusion. …
Learning Accurate And Robust Deep Visual Models,
2021
University of Central Florida
Learning Accurate And Robust Deep Visual Models, Yandong Li
Electronic Theses and Dissertations, 2020-2023
Over the last decade, we have witnessed the renaissance of deep neural networks (DNNs) and their successful applications in computer vision. There is still a long way to build intelligent and reliable machine vision systems, but DNNs provide a promising direction. The goal of this thesis is to present a few small steps along this road. We mainly focus on two questions: How to design label-efficient learning algorithms for computer vision tasks? How to improve the robustness of DNN based visual models? Concerning label-efficiency, we investigate a reinforced sequential model for video summarization, a background hallucination strategy for high-resolution image …
Modified Firearm Discharge Residue Analysis Utilizing Advanced Analytical Techniques, Complexing Agents, And Quantum Chemical Calculations,
2021
West Virginia University
Modified Firearm Discharge Residue Analysis Utilizing Advanced Analytical Techniques, Complexing Agents, And Quantum Chemical Calculations, William J. Feeney
Graduate Theses, Dissertations, and Problem Reports (ETD)
The use of gunshot residue (GSR) or firearm discharge residue (FDR) evidence faces some challenges because of instrumental and analytical limitations and the difficulties in evaluating and communicating evidentiary value. For instance, the categorization of GSR based only on elemental analysis of single, spherical particles is becoming insufficient because newer ammunition formulations produce residues with varying particle morphology and composition. Also, one common criticism about GSR practitioners is that their reports focus on the presence or absence of GSR in an item without providing an assessment of the weight of the evidence. Such reports leave the end-used with unanswered questions, …
Partial Adversarial Behavior Deception In Security Games,
2021
Singapore Management University
Partial Adversarial Behavior Deception In Security Games, Thanh H. Nguyen, Arunesh Sinha, He He
Research Collection School Of Computing and Information Systems
Learning attacker behavior is an important research topic in security games as security agencies are often uncertain about attackers’ decision making. Previous work has focused on developing various behavioral models of attackers based on historical attack data. However, a clever attacker can manipulate its attacks to fail such attack-driven learning, leading to ineffective defense strategies. We study attacker behavior deception with three main contributions. First, we propose a new model, named partial behavior deception model, in which there is a deceptive attacker (among multiple attackers) who controls a portion of attacks. Our model captures real-world security scenarios such as wildlife …
Analysis Of Github Pull Requests,
2020
Southern Methodist University
Analysis Of Github Pull Requests, Canon Ellis
Computer Science and Engineering Theses and Dissertations
The popularity of the software repository site GitHub has created a rise in the Pull Based Development Models' use. An essential portion of pull-based development is the creation of Pull Requests. Pull Requests often have to be reviewed by an individual to be approved and accepted into the Master branch of a software repository. The reviewing process can often be time-consuming and introduce a relatively high level of lost development time. This paper examines thousands of pull requests to understand the most valuable metadata of pull requests. We then introduce metrics in comparing the metadata of pull requests to understand …
Extending Import Detection Algorithms For Concept Import From Two To Three Biomedical Terminologies,
2020
New Jersey Institute of Technology
Extending Import Detection Algorithms For Concept Import From Two To Three Biomedical Terminologies, Vipina K. Keloth, James Geller, Yan Chen, Julia Xu
Publications and Research
Background: While enrichment of terminologies can be achieved in different ways, filling gaps in the IS-A hierarchy backbone of a terminology appears especially promising. To avoid difficult manual inspection, we started a research program in 2014, investigating terminology densities, where the comparison of terminologies leads to the algorithmic discovery of potentially missing concepts in a target terminology. While candidate concepts have to be approved for import by an expert, the human effort is greatly reduced by algorithmic generation of candidates. In previous studies, a single source terminology was used with one target terminology.
Methods: In this paper, we are extending …
Sybil Defense Using Efficient Resource Burning,
2020
University of New Mexico
Sybil Defense Using Efficient Resource Burning, Diksha Gupta
Computer Science ETDs
In 1993, Dwork and Naor proposed using computational puzzles, a resource burning mechanism, to combat spam email. In the ensuing three decades, resource burning has broadened to include communication capacity, computer memory, and human effort. It has become a well-established tool in distributed security. Due to the cost attached to utilizing resource burning mechanism, these have not been popularized in domains apart from cryptocurrency.
In this dissertation, we design efficient resource burning based Sybil defense techniques for permissionless systems. As a first step, we identify existing resource burning mechanisms in literature in Chapter 2. Additionally, we enumerate numerous open problems …
